We have harnesses.. they can do most anything. What is stopping us now? Scale. We need a harness that can track an ever expanding and accelerating sphere of effort while remaining legible to operators. and to get this done we have nothing but the cosmic railgun of self improvement.
Because suppose you have a harness, and all you care about is that it scales forever i.e. it is always getting better at writing more kinds of correct software quickly. And suppose you develop it in a non-self-improving way, bringing an external system like a human or another harness to bear. Eventually, the amount of work the external harness must track about its own design will exceed its scale. If this happens before the harness you're building is self improving, then you lose. If it happens after, then you can migrate to the new harness and continue scaling from there. So self improvement is a necessary condition.
But self improvement endpoints are strongly path dependent (I've noticed) so it seems worthwhile to pause and ask what we want, for this to work, so we can take the biggest shot possible.
I think the big scale flavored problems are:
* compute is expensive;
* operator attention is scarce
* verification is complicated
* crime is bad (downtime is also a crime)
I'm gonna make a big leap and say we want:
Accurate estimates: prevent crime by noticing danger, make work cheap by prioritizing it correctly, prioritize work by estimating the difficulty of its dependencies. How? work items must be estimated and risked with confidence and basis, and estimation error must be tracked and minimized over the system's lifetime.
Cheap verification: CI moves to a 'report dont proscribe' regime; agents run jobs and report to pipelines; the auto-run pipeline must go; they were built for an age of machine verification and expensive developer eyeballs... having an e2e run that takes an hour used to be totally fine but now it's a non-starter. Waiting till CI completes to throw inference at the test process wastes valuable CPU and critical-path time. But! Verification is more important than ever; so clearly we need lots of tests. That's why we let the agents decide what tests to run. This also gets them a first crack at any problems that come up, with caches warm and ready.
rent-paying UX: user facing info must be both semantically and visually dense and layered. UX is the primary channel between agent and user, and it is visual/interactive. The agent and user must aggressively converge on the same UX. This is the hardest problem, and I have to think about it more
Comments