← Essays

The Autonomy Asymptote

· Autonomy · 8 min read

CLAIM: Autonomy is not a model capability. It's the product of capability × feedback-loop quality, and software engineering falls first because it's the closest thing we have to a closed world.

Most days I hand a piece of my job to a machine. I run agent loops (an AI wired into a try, check, retry cycle) at work and on my own product, and over the past year I shipped acted.app, a cross-platform app where agents wrote much of the code under my direction. A year of watching them convinced me of something that sounds pessimistic and is actually the opposite: the model is almost never the bottleneck. The harness is. The harness is everything around the model: where it runs, what it touches, how it learns it was wrong.

Give a frontier model (the biggest, newest AI money can rent) a vague task in an environment with no feedback, and it flails like an intern with no onboarding. Give a mid-tier model a tight loop (run the code, read the error, see the test results, try again) and it grinds its way to correct. Weaker model, better loop, better outcome. The autonomy was never inside the model to begin with.

The last essay was about the gate: the checks that decide whether one action is allowed to have consequences. This one is about what a gate buys you over time, why the curve it puts you on flattens out, and why one profession is going to ride that curve further than any other.

Autonomy = capability × feedback

Two factors, multiplied. Capability is what the AI labs sell, and it improves on their schedule, not yours. Feedback-loop quality is what you build: can the agent act, observe a truthful consequence, and correct itself, quickly, safely, over and over? A sandbox (a sealed-off space where code runs without breaking anything real). A test suite that won't lie. A compiler that politely refuses broken programs. That factor is fully under your control, and it's where almost all the difference between "impressive demo" and "runs unattended" lives.

Strip away the vendor slide decks and every agent that works is the same four lines:

pseudo-code · where autonomy actually lives
while not done:
    action  = model.propose(task, context)   # capability: the labs' factor
    result  = sandbox.run(action)            # a consequence, not an opinion
    context = update(context, result)        # truthful, machine-readable
    done    = checks.verify(result)          # "done" someone had to define

Three of those four lines are harness, not model. That ratio is the post.

Multiplication also explains why the two factors don't trade evenly. Zero feedback times any capability is zero: the most brilliant model in the world, working blind, is a very expensive random walk. The reverse does not hold. A modest model in a loop that never lies climbs, one corrected attempt at a time, and ends up ahead of a much stronger model that had to guess.

This also explains the embarrassing gap everyone notices: agents crush coding benchmarks (the standardized exams used to score models) and then can't survive a week of real engineering unattended. Benchmarks ship with the loop pre-built: the task is defined, the check runs itself, the verdict is instant. Reality doesn't. The job of an engineer who works with agents, as far as I can tell, is to take messy reality and force it into benchmark shape. Define done, make done checkable, wire the consequences back in.

Counting nines

The path to full autonomy is an asymptote: a curve that keeps climbing toward a ceiling it never touches. Getting from 90% reliable to 99% doesn't cost what the first 90 did. It costs roughly ten times more, and the next nine costs ten times more again. Demos live at 90. Production lives at 99.9 for anything that touches money. Each nine is bought with unglamorous material: tighter verification, an agent allowed to touch exactly one repo (one codebase), a rollback (an automatic undo) that fires before anyone notices.

Most "we tried agents and it didn't work" stories I hear are teams who assumed the distance from 90 to 99.9 was a model upgrade. That distance is an engineering budget, and somebody has to approve it.

I paid that budget myself on acted.app. The import that turns a shared video link into a collection entry worked flawlessly in every demo I gave. Then production introduced it to the long tail of real links: tracking redirects, regional URL variants, posts deleted between save and import. Hardening that last sliver took longer than writing the feature, and none of it was clever. Each failure that reached production became a test case, so the next attempt, mine or an agent's, had to pass against the real long tail instead of my demo. That is what a nine looks like up close: not a smarter model, a longer list of things the loop refuses to let through.

diagram · each nine costs ten times the last
Each nine of reliability costs ten times the engineering effort of the last, and the curve never reaches 100 percent.100% (never arrives)90 · demo99 · pilot99.9 · productiontouches money1×10×100×engineering effort, log scalereliability

One thing the diagram hides: the curve is per task, not per agent. The trust ratchet from the last essay is exactly the mechanism that moves one kind of task along it. "Format the weekly report" graduates to unsupervised after five hundred green runs while "write a database migration" stays at the draft stage, and the same agent sits at two different points on the curve on the same day. Ask "how autonomous is our agent?" and the honest answer is a list: this task type at 99.9, that one at 90, this one nowhere near, because nobody has defined what done means for it yet.

Why software falls first

Here's what excites me. Perfection (actual, verified correctness) is only reachable in closed worlds: domains where the rules are explicit and the outcomes are checkable. Chess is closed, which is why machines reached superhuman play decades ago. Marketing strategy is open: "did it work?" takes a quarter and a debate.

Two things decide how closed a job is. How long the world takes to answer, and who can read the answer. Chess answers in milliseconds, and the answer is a number. A marketing plan answers next quarter, and the answer is a meeting.

diagram · how closed is the world
how closed is the worldautonomy arrives here firstchesssoftwarecompiles · tests · replaysbookkeepingdoes it balancecustomer supportdid they come backlegal draftingdid it hold upmarketing strategya quarter and a debatea machinea persona meetingsecondsminutesdaysmonthsa quarterhow long the world takes to answer →who can read the answer

Software engineering is the most closed open-world profession we have. Code compiles or it doesn't. Tests pass or they don't. Types check, and the same input replays to the same output in a sandbox, every time. Better still, the feedback is machine-readable: the agent can consume its own consequences without a human translating. No other knowledge profession comes close to that density of cheap, truthful feedback. Bookkeeping comes nearest, and it's no accident that it's the other place where agents already earn their keep.

So the conclusion writes itself: software engineering will be the first knowledge profession where machine autonomy approaches perfection. Not because code is easy, but because correctness in code is unusually checkable. Other professions will follow in order of how closed their feedback loops can be made. (Notice that's an engineering property, not a prestige property. Some very prestigious jobs have terrible feedback loops. I'll resist naming them.)

And "can be made" is doing real work in that sentence. Closedness is not fixed. Every profession that wants agents will spend the next few years making itself more checkable: structured output instead of prose, decisions split from executions, results a machine can grade. The professions that manage it move down the list. The ones that can't, or won't, keep their humans, and pay for them.

The open edge

The asymptote has a ceiling for a reason, and in software the reason is not the code. Tests only check what somebody thought to write down. A sandbox replays what somebody thought to simulate. Every closed loop is closed around a specification, and the specification is where the open world leaks back in. The bugs that survive at 99.9 are almost never "the code did the wrong thing". They are "the code did exactly what I asked, and I asked for the wrong thing". No amount of feedback catches that, because the feedback was built from the same misunderstanding.

That's the honest limit of the argument. Autonomy in software approaches perfection at executing a spec. It does not approach perfection at having the right spec, and the two are easy to confuse when the loops are humming and every run comes back green.

What's left uphill

If writing correct code stops being scarce, what is? Everything in the loop that defines it: deciding what "done" means, encoding that as checks, choosing what's worth building at all, and owning the consequences when the failure lived in the spec rather than the code. The work migrates from writing the answer to writing the question precisely enough that a machine can answer it.

In Ten Years Out I argued that execution disappears into the machines and direction survives. This is that argument run on a single profession at close range: the typing is execution, the spec is direction, and the loop is the line between them.

That's no demotion. Specification was always the hard part; we just hid it inside the typing. The typing is leaving. The hard part stays. On the good days, when the loops hum, it's the only part left on my desk. I wouldn't trade it back.

Thoughts, disagreements, better arguments? Reply below or on LinkedIn. I read everything. New essays show up in the RSS feed first.

Reply privately

Stored privately, used only to reply. Privacy.