Smart Is Cheap. Verified Is Valuable.
CLAIM: An agent's value is capped not by how smart it is, but by how cheaply its output can be verified. Design for verification first and intelligence becomes a commodity input.
Nobody boards a plane because the pilot seems clever. We board because of everything built to check the pilot: the license, the checklists, the co-pilot cross-checking every callout, the simulator visit every six months. Aviation stopped betting on brilliance decades ago and became boring, which is to say safe. With AI agents, somehow, we've spent three years doing the opposite: obsessing over how clever the model is and treating the checks as an afterthought. Then we act surprised when the pilot (this time the project kind, the small trial every agent starts life as) dies the week it meets real production data.
The math that kills these projects is not deep, which is what makes it so merciless. Say each step of an agent is right 95% of the time, which is genuinely good. Now chain twenty steps together, the way any real workflow does, and your end-to-end success rate is 0.9520 ≈ 36%. Roughly one run in three comes out clean. Reliability doesn't add across a chain. It multiplies. And multiplication is brutal to numbers below 1.
Better models raise the per-step number, sure. But you don't close the gap from 36% to production-grade with model upgrades. Even a model that gets every step right 99% of the time, which no benchmark on earth will promise you, still loses one twenty-step run in five. You close the gap with verification: machinery that catches the failures before they compound.
Two families of checks
Everything I build for verification falls into two buckets, and the order matters.
Deterministic verification is the gold standard: checks that are simply true or false, no vibes involved. Schema validation (does the output have exactly the shape and fields it promised?). Invariants ("debits equal credits", "the diff touches only files in /src"). And the one I lean on hardest, re-execution in a sandbox, a sealed environment where mistakes can't touch anything real: don't ask the model if the code works, run it and watch the test suite come back green or not. None of this is glamorous. All of it is bankable.
Probabilistic verification covers LLM judges (a second model grading the first one's work), ensembles, and self-critique. It's what you reach for when the deterministic option doesn't exist: tone, summary fidelity, judgment calls. It's useful and it's improving, but it inherits the very problem it's checking: it can be wrong, confidently. Use it as a filter, never as the final gate on anything expensive or irreversible.
The craft is in a move most teams miss: reshaping the task until deterministic checks become possible. Ask for typed, structured output instead of prose. Make the agent produce a diff (just the lines it changed) instead of a whole new file. Split "decide" from "execute": the agent drafts the refund, a separate step actually pays it. Verification isn't something you bolt on after the agent works. It's a design constraint that decides what the agent is even asked to do.
Hard to do, easy to check
Some work is hard to do and easy to check. That's where agents print money today. Generating a reconciliation (the proof that two sets of accounts agree) is hard; checking that the numbers balance is trivial. Writing code is hard; running the test suite is cheap. Wherever this asymmetry exists, you can deploy an imperfect model safely, because errors are caught for pennies. I started my career as an apprentice building Swiss finance software, where "hard to produce, cheap to check" is practically the industry's definition of a deliverable. The instinct transfers directly.
Some work is hard to do and hard to check: open-ended strategy, novel design. There, agent output is still expensive to trust, and humans stay firmly in the loop. This asymmetry is the first thing I look for when deciding whether a job is worth handing to an agent. It predicts success better than any benchmark score I've seen.
Verification must have teeth
One more distinction, because it separates the demos from the systems: verification has to be enforcement. A dashboard that shows you the agent went wrong yesterday is a confession, not a control. The check has to sit inside the loop with the authority to block: the action doesn't execute, the message doesn't send, the merge doesn't happen until the gates pass. Think CI/CD (the automated gates that stop broken code from ever shipping), but for everything the agent touches.
The model proposes. The harness disposes.
In pseudo-code, the whole philosophy fits on a napkin:
def gate(action: ProposedAction) -> Verdict:
# deterministic first: true or false, no vibes
if not schema_valid(action.output): return BLOCK("malformed")
if not invariants_hold(action.output): return BLOCK("debits != credits")
if not sandbox_replay(action).ok: return BLOCK("failed re-execution")
# probabilistic last: it can block or escalate, never approve alone
verdict = llm_judge(action)
if not verdict.ok: return BLOCK(verdict.reason)
if verdict.confidence < THRESHOLD: return ESCALATE(to=human)
return ALLOW # and only now does the side effect happen
This is how I run the AI pipeline behind acted.app: output that fails its schema check is retried or dropped, and no malformed plan ever reaches a user. Note what's not in there: no logging-and-hoping, no after-the-fact review queue. The return value of the gate is the permission. Everything upstream of it is just a proposal.
Trust is a ratchet, not a switch
The question I get whenever I show someone the gate: when do you get to loosen it? Fair question, because a gate that escalates everything to a human is just an elaborate way of doing the work yourself. The answer I've settled on is the one every team already uses for people. Nobody hands a new hire production access on day one. You read their first twenty pull requests line by line, then you skim the routine ones, and a few months in you only look up when something unusual happens. Autonomy is earned through a track record, and it is earned per kind of task, never wholesale.
Agents deserve the same contract, with one advantage: their track record is machine-readable. Every pass and every block the gate records is data. After five hundred green runs of "format the weekly report", that task type graduates: spot-checks replace full review and the agent executes without asking. The same agent, same day, still hands over every database migration as a draft, because on that task type its record is thin. And the ratchet turns both ways: one failed invariant on a graduated task and it drops back to supervised until it re-earns the slack. One more rule I hold to even though it costs me speed: trust never transfers across model upgrades. A new model is a new hire wearing the old one's badge.
The part that holds its price
There is a bigger reason I keep hammering on this, and it connects to the last essay, where I argued that execution disappears into the machines and direction survives. Verification is what direction looks like on a normal Tuesday. Deciding what "correct" means for a task, encoding that decision into a gate, and standing behind whatever the gate waves through: that is ownership, written down precisely enough for a machine to enforce. The execution below the gate already belongs to the agents. The gate itself is the job.
All of this is, frankly, boring engineering. Schemas, sandboxes, replay logs, gates. Nobody gives a keynote about it. But boring is exactly what production trust is made of, and right now it's the scarcest skill in the field. Intelligence keeps getting cheaper every quarter. Verified results are the part that holds its price.