Lauren Tan says she ships 2,000 pull requests to production a month. A pull request is a proposed change to a codebase, and 2,000 a month is roughly a hundred every working day. In a guide to her agent workflow published at the end of August, she credits one thing above the rest: a verification skill that lets an agent check its own work. “It can keep going until it succeeds at its task,” she writes, “because it can now close the loop without you being the bottleneck.” Her follow-up post raised the number. She reported that August closed with 2,462 of her changes in production.
I think she is right about verification, and I think her number points somewhere her guide does not go. Agents can now check whether their own work runs. They cannot answer for it, so the limit on what an organization can safely ship is no longer output but how much someone is willing to put their name on.
The case for agentic verification starts with arithmetic. Spread across the 21 working days of August, 2,462 changes is about 117 a day. A reviewer who read every one of them in an eight-hour day, and did nothing else, would get about four minutes per change. Four minutes is enough to glance at a diff. It is not enough to run the app, click through the new behavior, or notice the case nobody wrote a test for.
So Tan moves the checking to where it can scale. Her verification skill gives the agent the tools she would use by hand: a small command-line program that drives the app, takes screenshots, records performance traces and reports what happened. The agent makes a change, runs the app, looks at the result, and tries again until the checks pass. She treats the skill as critical infrastructure and suggests putting an on-call rotation on it.
It would be easy to caricature this as taking the human out. She does not say that. Her prompts end with requests like “show me a video and screenshots as proof” and “let me review before proceeding.” She calls herself the codebase’s “gardener and maintainer,” refactoring and adding checks while the wider team lands hundreds of changes a day. Her second post opens with what she calls “the art of supervising someone smarter than you, on a codebase you haven’t written yourself, and where humans can no longer fit the entire mental model of the codebase in their head.”
That is an honest description of the job. If a person has to be the checking layer, that person becomes the ceiling, and Tan’s answer is to make checking cheap enough that it stops being one. On that point I have no argument. A team shipping agent-written code without something like her verification loop is choosing to go slower and see less.
Checking is not answering for
A verification skill is built to answer one question: does this change do what it was asked to do? That question matters, and it can be automated. A second question arrives with every change headed to production: who answers for it if it turns out to matter?
The two feel like one because, for most of software’s history, the same person carried both. The engineer who wrote a change also tested it, and their name on the review made them the one to call when it broke. Agents pull the questions apart. Checking moves into the loop. Answering stays where it was, with a person or an institution, and it does not scale automatically because the checking did.
Tan’s own workflow shows the seam. The agent can produce a video as evidence that the new button behaves as specified. Someone still has to watch the video and decide that the evidence is enough. The loop can share the same faulty assumptions, incomplete checks, or missing edge case as the task it was given. Her pitch to non-engineers makes the seam wider: verification, she writes, lets people across an organization “validate that their changes actually work.” Confirming that a suite of checks passed and being able to answer for what a change does inside a system nobody can hold in their head are different acts. Checking is not answering for.
The same seam showed up in mathematics last month. On 21 September, OpenAI said a new internal model “has now resolved more than 100 long-standing open problems across most areas of mathematics.” It did not list the problems. In the same announcement it said nine mathematicians, including Timothy Gowers, Martin Hairer and Edward Witten, had formed an independent group hosted at the Institute for Advanced Study to advise on how the results should be assessed and released. Earlier in September, 25 Fields Medalists had signed an open letter warning that “the mass production at faster and faster pace of ‘true/false’ statements could destroy fertile ground instead of breathing life into new ideas.”
The group’s statement is precise about what it is: “Although we will give advice, we do not have decision making power at any AI company, and the responsibility for the decisions made by any company will rest with that company.” OpenAI drew the line from its side: “Importantly, the group will not be responsible for advising us on how to pace our internal progress on mathematics.”
Read those two sentences together. The rate of production is off the table. The reviewers can advise on release. Responsibility stays with the company. You could object that mathematics is the one field where this should not matter, because a proof can in principle be checked by software with no one vouching for it. For a fully formalized proof, mechanical checking can settle whether the argument follows within the formal system. OpenAI has not said that all of those claimed results have been formally checked. Even a machine-checked proof would not settle whether a result is significant, how it should be credited, or how a field absorbs a hundred of them at once, and that is the work the group was convened to do.
The budget, and what raises it
Call the amount of output a person or institution is willing and able to answer for the answerability budget. People and institutions set it: how many changes a named person can understand well enough to stand behind, and how much risk an organization will let that name carry. When generation was slow, the budget rarely showed. Writing code took long enough that understanding it came along with the work.
Now it shows. If a team’s output grows tenfold and the number of names that can honestly sign for it stays the same, one of three things happens. The signatures thin out into clicks. The work waits. Or unsigned work ships, and the organization learns later whose name it was under, usually the on-call engineer’s. The formal process can look identical in all three cases. There is still an approval button. What changes is whether pressing it means anything.
Tan’s second post reads to me as a guide to raising the budget. She asks agents to restate a problem “in your own words” before writing any code, so she can catch a misunderstanding early. She built a skill called /teach and uses it “whenever I want the agent to restate something so I can better understand and trust its work.” She warns that agents “still often state things confidently without backing it up with data or actually reading the code,” and her planning playbook tells agents that tests alone are not sufficient verification. She keeps each change “small, self-contained, and easily reviewed.”
That is where the budget can grow: explanation, smaller units of change, better evidence, and a record of why the code is the way it is. None of them makes a person think ten times faster. They reduce how much context that person has to reconstruct before making a consequential judgment. The unit of work becomes cheaper to understand.
It is also where the growth stops. None of those tools moves the name. An agent can produce strong evidence that a change behaves as specified. It cannot be the one who answers for it.
In Mona Lisas and Readymades I argued that authorship is someone taking responsibility for what a text claims. A change headed to production is a claim too: that it behaves as specified, and that someone knows what it will do. Agent Contracts asked who decides, what may go in, and what must come back from an agent’s job. The better the contract, the less of the surrounding system a human has to reconstruct every time they approve something. The contract conserves scarce human answerability. The answerability budget is the line in that job file an agent cannot fill in.
What I cannot resolve is the gap between two curves. Generation and checking improve on the schedule of models and tooling. The answerability budget improves on the schedule of human and institutional judgment: how fast someone learns a system, how far an organization trusts them, how much consequence customers and regulators will let attach to a person, a team, or a company. Tan’s methods raise it. So do smaller changes and better explanation. I do not see anything that makes it grow as fast as output.
Accountability may adapt. Institutions have assigned responsibility for other automated systems before, and some version of that will arrive here. The mathematicians’ arrangement shows what the early version looks like: advisers without power, a company that keeps the decision, and a pace no one outside the company controls. That arrangement can hold while the volume is modest. I do not know what it does with a hundred results at once, or a hundred changes a day. We know how to make generation cheaper. We are learning how to make checking cheaper. I have not seen a design that makes responsibility cheaper without making it meaningless.
For now the budget is still counted in names, and the output is counted in thousands.
Madam I’m Adam
Discover more from Adam Monago
Subscribe to get the latest posts sent to your email.
Leave a Reply