Autopilot Built Itself in a Day. Here Is What That Proves.

Autopilot, Part 3 of 3. A series for engineering leaders on running coding agents with nobody watching. Part 1: Every Prompt Assumes Someone Is There. Part 2: The Driver Is a Script, Not an Agent.
When I wrote the backlog item for the autopilot, I sized it honestly. Weeks, in four separable changes. My recommendation on the item itself was "do later".
On October 2nd the first of those four changes opened at 16:11 UTC. The fourth closed at 17:40. In between, four releases of AIDLC went to npm, v1.42.0 through v1.45.0, the first and last published 43 minutes apart. Each change went from requirements to tested code in under an hour, most in 40 to 50 minutes. They ran in parallel worktrees and queued for one release lane, the way Part 2 describes.
Every one of the four retrospectives opens with the same line: answers are the agent's own; no person was asked.
It's a great headline. I want to spend this post on why it's the least interesting part.
Speed was never the bottleneck
Look at where time and money actually went in the dry run that started all this, the one from Part 1. The work itself, requirements through testing, took 17 minutes. Deployment took 27. Of the $7.76 that item cost, $5.16 was spent getting it released.
And the release failed twice before it tagged. Neither failure had anything to do with the change. A stale dependency folder in the main checkout. Two small merges that the hand-written changelog forgot to declare.
That's the pattern I keep seeing with engineering teams adopting agents. Writing code got cheap. Everything around the code didn't: integration, release, review, the shared state between people and systems. If your plan for AI is "developers write code faster", you're optimizing the 17 minutes and ignoring the 27.
What broke, which is the useful part
I asked the agents to be honest in their retrospectives, and they were. A few things went wrong that day, and each one is a lesson for anyone running this at scale.
We built a cost cap on a day we couldn't measure cost. One of the driver's stop rules is a per-item spending limit. Meanwhile, for the autopilot work itself, the cost meter recorded nothing usable. One agent session held locks on several pieces of work at once, so spend couldn't be attributed to any of them. The lesson was already in the project's notes from August: one session per piece of work. Autonomy doesn't make old lessons optional. It makes them load-bearing.
Tests that passed and proved nothing. All 23 of the driver's tests passed on the first run. Then the agent deliberately planted defects to check the tests would catch them, and found a hole: the dry-run test never reached a GitHub write, so a bug that treated every GitHub call as read-only would have sailed through. A separate change had the same story: a deliberately broken sort passed the ordering test because the bad comparator happened to sort that fixture correctly. If your agents only report "tests green", you're not seeing the interesting number. Ask whether anyone checked the tests can fail.
A file nobody meant to commit. CI went red once because a generated file slipped in through git add -A in a fresh worktree. Same slip in two of the four changes. The driver's own agents would hit it on every item, so it was filed as a backlog item on the spot, and fixed and released as v1.45.1 within the hour.
A worktree that couldn't see its own config. One change edited project configuration, and a different branch's test broke two hours later, because a generator read config from the main checkout, not the worktree. Parallel agents surface every hidden assumption that "there is only one copy of the repo".
None of these are AI problems. They're the same integration problems human teams have always had, showing up faster and in parallel.
What it does not prove
Here's the honest part.
The driver has never run end to end. It shipped with two settings deliberately empty: what the agent is allowed to do on the machine, and the schedule that turns it on. Both are mine to set, and as of this writing I haven't. Agents built the autopilot. The autopilot has not yet built anything.
A human still started the sessions. Agents answered every question themselves, but a person launched the work and handed it the brief. That's the gap the driver closes, and it's untested in anger.
One project, one maintainer, one day. These are good numbers from a small, well-instrumented codebase with a lot of written process. I wouldn't extrapolate them to a 200-engineer monorepo without your own dry run first.
What it does prove is narrower and, I think, more useful. When the process is written down, the stops are listed, the defaults are logged and the dangerous steps are deterministic, agents can carry a well-specified backlog item from requirements to a passing release without anyone answering a question. The bottleneck moves. It doesn't disappear.
The trap that gets worse unattended
There's one risk I'd put above all the others, and it's not a rogue agent deleting production.
It's a loop that feeds itself.
Every finished piece of work ends with a retrospective, and every retrospective suggests follow-ups. My project measured what that does. In its first nine days it shipped 11 features and no bug fixes. Over the next five, one feature and seven bug fixes, most of them caused by earlier features that existed to watch the process rather than serve a user. One of those features cost $153.65 to build. Counting the two fixes it caused, its true cost was $219.78. Its own number understated it by about a third.
Now imagine that dynamic with no human in the loop and a scheduler firing every ten minutes.
That's why the autopilot's scope says, in so many words, that the driver may write to the inbox but only a human promotes. Follow-ups are capped at two per finished item; the rest are recorded as "observed, not queued". The backlog is capped at about a dozen raw ideas. These limits look like bureaucracy. They're the thing that keeps an autonomous system pointed at your priorities instead of its own.
The job changes, it doesn't go away
So what does an engineering leader actually do when agents can carry work from ticket to release?
- Triage becomes the job. What goes into the backlog, in what order, with what definition of success. The agent will build what you agreed. Be sure you agreed the right thing.
- Write the boundaries. Hard stops, stop rules, permissions. These are policy decisions, and they need an owner with a name.
- Review the decisions, not the keystrokes. A five-minute read of the decision log beats watching a session for an hour.
- Measure true cost. Include the fixes a feature causes later. Own cost lies.
- Earn autonomy in steps. We went dry run, then a written rule, then a deterministic picker, then generated release notes, then the driver. Each step had evidence from the one before. Skipping to the last step is how you end up with a very fast way to make a mess.
There's a lively argument right now about whether copilot beats autopilot. After this week, my answer is that autonomy isn't a product feature you buy. It's something a team earns one class of decision at a time, with logs to show for it.
I'll report back after the first night the driver runs on its own. If it parks everything, that will be a result too.
Missed the start? Part 1 covers the 13 stops, and Part 2 the loop and its stop rules.