Skip to content
← Back to blog
·8 min read

The Driver Is a Script, Not an Agent

aidevopsmanagement
The Driver Is a Script, Not an Agent

Autopilot, Part 2 of 3. A series for engineering leaders on running coding agents with nobody watching. Part 1: Every Prompt Assumes Someone Is There. Part 3: Autopilot Built Itself in a Day.

In Part 1 I showed how we got one agent to finish one piece of work with nobody answering its questions. That gets you exactly one item. The agent stops at the release boundary, writes a plan, and ends.

Something has to start the next one. Something has to open the pull request, wait for CI, merge, release, and pick the next item. Around the clock.

The fashionable answer is "another agent". An orchestrator agent that manages worker agents. I went the other way. The thing that runs the loop is a plain script. No model, no prompt, no judgment. And I'd recommend the same split to any engineering org thinking about this.

Creative work for agents, control loops for code

Here's the principle. An agent is good at the part of the job where there are many acceptable answers: how to structure a module, what to name a test, how to phrase a migration. A control loop is the part where there's exactly one right answer and getting it wrong is expensive: did CI pass, is a release already running, has this item cost too much, should we stop.

Put a model in charge of the second category and you've made your safety mechanism non-deterministic. You can't unit-test it properly, you can't reason about it at 3 AM, and you definitely can't explain it to an auditor.

So the driver is a script that runs on a timer. Every ten minutes a scheduler calls it, it moves each item in flight one step, saves its state, and exits. Nothing waits. Nothing holds a connection open for an hour. If the machine reboots mid-run, the next tick picks up where the last one left off.

The loop, in four steps

  1. Pick. Ask the backlog for the next ready item.
  2. Build. Create an isolated copy of the repo (a git worktree), and run the agent there, unattended, until it reaches the release boundary or parks the item.
  3. PR and CI. Push the branch, open the pull request. Green checks put it in the release queue. A red check relaunches the agent once with the failing output. A second failure parks the item with the reason.
  4. Merge and release. One at a time, never while another release is running. Merge, write the release notes, publish, close the work item. Then pick again.

At most two items build in parallel. Only one ever releases.

Each of those steps encodes a decision your organization should make on purpose. Let me go through the three that matter most.

Picking is not promoting

"What's next?" used to be answered by me reading the backlog, or by an agent guessing. The agreed order lived in a chat message.

Now each backlog item can carry a priority, and a command returns the next ready item: lowest rank among the items whose dependencies are all done, unranked items last, ties broken by file name. Same input, same answer, every time.

The important part is what that command can't do. It only reads. It cannot move an item into the backlog, change a rank, or add a dependency. The driver can park what it's building and nothing else.

That's deliberate. Deciding what counts as agreed work stays a person's job. An autonomous system that can also decide what to work on next will, sooner or later, feed itself. Every finished piece of work generates follow-ups, and an unattended loop that promotes its own follow-ups is a machine for growing scope. My own project has the scars: it once went 11 days with zero product backlog items moving while self-repair work got promoted in one to three days.

So the rule is simple. The driver may suggest. Only a human promotes.

The release is one lane

This one I learned the expensive way, on October 1st.

I had four agent sessions running in four separate worktrees. That part worked beautifully: no session touched another's files. Then they all finished and queued at the same door. One release was already in flight in the main checkout, visible only as an unpushed local tag. Its uncommitted files blocked the next release script. A version number was already taken. The machine's load average hit 67 and the test suite started timing out. That same day, two releases, v1.40.1 and v1.40.2, went out 30 minutes apart.

Worktrees parallelize the work. They do nothing for the release. A release is a shared, ordered, mostly irreversible thing: a version number, a tag, a published package, a deploy.

So the driver serializes it. One queue. It checks that no release workflow is running before starting one. And if anything fails mid-release, it does not retry:

A half-finished release is the one state a loop must not paper over.

Retries are great for flaky reads. For a release that may already be half published, a retry is how you get two broken versions instead of one. The driver halts the queue and leaves it for a person.

We also took the human out of the release paperwork. Every release used to need someone to choose the version bump and hand-write a changelog entry. Earlier that day a release had failed on exactly that: two small chore merges nobody declared. Now the bump comes from what was merged (feature means minor, fix means patch) and the entry is generated and dry-run checked. The first release to use it was also the first that day that didn't need a hand-fixed entry.

Write the stop rules before you buy the autonomy

Once no one is watching, the stop rules are your safety. These are ours:

  • A kill switch anyone can pull. A file named STOP in the repo's git directory halts every tick until someone removes it. No dashboard, no login. touch it and walk away.
  • A cost cap per item. An item that spends more than the configured amount is stopped and parked.
  • Two strikes on CI. One automatic retry with the failure attached. The second failure parks the item with the reason.
  • Back off on rate limits. If GitHub or the model provider says slow down, wait five minutes, doubling up to two hours. Detection is textual and crude on purpose. A false positive costs a few minutes.
  • Halt, never retry, during a release.

And a dry-run mode that prints every command it would run and only executes the reads. You want that before the first real night.

Two things the script refuses to decide

This is the part I'd want every engineering leader to steal.

The driver ships with two settings deliberately empty, and until both are set it does nothing at all.

What the agent is allowed to do on the machine. The agent command has no default. Which permission mode, which tools, which network access: that's a security decision about your environment, and a framework has no business making it for you. So it doesn't.

Turning it on. The schedule is yours to enable. The repo ships documentation for cron and launchd and not one line that installs either.

That's the right shape for accountability. The framework can be opinionated about process. The person who owns the environment owns the blast radius, and has to say so explicitly.

What to do Monday morning

  1. Draw your loop and color it. For each step in your path to production, mark whether it needs judgment (agent) or has one right answer (code). Anything in the second category that an agent currently does is a risk.
  2. Find your release lane. If two teams or two agents can release at once, you'll find out the way I did. Make it one queue.
  3. Separate suggest from promote. Whatever picks the next work item for an agent should be read-only. Promotion is a human decision with a name attached.
  4. Agree the stop rules first. Kill switch, cost cap, failure limit, backoff, release halt. If you can't state them, you're not ready to run unattended.
  5. Keep permissions and the on switch human. Explicitly, in writing, with an owner.

In Part 3 I'll share the honest field report: we estimated this whole thing at weeks, agents built it in about 90 minutes, and I'll go through what that does and does not prove.

The AIDLC changelog post for the driver is The backlog works itself. The ranking and release pieces are The next item is not a guess and The release writes its own entry.