Skip to content
← Back to blog
·7 min read

Every Prompt Assumes Someone Is There

aidevopsmanagement
Every Prompt Assumes Someone Is There

Autopilot, Part 1 of 3. A series for engineering leaders on running coding agents with nobody watching. Part 2: The Driver Is a Script, Not an Agent. Part 3: Autopilot Built Itself in a Day.

Every engineering leader I talk to has the same question this year. When can the agents just run? Not pair with a developer. Run. Pick up a ticket at 2 AM, build it, test it, open the pull request, and leave a clean trail for the morning.

The usual answer is "when the models get better". I think that answer is wrong, and I have the receipts.

On October 2nd I tried it on my own project. I took one small item from the AIDLC backlog and told an agent to build it with nobody answering. Same item, two machines, two separate runs. I asked each agent for one thing besides the code: write down every place where the instructions told you to stop and ask a person.

Each run came back with 13.

The model was fine. The process wasn't.

None of the 13 stops were the model failing. The model wrote the code, wrote the tests, and passed them. Every stop was the process assuming a human was sitting there.

Some were obvious. An entry menu that waits for a choice. A question protocol that says "never assume an answer". A requirements phase that ends with "the user has confirmed the requirements are complete".

Some were sneakier, and those are the ones that should worry you.

The confirmation that was never checked. That requirements sign-off? The agent treated it as met and moved on. Nothing stopped it, because the gate that enforces phase transitions doesn't check for a human confirmation. It can't. It's a sentence in a document. I had been reading that step as a control. It was a polite request.

The default that would have quietly cut a phase. The project's default template runs only implementation and testing. The prompt asked for requirements too. One agent noticed and picked a bigger template. If it had followed the default, the requirements phase would have vanished and nobody would have known until review.

Two agents, one instruction, two outcomes. One run linked the backlog item to the new work by hand. The other skipped that step and left the item orphaned. Same instructions, same model, different result. That's not a model problem either. It's an instruction that relies on judgment where it should be a command.

Put the two lists together and you get 20 distinct kinds of stop. That list is the honest size of the gap between "AI-assisted" and "autonomous" in a real codebase. I'd bet yours is longer.

Three outcomes, and nothing else

Here's what we did with the list. Every stop gets exactly one of three outcomes:

  1. Default and log. Take a named, safe default, write it down, keep going. The entry menu routes to what the prompt asked for. An open question takes the recommended option. The end of a phase continues into the next one.
  2. Hard stop. Write it down, park the item with the reason, release the lock, end the session. Try no other route to the same outcome.
  3. Release boundary. The normal end of an unattended run. The agent finishes testing, writes the release plan, and stops. Something else ships it.

Then one more rule, the one I'd put on a poster:

A situation no entry covers is a hard stop. Guessing is what this mode exists to prevent.

That flips the usual failure mode. Most agent setups fail open: anything you didn't anticipate, the agent improvises. This fails closed. The cost is that some runs park items a human would have waved through. I'll take that trade every time.

Where the line goes

The hard stops are the part worth arguing about with your team, so here are ours. An unattended agent always stops when the next step would:

  • cross a prohibition in the project's constitution,
  • delete or overwrite data the session did not create,
  • touch authentication, secrets, permissions or CI tokens,
  • or grow the scope beyond what the backlog item said success looks like.

That last one matters more than it looks. An agent that answers its own questions will happily answer them in the direction of more. "Should we also handle the JSON case?" Sure, why not. So both sign-offs, requirements and design, became scope tests. If the agent's own requirements outgrow the item, that's a hard stop, not a default.

Notice what's not on the list: code style, library choice, test structure, naming. Those are defaults. Your developers make those calls a hundred times a day without a meeting, and so can an agent, as long as it writes them down.

The decision log is the product

Every default the agent takes goes into one file per piece of work, with the source instruction, what it decided, and why that was safe. Every question it answered for itself carries a marker: (unattended default).

This is the part I'd push hardest with leadership. You don't need to watch an autonomous run. You need to be able to audit one in five minutes. A reviewer opens one file and sees every decision nobody made. That's a better trail than most human-driven changes leave behind.

It also turns trust into something you can measure. If a month of logs shows the agent's defaults were right 50 times out of 50 for a class of decision, that class can stay a default. If one went wrong, you've found a new hard stop, with evidence.

A rule, not a switch

The obvious way to build this is a flag: UNATTENDED=1, and every command checks it and behaves differently. We didn't do that, on purpose.

A switch in every command is a new code path in every command, forever. My project already learned that lesson the hard way. In its first nine days it shipped 11 features and zero bug fixes. In the next five, one feature and seven bug fixes. Most of the damage traced back to features that watched the lifecycle rather than served a user. Each one added a check every future change had to pass.

So unattended mode is one rule file the agent reads. The commands don't know it exists. The code changes were small: two commands for the steps we saw agents fumble, linking a backlog item and parking one. Those became deterministic, so no agent ever has to improvise them again.

Not everyone agrees, and that's fine

AWS's own AI-DLC methodology takes close to the opposite position: AI creates the plan, seeks clarification, and defers critical decisions to humans. Plenty of people argue copilot still beats autopilot in 2026 because background agents lose the human safety net.

I don't think they're wrong. I think they're answering a different question. "Should a human approve every stage?" is the wrong frame. The useful question is which decisions need a human, and you can't answer that until you've listed them. Most teams never have. The 13 stops were invisible to me until an agent tripped over every one of them.

What to do Monday morning

  1. Inventory your stops. Grep your agent instructions, runbooks and PR templates for "ask", "confirm", "wait", "approve". Each hit is a place your process assumes a person.
  2. Run one dry run. Pick a small, boring backlog item. Tell an agent to build it and log every stop instead of asking. It costs an afternoon and gives you a real number.
  3. Write your hard stops down. Four or five classes, agreed with your security and platform leads. Everything else gets a named default.
  4. Make the log non-negotiable. No decision log, no merge. That's the deal that makes autonomy reviewable.
  5. Turn fumbled judgment into commands. When two runs do the same step two ways, the step should not be left to judgment.

In Part 2 I'll cover what starts the agent, ships its work, and knows when to stop: a loop we deliberately did not hand to an agent.

The AIDLC changelog post behind this one, with the full list of stops, is Nobody is watching.