Why People Still Decide the Outcome of AI Delivery
When Claude Code builds inside our delivery harness, writing the code stops being the slow part. In a change we measured end to end, the moments that shaped the result were three business questions, found during planning and answered in three minutes after waiting almost a week. The business analyst, the QA analyst, and an experienced engineer are where the outcome is decided, and with the harness they account for most of the people’s time.
We recently measured a single change from the first prompt to the last merge, and wrote it up in full as Anatomy of an AI Change. Two complex, linked user stories on one of the booking platforms we build went from open questions to two merged pull requests in 15 hours 35 minutes, overnight included. Claude worked for about three hours of that, and the person driving it spent an estimated 25 to 50 minutes hands-on.
It would be easy to read those figures as Claude doing the work, and by volume of output it did. When we went back through the session, though, the decisions that changed what was built were all made by people. This post is about those decisions, and what they mean for anyone planning AI-augmented delivery.
Where did the decisive moments come from?
They came from planning, almost a week before the build. Our planning skill read the first story against its own discussion thread and the existing code, and found three problems:
- Two acceptance criteria still offered an option that the discussion thread had recorded as dropped.
- The existing update endpoints next to the new one checked that a user had the edit permission, but not which sites they were responsible for.
- Migration from the legacy system would have lost settings that staff had chosen by hand on the large majority of records.
The skill did not resolve any of these itself. Its rules are to ask before writing to the work-item system and never to edit a story’s acceptance criteria. So it raised a business analyst task with three questions, each set out with the evidence and the options.
The story then waited almost a week for answers. In the session, the person driving it was also acting as product owner, and the three decisions took three minutes.
Why did three minutes of decisions matter more than hours of code?
Each answer changed what was built, and none of them could have been worked out from the code:
| Question | Decision | What would have happened without it |
|---|---|---|
| Is the dropped option still in scope? | No | The model would most likely have built it, because the acceptance criteria still asked for it |
| Are thresholds a percentage or a number? | A number of places remaining | A guess either way, with tests written to match the guess |
| Are legacy settings carried across at migration? | Yes | Most migrated records would have arrived on the default, and staff would have lost the settings they chose |
In our experience this is where AI-built work goes wrong when nobody asks. The model is good at building what it is told, and the tests it writes alongside the code agree with the code. A contradictory requirement therefore becomes a feature that works, passes its tests, and does the wrong thing.
The three minutes were short because the questions arrived prepared, not because the answers were easy. Someone with the authority to decide still had to decide.
What does the experienced engineer actually do?
The person driving the session knows the codebase, the client, and our delivery process, and that knowledge sits behind every short answer in the prompt log. They sent 13 prompts across the whole session, and seven of them were ‘yes’. Each of those short answers carried a judgement:
- Approving test cases means knowing whether they test the right things.
- Choosing to fix a regression in the pull request rather than raise a bug is a judgement about scope and risk.
- Reviewing the code before it merges, including whether it fits the system’s architecture. Every pull request gets an expert human review.
- Merging is taking responsibility for the change.
Claude asked seven questions during the session rather than guessing. Our build skill has a rule to pause on ambiguity and ask, and the engineer’s job is to know the answer, or to know who does.
We should be open about one detail. The engineer reviewed the second story’s draft test cases in the session, which took six minutes. The first story’s were approved within seconds. All of them are reviewed again in the manual session.
The harness itself is the same expertise written down. Its rules and guard tests exist because an experienced team was caught out by each defect once. We would not expect the same result from the same tools in the hands of someone who could not judge the answers.
Why is the manual session the one step we would not automate?
Every approach we compared ends with a manual session of two to three hours per story. A QA analyst and a business analyst use the feature on the test environment and ask whether it makes sense to the people who will use it. The tester’s own pass happens within that session.
No acceptance criterion fully describes what makes a feature usable, so no automated test can confirm it. That is why the session stays the same size whichever way the software is built.
It also changes the shape of the work. We estimate the manual session is about a tenth of the people’s time when the same change is built by hand. With the harness it is about nine tenths. The hours saved fall on the developer and on test design, while the judgement-led work stays where it was.
For these two stories the session had not happened at the time of writing. The change is merged, but it is not signed off until people have used it.
Why is cognitive load the limit?
Given the figures, most of the productivity comes from parallel workstreams. The engineer’s hands-on time was under an hour across a change that took 15 hours 35 minutes end to end, while Claude worked for about three hours. That leaves an engineer free to drive several changes at once, each at its own pace.
How many they can run well depends on more than the hours available. Each workstream still needs its questions answered, its test cases judged, and its code reviewed, and every switch between them costs attention. Cognitive load, meaning how much one person can hold in mind and judge well at once, is the limit on top of time.
What does this mean if you are planning AI-augmented delivery?
The constraint has moved. When a build takes an afternoon, a week spent waiting for a decision dominates the timeline, and no tool shortens it. We find these are the things worth checking, whether you run delivery yourself or buy it from a supplier:
- Who raises the questions, and when? Ask to see how open requirements are surfaced before a build starts, and whether they arrive with evidence and options.
- Who answers them, and how quickly? Nominate someone with the authority to decide, and agree how fast decisions will come back.
- Is business analysis and QA time in the plan? A plan that shows developer hours falling and nothing else is missing the work that now takes most of the people’s time.
- Who reviews the code, and against what? Every pull request should get an expert human review that includes a check against the architecture.
- Who signs the work off? Sign-off should sit with people who have used the feature on a test environment, not with a green pipeline.
Our guide to the AI-augmented software development lifecycle covers where each role fits across the whole process. For why we keep people in the loop at every stage, see why we don’t let AI ship code unsupervised.