Skip to content
aisoftware-developmentqaclaude

Why People Still Decide the Outcome of AI Delivery

Matt Hammond 7 min read
Two colleagues discussing a plan at a whiteboard while a team works at desks behind them

When Claude Code builds inside our delivery harness, writing the code stops being the slow part. In a change we measured end to end, the moments that shaped the result were three business questions, found during planning and answered in three minutes after waiting almost a week. The business analyst, the QA analyst, and an experienced engineer are where the outcome is decided, and with the harness they account for most of the people’s time.

We recently measured a single change from the first prompt to the last merge, and wrote it up in full as Anatomy of an AI Change. Two complex, linked user stories on one of the booking platforms we build went from open questions to two merged pull requests in 15 hours 35 minutes, overnight included. Claude worked for about three hours of that, and the person driving it spent an estimated 25 to 50 minutes hands-on.

It would be easy to read those figures as Claude doing the work, and by volume of output it did. When we went back through the session, though, the decisions that changed what was built were all made by people. This post is about those decisions, and what they mean for anyone planning AI-augmented delivery.

Where did the decisive moments come from?

They came from planning, almost a week before the build. Our planning skill read the first story against its own discussion thread and the existing code, and found three problems:

  • Two acceptance criteria still offered an option that the discussion thread had recorded as dropped.
  • The existing update endpoints next to the new one checked that a user had the edit permission, but not which sites they were responsible for.
  • Migration from the legacy system would have lost settings that staff had chosen by hand on the large majority of records.

The skill did not resolve any of these itself. Its rules are to ask before writing to the work-item system and never to edit a story’s acceptance criteria. So it raised a business analyst task with three questions, each set out with the evidence and the options.

The story then waited almost a week for answers. In the session, the person driving it was also acting as product owner, and the three decisions took three minutes.

Why did three minutes of decisions matter more than hours of code?

Each answer changed what was built, and none of them could have been worked out from the code:

QuestionDecisionWhat would have happened without it
Is the dropped option still in scope?NoThe model would most likely have built it, because the acceptance criteria still asked for it
Are thresholds a percentage or a number?A number of places remainingA guess either way, with tests written to match the guess
Are legacy settings carried across at migration?YesMost migrated records would have arrived on the default, and staff would have lost the settings they chose

In our experience this is where AI-built work goes wrong when nobody asks. The model is good at building what it is told, and the tests it writes alongside the code agree with the code. A contradictory requirement therefore becomes a feature that works, passes its tests, and does the wrong thing.

The three minutes were short because the questions arrived prepared, not because the answers were easy. Someone with the authority to decide still had to decide.

What does the experienced engineer actually do?

The person driving the session knows the codebase, the client, and our delivery process, and that knowledge sits behind every short answer in the prompt log. They sent 13 prompts across the whole session, and seven of them were ‘yes’. Each of those short answers carried a judgement:

  • Approving test cases means knowing whether they test the right things.
  • Choosing to fix a regression in the pull request rather than raise a bug is a judgement about scope and risk.
  • Reviewing the code before it merges, including whether it fits the system’s architecture. Every pull request gets an expert human review.
  • Merging is taking responsibility for the change.

Claude asked seven questions during the session rather than guessing. Our build skill has a rule to pause on ambiguity and ask, and the engineer’s job is to know the answer, or to know who does.

We should be open about one detail. The engineer reviewed the second story’s draft test cases in the session, which took six minutes. The first story’s were approved within seconds. All of them are reviewed again in the manual session.

The harness itself is the same expertise written down. Its rules and guard tests exist because an experienced team was caught out by each defect once. We would not expect the same result from the same tools in the hands of someone who could not judge the answers.

Why is the manual session the one step we would not automate?

Every approach we compared ends with a manual session of two to three hours per story. A QA analyst and a business analyst use the feature on the test environment and ask whether it makes sense to the people who will use it. The tester’s own pass happens within that session.

No acceptance criterion fully describes what makes a feature usable, so no automated test can confirm it. That is why the session stays the same size whichever way the software is built.

It also changes the shape of the work. We estimate the manual session is about a tenth of the people’s time when the same change is built by hand. With the harness it is about nine tenths. The hours saved fall on the developer and on test design, while the judgement-led work stays where it was.

The manual session grows from about a tenth of people's time by hand to about nine tenths with the harness.At the midpoint estimates, the 4 to 6 hour manual session with a QA analyst and a business analyst is about 9% of people's time by hand, about 27% with Claude Code and no harness, and about 89% as delivered with the harness.Developer and QA analystManual session, QA analyst and BA (allowance)By hand, no AIManual session about 9% of 42 to 72.5 hClaude Code, no harnessManual session about 27% of 12.5 to 24 hAs deliveredManual session about 89% of 4.4 to 6.8 hHatched: estimated. Midpoint shares.
Shares of estimated people's time at the midpoint of each range. The manual session is an allowance in every approach. Source: Talk Think Do estimates.

For these two stories the session had not happened at the time of writing. The change is merged, but it is not signed off until people have used it.

Why is cognitive load the limit?

Given the figures, most of the productivity comes from parallel workstreams. The engineer’s hands-on time was under an hour across a change that took 15 hours 35 minutes end to end, while Claude worked for about three hours. That leaves an engineer free to drive several changes at once, each at its own pace.

How many they can run well depends on more than the hours available. Each workstream still needs its questions answered, its test cases judged, and its code reviewed, and every switch between them costs attention. Cognitive load, meaning how much one person can hold in mind and judge well at once, is the limit on top of time.

What does this mean if you are planning AI-augmented delivery?

The constraint has moved. When a build takes an afternoon, a week spent waiting for a decision dominates the timeline, and no tool shortens it. We find these are the things worth checking, whether you run delivery yourself or buy it from a supplier:

  • Who raises the questions, and when? Ask to see how open requirements are surfaced before a build starts, and whether they arrive with evidence and options.
  • Who answers them, and how quickly? Nominate someone with the authority to decide, and agree how fast decisions will come back.
  • Is business analysis and QA time in the plan? A plan that shows developer hours falling and nothing else is missing the work that now takes most of the people’s time.
  • Who reviews the code, and against what? Every pull request should get an expert human review that includes a check against the architecture.
  • Who signs the work off? Sign-off should sit with people who have used the feature on a test environment, not with a green pipeline.

Our guide to the AI-augmented software development lifecycle covers where each role fits across the whole process. For why we keep people in the loop at every stage, see why we don’t let AI ship code unsupervised.

Frequently asked questions

Does AI-augmented delivery still need a business analyst?
Yes. The model builds quickly from whatever it is given, so an unanswered or contradictory requirement turns into working code that does the wrong thing. A business analyst makes sure the questions are asked and answered by someone with the authority to decide, before the build starts.
What does a QA analyst do when AI writes the tests?
They judge whether the feature makes sense to use. Automated tests prove the code matches the acceptance criteria, but no criterion fully describes what makes a feature usable. In our process a QA analyst and a BA run a manual session on the test environment for every story, whichever way it was built.
Why does an experienced engineer need to drive an AI coding session?
Because the short answers carry the weight. Approving test cases, choosing whether to fix a regression or raise a bug, and merging a change are all judgements about scope and risk. Someone who cannot judge the answers gets the same speed with none of the safety.
What slows AI-augmented delivery down now?
Waiting on decisions. In the change we measured, the build took an afternoon and an evening, and the business questions it depended on had waited almost a week. The speed at which a client can answer well-prepared questions is now the main constraint on throughput.
How many workstreams can one engineer run with AI coding agents?
There is no fixed number; it depends on the engineer and the work. In the change we measured, the engineer's hands-on time was under an hour across 15 hours 35 minutes, which leaves room for several changes at once. Cognitive load sets the limit, because each workstream still needs decisions, test case judgement, and code review.

Ready to transform your software?

Let's talk about your project. Contact us for a free consultation and see how we can deliver a business-critical solution at startup speed.