Anatomy of an AI Change: How Long It Really Takes
AI isn't instant or free. We measured one real change end to end: 15h 35m elapsed, about three hours of Claude working, and about $71 of AI usage.
AI isn't instant and it isn't free. To estimate accurately, we need to know how long AI-built work really takes and what it costs, so we measured one real change from the first business question to the last merge. Two complex, linked user stories on a live booking platform took 15 hours 35 minutes, overnight included. Claude worked for about three hours of that, an engineer spent under an hour on it, and AI usage came to about $71. Waiting for business decisions took far longer than the build, and running tests took more of Claude's time than writing the code.
Two stories tend to get told about AI and software. In one, AI builds an application in minutes for next to nothing. In the other, it produces plausible code that falls apart in production. Neither matches what we see in our own delivery. We are getting very good outcomes, but they take real time, real money, and real people, and we want to understand exactly how much of each.
That matters because of a problem we set out in our Q2 2026 report: when productivity keeps changing, estimating the next piece of work is hard, and you cannot predict what you have not measured. We also promised a closer look at the harness, the rules, checks, and people that surround the model in our delivery. A single real change, followed from start to finish, lets us do both.
This issue follows two complex, linked user stories on one of the booking platforms we build for clients, through planning, build, quality assurance (QA), and merge, in one Claude Code session. We have anonymised the client and the detail of the work, and rounded the counts that could point to a particular codebase. The durations, the costs, and the shape of the work are as they happened.
A note on the numbers. Claude’s working time and the cost are measured. The engineer’s time is our estimate from the session. The comparisons with other ways of working are estimates too, and we say so wherever they appear.
What did we build?
We built a way for staff to control how a bookable item’s availability is shown to the public, for example as limited or full, with a different threshold for each type of item. A second story then extended it to two edge cases the first had left open.
It is ordinary product work, which is what makes it a useful measure. Behind the one new setting sat:
- changes to the database, and a migration to apply them
- more than a dozen of the database routines the application calls
- new admin screens and settings
- a change to the public availability feed
- a mapping that carries staff’s existing settings across from the legacy system being replaced
How long did it really take?
The session started mid-afternoon and ended with the second merge early the next morning, 15 hours 35 minutes later. Claude worked for about three hours of that. The rest of the time it was waiting: on the engineer, on the automated checks that run when code is submitted, or on the night.
Claude worked in bursts. The two builds were the long stretches, at about an hour each. Everything else, from turning the business answers into a plan to writing the test cases and running QA, took minutes. The first story was ready for review two and a half hours after the first question and merged half an hour later. The second followed in the evening and merged the next morning.
What did the engineer actually do?
The engineer did very little typing and a great deal of judging. The engineer driving the session sent 13 prompts across the whole day, and seven of them were ‘yes’. Claude asked seven questions of its own rather than guessing, and most were answered in under a minute.
We estimate the engineer’s hands-on time at 25 to 50 minutes. That went on reading Claude’s summaries, making three business decisions, reviewing a dozen draft test cases, and merging two pull requests.
Where does Claude’s time go?
Most of it goes on checking its own work. Running tests took more of Claude’s time than anything else, including writing code, and a real share went on the development environment itself when it misbehaved.
Where did the people make the difference?
By volume, Claude did most of the work. The decisions that shaped the result were made by people, and our process is built so that they have to be.
A week of waiting, then three minutes of decisions
Almost a week before the build, our planning step read the first story against its own discussion thread and the existing code, and found three problems. Two acceptance criteria still asked for an option the business had already dropped. The existing ways of updating those settings checked that a user had the edit permission, but not which sites they were responsible for. And migration would have lost settings that staff had chosen by hand on most records.
Planning did not guess at any of these. It raised them with a business analyst (BA) as three questions, each with the evidence and the options, and the story waited almost a week for answers. When they came, in the session itself, the three decisions took three minutes.
| Question | Decision | What it changed |
|---|---|---|
| Is the dropped option still in scope? | No | The acceptance criterion that named it was removed, and the work planned for it was dropped |
| Are thresholds a percentage or a number? | A number of places remaining | Confirmed the planned design |
| Are the legacy system’s settings carried across at migration? | Yes | Added the migration mapping, so staff keep the settings they chose |
Why does the engineer’s experience matter so much?
Every short answer the engineer gave carried a judgement. Approving test cases means knowing whether they test the right things. Choosing to fix a problem now rather than log it for later is a call about scope and risk. Merging is taking responsibility for the change. We would not expect the same result from the same tools in the hands of someone who could not judge the answers.
Every pull request is reviewed by an expert engineer before it merges, including a check that the change fits the system’s architecture. The automated checks run first, so the reviewer can concentrate on design and intent.
Sign-off stays with people too. Every approach we compared ends with a manual session of two to three hours per story, in which a QA analyst and a BA use the feature and ask whether it makes sense. It is the one step we would not try to automate, and at the time of writing it had not yet happened for these two stories.
What does the harness add to the model?
The harness is our team’s delivery experience written down so that the model applies it every time. It is the AI-first evolution of Codenative, the internal toolset we have built up over several years of research and development (R&D). A capable model writes code quickly on its own. The harness makes sure that code follows the system’s own patterns, is tested at several levels, and is checked against requirements that someone has confirmed.
On this project the harness includes:
- more than 30 rule files that Claude reads before writing each part of the system
- more than 20 written procedures for planning, building, and QA
- an independent automated review of every change
- around two dozen tests that fail the build when a known rule is broken
- a check that no acceptance criterion goes untested
Several of those rules exist because the team was once caught out by the defect they now prevent.
The table shows what that meant in this session. The right-hand column is our expectation of an unguided session, not a measurement.
| Area | What happened here | What we would expect without the harness |
|---|---|---|
| Requirements | The out-of-date criterion was found and settled before any code was written | The model builds the option the business had dropped |
| Security | The gap in the existing update functions was closed in the same change, and proven with tests | A permission check at the front door looks like enough, and nobody looks one layer deeper |
| Correctness | Reading the booking logic showed that a display setting could have stopped the last places being sold | The natural implementation reuses the displayed state, and the fault appears in production as lost sales |
| Traceability | Every acceptance criterion was covered by a named test | A pull request and a commit message |
Our harness engineering guide covers the mechanics for teams who want to build the same thing.
What went wrong along the way?
Most of what went wrong was the environment. The development container restarted twice during the session, and together with a hung local proxy that cost about 20 minutes of rebuilding and investigation, plus a burst of false test failures along the way.
QA caught one real problem before the second story merged. A test from an earlier story failed because the first story had introduced a new kind of history entry that the older test did not expect. It took about ten minutes to prove and seven to fix. Separately, the build checks did not run one set of front-end tests, so a small breakage showed up only after merge and was fixed the same evening. Adding those tests to the build checks closes the gap.
How is quality kept at this speed?
Quality comes from layers, each designed to catch a different kind of problem, and most of them automatic. The unit test suite of about 2,400 tests ran twice with no failures, and around 160 end-to-end tests passed in full.
What would the same change have taken otherwise?
We estimate the same scope would have taken 42 to 72.5 hours of people’s time by hand, or 12.5 to 24 hours with Claude Code and no harness. As delivered, it took 4.4 to 6.8 hours, most of which is the manual session still to come. All three are estimates against the scope that was delivered.
| Role | By hand, no AI | Claude Code, no harness | As delivered |
|---|---|---|---|
| Developer | 32.5 to 55.5 h | 5.5 to 11 h | 0.4 to 0.8 h |
| QA analyst, before the manual session | 5.5 to 11 h | 3 to 7 h | none |
| Manual session, QA analyst and BA | 4 to 6 h | 4 to 6 h | 4 to 6 h |
| People’s time in all | 42 to 72.5 h | 12.5 to 24 h | 4.4 to 6.8 h |
| Elapsed, including the manual session | 8 to 13 working days | 2.5 to 4 working days | 1 to 1.5 working days |
The estimates assume a mid-level to senior developer who already knows the codebase. Without the harness, we expect the extra time to go on restating the project’s patterns to the model, reworking what drifted, and doing the delivery administration and test design by hand.
In Q2 we reported proposals for large greenfield builds landing at 28 to 34% of their pre-AI cost. This change comes out at around 10% of the manual people time. We do not read that as a new headline: it is one well-specified change on a mature harness, and a programme-level figure averages across many features, including ones that are less well defined. None of these figures include what it has cost us to build the harness, which a full comparison would need to count. Our AI-augmented development ROI guide sets out how we think about that.
What does AI actually cost?
Both stories together cost about $71 at API list prices, by Claude Code’s own counter. Almost all of it was Claude Opus 5.5, with about $3 for Claude Sonnet 5 running the independent reviews. That is small against the people’s time, but it is not nothing, and it is worth understanding where it comes from.
Most of the cost is Claude re-reading, not writing. Each time Claude takes a step, it reads the whole conversation so far, the way you might reread a long email thread before replying. The longer a session runs, the more there is to reread, so the cost of each step climbs. Claude wrote about 385,000 tokens of output across the session, but re-read about 220 million tokens of context to do it.
The chart shows the conversation growing through the first build until Claude compacted it, summarising the history to start afresh. The session ran with a context window of one million tokens, which is how a single step could carry 785,000. Long pauses had a cost too. After the two longest waits, one of 81 minutes and the other overnight, Claude’s cached copy of the conversation had expired and had to be rebuilt, which costs more than reading it.
What does this tell us about predicting and optimising AI delivery?
The code is no longer where most of the time goes, so it is no longer the right thing to estimate from. This is one worked example rather than a benchmark, but it points to five things we want to understand better.
- Elapsed time is set by decisions, not by the build. The build took an afternoon and an evening, and the questions it depended on waited almost a week. Predicting delivery means predicting how quickly decisions come back, and the best way to speed that up is to put well-prepared questions to the business early.
- Claude’s time goes on checking, not typing. Running tests took more of Claude’s time than writing code, and environment problems cost about 20 minutes. Faster, more reliable test environments pay back directly in shorter sessions.
- Session length drives AI cost. The longest build carried the most history and accounted for over half of everything Claude re-read. Smaller pieces of work, earlier compaction, and fewer long pauses mid-session are levers we can test.
- Engineer attention is the scarce resource. Under an hour of hands-on time across a day-long change means one engineer can drive several changes in parallel. Time still matters, but cognitive load, meaning how many changes one person can hold in mind, judge, and review well, sets the tighter limit.
- One session is not a baseline. We want to measure more changes this way, across different kinds of work, so that our estimates rest on how long AI-built work actually takes rather than on how long we hope it will.
For the difference between AI as an editor helper and AI across the whole lifecycle, see our guide to AI-augmented vs AI-assisted development. Our next quarterly report follows in October. For everything else, browse the Talk Think Do guides hub.
Cite this report
Talk Think Do, "Anatomy of an AI Change: How Long It Really Takes," September 2026. https://talkthinkdo.com/ai-velocity-report/anatomy-of-an-ai-change/
Frequently asked questions
How long does a feature take to deliver with Claude Code and a delivery harness?
How much does an AI-delivered change cost in model usage?
Why measure AI-assisted software delivery?
What makes AI-built software hard to predict?
Does AI delivery remove the need for business analysts and testers?
Want to talk about what we're seeing?
Book a free 30-minute consultation. We will give you an honest assessment of your options.
Get each AI Velocity Report in your inbox