Skip to content

Field note 5 min read

Measuring an AI-native delivery uplift honestly

Everyone reports a productivity gain from AI-assisted engineering. Very few can show a baseline. Here is the instrumentation we insist on before claiming a number.

  • AI-native delivery
  • Engineering leadership
  • Measurement

Claims of 30%, 50%, even 10× productivity gains from AI-assisted engineering circulate freely, and almost none survive a follow-up question about methodology. This matters beyond credibility: if you cannot measure the uplift, you cannot tell which parts of the lifecycle to invest in next, and you will keep buying tools instead of changing how the work works.

We hold ourselves to a specific bar. A number is only quotable if there is a pre-period baseline, a stable unit of work, and an explanation for the confounders.

Pick outcome measures, then guard them

The temptation is to measure what the tools emit — lines accepted, suggestions taken, prompts run. These measure adoption, not value, and they are trivially gamed. The measures that matter sit at the delivery boundary:

  • Lead time from committed scope to production, per increment.
  • Throughput in stable units of work, with a fixed definition agreed before the trial.
  • Change failure rate and time to restore, so speed gains are not quietly funded by quality.
  • Review latency and rework rate, which is where AI-generated code most often shifts cost rather than removing it.

The last one deserves emphasis. A team can generate code faster and spend the saving on review and defect repair, netting nothing. Unless rework is instrumented, that outcome is indistinguishable from a genuine gain.

Establish the baseline before enthusiasm sets in

Baselines gathered after a rollout has started are contaminated by the Hawthorne effect and by selection: the squads that volunteer first are not representative. We take at least one full quarter of pre-period data across a spread of teams, and we write down the expected effect size in advance. Pre-registration is unfashionable in industry and it is the cheapest guard against reading noise as signal.

Name the confounders out loud

In every real programme, the AI rollout coincides with other changes — a reorganisation, a new platform, a hiring wave, a quieter quarter. Attribution is genuinely hard. Two things help. First, staggered adoption across comparable teams gives you an internal comparison group rather than only a before-and-after. Second, decomposing the gain by lifecycle stage makes implausible attributions visible: if the uplift is concentrated in test authoring and scaffolding, that is a coherent story; if it is spread evenly across activities AI barely touches, something else is driving it.

What the honest answer usually looks like

Where the discipline holds, the sustained figure we see across squads sits in the region of 30–40% on delivery throughput — meaningful, compounding, and considerably less dramatic than the headline numbers in circulation. It is concentrated in specification and story generation, scaffolding, test creation, migration work and review assistance. It is close to zero on genuinely novel design work, and it can go negative in domains where the model’s training data misleads confidently.

That shape is more useful than a bigger number, because it tells you where to push next. Which is the entire point of measuring.

The organisational part

The measurement discipline is also what makes the change stick. Teams accept a new way of working when they can see it working, in their own numbers, rather than being told about someone else’s success. Publishing the uplift by stage — including the stages where it did nothing — buys more credibility than any mandate.

From note to engagement

Recognise this problem in your own organisation?

A 45-minute technical conversation. Bring the constraint you are stuck on and you will leave with an honest read on whether it is tractable, what it would take to move it, and how we would sequence the work.

or write directly — hello@bluestreaklabs.com