Disciplines → 6 of 6

Value & Measurement

Evidence, not anecdote: what AI adoption changes in flow, quality and cost, measured for teams against a baseline.

Overview

Value & Measurement decides whether AI adoption is working. It sets a baseline before adoption, compares AI-assisted with non-assisted work, counts the full cost of AI use and reports the result honestly, including where AI made work slower or worse.

Claims that AI makes teams faster are easy to make and hard to verify. Generating code quickly is not the same as delivering change quickly: time saved writing can reappear as review queues, rework and incidents. This discipline measures the whole flow, from accepted idea to working change in production, so that decisions to expand, hold or withdraw AI use rest on evidence. It measures teams and value streams, never individuals.

The established delivery measures give the outcome frame: lead time for changes, deployment frequency, change failure rate and time to restore service. Around them sit the measures that show where AI helps or hurts: cycle time, rework rate, defect escape rate, review burden and cost.

What changes in existing practice

PracticeTodayWith Value & Measurement
Engineering metrics dashboardsDelivery measures for the whole teamThe same measures split by AI-assisted vs not, plus review burden and rework; aggregated to team level, with no per-person views
Business cases and benefit trackingBenefits estimated once, rarely checkedA value hypothesis written at Conceive, with a baseline, and tested at Learn
Finance and chargebackLicences budgeted as a tooling lineTotal cost of AI use allocated to the value streams that incur it
Portfolio reviewsProgress against planAI initiatives reviewed on measured outcomes and cost; expand, hold or stop decided on evidence

Key practices

Write the value hypothesis at Conceive

Every AI initiative states what it expects to change, by how much, for which value stream, and how that will be measured. The hypothesis is written at Conceive alongside the business case, and tested at Learn against the measured result. An initiative that cannot name its measure is not ready to start.

Take a baseline before adoption

Measure cycle time, rework rate, defect escape rate and review burden for the affected teams before AI tools arrive, over a period long enough to cover normal variation. Without a baseline, any later number is uninterpretable and every trend becomes a story told after the fact.

Compare like with like

Compare AI-assisted with non-assisted work at team or value-stream level, using the provenance already recorded on each change. Control for what else changed: team composition, work mix, release calendar, a new platform. Attributing a result to AI is an inference; record the method and the confounders with the result, and have someone outside the delivering team review both.

Count the total cost

The licence is the smallest visible part of the cost. Total cost of AI use includes licences, usage charges, reviewer time, rework, incidents traced to AI-assisted changes and the enablement effort to train people and maintain context assets. Report cost per change including review time, so a faster author with a slower reviewer does not look like a gain.

Watch review burden

AI shifts effort from writing to checking. Track reviewer hours per change and the time changes wait in review queues. Rising review burden is often the first sign that generated output exceeds what teams can verify, and it predicts rework and escaped defects before they show up in production.

Report benefits and costs together

Publish what improved, what did not and where AI slowed work down. A result that shows no gain, or a loss, is a valid finding that saves money and protects quality. Expand delegation only when the evidence supports it, and pull it back when the evidence turns.

Controls by risk tier

TierPrevent (Guardrails)Prove (Audit Trail)Detect (Self-Monitoring)
LowBaseline in place before a pilot starts; team-level aggregation onlyMetric definitions and baseline versioned with the value hypothesisRework and pipeline failure trends, AI-assisted vs not
MediumDelegation expanded only after a completed comparison periodComparison method, confounders and result recorded and reviewed at LearnReview burden and change failure rate against agreed thresholds
HighExpansion requires independent review of the evidence and risk sign-offEvidence pack with total cost, attribution reasoning and reviewerDefect escapes and incidents tracked for a defined window after each expansion
CriticalNo expansion of delegation on productivity grounds; AI stays advisoryDocumented rationale and second-line sign-off for any change in AI useAny measured quality regression triggers review

Maturity path

  1. Ad hoc — value claimed from anecdote and licence counts; no baseline exists.
  2. Experimenting — pilots report usage and satisfaction; outcomes are estimated, not measured.
  3. Managed — baselines taken before adoption; value hypotheses written and tested; AI-assisted vs not compared at team level.
  4. Governed — total cost reported per value stream; portfolio decisions to expand or stop rest on reviewed evidence.
  5. Optimised — delegation and spend adjusted continuously from flow, quality and cost data.

Assess your organisation →

Measures

Anti-patterns

Regulatory anchors

The EU AI Act lists AI systems used to monitor or evaluate the performance and behaviour of workers as high-risk (Annex III, point 4). Scoring individual engineers from AI usage or output data moves towards that category and its obligations, which is one more reason to measure teams and flow, not people. ISO/IEC 42001 expects an AI management system to evaluate its performance (clause 9) and improve continually (clause 10); baselines, tested value hypotheses and reviewed evidence supply both.