Disciplines → 6 of 6
Evidence, not anecdote: what AI adoption changes in flow, quality and cost, measured for teams against a baseline.
Value & Measurement decides whether AI adoption is working. It sets a baseline before adoption, compares AI-assisted with non-assisted work, counts the full cost of AI use and reports the result honestly, including where AI made work slower or worse.
Claims that AI makes teams faster are easy to make and hard to verify. Generating code quickly is not the same as delivering change quickly: time saved writing can reappear as review queues, rework and incidents. This discipline measures the whole flow, from accepted idea to working change in production, so that decisions to expand, hold or withdraw AI use rest on evidence. It measures teams and value streams, never individuals.
The established delivery measures give the outcome frame: lead time for changes, deployment frequency, change failure rate and time to restore service. Around them sit the measures that show where AI helps or hurts: cycle time, rework rate, defect escape rate, review burden and cost.
| Practice | Today | With Value & Measurement |
|---|---|---|
| Engineering metrics dashboards | Delivery measures for the whole team | The same measures split by AI-assisted vs not, plus review burden and rework; aggregated to team level, with no per-person views |
| Business cases and benefit tracking | Benefits estimated once, rarely checked | A value hypothesis written at Conceive, with a baseline, and tested at Learn |
| Finance and chargeback | Licences budgeted as a tooling line | Total cost of AI use allocated to the value streams that incur it |
| Portfolio reviews | Progress against plan | AI initiatives reviewed on measured outcomes and cost; expand, hold or stop decided on evidence |
Every AI initiative states what it expects to change, by how much, for which value stream, and how that will be measured. The hypothesis is written at Conceive alongside the business case, and tested at Learn against the measured result. An initiative that cannot name its measure is not ready to start.
Measure cycle time, rework rate, defect escape rate and review burden for the affected teams before AI tools arrive, over a period long enough to cover normal variation. Without a baseline, any later number is uninterpretable and every trend becomes a story told after the fact.
Compare AI-assisted with non-assisted work at team or value-stream level, using the provenance already recorded on each change. Control for what else changed: team composition, work mix, release calendar, a new platform. Attributing a result to AI is an inference; record the method and the confounders with the result, and have someone outside the delivering team review both.
The licence is the smallest visible part of the cost. Total cost of AI use includes licences, usage charges, reviewer time, rework, incidents traced to AI-assisted changes and the enablement effort to train people and maintain context assets. Report cost per change including review time, so a faster author with a slower reviewer does not look like a gain.
AI shifts effort from writing to checking. Track reviewer hours per change and the time changes wait in review queues. Rising review burden is often the first sign that generated output exceeds what teams can verify, and it predicts rework and escaped defects before they show up in production.
Publish what improved, what did not and where AI slowed work down. A result that shows no gain, or a loss, is a valid finding that saves money and protects quality. Expand delegation only when the evidence supports it, and pull it back when the evidence turns.
| Tier | Prevent (Guardrails) | Prove (Audit Trail) | Detect (Self-Monitoring) |
|---|---|---|---|
| Low | Baseline in place before a pilot starts; team-level aggregation only | Metric definitions and baseline versioned with the value hypothesis | Rework and pipeline failure trends, AI-assisted vs not |
| Medium | Delegation expanded only after a completed comparison period | Comparison method, confounders and result recorded and reviewed at Learn | Review burden and change failure rate against agreed thresholds |
| High | Expansion requires independent review of the evidence and risk sign-off | Evidence pack with total cost, attribution reasoning and reviewer | Defect escapes and incidents tracked for a defined window after each expansion |
| Critical | No expansion of delegation on productivity grounds; AI stays advisory | Documented rationale and second-line sign-off for any change in AI use | Any measured quality regression triggers review |
The EU AI Act lists AI systems used to monitor or evaluate the performance and behaviour of workers as high-risk (Annex III, point 4). Scoring individual engineers from AI usage or output data moves towards that category and its obligations, which is one more reason to measure teams and flow, not people. ISO/IEC 42001 expects an AI management system to evaluate its performance (clause 9) and improve continually (clause 10); baselines, tested value hypotheses and reviewed evidence supply both.