Measuring without illusions 📏
The AI metrics that are easiest to collect are often the least useful: prompts, tokens, accepted suggestions. They measure activity. Value requires a baseline and two families of signals: operational outcome and the capability the system retains after the pilot. The rest is dashboard aesthetics. Many pilots with no measurable return start here: nobody measures because nobody ever decided what to look at.
A dual metrics system ⚖️
| Operational | Capability |
|---|---|
| Lead time, deploy frequency, MTTR, change failure rate | Teams able to repeat the workflow without an expert |
| Triage time, errors avoided, deflected work | Curated data, clear policies, explicit ownership |
| Time-to-value and incremental revenue | Quality of feedback, evals and learning |
The DORA Four Keys are a good lingua franca for a software pipeline: they don’t automatically attribute every change to AI, but they show whether the flow has really improved. Always add a guardrail: speed without quality is just a scheduled incident.
flowchart LR
A[Baseline] --> B[Bounded pilot]
B --> C[Operational metrics]
B --> D[Quality guardrail]
C --> E[Decision]
D --> E
The Four Keys are no longer enough 🧩
With cheaply generated code, deployment frequency and lead time become misleading: they rise even when the system gets worse. Faros AI data is the textbook case: more tasks completed, but also more bugs, more incidents per PR, much longer reviews and much more rewritten code. If you only measure delivery speed, you see a success. If you measure the whole flow, you see debt piling up.
Among the Four Keys, MTTR remains the most reliable. DORA alone, however, is not enough: complete it with frameworks such as SPACE (satisfaction, performance, activity, communication, efficiency) or DX Core 4, and with AI-specific metrics such as the share of code written by agents, code churn and bug attribution between AI and humans (tagged in CI, not by gut feeling). The rule stays: few metrics, well defined, driving a decision.
The shapes of value 💶
Count avoided costs, reallocated hours, eliminated work, prevented errors and risks, time to value and incremental revenue. Tokens are a cost to govern, not proof that the pilot worked: a non-trivial share of consumption is waste (redundant contexts, pointless retries, output never used) and should be optimized separately from the business case.
Total cost must be computed over a multi-year horizon and include what is usually forgotten: recurring tokens, retrieval infrastructure, observability, human review and rework. The license is the most visible part and the least decisive: what moves the real cost is the quality of what you produce.
Baseline, control and a decision made in advance 🎯
The pre-AI baseline is the most critical and most often missing element. Without it there is no result: there is a before and after with no starting point. Three practical rules:
- Control group, not just before/after. Compare a group using AI with one that doesn’t (or barely does). It is the only way to filter out context variables.
- Measure at team level, never individual rankings. Individual metrics with AI turn into downward games: whoever doesn’t delegate looks slow, whoever delegates too much looks fast. The team is the unit of measure.
- Decision written before seeing results. Criteria to expand, correct or stop are fixed at the start of the pilot. Cherry-picking means defining metrics after the fact: banning it is a governance matter, not a matter of goodwill.
Horizon matters too: immediate value and capability building are two different conversations. Measuring first-year training with ROI metrics from years later is a sure way to shut down an initiative that was building the future.
Anti-metrics and “phantom productivity” 🚫
- Tool-usage fetish: “almost every team uses AI” is not a result, it is a cost. Put impact metrics on the value stream.
- Phantom productivity: hours saved but never recaptured. If you don’t specify where the hours go, the saved time evaporates.
- Cognitive debt neglect: hours of review and refactoring of AI output that never reach the books.
- Incomplete TCO: counting only the license and discovering token, retrieval and observability costs later.
- Cherry-picking and no baseline: two faces of the same coin — the metric defined after the result.
A reasonable scorecard keeps few active metrics, five to seven at most: delivery (DORA, watching rework and churn), experience (SPACE/DX), AI-specific (such as the share of PRs without review, to keep at zero), cost tracked separately, and a business outcome tied to OKRs or P&L. Golden rule: every KPI must drive a decision — if it doesn’t change the action, drop the KPI: it is theater.
Checklist ✅
- Baseline collected before rollout.
- One outcome metric and one quality guardrail.
- Total cost, including review and rework.
- Decision written before seeing results.
- Comparison group or explicit estimate of variables.
- Few metrics, each tied to a decision.
Connect the picture to FOCUS and lightweight governance: measuring does not replace judgment, it stops judgment from turning into narrative.