Skip to content

Measuring without illusions 📏

The AI metrics that are easiest to collect are often the least useful: prompts, tokens, accepted suggestions. They measure activity. Value requires a baseline and two families of signals: operational outcome and the capability the system retains after the pilot. The rest is dashboard aesthetics. Many pilots with no measurable return start here: nobody measures because nobody ever decided what to look at.

A dual metrics system ⚖️

Operational Capability
Lead time, deploy frequency, MTTR, change failure rate Teams able to repeat the workflow without an expert
Triage time, errors avoided, deflected work Curated data, clear policies, explicit ownership
Time-to-value and incremental revenue Quality of feedback, evals and learning

The DORA Four Keys are a good lingua franca for a software pipeline: they don’t automatically attribute every change to AI, but they show whether the flow has really improved. Always add a guardrail: speed without quality is just a scheduled incident.

    flowchart LR
    A[Baseline] --> B[Bounded pilot]
    B --> C[Operational metrics]
    B --> D[Quality guardrail]
    C --> E[Decision]
    D --> E
  

The Four Keys are no longer enough 🧩

With cheaply generated code, deployment frequency and lead time become misleading: they rise even when the system gets worse. Faros AI data is the textbook case: more tasks completed, but also more bugs, more incidents per PR, much longer reviews and much more rewritten code. If you only measure delivery speed, you see a success. If you measure the whole flow, you see debt piling up.

Among the Four Keys, MTTR remains the most reliable. DORA alone, however, is not enough: complete it with frameworks such as SPACE (satisfaction, performance, activity, communication, efficiency) or DX Core 4, and with AI-specific metrics such as the share of code written by agents, code churn and bug attribution between AI and humans (tagged in CI, not by gut feeling). The rule stays: few metrics, well defined, driving a decision.

The shapes of value 💶

Count avoided costs, reallocated hours, eliminated work, prevented errors and risks, time to value and incremental revenue. Tokens are a cost to govern, not proof that the pilot worked: a non-trivial share of consumption is waste (redundant contexts, pointless retries, output never used) and should be optimized separately from the business case.

Total cost must be computed over a multi-year horizon and include what is usually forgotten: recurring tokens, retrieval infrastructure, observability, human review and rework. The license is the most visible part and the least decisive: what moves the real cost is the quality of what you produce.

Baseline, control and a decision made in advance 🎯

The pre-AI baseline is the most critical and most often missing element. Without it there is no result: there is a before and after with no starting point. Three practical rules:

  1. Control group, not just before/after. Compare a group using AI with one that doesn’t (or barely does). It is the only way to filter out context variables.
  2. Measure at team level, never individual rankings. Individual metrics with AI turn into downward games: whoever doesn’t delegate looks slow, whoever delegates too much looks fast. The team is the unit of measure.
  3. Decision written before seeing results. Criteria to expand, correct or stop are fixed at the start of the pilot. Cherry-picking means defining metrics after the fact: banning it is a governance matter, not a matter of goodwill.

Horizon matters too: immediate value and capability building are two different conversations. Measuring first-year training with ROI metrics from years later is a sure way to shut down an initiative that was building the future.

Anti-metrics and “phantom productivity” 🚫

  • Tool-usage fetish: “almost every team uses AI” is not a result, it is a cost. Put impact metrics on the value stream.
  • Phantom productivity: hours saved but never recaptured. If you don’t specify where the hours go, the saved time evaporates.
  • Cognitive debt neglect: hours of review and refactoring of AI output that never reach the books.
  • Incomplete TCO: counting only the license and discovering token, retrieval and observability costs later.
  • Cherry-picking and no baseline: two faces of the same coin — the metric defined after the result.

A reasonable scorecard keeps few active metrics, five to seven at most: delivery (DORA, watching rework and churn), experience (SPACE/DX), AI-specific (such as the share of PRs without review, to keep at zero), cost tracked separately, and a business outcome tied to OKRs or P&L. Golden rule: every KPI must drive a decision — if it doesn’t change the action, drop the KPI: it is theater.

Checklist ✅

  • Baseline collected before rollout.
  • One outcome metric and one quality guardrail.
  • Total cost, including review and rework.
  • Decision written before seeing results.
  • Comparison group or explicit estimate of variables.
  • Few metrics, each tied to a decision.

Connect the picture to FOCUS and lightweight governance: measuring does not replace judgment, it stops judgment from turning into narrative.

Last updated on