KPIs and the Three-Tier ROI Stack
Prove value, avoid vanity metrics
- Apply the three-tier ROI stack — adoption/usage, workflow efficiency, business/P&L impact — and explain why stopping at tier one is the most common trap
- Identify vanity metrics (seats, logins, tokens, suggestion-acceptance rate) and explain how they can be inversely correlated with real value
- Insist on a baseline before spend and on quality metrics (CSAT/NPS), using the Klarna reversal as the cautionary case
- Frame '% of EBIT attributable to AI' as the high-performer line and set realistic ROI timelines (typically 2–4 years), citing live sources
- Run a disciplined pilot review that kills or reshapes any initiative that does not move a defined number
- Ask the right boardroom questions that separate proven AI value from expensive theater
Almost every company now uses AI, yet most cannot point to a dollar of value from it — because they measure activity instead of impact. This lesson gives you a rigorous, three-tier way to measure AI value (adoption, then workflow efficiency, then P&L), names the vanity metrics that flatter dashboards while value quietly stalls, and equips you to baseline before you start, track quality not just volume, and kill or reshape any pilot that won't move a number.
- 1The headline: everyone has AI, almost nobody can prove value
- 2The three-tier ROI stack
- 3Vanity metrics: the numbers that mislead
- 4Baseline first, and track quality — not just volume
- 5The high-performer line: % of EBIT attributable to AI — and a realistic clock
- 6The discipline of the kill-or-reshape decision
- 7The right questions to ask
The headline: everyone has AI, almost nobody can prove value
Start with the uncomfortable fact every leader should carry into the boardroom: adoption is near-universal, but measured value is rare. The differentiator between the winners and everyone else is not the model — it is whether leadership measures the right things and acts on them.
The 2025–2026 evidence converges hard:
| Finding | Source (year) |
|---|---|
| ~88% of organizations use AI in ≥1 function, but only ~39% report enterprise-level EBIT impact — and most of those attribute under 5% of EBIT to it | McKinsey, State of AI (Nov 2025) |
| ~95% of enterprise gen-AI pilots delivered no measurable P&L return; only ~5% reached rapid value. Root cause is the learning/integration gap, not model quality | MIT NANDA, GenAI Divide (Jul 2025) |
| Only ~5% of firms are "future-built" (value at scale); ~60% are "laggards" with minimal gains despite spending | BCG, Widening AI Value Gap (Sept 2025) |
"Everyone has AI; almost nobody has AI value yet."
The gap is not a technology problem — it is a measurement-and-management problem. When pilots don't move EBIT, the cause is rarely the model; it is the absence of a baseline, a workflow redesign, and a disciplined number to hit. McKinsey's own analysis found that tracking well-defined KPIs is the single attribute most correlated with AI adoption — measurement is not the scoreboard after the game, it is part of how you win. These are time-stamped findings from a fast-moving field; treat the percentages as directional and re-verify them live before you quote them.
Key insight
The reframe
Measurement is not the report you write after the value arrives. Disciplined, well-defined KPIs are how value gets created — McKinsey found KPI tracking is the top driver of adoption, and adoption is the on-ramp to impact. If you can't measure it, you can't manage it, fund it, or defend it to the board.
The three-tier ROI stack
The core framework of this lesson is a simple ladder. AI value reveals itself in three tiers, and you must climb all three. Most organizations proudly report tier one and never get to tier three — which is exactly why their AI looks busy but their P&L doesn't move.
| Tier | What it measures | Example metrics | The question it answers |
|---|---|---|---|
| 1 · Adoption / usage | Are people actually using it? | Active users, usage frequency, % of eligible workforce, adoption velocity | "Is it switched on?" |
| 2 · Workflow efficiency | Is work getting faster, higher-throughput, better quality? | Time saved per task, throughput, handling-time reduction, error/defect reduction, CSAT/NPS lift | "Is the work itself better?" |
| 3 · Business / P&L impact | Did it change the financials? | Revenue uplift, cost reduction, and ultimately % of EBIT attributable to AI | "Did it move money?" |
The discipline is to connect the tiers, not to stop at one. Adoption without efficiency is just a tool nobody benefits from. Efficiency without P&L impact is a productivity claim with no proof it reached the bottom line — time "saved" that is never reinvested or removed from cost is value that evaporates. The leadership move is to define, before you spend, what a win looks like at tier three and trace the line back down to the tier-one behaviors that produce it.
A caution from the research: be deliberate about where you chase tier-three value. MIT NANDA found more than half of gen-AI budgets flow to sales and marketing tools, while some of the largest near-term ROI sat in unglamorous back-office automation. The flashiest use case is not always the one that moves the most money.
Watch out
Where leaders get it wrong
They declare victory at tier one. "90% of the team logged in this month" feels like progress and proves nothing about value. A dashboard full of green tier-one metrics with no tier-three number attached is the single most common pattern behind the ~95% of pilots that show no P&L return (MIT NANDA, 2025).
Tip
The leadership move
For every funded initiative, write the tier-three number first — the specific revenue, cost, or EBIT figure it must move — then work down to the tier-two and tier-one metrics that should predict it. Value flows up the stack; accountability is designed down from the top.
Vanity metrics: the numbers that mislead
Some metrics are not just weak — they are actively misleading, because they can move in the opposite direction of value. These are the numbers a vendor or an over-eager team will put on a slide precisely because they always go up.
| Vanity metric | Why it misleads |
|---|---|
| Seat licenses sold | Counts what you bought, not what created value. A common failure: blanket per-seat licensing on day one, with 30–40% of seats unused within 90 days. |
| Logins / active users | Measures presence, not productivity or P&L. People can log in daily and produce nothing. |
| Token consumption | Counts cost, not benefit. More tokens burned is more spend, not more value. |
| Suggestion-acceptance rate | Counts agreement with the tool, not whether the output was correct or made anything better. |
The most counter-intuitive trap is that these can be inversely correlated with value. Consider agentic AI: as work moves from one-human-one-seat toward one orchestrator running many agents, a more valuable, more automated workflow may show fewer seats and fewer human logins. If "seats" or "logins" is your success metric, your dashboard will fall exactly as your value rises. You would be punishing your best transformation.
The fix is to anchor every vanity metric to a tier-two or tier-three outcome. Seats only matter if utilization converts to time saved; time saved only matters if it converts to cost out or revenue in. Treat tier-one metrics as diagnostics (is anyone using this?) — never as the scoreboard.
Key insight
The reframe
A rising vanity metric is not evidence of value — and a falling one is not evidence of failure. As agents replace seats, the most valuable workflows can post the lowest seat and login counts. Judge initiatives by what reaches the P&L, not by what's easy to count.
Baseline first, and track quality — not just volume
Two disciplines separate credible measurement from theater: baseline before you start, and track quality, not just volume.
Baseline before you spend. If you don't capture today's handling time, resolution rate, error rate, CSAT, and cost before you turn AI on, you can never prove the change was AI's doing. "Vibe-based" spending — funding AI on fear-of-missing-out with no baseline and no hypothesis — is one of the most-cited reasons ROI goes unproven (Worklytics, 2025). A pre-registered baseline turns a debate about whether AI "feels" helpful into a measurable before-and-after.
Track quality, not just volume. Volume metrics (chats handled, tickets closed, content generated) are seductive because AI makes them spike. But volume can mask a quality collapse — and quality is what customers and revenue actually respond to. Always pair every volume metric with a quality metric: containment/deflection alongside CSAT/NPS, throughput alongside error/defect rate, speed alongside accuracy. This is the Klarna lesson (next callout), and it is the difference between a number that flatters you and a number that protects you.
Example
The Klarna reversal — the canonical cautionary tale
In 2024 Klarna's AI assistant handled ~two-thirds of customer-service chats (~2.3M conversations in month one), cut resolution time from ~11 minutes to under 2, did the work of ~700 agents, and was tied to a projected ~$40M profit improvement (Klarna press, Feb 2024 — vendor-sourced). On volume metrics it looked like a triumph. Then quality slipped: CSAT/NPS deteriorated, and by mid-2025 Klarna reversed course and rehired humans. The CEO's own framing: they "focused too much on efficiency and cost," the lower quality was "not sustainable." The volume metrics masked a quality collapse — for nuanced, empathetic work, augment-and-blend beat wholesale replacement. (Verify current status live.)
Watch out
Where leaders get it wrong
They ship a pilot, watch the volume chart go up and to the right, and declare success — with no baseline to compare against and no quality metric to catch the downside. By the time CSAT shows up in churn, the damage is done. Never let a volume number travel to the board without its paired quality number.
The high-performer line: % of EBIT attributable to AI — and a realistic clock
At the top of the stack sits the metric that separates serious programs from busy ones: the percent of EBIT attributable to AI. McKinsey uses precisely this line to identify high performers — the firms that can attribute a meaningful, growing share of earnings to AI, rather than gesturing at productivity anecdotes. As of late 2025, even among the ~39% of firms reporting any enterprise EBIT impact, most attribute under 5% of EBIT to AI (McKinsey, State of AI, Nov 2025) — which tells you both how early this is and how distinctive it is to be on the right side of that line.
Set a realistic clock. AI returns take longer than typical technology investments. Most organizations reach satisfactory ROI in roughly 2–4 years; only about 6% see it in under a year (Deloitte, AI ROI Paradox). Promising board-level EBIT impact in two quarters is how you manufacture the disappointment that gets a good program cancelled. The leadership move is to commit to the 2–4-year arc on tier three while demanding near-term movement on tiers one and two as leading indicators — proof you're on the path even before the EBIT lands.
Both the EBIT-share figure and the 2–4-year timeline are volatile, dated findings. Teach the durable concept — EBIT attribution is the high-performer line; AI ROI is a multi-year arc — and re-verify the exact numbers at the live sources before you cite them.
Tip
The leadership move
Make '% of EBIT attributable to AI' a standing line in your AI review — even if it reads "~0%" today. Naming the metric forces the organization to build the attribution discipline now, so that when value lands you can prove it. What gets measured at the top gets managed all the way down.
Watch out
Where leaders get it wrong
They under-promise on usage and over-promise on payback — pledging fast EBIT impact, then killing the program when 2–4-year returns don't show up in two quarters. Mismatched timelines, not weak technology, sink many otherwise-sound initiatives. Set the clock honestly up front.
The discipline of the kill-or-reshape decision
Measurement without consequence is decoration. The hardest — and most valuable — leadership act is the scaling decision: at the end of a pilot, every initiative either scales, gets reshaped, or gets killed. What you must not allow is pilot purgatory — the endless proof-of-concept that never faces a verdict. The share of companies abandoning most AI projects rose to roughly 42% in 2025, up from ~17% in 2024 (cited via Worklytics/S&P; trace to primary before quoting); the lesson is not to avoid killing pilots — it is to kill them deliberately and early, on evidence, rather than letting them drift.
Use a simple decision rule tied directly to the three-tier stack:
| Verdict | When to choose it | What it looks like |
|---|---|---|
| Scale | Moves a defined tier-two/tier-three number, with quality holding | Fund the 70% (people, process, workflow redesign); set EBIT-linked targets |
| Reshape | Promising signal but the workflow, scope, or data is wrong | Redesign the workflow around the tool — don't bolt it onto a broken process — and re-baseline |
| Kill | No movement on a real number after a fair, baselined trial | Stop the spend, capture the lesson, redeploy the budget |
The positive counter-example to pilot purgatory is JPMorgan, which runs 450+ AI use cases in production with benefits growing ~30–40% year over year — the hallmark of scoped, KPI-anchored deployment where each use case has to earn its place against a number. The contrast with the ~95% of pilots showing no P&L return is not luck; it is the discipline of measuring, then deciding.
Tip
The leadership move
Put a scaling decision on the calendar when you launch the pilot, not after it has quietly run for a year. Pre-commit to the number it must move and the date you'll judge it. "Kill, reshape, or scale — decide by Q3" is a sentence that ends pilot purgatory before it starts.
Example
Scoped and measured beats sprawling and vague
JPMorgan's 450+ production use cases, each anchored to a number and compounding ~30–40% in benefits per year, exemplify the alternative to pilot sprawl. The discipline isn't running fewer experiments — it's giving every experiment a number to hit and a date to be judged.
The right questions to ask
You don't need to build the dashboard — you need to interrogate it. These are the questions that separate a measured program from a flattering one. Keep them on a card for every AI review.
On any metric you're shown:
- Which tier is this — adoption, efficiency, or P&L? (If everything on the slide is tier one, push back.)
- Could this number go up while value goes down? (The vanity-metric test.)
- What baseline is it measured against? Did we capture "before"?
On quality and risk:
- For every volume metric, where is its paired quality metric (CSAT/NPS, error rate, accuracy)?
- If this is "90% accurate," accurate how — and who owns the failing 10%?
On the P&L line:
- What % of EBIT do we attribute to AI today, and what's the trajectory?
- Is our timeline honest — are we judging a 2–4-year payback on a two-quarter clock?
On the verdict:
- Which pilots haven't moved a number, and when is their kill-or-reshape decision?
- When we report time saved, has it actually been reinvested or removed from cost — or did it just evaporate?
Asking these consistently does something subtle but powerful: it trains the organization to bring you tier-three thinking by default, because they know tier-one theater won't survive the meeting.
Key insight
The reframe
Your job is not to read the dashboard — it's to audit it. The most valuable thing a leader contributes to AI measurement is a short, repeatable set of questions that makes vanity metrics and unbaselined claims uncomfortable to present. Govern the metrics and you govern the value.
Try it: Build a three-tier ROI scorecard and pass two pilots through the kill-or-reshape gate
A strategic exercise — no tools, no code, just the discipline. 1) Pick two real (or realistic) AI initiatives in your organization — ideally one customer-facing (e.g. support assist) and one back-office (e.g. finance reporting). 2) Build a three-tier scorecard for each. In a simple table, fill one row per tier: Tier 1 (adoption — e.g. % of eligible workforce using it weekly), Tier 2 (workflow efficiency — name a time/throughput metric AND its paired quality metric, e.g. handle time + CSAT), Tier 3 (P&L — the specific revenue, cost, or EBIT number it must move). Write the tier-three number first. 3) Flag the vanity metrics. List any seats/logins/tokens/acceptance-rate numbers currently being reported and write one sentence on how each could rise while value falls. 4) Define the baseline. For each metric, state what 'before' value you would capture, and admit honestly where you have no baseline today. 5) Set an honest clock. Mark which tier-three results you'd expect inside a year vs. 2–4 years, and which tier-1/2 metrics are your leading indicators. 6) Run the gate. For each initiative, write the date and the threshold for its kill-or-reshape-or-scale decision — the exact number it must move by when. 7) Draft three boardroom questions you will ask at the next AI review to keep tier-one theater from passing as value. Deliverable: a one-page scorecard per initiative plus your three questions. This builds the instinct the rest of your AI program depends on — that every initiative carries a number, a baseline, a quality guardrail, and a verdict date.
Key takeaways
- 1AI's defining 2025–2026 reality is a measurement-and-management gap: ~88% of firms use AI but only ~39% report enterprise EBIT impact and ~95% of pilots show no P&L return — value stalls on management, not models (McKinsey & MIT NANDA, 2025).
- 2Use the three-tier ROI stack — (1) adoption/usage → (2) workflow efficiency (time, throughput, error/quality) → (3) business impact (revenue, cost, EBIT) — and never stop at tier one; define the tier-three number first and trace it down.
- 3Vanity metrics (seats, logins, tokens, suggestion-acceptance rate) can be inversely correlated with value — as agents replace seats, the most valuable workflows can post the lowest counts; treat tier-one numbers as diagnostics, not the scoreboard.
- 4Baseline before you spend and pair every volume metric with a quality metric (CSAT/NPS) — the Klarna reversal shows how strong volume numbers masked a quality collapse that forced a U-turn.
- 5'% of EBIT attributable to AI' is McKinsey's high-performer line; set an honest clock, because AI ROI typically takes 2–4 years (only ~6% under a year, per Deloitte) — demand near-term tier-1/2 movement as leading indicators.
- 6End pilot purgatory: schedule a kill-or-reshape-or-scale decision at launch and kill or reshape anything that won't move a defined number — scoped, KPI-anchored deployment (e.g. JPMorgan's 450+ use cases) is what beats the 95% failure rate.
Quiz
Lock in what you learned
Check your understanding
0 / 4 answered
1.Your AI program review opens with a slide showing seat licenses purchased, monthly logins, and tokens consumed — all trending up. What is the most accurate executive read of this dashboard?
2.What is the central lesson executives should draw from the Klarna customer-service AI case?
3.A division head proposes judging a new AI initiative purely on 'percent of EBIT attributable to AI' within the first two quarters. What is the soundest objection?
4.Six months in, a pilot shows enthusiastic usage but has not moved any tier-two or tier-three number against its baseline. What is the disciplined leadership response?
Go deeper
Hand-picked sources to keep learning
Source for the ~88% adoption / ~39% EBIT-impact gap, '% of EBIT attributable to AI' as the high-performer line, and KPI tracking as the top adoption driver. Refreshes ~annually — re-verify figures.
Source for ~95% of pilots showing no measurable P&L return and the 'learning gap'; also the finding that >50% of budgets go to sales/marketing while back-office held larger near-term ROI.
Source for the realistic ROI timeline: most firms reach satisfactory AI ROI in 2–4 years, only ~6% in under a year. Verify current figures live.
The original (vendor-sourced) volume numbers — ~700 agents, 11 min → <2 min, ~$40M projected uplift — that later masked the quality decline.
The other side of the Klarna case: CSAT/NPS decline and the mid-2025 rehiring of humans — the 'track quality, not just volume' cautionary tale.
Source for measuring-activity-not-value anti-patterns and the ~42%-in-2025 (vs ~17% in 2024) project-abandonment figure. Trace abandonment number to Gartner/S&P primary before quoting.