Agentic AI AcademyAgentic AI Academy

KPIs and the Three-Tier ROI Stack

Prove value, avoid vanity metrics

Intermediate 13 minDecision-maker
What you'll be able to do
  • Apply the three-tier ROI stack — adoption/usage, workflow efficiency, business/P&L impact — and explain why stopping at tier one is the most common trap
  • Identify vanity metrics (seats, logins, tokens, suggestion-acceptance rate) and explain how they can be inversely correlated with real value
  • Insist on a baseline before spend and on quality metrics (CSAT/NPS), using the Klarna reversal as the cautionary case
  • Frame '% of EBIT attributable to AI' as the high-performer line and set realistic ROI timelines (typically 2–4 years), citing live sources
  • Run a disciplined pilot review that kills or reshapes any initiative that does not move a defined number
  • Ask the right boardroom questions that separate proven AI value from expensive theater
At a glance

Almost every company now uses AI, yet most cannot point to a dollar of value from it — because they measure activity instead of impact. This lesson gives you a rigorous, three-tier way to measure AI value (adoption, then workflow efficiency, then P&L), names the vanity metrics that flatter dashboards while value quietly stalls, and equips you to baseline before you start, track quality not just volume, and kill or reshape any pilot that won't move a number.

  1. 1The headline: everyone has AI, almost nobody can prove value
  2. 2The three-tier ROI stack
  3. 3Vanity metrics: the numbers that mislead
  4. 4Baseline first, and track quality — not just volume
  5. 5The high-performer line: % of EBIT attributable to AI — and a realistic clock
  6. 6The discipline of the kill-or-reshape decision
  7. 7The right questions to ask

The headline: everyone has AI, almost nobody can prove value

Start with the uncomfortable fact every leader should carry into the boardroom: adoption is near-universal, but measured value is rare. The differentiator between the winners and everyone else is not the model — it is whether leadership measures the right things and acts on them.

The 2025–2026 evidence converges hard:

FindingSource (year)
~88% of organizations use AI in ≥1 function, but only ~39% report enterprise-level EBIT impact — and most of those attribute under 5% of EBIT to itMcKinsey, State of AI (Nov 2025)
~95% of enterprise gen-AI pilots delivered no measurable P&L return; only ~5% reached rapid value. Root cause is the learning/integration gap, not model qualityMIT NANDA, GenAI Divide (Jul 2025)
Only ~5% of firms are "future-built" (value at scale); ~60% are "laggards" with minimal gains despite spendingBCG, Widening AI Value Gap (Sept 2025)

"Everyone has AI; almost nobody has AI value yet."

The gap is not a technology problem — it is a measurement-and-management problem. When pilots don't move EBIT, the cause is rarely the model; it is the absence of a baseline, a workflow redesign, and a disciplined number to hit. McKinsey's own analysis found that tracking well-defined KPIs is the single attribute most correlated with AI adoption — measurement is not the scoreboard after the game, it is part of how you win. These are time-stamped findings from a fast-moving field; treat the percentages as directional and re-verify them live before you quote them.

Key insight

The reframe

Measurement is not the report you write after the value arrives. Disciplined, well-defined KPIs are how value gets created — McKinsey found KPI tracking is the top driver of adoption, and adoption is the on-ramp to impact. If you can't measure it, you can't manage it, fund it, or defend it to the board.

The three-tier ROI stack

The core framework of this lesson is a simple ladder. AI value reveals itself in three tiers, and you must climb all three. Most organizations proudly report tier one and never get to tier three — which is exactly why their AI looks busy but their P&L doesn't move.

TierWhat it measuresExample metricsThe question it answers
1 · Adoption / usageAre people actually using it?Active users, usage frequency, % of eligible workforce, adoption velocity"Is it switched on?"
2 · Workflow efficiencyIs work getting faster, higher-throughput, better quality?Time saved per task, throughput, handling-time reduction, error/defect reduction, CSAT/NPS lift"Is the work itself better?"
3 · Business / P&L impactDid it change the financials?Revenue uplift, cost reduction, and ultimately % of EBIT attributable to AI"Did it move money?"

The discipline is to connect the tiers, not to stop at one. Adoption without efficiency is just a tool nobody benefits from. Efficiency without P&L impact is a productivity claim with no proof it reached the bottom line — time "saved" that is never reinvested or removed from cost is value that evaporates. The leadership move is to define, before you spend, what a win looks like at tier three and trace the line back down to the tier-one behaviors that produce it.

A caution from the research: be deliberate about where you chase tier-three value. MIT NANDA found more than half of gen-AI budgets flow to sales and marketing tools, while some of the largest near-term ROI sat in unglamorous back-office automation. The flashiest use case is not always the one that moves the most money.

Watch out

Where leaders get it wrong

They declare victory at tier one. "90% of the team logged in this month" feels like progress and proves nothing about value. A dashboard full of green tier-one metrics with no tier-three number attached is the single most common pattern behind the ~95% of pilots that show no P&L return (MIT NANDA, 2025).

Tip

The leadership move

For every funded initiative, write the tier-three number first — the specific revenue, cost, or EBIT figure it must move — then work down to the tier-two and tier-one metrics that should predict it. Value flows up the stack; accountability is designed down from the top.

Vanity metrics: the numbers that mislead

Some metrics are not just weak — they are actively misleading, because they can move in the opposite direction of value. These are the numbers a vendor or an over-eager team will put on a slide precisely because they always go up.

Vanity metricWhy it misleads
Seat licenses soldCounts what you bought, not what created value. A common failure: blanket per-seat licensing on day one, with 30–40% of seats unused within 90 days.
Logins / active usersMeasures presence, not productivity or P&L. People can log in daily and produce nothing.
Token consumptionCounts cost, not benefit. More tokens burned is more spend, not more value.
Suggestion-acceptance rateCounts agreement with the tool, not whether the output was correct or made anything better.

The most counter-intuitive trap is that these can be inversely correlated with value. Consider agentic AI: as work moves from one-human-one-seat toward one orchestrator running many agents, a more valuable, more automated workflow may show fewer seats and fewer human logins. If "seats" or "logins" is your success metric, your dashboard will fall exactly as your value rises. You would be punishing your best transformation.

The fix is to anchor every vanity metric to a tier-two or tier-three outcome. Seats only matter if utilization converts to time saved; time saved only matters if it converts to cost out or revenue in. Treat tier-one metrics as diagnostics (is anyone using this?) — never as the scoreboard.

Key insight

The reframe

A rising vanity metric is not evidence of value — and a falling one is not evidence of failure. As agents replace seats, the most valuable workflows can post the lowest seat and login counts. Judge initiatives by what reaches the P&L, not by what's easy to count.

Baseline first, and track quality — not just volume

Two disciplines separate credible measurement from theater: baseline before you start, and track quality, not just volume.

Baseline before you spend. If you don't capture today's handling time, resolution rate, error rate, CSAT, and cost before you turn AI on, you can never prove the change was AI's doing. "Vibe-based" spending — funding AI on fear-of-missing-out with no baseline and no hypothesis — is one of the most-cited reasons ROI goes unproven (Worklytics, 2025). A pre-registered baseline turns a debate about whether AI "feels" helpful into a measurable before-and-after.

Track quality, not just volume. Volume metrics (chats handled, tickets closed, content generated) are seductive because AI makes them spike. But volume can mask a quality collapse — and quality is what customers and revenue actually respond to. Always pair every volume metric with a quality metric: containment/deflection alongside CSAT/NPS, throughput alongside error/defect rate, speed alongside accuracy. This is the Klarna lesson (next callout), and it is the difference between a number that flatters you and a number that protects you.

Example

The Klarna reversal — the canonical cautionary tale

In 2024 Klarna's AI assistant handled ~two-thirds of customer-service chats (~2.3M conversations in month one), cut resolution time from ~11 minutes to under 2, did the work of ~700 agents, and was tied to a projected ~$40M profit improvement (Klarna press, Feb 2024 — vendor-sourced). On volume metrics it looked like a triumph. Then quality slipped: CSAT/NPS deteriorated, and by mid-2025 Klarna reversed course and rehired humans. The CEO's own framing: they "focused too much on efficiency and cost," the lower quality was "not sustainable." The volume metrics masked a quality collapse — for nuanced, empathetic work, augment-and-blend beat wholesale replacement. (Verify current status live.)

Watch out

Where leaders get it wrong

They ship a pilot, watch the volume chart go up and to the right, and declare success — with no baseline to compare against and no quality metric to catch the downside. By the time CSAT shows up in churn, the damage is done. Never let a volume number travel to the board without its paired quality number.

The high-performer line: % of EBIT attributable to AI — and a realistic clock

At the top of the stack sits the metric that separates serious programs from busy ones: the percent of EBIT attributable to AI. McKinsey uses precisely this line to identify high performers — the firms that can attribute a meaningful, growing share of earnings to AI, rather than gesturing at productivity anecdotes. As of late 2025, even among the ~39% of firms reporting any enterprise EBIT impact, most attribute under 5% of EBIT to AI (McKinsey, State of AI, Nov 2025) — which tells you both how early this is and how distinctive it is to be on the right side of that line.

Set a realistic clock. AI returns take longer than typical technology investments. Most organizations reach satisfactory ROI in roughly 2–4 years; only about 6% see it in under a year (Deloitte, AI ROI Paradox). Promising board-level EBIT impact in two quarters is how you manufacture the disappointment that gets a good program cancelled. The leadership move is to commit to the 2–4-year arc on tier three while demanding near-term movement on tiers one and two as leading indicators — proof you're on the path even before the EBIT lands.

Both the EBIT-share figure and the 2–4-year timeline are volatile, dated findings. Teach the durable concept — EBIT attribution is the high-performer line; AI ROI is a multi-year arc — and re-verify the exact numbers at the live sources before you cite them.

Tip

The leadership move

Make '% of EBIT attributable to AI' a standing line in your AI review — even if it reads "~0%" today. Naming the metric forces the organization to build the attribution discipline now, so that when value lands you can prove it. What gets measured at the top gets managed all the way down.

Watch out

Where leaders get it wrong

They under-promise on usage and over-promise on payback — pledging fast EBIT impact, then killing the program when 2–4-year returns don't show up in two quarters. Mismatched timelines, not weak technology, sink many otherwise-sound initiatives. Set the clock honestly up front.

The discipline of the kill-or-reshape decision

Measurement without consequence is decoration. The hardest — and most valuable — leadership act is the scaling decision: at the end of a pilot, every initiative either scales, gets reshaped, or gets killed. What you must not allow is pilot purgatory — the endless proof-of-concept that never faces a verdict. The share of companies abandoning most AI projects rose to roughly 42% in 2025, up from ~17% in 2024 (cited via Worklytics/S&P; trace to primary before quoting); the lesson is not to avoid killing pilots — it is to kill them deliberately and early, on evidence, rather than letting them drift.

Use a simple decision rule tied directly to the three-tier stack:

VerdictWhen to choose itWhat it looks like
ScaleMoves a defined tier-two/tier-three number, with quality holdingFund the 70% (people, process, workflow redesign); set EBIT-linked targets
ReshapePromising signal but the workflow, scope, or data is wrongRedesign the workflow around the tool — don't bolt it onto a broken process — and re-baseline
KillNo movement on a real number after a fair, baselined trialStop the spend, capture the lesson, redeploy the budget

The positive counter-example to pilot purgatory is JPMorgan, which runs 450+ AI use cases in production with benefits growing ~30–40% year over year — the hallmark of scoped, KPI-anchored deployment where each use case has to earn its place against a number. The contrast with the ~95% of pilots showing no P&L return is not luck; it is the discipline of measuring, then deciding.

Tip

The leadership move

Put a scaling decision on the calendar when you launch the pilot, not after it has quietly run for a year. Pre-commit to the number it must move and the date you'll judge it. "Kill, reshape, or scale — decide by Q3" is a sentence that ends pilot purgatory before it starts.

Example

Scoped and measured beats sprawling and vague

JPMorgan's 450+ production use cases, each anchored to a number and compounding ~30–40% in benefits per year, exemplify the alternative to pilot sprawl. The discipline isn't running fewer experiments — it's giving every experiment a number to hit and a date to be judged.

The right questions to ask

You don't need to build the dashboard — you need to interrogate it. These are the questions that separate a measured program from a flattering one. Keep them on a card for every AI review.

On any metric you're shown:

  • Which tier is this — adoption, efficiency, or P&L? (If everything on the slide is tier one, push back.)
  • Could this number go up while value goes down? (The vanity-metric test.)
  • What baseline is it measured against? Did we capture "before"?

On quality and risk:

  • For every volume metric, where is its paired quality metric (CSAT/NPS, error rate, accuracy)?
  • If this is "90% accurate," accurate how — and who owns the failing 10%?

On the P&L line:

  • What % of EBIT do we attribute to AI today, and what's the trajectory?
  • Is our timeline honest — are we judging a 2–4-year payback on a two-quarter clock?

On the verdict:

  • Which pilots haven't moved a number, and when is their kill-or-reshape decision?
  • When we report time saved, has it actually been reinvested or removed from cost — or did it just evaporate?

Asking these consistently does something subtle but powerful: it trains the organization to bring you tier-three thinking by default, because they know tier-one theater won't survive the meeting.

Key insight

The reframe

Your job is not to read the dashboard — it's to audit it. The most valuable thing a leader contributes to AI measurement is a short, repeatable set of questions that makes vanity metrics and unbaselined claims uncomfortable to present. Govern the metrics and you govern the value.

Try it: Build a three-tier ROI scorecard and pass two pilots through the kill-or-reshape gate

A strategic exercise — no tools, no code, just the discipline. 1) Pick two real (or realistic) AI initiatives in your organization — ideally one customer-facing (e.g. support assist) and one back-office (e.g. finance reporting). 2) Build a three-tier scorecard for each. In a simple table, fill one row per tier: Tier 1 (adoption — e.g. % of eligible workforce using it weekly), Tier 2 (workflow efficiency — name a time/throughput metric AND its paired quality metric, e.g. handle time + CSAT), Tier 3 (P&L — the specific revenue, cost, or EBIT number it must move). Write the tier-three number first. 3) Flag the vanity metrics. List any seats/logins/tokens/acceptance-rate numbers currently being reported and write one sentence on how each could rise while value falls. 4) Define the baseline. For each metric, state what 'before' value you would capture, and admit honestly where you have no baseline today. 5) Set an honest clock. Mark which tier-three results you'd expect inside a year vs. 2–4 years, and which tier-1/2 metrics are your leading indicators. 6) Run the gate. For each initiative, write the date and the threshold for its kill-or-reshape-or-scale decision — the exact number it must move by when. 7) Draft three boardroom questions you will ask at the next AI review to keep tier-one theater from passing as value. Deliverable: a one-page scorecard per initiative plus your three questions. This builds the instinct the rest of your AI program depends on — that every initiative carries a number, a baseline, a quality guardrail, and a verdict date.

Key takeaways

  1. 1AI's defining 2025–2026 reality is a measurement-and-management gap: ~88% of firms use AI but only ~39% report enterprise EBIT impact and ~95% of pilots show no P&L return — value stalls on management, not models (McKinsey & MIT NANDA, 2025).
  2. 2Use the three-tier ROI stack — (1) adoption/usage → (2) workflow efficiency (time, throughput, error/quality) → (3) business impact (revenue, cost, EBIT) — and never stop at tier one; define the tier-three number first and trace it down.
  3. 3Vanity metrics (seats, logins, tokens, suggestion-acceptance rate) can be inversely correlated with value — as agents replace seats, the most valuable workflows can post the lowest counts; treat tier-one numbers as diagnostics, not the scoreboard.
  4. 4Baseline before you spend and pair every volume metric with a quality metric (CSAT/NPS) — the Klarna reversal shows how strong volume numbers masked a quality collapse that forced a U-turn.
  5. 5'% of EBIT attributable to AI' is McKinsey's high-performer line; set an honest clock, because AI ROI typically takes 2–4 years (only ~6% under a year, per Deloitte) — demand near-term tier-1/2 movement as leading indicators.
  6. 6End pilot purgatory: schedule a kill-or-reshape-or-scale decision at launch and kill or reshape anything that won't move a defined number — scoped, KPI-anchored deployment (e.g. JPMorgan's 450+ use cases) is what beats the 95% failure rate.

Quiz

Lock in what you learned

Check your understanding

0 / 4 answered

1.Your AI program review opens with a slide showing seat licenses purchased, monthly logins, and tokens consumed — all trending up. What is the most accurate executive read of this dashboard?

2.What is the central lesson executives should draw from the Klarna customer-service AI case?

3.A division head proposes judging a new AI initiative purely on 'percent of EBIT attributable to AI' within the first two quarters. What is the soundest objection?

4.Six months in, a pilot shows enthusiastic usage but has not moved any tier-two or tier-three number against its baseline. What is the disciplined leadership response?

Go deeper

Hand-picked sources to keep learning