Proof Points and Cautionary Tales
Klarna, JPMorgan, and the both-sides case
- Be able to summarize the Klarna case as a both-sides story — the efficiency win and the quality-driven reversal — and explain what it teaches
- Be able to contrast wholesale replacement with an augment-and-blend approach for nuanced, empathetic work
- Be able to use JPMorgan's scoped, KPI-anchored portfolio as a contrast model for how to deploy AI at scale
- Be able to explain why customer support is the best-proven quick win, and which metrics must be tracked alongside volume
- Be able to read a vendor or press case study with a critical, both-sides lens and ask the right verifying questions
This lesson grounds the value discussion in real cases — the wins and the reversals — so you learn to read AI case studies critically instead of chasing headlines. Klarna is the canonical 'both-sides' story: a spectacular efficiency win that was partly reversed once quality fell, while JPMorgan models the disciplined, scoped, KPI-anchored alternative. The durable takeaway: customer support is the best-proven quick win, but only if you track quality (CSAT/NPS), not just volume.
- 1Why read cases critically
- 2Klarna, chapter one: the efficiency win
- 3Klarna, chapter two: the quality-driven reversal
- 4JPMorgan: the scoped, KPI-anchored contrast
- 5How to read any case study
Why read cases critically
Every AI vendor, consultancy, and press release offers a number that sounds like proof: handled two-thirds of chats, 55% faster, $40M uplift. Numbers like these are how budgets get approved and how flagship programs get launched — and they are also how programs quietly fail, because the headline almost never tells you what the failing fraction looked like or whether the win held up over time.
The single most valuable habit a leader can build in this module is to read a case study the way an auditor reads a forecast: What exactly was measured? Over what window? Who is the source — and what do they have to gain? And crucially, what happened next? Most published cases are a snapshot near launch, when the metrics look best. The interesting story is usually the second chapter.
This lesson works through two cases that have a documented second chapter. Klarna is the canonical both-sides story — a genuine efficiency win followed by a quality-driven course correction. JPMorgan is the contrast: a scoped, KPI-anchored portfolio that grows because every use case is tied to a number, not a narrative. Read together, they teach the discipline the headlines can't.
Key insight
The reframe
A case study is a marketing artifact until you find its second chapter. The win at launch is rarely the whole story; the durable lesson is in what the company did six to twelve months later.
Watch out
Where leaders get it wrong
Treating a vendor's launch-day metric as a settled outcome. Most published AI wins are measured at the moment they look best — near launch, on volume, before quality and durability are known.
Klarna, chapter one: the efficiency win
In February 2024, the fintech Klarna announced one of the most-cited AI deployments in the enterprise. Its AI customer-service assistant, built on a large language model, took on the front line of support — and the launch numbers were genuinely impressive.
| What Klarna reported (Feb 2024) | Figure |
|---|---|
| Share of customer-service chats handled | ~two-thirds (~2.3M conversations in month one) |
| Equivalent human workload | ~700 full-time agents |
| Average resolution time | 11 minutes → under 2 minutes |
| Repeat inquiries | ~25% fewer |
| Projected profit improvement (2024) | ~$40M |
| Reach | 23 markets, 35+ languages |
On the surface this is the textbook quick win, and it lines up with the strongest independent evidence we have. A landmark study by Brynjolfsson, Li and Raymond of 5,179 real support agents (NBER, 2023) found a gen-AI assistant raised issues resolved per hour by ~14% on average and ~34% for novice agents — exactly the kind of lift customer support is built to deliver: repeatable queries, a structured knowledge base, and a clear success metric. Customer support is, on the evidence, the best-proven place to start.
The trap is to stop reading here. Klarna's own figures are vendor-sourced — a press release, not an audited result — and every metric in that table is a volume or speed metric. None of them measures whether customers were actually well served.
Example
The independent evidence behind the win
Brynjolfsson, Li & Raymond, 'Generative AI at Work' (NBER, 2023), tracked 5,179 customer-support agents: a gen-AI assistant lifted issues-resolved-per-hour ~14% on average and ~34% for the least-experienced agents, and improved customer sentiment and retention by spreading top performers' know-how. Support is the best-proven quick win — re-verify the figures live.
Watch out
Every number here is volume or speed
Chats handled, resolution time, repeat inquiries, projected profit — all efficiency metrics. Not one of them tells you whether the answer was good. That gap is the whole story.
Klarna, chapter two: the quality-driven reversal
By mid-2025, the story had a second chapter the original headlines never anticipated. After leaning hard into automation and reducing its human workforce, Klarna found that customer satisfaction declined — and the company began rehiring humans to handle the work the AI did poorly. CEO Sebastian Siemiatkowski put it plainly:
"We focused too much on efficiency and cost… it led to lower quality… not sustainable."
This is why Klarna is the canonical both-sides case. Nothing in chapter one was false — the AI really did handle enormous volume, fast and cheaply. The failure was that the metrics chosen to declare victory (volume, speed, cost) were inversely informative about the thing that actually mattered (quality, trust, the nuanced and empathetic cases). The volume metrics looked spectacular precisely while the quality the business depended on was eroding underneath them.
The durable lesson is not "AI support doesn't work." It plainly does, for a large share of cases. The lesson is about what you measure and how far you push replacement:
- Volume metrics can mask a quality collapse. A system that resolves more tickets faster can be simultaneously making customers less happy. Track CSAT/NPS alongside throughput, and baseline both before you start.
- Augment-and-blend beats wholesale replacement for nuanced work. The empathetic, ambiguous, high-stakes conversations are exactly where the jagged frontier bites. The winning pattern is AI for the repeatable bulk, humans kept in the loop for the cases that need judgment — not a headcount-elimination target.
Tip
The leadership move
Before any support-automation pilot, define a quality floor — a CSAT or NPS line you will not cross — and instrument it from day one. If quality dips below the floor, the pilot fails, no matter how good the volume numbers look. Make quality a gate, not a footnote.
Key insight
Augment, don't amputate
The reversal wasn't 'AI vs. humans' — it was 'efficiency-only vs. efficiency-plus-quality.' Blend: let AI carry the repeatable volume and keep humans on the nuanced, empathetic, high-stakes tail. That is the sustainable design.
JPMorgan: the scoped, KPI-anchored contrast
If Klarna shows the danger of one big bet measured on volume, JPMorgan shows the opposite discipline. The bank has put 450+ AI use cases into production, with reported benefits growing on the order of 30–40% year over year — and the structural lesson is not the count but the shape of the portfolio.
Each use case is scoped and anchored to a specific KPI rather than launched as a flagship transformation hoping to move everything at once. This is the same conclusion the broader evidence keeps reaching: McKinsey's work on capturing agentic value argues that winners concentrate on a focused set of high-value use cases, embed them in redesigned workflows, and scale with governance — rather than running scores of disconnected pilots or one headline-grabbing replacement.
| Klarna (chapter one) | JPMorgan model | |
|---|---|---|
| Shape of the bet | One large, visible replacement | A portfolio of many scoped use cases |
| What success was tied to | Volume, speed, cost | A specific KPI per use case |
| Quality measured? | Not at launch | Anchored per use case |
| What it teaches | How a win can reverse | How value compounds when it's measured |
The contrast is the point. Klarna's chapter one is not a story of incompetence — it's a story of measuring the wrong thing and pushing replacement too far. JPMorgan's portfolio compounds precisely because each piece is small enough to measure honestly and is tied to a number a leader can defend.
Example
The disciplined model
JPMorgan reports 450+ AI use cases in production with benefits growing ~30–40% year over year — an example of scoped, KPI-anchored deployment where each use case earns its place by moving a specific number, not by generating a headline. (Figures are reported, not audited — verify live.)
Tip
The leadership move
Ask of every proposed AI initiative: 'Which single KPI does this move, and what's the baseline?' If there isn't a clear answer, it's a demo, not a use case. Concentrate on 2–3 high-value bets you can measure, and let the winners compound.
How to read any case study
You will be handed AI case studies constantly — by vendors selling a platform, by consultants selling a transformation, by your own teams seeking budget. Almost all of them are real and incomplete. The skill is not cynicism; it's a short, repeatable interrogation that separates a durable result from a launch-day snapshot.
The both-sides checklist — five questions for any AI case study:
- Source and incentive. Who published this, and what do they gain if you believe it? A vendor press release and a peer-reviewed field study are not the same evidence.
- What was measured. Is the headline a volume/speed/cost metric or a quality/outcome metric? If quality isn't in the story, assume it wasn't tracked.
- The failing fraction. If it's "90% accurate" or "two-thirds handled," what do the other 10% / one-third look like, and who is accountable for them?
- The time window. Is this a launch snapshot or a sustained result? Ask what the numbers looked like six to twelve months later — the second chapter.
- Transferability. Does this case share your data quality, your workflow, your regulatory exposure? A win in one context is a hypothesis, not a guarantee, in yours.
Applied to our two cases: Klarna passes question 1 only weakly (vendor-sourced), fails questions 2 and 4 at launch (volume-only, snapshot) — and that's exactly where the reversal came from. JPMorgan's portfolio is built to pass question 2 (a KPI per use case) and question 4 (benefits tracked year over year). The checklist would have predicted both outcomes.
Watch out
Vendor cases are marketing-flavored
A customer success story exists to sell the next deal. Treat its numbers as a hypothesis to verify, not a result to copy. The more impressive the single number, the harder you should look for what it leaves out.
Key insight
The one question that does the most work
'What happened next?' Forces the snapshot into a trajectory and surfaces the second chapter — the quality dip, the reversal, the durability — that the launch-day headline was never going to mention.
Try it: Run the both-sides checklist on a real case study
Goal: build the instinct to read AI case studies critically instead of acting on the headline. 1) Pick a case. Choose one AI case study relevant to your business — a vendor success story, a consultancy report, or an internal pilot's results deck. 2) Run the five questions. For each, write one or two sentences: (a) Source and incentive — who published it and what do they gain if you believe it? (b) What was measured — is the headline a volume/speed/cost metric or a quality/outcome metric, and was quality tracked at all? (c) The failing fraction — if it's '90% accurate' or 'two-thirds handled,' what do the failures look like and who is accountable? (d) The time window — is this a launch snapshot or a sustained result, and what (if anything) is reported 6–12 months later? (e) Transferability — does this case share your data quality, workflow, and regulatory exposure? 3) Write the verdict. In three sentences: Would you fund a program based on this case as written? What single additional data point would most change your answer? And which metric would you insist on instrumenting from day one if you ran this yourself? 4) Bonus: draft the one quality-floor metric (e.g., a CSAT or NPS line) you would make a hard gate for the pilot — the number that fails the pilot no matter how good the volume looks. Deliverable: a one-page critical read of the case using the checklist — the artifact you would bring to a steering-committee discussion.
Key takeaways
- 1Klarna is the canonical both-sides case: a real efficiency win (AI handled ~2/3 of chats, ~700 FTE-equivalent, ~$40M projected uplift) that was partly reversed by mid-2025 when CSAT fell and the company rehired humans — verify the figures live.
- 2The Klarna lesson is about measurement and reach: volume and speed metrics masked a quality collapse, and augment-and-blend beats wholesale replacement for nuanced, empathetic work.
- 3Customer support is the best-proven quick win (NBER: ~14% average, ~34% for novices), but quality metrics like CSAT/NPS must be tracked and baselined alongside volume — and treated as a gate, not a footnote.
- 4JPMorgan models the contrast: 450+ scoped, KPI-anchored use cases with benefits growing ~30–40% year over year — value compounds when each use case moves a defined number rather than chasing a headline.
- 5Vendor and press case studies are marketing-flavored snapshots; read them with a both-sides lens and always ask 'what happened next?'
- 6Use a five-question checklist on every case: source and incentive, what was measured, the failing fraction, the time window, and transferability to your context.
Quiz
Lock in what you learned
Check your understanding
0 / 4 answered
1.Why is Klarna described as the canonical 'both-sides' AI case study?
2.What is the core lesson of the Klarna reversal for how leaders should run a support-automation pilot?
3.What makes JPMorgan's reported AI approach a useful contrast to Klarna's launch story?
4.A vendor shows you a case study claiming its AI 'resolves 90% of tickets and cut handle time by half.' Using the both-sides lens, which question does the most to reveal whether this is a durable result?
Go deeper
Hand-picked sources to keep learning
Chapter one: the original (vendor-sourced) press release with the ~700-FTE, 11-min-to-2-min, ~$40M figures. Read it as marketing — the launch snapshot.
Chapter two: the mid-2025 reversal after CSAT fell. The 'what happened next' that makes Klarna the canonical both-sides case.
The independent field study (5,179 agents) behind support being the best-proven quick win: ~14% average and ~34% novice productivity gains.
Why winners concentrate on a focused set of high-value use cases in redesigned workflows rather than scores of disconnected pilots — the discipline behind the JPMorgan-style portfolio.
Context for why most pilots show no measurable P&L return — a 'learning gap' of workflow and measurement, not model quality. Verify figures live.
The adoption-vs-value backdrop: most volatile adoption, EBIT-impact, and agent-scaling figures live here — re-verify before teaching.