Augmentation Beats Replacement
AI lifts your novices most — the evidence
- Be able to summarize the Brynjolfsson, Li & Raymond support-agent study and what its +14% overall / +34% novice / ~0 expert pattern means for the workforce
- Be able to cite the GitHub Copilot experiment as corroborating evidence and explain why less-experienced workers gain most
- Be able to explain the knowledge-transfer mechanism — how AI compresses skill gaps by encoding top performers' tacit know-how
- Be able to separate a task-level speedup from genuine organizational throughput, and name what eats the difference
- Be able to make the strategic case for designing roles around human + AI rather than treating AI as a replacement play
The strongest field evidence on AI at work points one direction: it lifts your less-experienced people the most and spreads your best people's know-how across the team. This lesson turns two rigorous studies — a 5,179-agent customer-support trial and a GitHub Copilot coding experiment — into a strategy choice: build for human + AI augmentation, not headcount replacement. It also flags the trap that sinks ROI cases — confusing a task-level speedup with org-level throughput.
- 1The headline: AI lifts your novices most
- 2Evidence #1: 5,179 support agents
- 3Evidence #2: the GitHub Copilot experiment
- 4The mechanism: AI as a knowledge-transfer machine
- 5The caveat that sinks ROI cases: task speed ≠ throughput
- 6The strategic choice: design for human + AI
The headline: AI lifts your novices most
If you remember one finding from the research on AI in the workplace, make it this: the biggest, most reliable productivity gains from AI go to your least-experienced people. That single fact should reshape how a leader thinks about where to deploy AI and what it is for.
The intuition runs the other way for most executives. We assume the expert — the person who already knows the most — will squeeze the most out of a powerful new tool. The field evidence says the opposite. A novice paired with AI starts to perform like a mid-level employee; the expert, who already operates near the top of the task, has little headroom to gain.
This reframes the entire conversation. AI is not primarily a tool for replacing people. It is, on the best current evidence, a tool for raising the floor — pulling your weakest performers up toward your average, and your average up toward your best. That is an augmentation story, and it is a far more defensible bet than a replacement story, because it grows your existing capacity rather than betting that the technology can do a whole job alone.
The leadership question shifts from "how many roles can AI replace?" to "how much can AI raise the capability of the people I already have?"
Key insight
The reframe
Replacement asks 'what can we cut?' Augmentation asks 'how high can we lift everyone?' The evidence in this lesson supports the second question — and the second question is the one with a stronger, lower-risk business case.
Evidence #1: 5,179 support agents
The cleanest, largest piece of evidence comes from a study by Brynjolfsson, Li & Raymond, Generative AI at Work (NBER, 2023), run across 5,179 customer-support agents at a software firm. The agents were given a generative-AI assistant that suggested responses in real time, and their performance was measured against agents working without it.
The results (as of the 2023 study — verify the live source before quoting):
| Group | Effect of the AI assistant on issues resolved per hour |
|---|---|
| All agents (average) | +14% |
| Novice / low-skilled agents | +34% |
| Experienced / high-skilled agents | ~0% (little to no gain) |
Three things make this study unusually trustworthy for executives. First, it is large and real — over five thousand agents doing actual paid work, not a lab simulation. Second, it measured a hard business outcome — issues resolved per hour — not a soft proxy like "satisfaction with the tool." Third, it found two bonus effects beyond raw speed: customer sentiment improved, and employee retention rose, because the job got less frustrating and newer agents ramped faster.
The shape of the result — big gains for novices, near-zero for experts — is the augmentation thesis in one chart. AI did not replace the support team. It made the less-experienced half of the team dramatically more productive.
Example
Why the experts barely moved
Experienced agents already knew the best answers and the fastest paths, so the AI's suggestions rarely told them anything new. The novices, by contrast, were handed the expert playbook in real time — which is exactly why their numbers jumped 34%.
Tip
The leadership move
When piloting AI assist, point it at functions with a wide skill spread and a steady stream of new or junior staff (support, sales development, claims, onboarding). That is where the +34% lives — not in your team of seasoned experts.
Evidence #2: the GitHub Copilot experiment
A second study from an entirely different domain — software engineering — found the same pattern, which is what makes the augmentation thesis credible rather than a one-off.
In a controlled experiment, GitHub had developers complete a coding task (building an HTTP server) with and without its Copilot assistant. The developers using Copilot finished ~55.8% faster (the study reported a wide confidence interval, roughly 21–89%, as of 2023 — verify live), and a higher share completed the task at all (78% vs. 70%). As with the support study, the largest gains went to the less-experienced developers.
Two different functions, two different research teams, two different kinds of work — and the same directional finding: AI augments the less-experienced most, and the gains are real and measurable on tasks that sit inside the technology's frontier. When independent studies in unrelated domains converge, a leader can treat the pattern as a planning assumption rather than a vendor anecdote.
Watch out
Where leaders over-read this
A 55.8% speedup on one well-defined task is not a 55.8% increase in your engineering org's output. It is a measurement of one isolated task, under ideal conditions, with no review, integration, or rework counted. We unpack exactly why in the throughput section — do not put the raw number in a board ROI case.
The mechanism: AI as a knowledge-transfer machine
Why do novices gain most? The answer is the strategically useful part, because it tells you what AI is actually doing inside your organization.
The AI assistant in the support study had effectively absorbed the patterns of the best agents' answers and was surfacing them, in the moment, to everyone. A new hire who would have taken months to learn the right phrasing, the right troubleshooting path, and the right tone was handed that tacit knowledge instantly. In effect, AI compressed the skill gap — it moved know-how that normally lives in your top performers' heads out to the rest of the team.
This is the deepest reason augmentation beats replacement as a frame: the value is not just "a tool that types faster." It is a mechanism that diffuses your best people's expertise across your whole workforce. Top performers' know-how is usually your most valuable, least transferable asset — locked in individuals, lost when they leave. AI turns some of that tacit knowledge into something the whole team can draw on.
| Without AI | With AI as knowledge transfer |
|---|---|
| Expertise lives in a few heads | Expertise is surfaced to everyone, in the moment |
| New hires ramp slowly, by osmosis | New hires get the expert playbook on day one |
| Top performers' know-how leaves when they do | Some of it is captured and reused |
| Skill gap between best and worst is wide | Skill gap compresses — the floor rises |
Reframed for the board: AI's clearest near-term value is spreading the know-how you already have, not acquiring know-how you don't.
Key insight
The asset you're activating
Your best people's tacit knowledge is real capital sitting on your balance sheet that you can't see. The augmentation play is to capture and distribute it — which is also why your data and your experts' patterns matter more than the model you license.
The caveat that sinks ROI cases: task speed ≠ throughput
Here is where most AI business cases quietly fall apart, and where a sharp leader earns their seat: a task-level speedup is not the same as organizational throughput.
The studies measured one task in near-ideal conditions. A 55.8% faster coding task or a 14% lift in issues-per-hour is a true result — for that task. But the value your organization actually banks is the work that flows all the way through to a shipped, correct, integrated outcome. Between the fast task and the realized value sit several things the studies did not count:
- Review. Faster drafts still need human review — and AI output, which never says "I don't know," often needs more scrutiny, not less.
- Integration. A function that's quicker in isolation can stall at the handoff if the next step in the workflow wasn't redesigned around it.
- Rework. Output that looks done but is subtly wrong creates downstream rework that can erase the upfront time saved.
- The bottleneck may be elsewhere. Speeding up a step that wasn't the constraint produces no net throughput gain at all.
This is why so many organizations report local speedups that never show up in the financials. The honest executive position is: treat task-level numbers as evidence the capability is real, then ask the harder question — does this speedup survive review, integration, and rework to become value the business can actually count?
Watch out
Where leaders get it wrong
Taking a vendor's task-level benchmark ("55% faster!") and multiplying it across the headcount to project savings. That math is almost always wrong because it ignores review, integration, rework, and the real bottleneck. Demand end-to-end, outcome-level measurement before you commit to a number.
Tip
The question to ask
When someone presents an AI productivity number, ask: 'Is that a task-level speedup or a measured change in what we actually ship? What did you count between the faster task and the finished, correct outcome?'
The strategic choice: design for human + AI
Put the evidence together and a clear strategy choice emerges. The augmentation case rests on three durable findings: AI lifts the less-experienced most, it works by transferring expert know-how across the team, and its task-level gains only become value when you redesign the work around them. None of those findings supports a wholesale replacement play.
So the strategic posture is to design roles around human + AI, not to swap one for the other:
| Replacement framing | Augmentation framing | |
|---|---|---|
| Goal | Cut headcount, automate the job | Raise the capability of every person |
| Where value comes from | Labor cost removed | Floor raised, expertise spread, faster ramp |
| Risk profile | High — bets the tech can do a whole job alone | Lower — grows capacity you already have |
| What you redesign | The org chart | The workflow and the role |
| Evidence support | Thin for nuanced work | Strong (NBER, GitHub) |
Replacement is the riskier bet on every axis, and — as the canonical Klarna case showed — it can reverse: Klarna replaced roughly 700 customer-service roles with AI in 2023, looked excellent on volume metrics, then saw customer-satisfaction quality deteriorate and rehired humans by mid-2025, with its CEO conceding the company had focused too much on cost at the expense of quality. The augment-and-blend approach the evidence supports would likely have avoided that round trip.
The practical implication: aim AI at lifting people, redesign the role and the workflow around the human-plus-AI pairing, keep a human accountable for judgment and quality, and measure the outcome, not the task.
Example
Klarna: the both-sides case
Klarna's AI handled about two-thirds of support chats and cut resolution time from ~11 minutes to under 2 (vendor-reported, Feb 2024) — then CSAT fell and it rehired humans by mid-2025. The lesson isn't 'AI failed'; it's that wholesale replacement of nuanced, empathetic work backfired where augment-and-blend would have held. Verify the live sources before citing the figures.
Try it: Map your augmentation opportunity (one page)
A strategic exercise, no technology required. The goal is to turn this lesson's evidence into a deployment decision for your own organization.
1) Pick three functions with a wide spread between your best and weakest performers and a steady flow of new or junior staff (candidates: customer support, sales development, claims, onboarding, junior analysts). These are where the '+34% for novices' effect lives.
2) For each, fill a four-column row: (a) the repetitive, knowledge-heavy task; (b) where your top performers' tacit know-how currently lives and how new hires learn it today; (c) the augmentation hypothesis — what would a novice-plus-AI pairing look like, and which expertise would AI surface; (d) the honest skeptic's column — what review, integration, or rework would have to happen for a task speedup to become real throughput.
3) Write the two measurement questions you will demand of any pilot in these areas: one that distinguishes a task-level speedup from an outcome-level result, and one that tracks quality (e.g., CSAT/NPS, error rate) so you don't repeat Klarna's volume-over-quality mistake.
4) Make the call. For each function, state in one sentence whether you'd pursue augmentation (human + AI, role redesigned) or hold — and why. Note explicitly any place you were tempted toward replacement and what the evidence in this lesson says about that temptation.
Deliverable: a single page you could put in front of your leadership team that names where to aim AI to raise the floor, and how you'll know if it actually worked.
Key takeaways
- 1The strongest field evidence shows AI's biggest gains go to less-experienced workers: in a 5,179-agent support study, +14% issues/hour overall, +34% for novices, and ~0 for experts (Brynjolfsson, Li & Raymond, NBER 2023 — verify live).
- 2An independent GitHub Copilot experiment found a coding task ~55.8% faster with the largest gains for less-experienced developers — convergent evidence from a different domain.
- 3The mechanism is knowledge transfer: AI surfaces top performers' tacit know-how to the whole team in the moment, compressing the skill gap and raising the floor.
- 4A task-level speedup is not organizational throughput — review, integration, rework, and the real bottleneck eat the difference, which is why local gains often never reach the financials.
- 5The strategic choice is to design for human + AI augmentation, not replacement: it's lower-risk, better-evidenced, and avoids the quality collapse that forced Klarna to reverse course and rehire humans.
- 6Demand outcome-level, end-to-end measurement before committing to an ROI number — never multiply a vendor's task-level benchmark across your headcount.
Quiz
Lock in what you learned
Check your understanding
0 / 4 answered
1.In the Brynjolfsson, Li & Raymond study of 5,179 customer-support agents, which group benefited MOST from the AI assistant, and why does that matter strategically?
2.An executive sees that a GitHub study found developers completed a coding task ~55.8% faster with Copilot, and proposes projecting a 55.8% reduction in the engineering department's costs. What is the central flaw?
3.The lesson describes AI as a 'knowledge-transfer machine.' What does this mean, and why is it strategically valuable?
4.Based on the evidence in this lesson, what is the soundest strategic posture toward AI in the workforce?
Go deeper
Hand-picked sources to keep learning
The 5,179-agent customer-support study: +14% issues/hour overall, +34% for novices, ~0 for experts. Re-verify figures before quoting.
The controlled experiment behind the ~55.8% faster coding-task figure, with the largest gains for less-experienced developers.
Companion evidence on where AI helps vs. hurts (758 BCG consultants) — context for calibrated, task-by-task augmentation.
The canonical replacement-reversal case: ~700 support roles replaced in 2023, humans rehired by mid-2025 after quality fell.
The original vendor-sourced claims (2.3M chats, ~700 FTE-equivalent, 11 min → <2 min). Treat as marketing-flavored; verify live.