The Jagged Frontier
AI is unevenly brilliant and the edge is invisible
- Explain the jagged frontier — why AI capability is uneven across tasks and the edge between 'good at this' and 'bad at this' is invisible
- Cite the BCG/HBS field experiment to show AI's effect inside vs. outside the frontier, and where leaders can re-verify the numbers
- Reframe the core risk from 'weak AI' to 'over-trusted AI' — the danger of confident output where it is silently wrong
- Apply calibrated trust: decide task-by-task where AI helps, where it hurts, and where a human check must sit before AI acts
- Distinguish the two human-AI collaboration patterns — centaurs and cyborgs — and choose configurations deliberately rather than making a binary call
- Lead the right conversation in your organization: interrogate accuracy claims and keep accountability human
AI is not uniformly good or bad — it is sharply brilliant inside an invisible boundary called the jagged frontier, and it degrades, often confidently, just outside it. This lesson uses the landmark BCG/HBS consultant experiment to make that concrete, then draws the leadership lesson: not 'use AI / don't,' but calibrated, task-by-task trust, with the human keeping final judgment.
- 1The most important nuance: the edge is invisible
- 2The experiment that made it concrete
- 3Why AI fails the way it does
- 4Calibrated trust: the skill that replaces the on/off switch
- 5Centaurs and cyborgs: two ways to work with AI
- 6The conversation a leader must lead
The most important nuance: the edge is invisible
If you take one idea from this course about what AI can and cannot do, make it this: AI capability is unevenly good across tasks, and the boundary between what it does well and what it does badly is invisible.
We instinctively grade tools the way we grade people — "this one is strong" or "this one is weak," a single dial of competence. AI does not behave that way. On one task it produces work that rivals your best people. On a neighboring task that looks just as easy, it fails — and, crucially, it fails with the same fluent confidence it shows when it succeeds. Researchers named this shape the jagged frontier: a capability map with sharp peaks and deep valleys sitting right next to each other, with no warning label marking where one ends and the other begins.
The risk is not weak AI. The risk is over-trusting AI exactly where it is silently bad.
That single reframe changes the leadership question. The naive debate — "should we use AI or not?" — is the wrong altitude. The real question is "on which tasks does AI lift us, on which does it quietly drag us down, and how would we even tell the difference?" A leader who understands the jagged frontier stops issuing blanket mandates and starts governing trust task by task.
Key insight
The reframe
Most leaders picture AI as a single competence dial — better or worse than a human. The truth is a jagged map: brilliant and incompetent on adjacent, equally-easy-looking tasks, with no visible line between them. Govern the line, not the dial.
Watch out
Where leaders get it wrong
Treating 'use AI' as a one-time yes/no decision. A blanket 'yes' pushes the tool onto tasks where it silently degrades quality; a blanket 'no' forfeits real, proven gains. Both miss the point: the answer is task-by-task, not organization-wide.
The experiment that made it concrete
The frontier moved from intuition to evidence in a 2023 field experiment run by researchers at Harvard Business School and BCG (Dell'Acqua, Mollick, Lakhani and colleagues). They gave 758 BCG consultants real consulting-style tasks, with some randomly given access to a frontier AI model and some not. Then they split the tasks into two groups: ones that sat inside the AI's capability frontier, and ones engineered to fall outside it.
The results are the cleanest illustration of the jagged frontier we have:
| Task INSIDE the frontier | Task OUTSIDE the frontier | |
|---|---|---|
| Speed | ~25% faster | — |
| Quality | ~40% higher quality | — |
| Completion | +12% more tasks completed | — |
| Correctness vs. peers without AI | Higher | ~19 percentage points LESS likely to be correct |
| Leadership read | A genuine performance lift | A confident, invisible drag |
Same people. Same tool. Same apparent difficulty. Opposite effect — depending only on which side of an invisible line the task fell. Inside the frontier, AI was a powerful accelerant. Outside it, consultants using AI were about 19 percentage points less likely to reach the correct answer than colleagues working without it — because the AI gave fluent, plausible, wrong answers, and capable professionals trusted them.
These figures are time-stamped findings from one landmark study, not permanent laws — re-verify them at the live source before you quote them in a board meeting.
Example
BCG/HBS 'Navigating the Jagged Technological Frontier' (2023)
758 BCG consultants; inside the frontier AI users were ~25% faster, ~40% higher quality, +12% completion; outside it they were ~19 percentage points LESS likely to be correct than non-AI peers. Source: HBS AI Institute and SSRN 4573321 (published in Organization Science). Re-verify the numbers before teaching them.
Key insight
Why 'silently' is the dangerous word
If AI failed loudly — error messages, blank outputs — no one would be misled. The damage comes precisely because outside-the-frontier failures arrive looking exactly like inside-the-frontier successes: confident, fluent, well-formatted. The packaging is identical; only the substance is wrong.
Why AI fails the way it does
You don't need the math to govern this, but you do need the intuition, because it explains why the failures are confident rather than obvious.
The text engine under most generative AI is a large language model — think of it as autocomplete on steroids. It is trained to predict "what plausibly comes next," which is exactly why it is so fluent and articulate. But it is predicting plausible text, not retrieving verified fact. So when a task drifts outside what it reliably knows, it does not stop and say "I'm not sure." It keeps producing confident, well-formed output — it simply makes things up. The industry word for this is hallucination, and it is an inherent property of how these systems work, not a bug awaiting a patch.
A useful analogy for the boardroom:
Generative AI is a brilliant, fast, eager intern who has read almost everything, has no memory of your company, never says "I don't know," and occasionally invents things with total confidence. You delegate to it — and you check its work.
That single image carries the whole governance lesson. You would never let a first-week intern send a regulatory filing or approve a loan unsupervised, no matter how articulate they sounded. AI deserves the same posture: delegate generously, verify deliberately, and put a human check wherever a wrong answer is load-bearing.
Tip
The leadership move
Anchor your organization on one phrase: 'It never says I don't know.' Once people internalize that AI will not flag its own uncertainty, they stop expecting the tool to warn them — and start designing the human check that the tool cannot provide for itself.
Calibrated trust: the skill that replaces the on/off switch
The opposite of a blanket AI mandate is not caution — it is calibrated trust: deciding, task by task, how much to rely on AI based on two things you can assess even without technical depth — how reliable AI is on this kind of task, and how costly a wrong answer would be.
A simple 2x2 makes it usable on Monday morning. Sort any candidate task on how load-bearing a wrong answer is (low to high stakes) against how well-suited the task is to AI (a fluent-but-bounded drafting job vs. novel judgment):
| Low stakes (a wrong answer is cheap) | High stakes (a wrong answer is costly) | |
|---|---|---|
| Well-suited to AI (drafting, summarizing, triage) | Let AI run — lightweight review | AI drafts, human approves — verification before it counts |
| Poorly-suited to AI (novel strategy, final judgment) | Use AI to expand options, not decide | Keep human-led — AI as a sounding board at most |
The enemy of calibrated trust is automation bias — the well-documented human tendency to over-trust confident, automated output and switch off our own scrutiny. The more fluent the answer, the stronger the pull. Calibrated trust is the deliberate counterweight: you decide in advance where the human check sits, rather than discovering after the fact that everyone quietly deferred to the machine.
Use AI to expand your options, not to make your choice. Final judgment and accountability stay human — not as a courtesy, but because the tool cannot be held accountable and cannot tell you when it is wrong.
Watch out
Automation bias is the quiet failure mode
The jagged-frontier damage doesn't come from people who distrust AI — it comes from capable professionals who trust a confident answer and stop checking. Assume your best people are the most exposed, because they're the ones moving fast and least likely to second-guess a polished output.
Tip
The leadership move
Don't ask 'are we using AI?' Ask 'where have we decided AI can run on its own, where must a human approve before it counts, and who owns that decision?' Calibrated trust is a governance choice you make explicitly — or one that gets made for you by default.
Centaurs and cyborgs: two ways to work with AI
The same researchers observed that the people who got the most from AI did not all work the same way. They fell into two patterns — useful labels for any leader designing how their teams collaborate with AI.
| Centaur | Cyborg | |
|---|---|---|
| Pattern | Cleanly divides the work between human and AI | Interweaves human and AI throughout the task |
| Mental image | A clear handoff line — AI does its half, human does theirs | A constant back-and-forth, blended turn by turn |
| Best when | The frontier is roughly known — you can assign AI the tasks it's good at | The frontier is fuzzy — you stay in the loop to catch errors as they appear |
| Example | "AI drafts the first version; I write the strategic recommendation" | "I co-write each paragraph with AI, editing as I go" |
The leadership takeaway is not to crown one pattern the winner. It is to recognize that "how should humans and AI work together here?" is a real design decision — and to make it deliberately, task by task, instead of defaulting to "everyone just uses the chatbot however." Centaur thinking suits tasks where you can confidently map the frontier and hand AI a bounded slice. Cyborg thinking suits messier tasks where the frontier is uncertain and a human staying interwoven is your error-catching mechanism.
Either way, the human is present by design — dividing or interweaving, never absent. That presence is what turns the jagged frontier from a hidden liability into a managed one.
Example
Centaurs vs. cyborgs (HBS, 2023)
From the same 758-consultant study: 'centaurs' split tasks between human and AI along a clear line; 'cyborgs' blend their workflow with AI continuously. The lesson for leaders is to choose the configuration that fits the task, not to make a binary 'use AI / don't' call.
The conversation a leader must lead
Calibrated trust only becomes real when it shows up in how your organization talks about AI. Three habits turn the concept into practice.
1. Interrogate every accuracy claim. When a vendor or a team says "it's 90% accurate," never accept the headline. Ask the questions that reveal the shape of the risk:
- "Accurate in what way — and what do the failing 10% actually look like?"
- "Where is a wrong answer load-bearing, and what human check sits before it acts?"
- "Who is accountable for the failures?"
The shape of the errors matters far more than the headline number. Ten percent of harmless typos is nothing; ten percent of confidently wrong medical, legal, or financial conclusions is a crisis.
2. Use AI to expand options, not make the choice. AI is superb at generating drafts, alternatives, and syntheses — and weak at novel, accountable judgment. Research has found that when asked for strategic advice, models tend to return generic, fashionable "trendslop" rather than genuine insight, and can subtly flatter the person asking. Treat AI as a tireless option-generator and keep the decision human.
3. Keep accountability human and explicit. Every consequential AI-assisted output needs a named human owner who can stand behind it. The tool cannot be held to account, cannot explain itself to a regulator, and cannot feel the weight of a wrong call. That is — and should remain — a person's job.
Tip
The leadership move
Make 'Accurate how? What do the failures look like? Who owns them?' a standing question in every AI vendor pitch and pilot review. Modeling that interrogation in the room teaches your whole organization to stop trusting headline accuracy numbers.
Watch out
Where leaders get it wrong
Asking AI to make the strategic call. Studies find models hand back generic 'trendslop' and can quietly flatter the asker. Use AI to widen the set of options you consider — then make the decision yourself, and own it.
Try it: Map your team's jagged frontier
Goal: turn calibrated trust from a concept into a one-page artifact you could take into a leadership meeting. 1) Pick a function you own (e.g. marketing, customer support, finance) and list 8–10 recurring tasks your team actually does. 2) Plot each task on the 2x2 from this lesson: 'how load-bearing is a wrong answer?' (low → high stakes) against 'how well-suited is this to AI?' (fluent-but-bounded drafting → novel judgment). Use a simple table — one row per task, two columns for the two axes, plus a 'verdict' column. 3) Assign a verdict to each: 'let AI run,' 'AI drafts / human approves,' or 'keep human-led.' For every task in the high-stakes column, write the specific human check that must sit before AI output counts, and name who owns it. 4) Choose a collaboration pattern (centaur = divide, or cyborg = interweave) for your three highest-value tasks, and say in one line why that pattern fits. 5) Write the interrogation script: three questions you will ask the next time a team or vendor claims an AI tool is 'X% accurate' (accurate how? what do the failures look like? who owns them?). 6) Identify your blind spot: name one task where your organization is probably already over-trusting AI on the wrong side of the frontier, and the first thing you'll do about it. Deliverable: a single page — the task map, the human-approval gates with named owners, and your interrogation script — that makes 'where AI helps vs. hurts here' an explicit, governed decision instead of a default.
Key takeaways
- 1AI capability is jagged, not uniform: it is brilliant inside an invisible frontier and degrades just outside it — and the two failures look identical because both arrive confident and fluent.
- 2The real risk is over-trust, not weakness: in the BCG/HBS 2023 study of 758 consultants, AI users were ~25% faster, ~40% higher quality and +12% completion inside the frontier, but ~19 percentage points LESS likely to be correct outside it (re-verify at the live source).
- 3AI fails silently because large language models predict plausible text, not verified fact — they never say 'I don't know,' so the human must supply the check the tool cannot.
- 4Replace the on/off switch with calibrated trust: decide task-by-task using stakes (how load-bearing a wrong answer is) against AI suitability, and place a human approval gate wherever a wrong answer counts.
- 5Centaurs divide work between human and AI; cyborgs interweave it — choose the collaboration pattern deliberately per task rather than making a binary 'use AI / don't' call.
- 6Interrogate accuracy claims ('Accurate how? What do the failures look like? Who owns them?'), use AI to expand options not make the choice, and keep final judgment and accountability human.
Quiz
Lock in what you learned
Check your understanding
0 / 4 answered
1.What does the 'jagged frontier' most precisely describe about today's AI?
2.In the 2023 BCG/HBS field experiment with 758 consultants, what happened on tasks that fell OUTSIDE the AI's frontier?
3.A team reports that a new AI tool is '92% accurate.' What is the strongest leadership response?
4.Which statement best captures the difference between 'centaurs' and 'cyborgs' as human-AI collaboration patterns?
Go deeper
Hand-picked sources to keep learning
The landmark 758-consultant field experiment: inside vs. outside the frontier, and centaurs vs. cyborgs. Re-verify the headline figures here before quoting them.
The complete study (published in Organization Science) with methodology and the inside/outside-the-frontier results.
Brynjolfsson, Li & Raymond: AI helps most where it is genuinely good — +14% on average, +34% for novices — useful context for which tasks sit inside the frontier.
On interrogating accuracy claims, using AI to expand options not make the choice, and where to move fast vs. prioritize control.
Why AI is weak at novel, accountable strategic judgment and can subtly flatter the asker — the case for keeping decisions human.
Practical framing of calibrated trust, automation bias, and the right questions for non-technical leaders.