Agentic AI AcademyAgentic AI Academy

Genuinely Good vs. Genuinely Risky

Where value is real and where leaders over-trust

Beginner 12 minDecision-maker
What you'll be able to do
  • Be able to name the four task types where today's AI reliably creates value and explain why each works
  • Be able to identify the categories where AI is over-trusted — reliability, novel strategy, and final judgment — and why
  • Be able to apply the 'jagged frontier' idea to decide whether a given task is inside or outside AI's competence
  • Be able to recognize 'trendslop' and sycophancy when AI is used for strategic advice
  • Be able to separate a polished demo from a production-ready capability when evaluating a vendor or pilot
  • Be able to draw the line that keeps accountability and final judgment human while still capturing AI's upside
At a glance

Today's AI is not uniformly capable — it is brilliant at some tasks and quietly terrible at others, with an invisible boundary between the two. This lesson gives you a durable map of where AI value is genuinely real (drafting, synthesis, triage, augmenting less-experienced staff, search over your own documents) versus where leaders systematically over-trust it (factual reliability, novel strategy, anything requiring accountability). Learn the move and the warning so you deploy AI where it pays and guard the places it silently fails.

  1. 1The reframe: capability is jagged, not uniform
  2. 2Where AI is genuinely good now
  3. 3Where leaders over-trust it
  4. 4The good / risky map at a glance
  5. 5A great demo is not a production capability
  6. 6The questions that keep you calibrated

The reframe: capability is jagged, not uniform

The instinct most leaders bring to AI is to ask "is it good or not?" — as if capability were a single dial. That question is the trap. AI is unevenly good across tasks, and the boundary between what it does brilliantly and what it does badly is invisible. Researchers call this the jagged (technological) frontier.

The evidence is precise. In the BCG/HBS field experiment (Dell'Acqua, Mollick, Lakhani et al., 2023; 758 BCG consultants), on tasks inside the frontier, consultants using AI were ~25% faster, produced ~40% higher-quality work, and completed ~12% more tasks. But on tasks outside the frontier — equally hard-looking, just unsuited to the tool — AI users were ~19 percentage points less likely to reach the correct answer than peers working without AI.

Read that again: the danger is not that AI is weak. The danger is that it is confidently, fluently wrong in exactly the places you can't see the edge. The leadership job, therefore, is not "adopt AI" or "avoid AI." It is calibration — knowing which side of the jagged line a task sits on, and deploying or guarding accordingly. The rest of this lesson is a map of the two sides.

Key insight

The question to retire

"Is the AI good?" is the wrong question — it implies a single dial. The right question is "Is this task inside or outside the frontier?" The same tool that makes one task 40% better makes the next one 19 points worse. Capability is jagged, not a slider.

Where AI is genuinely good now

Strip away the hype and there is a stable, evidence-backed list of where today's AI reliably earns its keep. Notice the common thread: these are bounded, verifiable tasks with a human in the loop and fast feedback — not open-ended judgment calls.

StrengthWhat it looks like in your businessWhy it works
Drafting & summarizingFirst drafts of emails, reports, proposals; condensing long documents and meetings into key pointsLanguage synthesis is the core competence of an LLM; a human edits the draft, so errors get caught
Classification & triageTagging, sorting, routing tickets, flagging documents for review at scalePattern recognition over high volume; the cost of a single miss is low and recoverable
Augmenting less-experienced workersFront-line staff getting expert-level suggestions in the momentAI spreads top performers' know-how to everyone else (the augmentation effect)
Search over your own documents (RAG)A grounded assistant that answers from your approved, current knowledge baseAnchoring answers to your documents cuts hallucination and adds traceability

The augmentation evidence is the strongest single data point a leader can cite. In Brynjolfsson, Li & Raymond's "Generative AI at Work" (NBER, 2023; 5,179 customer-support agents), a GenAI assistant raised issues resolved per hour by ~14% on average and ~34% for novices, with near-zero effect for experts — it lifted the bottom of the distribution by transferring the best workers' know-how. The pattern repeats in software: a GitHub Copilot randomized trial found developers finished a coding task ~55.8% faster, with the largest gains for less-experienced developers.

Tip

The leadership move: start where value is proven

Sequence your first deployments into the proven-good quadrant — support assist, internal knowledge search, drafting, coding assist — where the task is repeatable, the success metric is clear, and a human reviews the output. You build credibility and adoption momentum before you reach for harder, higher-risk bets.

Example

Customer support: the canonical quick win — read both halves

Klarna's AI assistant (Feb 2024, vendor-sourced) handled ~2/3 of chats in month one (~2.3M conversations), cut resolution time from 11 minutes to under 2, and was credited with ~700 FTE-equivalent and a ~$40M projected 2024 profit uplift. It is the textbook 'good-now' use case — and the textbook cautionary tale: by mid-2025 Klarna rehired humans after CSAT declined, the CEO noting they 'focused too much on efficiency and cost… lower quality… not sustainable.' Good at the bounded task; over-trusted as a wholesale replacement.

Where leaders over-trust it

Now the other side of the frontier — the places AI fails, and worse, fails confidently. These are not edge cases; they are the predictable failure categories every leader should be able to name.

1. Factual reliability — it never says "I don't know." Because an LLM predicts plausible next words rather than retrieving verified facts, it produces fluent output that can be invented or wrong, and it will not flag its own uncertainty. Hallucination is an inherent property, not a fixable bug. The leadership question is where in our process is a wrong answer load-bearing — legal, medical, financial, safety — and what human check sits before it acts.

2. Tasks outside the jagged frontier. As the BCG/HBS study showed, on unsuited tasks AI users did worse than people with no AI at all — because the tool looks equally confident on both sides of the line. The risk is automation bias: humans over-trusting confident-sounding output.

3. Genuine reasoning and novel strategy. Ask an LLM for strategic advice and you tend to get generic, fashionable, hedge-everything answers — what HBR (2026) memorably labeled "trendslop" — and the models can subtly flatter or manipulate the user toward whatever they seem to want to hear. AI is excellent at expanding your options; it is poor at making the choice.

4. Anything requiring accountability, stakeholder nuance, or final judgment. When a decision affects people's rights, money, or trust, someone must own the outcome. A model cannot be accountable; it cannot be held responsible to a board, a regulator, or a customer. That seat stays human.

Watch out

Where leaders get it wrong

The failure mode is not using AI on hard tasks — it's using it on tasks that look easy but sit outside the frontier, then trusting the confident answer. 'It's 90% accurate' is never enough. The right follow-up is: 'Accurate in what way, what do the failing 10% look like, and who is accountable for them?'

Example

Trendslop and the reliability tax — two named cautions

On strategy: HBR (2026) researchers who asked LLMs for strategic advice got 'trendslop' — generic, buzzword-laden recommendations that flatter rather than challenge. On reliability: trackers of AI-fabricated legal citations had logged well over a thousand cases worldwide by early 2026 (per Scientific American, 2026), and even an elite firm (Sullivan & Cromwell) apologized for erroneous, invented citations — fluent, confident, and entirely made up. (Running tally is volatile; verify live.)

The good / risky map at a glance

Carry this one table into your next portfolio review. The left column is where to deploy with confidence; the right is where to guard — slow down, add a human check, or keep the decision human entirely.

Deploy with confidence (inside the frontier)Guard carefully (outside / over-trusted)
Drafting, summarizing, rewritingAnything where a wrong fact is load-bearing (legal, medical, financial, safety)
Classification, tagging, triage, routingNovel strategy and genuine reasoning (beware 'trendslop')
Augmenting front-line and less-experienced staffTasks that look easy but sit outside the frontier
Grounded search over your approved documents (RAG)Decisions needing accountability, stakeholder nuance, final judgment
Coding assistance with tests/reviewHigh-stakes autonomous action with no human check

The organizing principle behind the whole table: automation for execution, humans for judgment. Use AI to do the bounded, repeatable, checkable work — and to expand the set of options a human considers — while the consequential choice, and the accountability for it, stay with a named person.

Tip

The leadership move: ask 'which side of the line?'

Before approving any AI use case, ask one question of the team: is this task on the deploy side or the guard side of the frontier? If guard side, what is the human check, and who owns the outcome? Make that a standing question in your governance and pilot reviews.

A great demo is not a production capability

The single most expensive misjudgment in this space is mistaking a polished demonstration for a production-ready system. A demo is a curated best case. It hides the parts that actually determine whether AI creates value: the messy data, the integration into real workflows, the governance, the monitoring, and the long tail of failure cases that never appear on stage.

This is not a hunch — it is the dominant finding of 2025-2026. The MIT NANDA "GenAI Divide" study (2025) found that roughly 95% of enterprise GenAI pilots delivered no measurable P&L return, and the root cause was not model quality but a "learning gap" — failure to integrate AI into workflows, structures, and culture. McKinsey's parallel framing (the "gen AI paradox") is that most firms use GenAI yet most report no material P&L impact, because value sits in shallow demos and copilots rather than redesigned work.

The leadership translation: when you see an impressive demo, the right reaction is curiosity about everything it didn't show — the error rate on your real data, the integration cost, the human checks, and who is accountable when it is wrong. The headline capability is the easy 20%; the unglamorous 80% (data, workflow, governance) is where value is actually won or lost.

Watch out

The demo trap

Demos are engineered to succeed. Treating 'it worked in the demo' as 'it's production-ready' is how organizations join the ~95% of pilots that move no P&L number. Always ask to see the failure cases, the real-data error rate, and the integration and governance plan — not just the highlight reel.

The questions that keep you calibrated

You don't need to know how a model works to govern it well. You need a short, repeatable set of questions that force the right side of the frontier into view. Use these in any vendor meeting, pilot review, or board discussion.

On the claim: "It's 90% accurate" → Accurate in what way? What do the failing 10% look like? Who is accountable for them?

On the task: Is this inside the frontier (bounded, verifiable, checkable) or outside it (novel, high-stakes, judgment-heavy)?

On reliability: Where in this process is a wrong answer load-bearing — and what human check sits before it acts?

On strategy use: Are we using AI to expand our options, or letting it make the choice? Is this advice specific to us, or generic 'trendslop'?

On the demo: What did the demo not show — the real-data error rate, the integration cost, the failure cases?

On accountability: Who owns this outcome when the AI is wrong?

The discipline behind all six is the same: use AI to expand options, not to make the choice, and never let a confident answer substitute for a calibrated one. Final judgment and accountability remain human — that is not a limitation of today's tools to be engineered away; it is the line that makes delegating to AI safe.

Key insight

Calibrated trust beats blanket trust or blanket fear

The leaders who win are neither AI evangelists who trust every output nor skeptics who ban it. They practice calibrated trust: high trust on the deploy side of the frontier, deliberate friction and human checks on the guard side. The map and the six questions are how you operationalize that calibration.

Try it: Build your AI 'Deploy vs. Guard' map

Goal: turn the jagged-frontier idea into a one-page decision tool you actually use. 1) List 8-10 candidate AI use cases from across your business — pull from current pilots, vendor pitches, and ideas your teams have raised (e.g. 'draft customer emails,' 'summarize board packs,' 'screen job applicants,' 'recommend a pricing strategy,' 'triage support tickets,' 'answer staff HR questions from policy docs'). 2) Draw a two-column map — 'Deploy with confidence' and 'Guard carefully' — and place each use case on one side, using the test: is the task bounded, verifiable, and human-checked (deploy) or novel, high-stakes, and judgment-heavy (guard)? 3) For every 'Guard' item, write two things: where a wrong answer is load-bearing, and the named human who owns the outcome. 4) Pick the single most over-trusted item on your list — the one that looks easy but sits outside the frontier — and write the one sentence you'll say in the next meeting to slow it down. 5) Pressure-test one vendor claim: take a real 'X% accurate' or demo claim you've heard and write out the answers to 'accurate in what way? what do the failures look like? who owns them?' 6) Write three sentences capturing what surprised you: which use case you misjudged, where your organization is over-trusting AI today, and one place you're under-using it. Bring the one-page map to your next governance or portfolio review — it is the calibration tool this lesson is designed to produce.

Key takeaways

  1. 1AI capability is jagged, not uniform: the same tool can make one task ~40% better (inside the frontier) and another ~19 points worse (outside it) — and the boundary is invisible, so the real risk is over-trusting confident output where it is silently wrong.
  2. 2Genuinely good now: drafting and summarizing, classification and triage, augmenting less-experienced workers (NBER: +14% overall, +34% for novices), and grounded search over your own documents via RAG — all bounded, verifiable, human-in-the-loop tasks.
  3. 3Genuinely risky: factual reliability (it never says 'I don't know'), tasks outside the frontier, novel strategy ('trendslop' and flattery), and anything needing accountability, stakeholder nuance, or final judgment.
  4. 4A polished demo is a curated best case, not a production capability — ~95% of enterprise GenAI pilots delivered no measurable P&L return (MIT NANDA, 2025), and the cause was workflow integration, not model quality.
  5. 5The organizing principle is automation for execution, humans for judgment: deploy AI to do the checkable work and expand options, but keep the consequential choice and its accountability with a named person.
  6. 6Govern with six repeatable questions — on the claim, the task, reliability, strategy use, the demo, and accountability — to force calibrated trust instead of blanket trust or blanket fear.

Quiz

Lock in what you learned

Check your understanding

0 / 4 answered

1.What does the 'jagged frontier' concept (from the BCG/HBS field experiment) teach a leader about today's AI?

2.Which of the following is the BEST-supported 'genuinely good now' use of AI, and why?

3.A vendor gives a flawless demo of an AI tool. What is the most important thing a leader should conclude?

4.Researchers in HBR (2026) asked LLMs for strategic advice and got 'trendslop.' What is the executive lesson?

Go deeper

Hand-picked sources to keep learning