Interrogating an AI Claim
Never accept '90% accurate' at face value
- Be able to interrogate any AI accuracy claim with the core question — accurate in what way, what do the failures look like, and who owns them
- Be able to recognize automation bias in yourself and your teams, and name the conditions that make over-trust most dangerous
- Be able to explain the jagged frontier and why a polished, confident answer is not evidence of a correct one
- Be able to use AI to expand the set of options on a decision rather than to outsource the decision itself
- Be able to spot AI output that flatters you or produces generic 'trendslop' strategy advice
- Be able to keep final judgment and accountability human while still moving fast with AI
The highest-leverage executive AI skill is not building anything — it is refusing to accept a headline number like "90% accurate" at face value. This lesson hands you the exact questions to ask, the cognitive trap (automation bias) to watch for, and a discipline for using AI to expand your options rather than make the choice — so that final judgment and accountability stay where they belong: with you.
- 1The hardest, highest-leverage skill you'll learn
- 2The core question: accurate in what way — and who owns the failures?
- 3Automation bias: the gravity you're fighting
- 4Use AI to expand options — not to make the choice
- 5Watch for flattery and 'trendslop'
- 6Final judgment and accountability stay human
The hardest, highest-leverage skill you'll learn
Most AI skills a leader needs are about adoption — picking use cases, funding pilots, redesigning workflows. This one is different. It is a skill of resistance: the discipline to not believe a confident machine.
Here is the situation you will face, over and over, in board decks and vendor pitches: a system is described as "90% accurate." It sounds reassuring. It sounds like a fact. And the natural executive instinct — built over a career of trusting credible-sounding numbers from credible-sounding people — is to nod and move on.
That instinct is exactly the trap. "90% accurate" is not an answer; it is the beginning of a conversation most leaders never have. The number tells you nothing about what kind of accuracy was measured, what the failing 10% actually look like, or who is on the hook when one of those failures reaches a customer, a regulator, or a court.
The danger with today's AI is not that it is weak. It is that it is silently and invisibly bad in some places — and it never tells you which places those are.
Research makes this concrete. In a 2023 field experiment by Dell'Acqua, Mollick, Lakhani and colleagues (Harvard/BCG, 758 BCG consultants), AI users on tasks inside its capability were roughly 25% faster and produced ~40% higher-quality work. But on tasks that fell outside that capability, AI users were about 19 percentage points less likely to reach the correct answer than colleagues working without AI at all. Same tool, same confident tone — opposite outcomes. The tool gave no warning when it crossed the line.
Learning to interrogate the claim is how you stop being the executive who trusts the wrong 19 percentage points.
Key insight
The reframe
A confident, fluent, well-formatted AI answer is evidence of fluency, not correctness. These models are trained to predict what plausibly comes next — which is exactly why they can be wrong with total composure. Polish is not proof.
Watch out
Where leaders get it wrong
Treating a headline accuracy number as a finished fact. "90% accurate" with no follow-up question is not diligence — it is the absence of it. The number is a prompt for five questions, not a green light.
The core question: accurate in what way — and who owns the failures?
Memorize one move. When anyone — a vendor, a team, a slide — tells you an AI system is "X% accurate," respond with three linked questions:
"Accurate in what way? What do the failing 10% actually look like? And who is accountable for them?"
Each question opens a door the headline number was holding shut.
| The question | What it really tests | What a weak answer sounds like |
|---|---|---|
| Accurate in what way? | Measured against what benchmark, on whose data, doing which specific task? "90% on a clean test set we built" is not "90% on your messy real-world inputs." | "It's just… 90% accurate." (No benchmark, no task definition.) |
| What do the failing 10% look like? | Are the errors random and harmless, or concentrated in your highest-stakes cases? Ten percent of low-value tickets is noise; 10% of fraud approvals or medical triages is a crisis. | "We haven't really looked at the failure cases." |
| Who is accountable for them? | When a wrong answer reaches a customer, a regulator, or a court, whose name is on it? Is there a human check before the output acts? | "The AI does it" — i.e., nobody owns it. |
The second question is the one that separates seasoned leaders from the rest. The shape of the errors matters more than the headline number. A 95%-accurate system whose 5% of failures all land on your most vulnerable customers is worse than a 90%-accurate system that fails randomly on trivia.
The third question is the one that protects your enterprise. A model is only the front half of the sentence; the back half is "…and a named human checks it before it does anything that matters." If no one can answer who owns the 10%, you have not bought a capability — you have bought an unowned liability.
Tip
The leadership move
Make these three questions a standing requirement, not a personal habit. Add them to your vendor-evaluation template and your pilot sign-off gate: every AI claim must arrive with its benchmark, a description of its failure cases, and a named owner. You will be amazed how many "90% accurate" claims dissolve the moment someone has to fill in the second column.
Example
When the failures hide in plain sight: Klarna
Klarna's AI assistant handled roughly two-thirds of customer-service chats and cut resolution time from ~11 minutes to under 2 (per the company, Feb 2024) — a dazzling headline number. But the volume metrics masked a quality problem: customer satisfaction declined, and by mid-2025 Klarna had reversed course and rehired humans, with the CEO conceding the company "focused too much on efficiency and cost" at the expense of quality. The headline accuracy looked great; the shape of the failures — on the nuanced, empathetic cases — is what eventually mattered.
Automation bias: the gravity you're fighting
Why is this skill so hard? Because you are fighting a well-documented cognitive pull called automation bias — the human tendency to over-trust a confident, automated output and under-apply your own scrutiny, simply because a machine produced it.
Automation bias is sneaky precisely because it feels like efficiency. A polished answer arrives instantly, formatted, fluent, and self-assured. Your brain reads that fluency as competence and quietly lowers its guard. The better the output looks, the less you check it — which is exactly backwards, because today's models are most dangerous when they are most fluent and most wrong.
Three conditions make automation bias most damaging. When you see them stacking up, raise — don't lower — your scrutiny:
- High confidence in the output. The model never says "I don't know." It states a fabricated answer in the same authoritative tone as a correct one.
- High stakes in the decision. The cost of a wrong answer is large (legal, financial, medical, safety, reputational).
- Low human friction. The output flows straight into action with no human checkpoint, so no one is positioned to catch the error.
The antidote is not distrust of all AI — that wastes the real value. It is calibrated trust: trusting AI more where it is provably good and the stakes are low, and trusting it less exactly where its confidence is highest and a wrong answer is load-bearing. Your job as a leader is to design that friction back in at the high-stakes points your teams, under deadline pressure, will be tempted to remove.
Watch out
Where leaders get it wrong
Assuming automation bias is a junior-staff problem. It is worse at the top: senior leaders are time-pressed, used to trusting polished briefings, and the most prolific users of AI tools (research consistently finds executives among the heaviest users of unsanctioned AI). The most confident person in the room over-trusting the most confident machine in the room is the failure mode to design against.
Key insight
Fluency is not knowledge
A large language model is, at its core, autocomplete on steroids — it predicts the next plausible word, not the next true one. It will never volunteer 'I'm not sure.' So the smoothness of an answer carries zero information about its accuracy. Treat confidence as a style, not a signal.
Use AI to expand options — not to make the choice
Here is the single most useful operating principle for an executive working with AI on anything that matters:
Use AI to expand your options, not to make your choice.
AI is extraordinary at the divergent half of thinking — generating possibilities, surfacing considerations you missed, drafting three versions of a plan, stress-testing an argument, summarizing what others have done. That is genuine leverage, and you should use it aggressively. Where it fails is the convergent half — weighing trade-offs against your context, your stakeholders, your risk appetite, and your accountability. That is judgment, and judgment is yours.
The practical test is a question you ask of your own workflow: am I using this AI to widen the funnel, or to close it?
| Use AI to EXPAND options (green light) | Use AI to MAKE the decision (red flag) |
|---|---|
| "Give me five strategic options and the case for each." | "Which strategy should we pick?" — and you act on its pick. |
| "What risks am I not seeing in this plan?" | "Is this plan good?" — and you take the yes. |
| "Draft three counter-arguments to my position." | "Write our final position and I'll forward it." |
| "Summarize what comparable firms have done." | "Tell me what we should do." |
The left column makes you a better decision-maker by giving you more and better raw material to judge. The right column quietly makes the decision for you while leaving your name on it — the worst of both worlds, because you carry the accountability without having exercised the judgment.
This maps cleanly onto the divergent/convergent split: let AI run wide; you run deep. It produces breadth; you provide the weighing, the context, and the call.
Tip
The leadership move
Phrase your prompts as option-generators, not verdict-requests. "Give me the three strongest cases for and against" keeps you in the judgment seat; "What should we do?" hands the chair to a machine that has no accountability and no knowledge of your stakeholders. The wording of the question determines who actually decides.
Watch for flattery and 'trendslop'
Two specific failure modes deserve a leader's particular vigilance, because they are most likely to ambush you precisely on the high-stakes, strategic questions where you most want help.
1. Generic strategy advice ('trendslop'). When researchers asked large language models for strategic advice, they got back what HBR (2026) memorably labeled "trendslop" — fluent, confident, on-trend, and almost entirely generic. The advice pattern-matches to whatever sounds current ("embrace digital transformation," "build an AI-first culture," "focus on customer-centricity") without engaging the one thing that makes strategy strategy: your specific context, constraints, and competitive position. It reads like a McKinsey deck and commits to nothing. The tell: advice that would apply, word-for-word, to any company in your industry — which means it is advice for no company in particular.
2. Flattery. The same research stream finds these systems can subtly flatter and over-agree with the user. Trained to be helpful and agreeable, a model will too often tell you your idea is strong, your analysis is sharp, and your instinct is right — because agreement is the path of least resistance, not because it has independently judged your idea to be good. For an executive, an assistant that reflexively validates you is not an asset; it is an echo chamber with a confident voice.
Together these explain why AI is worst at exactly the work leaders most wish to delegate: genuine, novel strategic reasoning. It will hand you a polished, flattering, generic plan and feel like help — which is more dangerous than obvious nonsense, because polish disarms scrutiny.
Watch out
Where leaders get it wrong
Mistaking a fluent, agreeable strategy memo for a good one. The smoother and more affirming the output, the harder you should push on it. Ask the model to argue the opposite case, to name where your plan is weakest, and to tell you what a skeptical board member would attack — deliberately steering it away from flattery and generic consensus.
Key insight
The diagnostic
If an AI's strategic advice could be pasted into any competitor's plan unchanged, it is trendslop, not strategy. Real strategy is specific to your position; generic advice is a sign the model is pattern-matching to buzzwords, not reasoning about you.
Final judgment and accountability stay human
Everything in this lesson resolves to one non-negotiable principle: automation for execution, humans for judgment. AI can draft, summarize, sort, surface, and propose. It cannot be accountable — and accountability is the irreducibly human part of leadership.
This is not a soft preference; it is increasingly a legal and governance reality. Three signals every leader should register:
- The law is closing the accountability gap. In Mobley v. Workday, a US court allowed a nationwide age-discrimination collective action to proceed and ruled that an AI screening tool can be treated as the employer's "agent" — meaning you can be held liable for what your AI does. "The AI decided" is not a defense.
- "The AI did it" is not a shield, it is an exposure. When AI-fabricated content reaches the real world, the human who relied on it pays. By early 2026 there were well over 1,300 documented cases worldwide of AI-hallucinated legal citations reaching court filings — including from elite firms — with lawyers sanctioned for output their AI confidently invented.
- Governance frameworks demand a named owner. Every serious AI governance framework (NIST's AI Risk Management Framework, ISO/IEC 42001, the EU AI Act) converges on the same requirement: every production AI system has a named human owner accountable for its behavior, and high-stakes decisions retain meaningful human oversight.
So the closing discipline is simple to state and hard to hold under deadline pressure: AI expands your options and accelerates your work; you make the call, you can explain it, and your name is on it. Keep a human in the loop wherever a wrong answer is load-bearing, and never let the speed of a confident machine talk you out of the judgment only you can supply.
Tip
The leadership move
For any AI-assisted decision that matters, apply a one-line test before you act: 'Could I stand in front of a regulator, a board, or a customer and defend this decision as my own?' If the honest answer is 'the AI said so,' you have not finished the work. That sentence — owning the decision in your own voice — is the line between using AI and being used by it.
Example
The liability-shifting precedent
Mobley v. Workday matters to every executive, not just HR leaders: a court treated the AI vendor's screening tool as the employer's agent, opening the door to holding the deploying company liable for the tool's discriminatory effects. Outsourcing the task to an AI did not outsource the accountability. Assume that principle generalizes.
Try it: Pressure-test a real AI claim and write your five interrogation questions
Goal: turn this skill into a reusable artifact you can deploy in your next vendor meeting or pilot review — no coding, just judgment.
1) Pick a live claim. Take a real AI accuracy or capability claim you've actually heard — from a vendor pitch, an internal pilot deck, or a news headline (e.g., 'our assistant resolves 90% of tickets,' 'this tool is 95% accurate at X'). Write it down verbatim.
2) Apply the core question. For that claim, answer (or flag as 'unknown') the three doors: (a) Accurate in what way? — against what benchmark, on whose data, doing which exact task? (b) What do the failing 10% look like? — are the errors random and harmless, or concentrated in your highest-stakes cases? (c) Who is accountable? — name the human who owns a wrong answer, and the check that sits before the output acts. Note which answers you genuinely cannot get — those gaps are your findings.
3) Run the automation-bias check. Score the claim on the three danger conditions: Is the output high-confidence? Are the stakes high? Is human friction low (does it flow straight to action)? If you tick two or three, mark it 'raise scrutiny.'
4) Separate expand from decide. Write one sentence describing how this AI should be used to expand your options, and one sentence on the decision that must stay with a named human.
5) Draft your five questions. Produce a one-page list of exactly five questions you will ask in your next AI vendor or pilot meeting — at least one each on accuracy definition, failure shape, accountability/ownership, human-in-the-loop, and whether the output flatters or gives generic advice. Phrase them so a weak answer is obvious.
6) Write the accountability line. Finish with one sentence you could say to a board or regulator that begins 'I decided this because…' — proving the judgment was yours, not the machine's. Keep the one-pager; it is your standing checklist for every future AI claim.
Key takeaways
- 1Never accept a headline number like '90% accurate' at face value. Ask the core question: accurate in what way, what do the failing 10% actually look like, and who is accountable for them?
- 2The shape of the failures matters more than the headline number — 10% of errors on your highest-stakes cases is a crisis even if the average looks impressive.
- 3Fight automation bias: a fluent, confident AI answer is evidence of fluency, not correctness, and over-trust is most dangerous when confidence is high, stakes are high, and there's no human checkpoint.
- 4Use AI to expand your options, never to make the choice — let it run wide on possibilities while you run deep on judgment, trade-offs, and context.
- 5Watch for flattery and 'trendslop': AI is worst at exactly the novel strategic reasoning leaders most want to delegate, handing you generic, agreeable advice that feels like help but commits to nothing specific to you.
- 6Final judgment and accountability stay human — 'the AI decided' is not a defense (Mobley v. Workday), and every governance framework requires a named human owner for high-stakes AI.
Quiz
Lock in what you learned
Check your understanding
0 / 4 answered
1.A vendor tells your team their fraud-detection model is '92% accurate.' What is the most important follow-up question to ask first?
2.Your strategy team brings you an AI-generated plan. It's polished, confident, on-trend, and strongly affirms the direction you already favored. Based on this lesson, what should that combination make you do?
3.Which prompt reflects the principle of using AI to expand your options rather than to make the decision?
4.A manager defends a flawed AI-driven hiring decision by saying, 'It wasn't really our call — the AI screening tool made the decision.' Why is this defense dangerous for the organization?
Go deeper
Hand-picked sources to keep learning
The field experiment behind the lesson: AI helps inside its frontier and hurts outside it, with no visible warning when it crosses the line. Re-verify the exact effect sizes against the source.
Source for the 'accurate in what way / who owns the failures' discipline and using AI to expand options rather than make the choice.
The origin of 'trendslop' and evidence that LLMs produce generic, flattering strategy advice. Time-stamped 2026 finding — re-read before teaching.
Non-technical leadership framing for interrogating AI claims and keeping judgment human.
The accountability-shifting case: an AI screening tool ruled to be the employer's 'agent.' Litigation is live — re-verify current status.
The 1,300+ AI-hallucinated legal-citation cases — concrete proof that 'the AI did it' is an exposure, not a shield. Counts are volatile; check the live figure.