The Right Questions to Ask
A leader's question bank for vendors and teams
- Interrogate any AI performance claim with the 'accurate how, what do the failures look like, who owns them?' triad instead of accepting a headline number
- Ask the strategy questions that separate a tool rollout from a workflow redesign, and use AI to expand options rather than make the choice
- Diagnose adoption health by asking whether the organization is empowering line managers or hoping a central lab will fix it
- Probe data readiness — clean, governed, accessible — as the precondition that gates every AI ambition
- Pressure-test any system that acts with the human-in-the-loop and halt/rollback questions that govern autonomy
- Assemble a personal, reusable question bank for vendors, teams, and pilot reviews
The most valuable thing a non-technical leader can carry into any AI meeting is not an answer but a question — the kind that turns a polished demo into a defensible decision. This lesson consolidates the executive's interrogation skill into a reusable question bank across five fronts: claims, strategy, adoption, data, and agent autonomy. You leave with a one-page set of questions that surface truth from vendors, teams, and pilots before money or reputation is on the line.
- 1Why the question is the leader's real tool
- 2On a claim: interrogate the number
- 3On strategy: tool rollout or workflow redesign?
- 4On adoption: empowering managers or hoping a lab fixes it?
- 5On data: clean, governed, and accessible enough?
- 6On agents: human-in-the-loop, and can we halt?
- 7Your one-page question bank
Why the question is the leader's real tool
A non-technical executive will never out-argue a vendor on model architecture, and shouldn't try. The leadership skill is different and more durable: asking the question that makes the truth surface on its own. A confident demo, a slide that says "90% accurate," a roadmap full of "agentic" — none of these survive contact with the right follow-up. Your job is to ask it.
This matters because the evidence says most AI value is lost not in the lab but in the boardroom and the meeting room — in decisions made on a polished surface without interrogating what's underneath. Recall the headline numbers from this course: roughly 95% of enterprise generative-AI pilots delivered no measurable P&L return (MIT NANDA, State of AI in Business 2025), and McKinsey's State of AI 2025 found that about 51% of organizations had at least one negative AI incident in the prior year. Behind most of those is a question that nobody asked.
The questions in this lesson are organized on five fronts, mirroring the five places leaders are most often misled:
| Front | The question defends against |
|---|---|
| On a claim | Believing a metric without knowing how it fails or who owns the failures |
| On strategy | Funding a tool rollout when the value needs a workflow redesign |
| On adoption | Hoping a central lab fixes what only empowered managers can |
| On data | Building on data too messy, ungoverned, or hidden to feed the system |
| On agents | Letting software that acts run without a way to supervise or stop it |
Use AI — and use these questions — to expand your options, not to make your choice. Final judgment and accountability stay human.
Key insight
The reframe
You are not being tested on whether you know the answer. You are valuable precisely because you ask the question a technical team, eager to ship, or a vendor, eager to sell, would rather you didn't. The question is the leadership move; the answer belongs to the people you're holding accountable.
Watch out
Where leaders get it wrong
The most expensive failure mode is silent agreement — nodding at a demo because asking feels like exposing ignorance. The opposite is true. A precise, naive-sounding question ("what do the failures look like?") is the single most senior thing you can do in an AI meeting.
On a claim: interrogate the number
Every AI conversation eventually produces a number — "90% accurate," "cuts handling time by 40%," "resolves two-thirds of tickets." Treating that number as a fact is the most common, and most costly, mistake a leader makes. A headline metric is the beginning of a conversation, not the end of one.
The interrogate-the-claim triad breaks any number open:
"Accurate in what way — what do the failing cases look like — and who is accountable for them?"
- "Accurate in what way?" Accuracy on which task, measured against what benchmark, on data that looks like ours? A model can be 90% accurate on a vendor's curated test set and far worse on your messy, real-world inputs — the jagged-frontier problem, where capability is unevenly good and the boundary is invisible.
- "What do the failing cases look like?" A 10% error rate spread evenly across low-stakes cases is a different animal from a 10% rate concentrated on your highest-value, highest-risk decisions. The shape of the failure matters more than its size.
- "Who owns the failures?" When the model is wrong, who catches it, who is accountable, and what does it cost? If the answer is "nobody, until a customer complains," you have found your real risk.
This is the antidote to automation bias — the human tendency to over-trust confident-sounding output. Remember the dossier's example of generic AI strategy advice (HBR's "trendslop", 2026): fluent, plausible, and hollow. Confidence is not correctness.
Example
Klarna: the number that hid the truth
Klarna's AI assistant handled roughly two-thirds of customer-service chats (~2.3M conversations in month one) and cut resolution time from ~11 minutes to under 2 — a spectacular volume metric. But the failing cases were the nuanced, empathetic ones, and customer satisfaction quietly declined. By mid-2025 Klarna reversed course and rehired humans, the CEO conceding it had "focused too much on efficiency and cost." The volume number was real; the question "what do the failing cases look like, and who owns them?" would have surfaced the quality collapse far earlier. (Vendor- and press-sourced, 2024–2025 — verify current status.)
Tip
The leadership move
Never let a metric stand alone in the minutes. Write next to every claimed number: the task it was measured on, what the failures look like, and the named owner of those failures. If any of the three is blank, the decision isn't ready.
On strategy: tool rollout or workflow redesign?
The strategy questions decide whether your AI investment produces a step change or a rounding error. The dossier is blunt on this: there is no standalone "AI strategy" — only a business strategy strengthened by AI. The leaders who make "doing AI" the goal generate diffuse pilots with no P&L impact.
Three questions cut to the heart of it:
- "Where can we move fast, and where must we prioritize control?" Not every use case carries the same risk. A marketing-copy drafter and a loan-approval model belong on opposite ends of the speed-vs-control dial. Ask which dial position each initiative sits at — and whether anyone has consciously chosen it.
- "Are we using AI to expand our options, or to make the decision for us?" The right posture is AI as an option-generator that a human judges, not an oracle that a human rubber-stamps. The moment a system is making the choice rather than widening the menu, accountability has quietly slipped.
- "Is this a tool rollout or a workflow redesign?" This is the single most important strategy question, because it predicts whether value will materialize.
That third question rests on the strongest finding in the research. McKinsey's State of AI 2025 found workflow redesign has the single biggest effect on the EBIT impact an organization gets from AI — yet most firms simply bolt AI onto legacy processes. This is the gen AI paradox: ~80% of firms use generative AI, yet ~80% report no material bottom-line impact, because value sits in deep, redesigned vertical workflows, not shallow horizontal copilots. A copilot makes an existing task faster (incremental); an agentic, redesigned workflow does the task end-to-end (a step change). If the plan in front of you is "give everyone a copilot," you are buying speed, not transformation.
Key insight
The reframe: 10-20-70
BCG's 10-20-70 rule holds that ~10% of AI effort is the algorithm, ~20% the technology and data, and ~70% the people and process — adoption, workflow redesign, reskilling, incentives. The strategy question "tool rollout or workflow redesign?" is really asking: are we funding the 70% we actually own, or just the 10% the vendor sells?
Watch out
Where leaders get it wrong
Launching the flagship AI transformation in a support function (HR, legal, finance) to cut cost, rather than in a core revenue or customer function where ~75% of generative-AI value concentrates (McKinsey). The wrong question — "where can we save money?" — points you at the wrong workflow.
On adoption: empowering managers or hoping a lab fixes it?
The decisive adoption question is deceptively simple:
"Are we empowering our line managers to make this work — or are we hoping a central lab will fix adoption for us?"
The research points hard at the first half. The executive-versus-manager reality gap is one of the most under-appreciated reasons pilots stall: senior leaders see strategic advantage, while middle managers confront AI's flaws in real workflows, often unsupported (HBR, 2026). When executives mandate adoption top-down and a central team "owns" it, the people who actually have to change how their teams work are left out of the design — and quietly resist.
The operating-model answer the evidence favors is hub-and-spoke: a lean central hub sets standards, guardrails, and platforms, while business-unit "spokes" own execution and adoption locally. Firms that successfully scale AI are markedly more likely to use this structure than a fully centralized or fully decentralized model (Dataiku/IBM, cited in the dossier). The central lab is an enabler, not a gatekeeper — and never a substitute for engaged managers.
A short adoption-diagnostic question set:
| Question | What a weak answer reveals |
|---|---|
| Were the line managers who own this workflow in the room when it was designed? | Top-down mandate; predictable resistance |
| What admin burden does this add to a manager's day, and how are we offsetting it? | Adoption taxed, not enabled |
| What feedback channel turns frontline friction into fixes? | No learning loop; pilot purgatory ahead |
| Are leaders, including me, visibly using this themselves? | Disengaged leadership — a known value-killer |
That last row matters more than it looks: BCG's AI Radar found deeply engaged C-suite teams are roughly 12× more likely to be top-tier AI value creators. Championing is not delegation — it is role-modeling.
Tip
The leadership move
Before approving a rollout, insist that the line managers who own the affected workflow have co-designed it and have a standing channel to report what's broken. Adoption is won in their daily reality, not in a central lab's roadmap.
Watch out
Where leaders get it wrong
Outsourcing adoption to a Center of Excellence and treating it as 'handled.' A CoE can set guardrails and platforms; it cannot reach into a hundred teams and change how work gets done. That is the spokes' job — and yours, by role-modeling use.
On data: clean, governed, and accessible enough?
Every AI ambition eventually runs into the same wall, so ask about it before you fund the ambition:
"Is our data clean, governed, and accessible enough to feed this — honestly?"
Data is the one AI advantage that compounds and the one rivals cannot simply buy. But the same data is a liability when it is siloed and ungoverned, because — as the dossier puts it — "AI agents cannot query what they cannot find." Roughly 80% of the real work in an AI initiative is data, governance, and workflow integration, not the model (MIT Sloan, via the dossier). The model is the easy 20%.
Break the data question into its three load-bearing words:
- Clean — Is the data accurate, current, and consistent, or will "garbage in, garbage out" surface as confident, wrong answers? Machine learning learns from data, so data quality is the literal raw material.
- Governed — Do we know what data this system can access, where it flows (including to third-party model providers), and who approved it? This is the first risk-taxonomy question every board must be able to ask.
- Accessible — Can the system actually reach the right data when it needs it, or is it locked in silos and legacy systems? Eight in ten organizations cite data and legacy-integration limits as the gate on scaling agents.
The payoff for getting this right is the data flywheel: every interaction generates feedback data that makes the system better, which compounds into a moat. But a flywheel built on dirty, ungoverned, inaccessible data just spins faster in the wrong direction.
Key insight
The reframe
The board's job on data is three verbs: govern it, protect it, activate it. A leader who asks 'is our data ready?' before 'which model should we buy?' is sequencing the work correctly — because the model is a commodity and the proprietary, well-governed data is the only durable moat.
Example
When data access is the whole story
The dossier's recurring failure pattern is the pilot that dazzles on a curated demo dataset, then stalls in production because the live data is siloed across legacy systems the agent can't reach or trust. The 'learning gap' that MIT NANDA blames for ~95% of failed pilots is largely a data-and-integration gap — not a model-quality gap.
On agents: human-in-the-loop, and can we halt?
When a system doesn't just answer but acts — books, pays, emails, approves, transacts — the questions change in kind, not just degree. An agent that takes actions toward a goal needs governance precisely because it does things. The two non-negotiable questions:
"What is our human-in-the-loop policy for this system — and can we halt it and roll it back?"
Ground this in the autonomy delegation test: delegate a decision to a system only when the data is clear, the rules are defined, and the cost of waiting for a human is greater than the cost of a wrong autonomous action. The governing principle is automation for execution, humans for judgment. Autonomy is a dial, not a switch — and that single dial sets the value, the risk, and the governance required.
The agent question set every leader should carry:
| Question | Why it matters |
|---|---|
| Is a human in the loop (approving each action) or on the loop (monitoring, ready to intervene)? | Sets the supervision model to the stakes |
| For which actions is human approval mandatory before execution? | High-impact actions must not be fully autonomous |
| Can we halt the system and roll back what it did? | The single most important agent control |
| What is the blast radius if it's wrong or hijacked? | More autonomy + more tool access = bigger damage |
| Is this a genuine agent, or 'agent washing'? | Gartner: only ~130 of thousands of 'agentic' vendors are real |
The security dimension is real and board-relevant. Prompt injection — hidden instructions buried in a document, email, or web page the agent reads — can hijack an agent with tool access into exfiltrating data or making unauthorized transactions (OWASP Top 10 for Agentic Applications, 2026). The mitigations are conceptual but askable: least-privilege tool access, sandboxing, input validation, and human approval for high-impact actions. And the liability is not hypothetical — in Mobley v. Workday, a court treated an AI screening tool as the employer's "agent," a precedent that shifts responsibility for what the software does back onto the organization deploying it.
Watch out
Where leaders get it wrong
Believing 'agentic' on a vendor slide. Gartner predicts >40% of agentic AI projects will be canceled by end of 2027 — on cost, unclear value, and weak risk controls — and warns that most self-styled 'agentic' vendors are rebranded chatbots ('agent washing'). The anti-hype test: a real agent shows autonomous reasoning, tool orchestration, and persistent context. (Gartner, 2025 — verify live.)
Tip
The leadership move
Make 'can we halt it and roll it back?' a gating question for any system that acts. If the team can't demonstrate a kill switch and a rollback path — and name the human who owns them — the agent is not ready for production, no matter how good the demo looked.
Your one-page question bank
Pull the five fronts together into a single page you can carry into any vendor pitch, pilot review, or strategy meeting. Lead with the front that matters most for the decision in front of you.
| Front | The questions to ask |
|---|---|
| On a claim | Accurate in what way, on data like ours? What do the failing cases look like? Who is accountable for them? |
| On strategy | Where can we move fast vs. prioritize control? Are we expanding our options or outsourcing the choice? Is this a tool rollout or a workflow redesign? |
| On adoption | Did the line managers who own this workflow help design it? What feedback loop fixes frontline friction? Am I, the leader, visibly using it? |
| On data | Is our data clean, governed, and accessible enough to feed this? Where does the data flow, and who approved it? |
| On agents | Human in-the-loop or on-the-loop? Which actions need approval? Can we halt and roll back? What's the blast radius? Is it a real agent or 'agent washing'? |
Three cross-cutting questions belong on every front, because they expose the answers people most want to avoid giving:
- "In P&L terms, not demo terms — how will we know this worked?"
- "Are we buying or building, and why?" (The dossier's evidence favors a default-to-buy posture for most use cases — verify the specific figures live.)
- "Who is the single named owner accountable for this — and can they explain it to a regulator?"
A question bank is only as good as the discipline of using it. The habit to build: never approve an AI investment, sign a vendor, or green-light a pilot until each relevant front has a real answer — not a confident deflection.
Key insight
The reframe: questions are a governance control
Diffuse accountability is the clearest 'leaders get it wrong' signal in the research — only ~28% of organizations say the CEO owns AI governance and ~17% the board (McKinsey, 2025). A standing question bank, asked consistently, is itself a lightweight governance control: it forces a named owner and a real metric onto every decision before money moves.
Try it: Build and pressure-test your AI question bank
Goal: leave with a one-page question bank you will actually use, and prove it by running it against a real decision. 1) Pick a live decision. Choose an AI initiative your organization is considering, piloting, or being pitched — a vendor tool, an internal pilot, or an agentic workflow. 2) Draft your one page. For each of the five fronts — claim, strategy, adoption, data, agents — write the 2–3 sharpest questions you would ask, adapting the lesson's bank to your context. Add the three cross-cutting questions (P&L not demos; buy or build; single named owner). 3) Stress-test it. Run your questions against the chosen initiative and write the honest answer to each — or mark it 'no real answer yet.' Every blank is a risk you just surfaced. 4) Find the weakest front. Identify which of the five fronts has the thinnest answers; that is where this initiative is most likely to fail. 5) Assign owners. For each front, name the single person who should be accountable for the answer (yourself, a team lead, the vendor). If you can't name one, that is the finding. 6) Write the verdict (half a page): Given the answers and the gaps, would you approve, pause for more information, or decline — and which one question, if it had a better answer, would most change your decision? This converts the interrogation skill into a reusable governance habit: no AI investment approved until each front has a real answer and a named owner.
Key takeaways
- 1The non-technical leader's most powerful tool is the right question, not the right answer — it makes the truth surface from vendors, teams, and pilots before money or reputation is at stake.
- 2On a claim, run the triad: 'accurate in what way, what do the failing cases look like, and who owns them?' A headline number is the start of a conversation, not a fact — the Klarna reversal is the canonical lesson.
- 3On strategy, the decisive question is 'tool rollout or workflow redesign?' — McKinsey finds workflow redesign has the single biggest effect on EBIT impact, and AI should expand options, not make the choice.
- 4On adoption, ask whether you are empowering line managers (hub-and-spoke, co-design, role-modeling) or hoping a central lab will fix adoption for you — the executive-versus-manager gap stalls pilots.
- 5On data, ask honestly whether it is clean, governed, and accessible — ~80% of the real AI work is data, governance, and integration, and 'agents cannot query what they cannot find.'
- 6On agents, demand a clear human-in-the-loop policy plus the ability to halt and roll back; the more autonomy and tool access, the bigger the blast radius — and beware 'agent washing.'
Quiz
Lock in what you learned
Check your understanding
0 / 4 answered
1.A vendor's slide says their AI is '92% accurate.' Using the interrogate-the-claim triad, which response best protects you as a leader?
2.Which strategy question is most predictive of whether an AI initiative will produce a step change in value rather than a rounding error?
3.Your AI program is stalling in adoption despite executive enthusiasm. Which question best targets the most common root cause the research identifies?
4.A team proposes an autonomous agent that can send emails, issue refunds, and update customer records without per-action approval. What is the single most important governance question to ask before approving it?
Go deeper
Hand-picked sources to keep learning
Source for workflow redesign as the biggest EBIT driver, the ~51% negative-incident figure, and CEO/board governance ownership. Time-stamped 2025 — re-verify figures before quoting.
The ~95%-of-pilots-show-no-P&L-return finding and the 'learning gap' root cause behind it. Dated Aug 2025 — verify live.
Foundation for the interrogate-the-claim skill and using AI to expand options rather than make the choice.
The canonical 'volume metric hid a quality collapse' case used in the On-a-Claim section. Status evolving — verify current state.
Source for the agent-washing warning and the anti-hype test for genuine agents. Predictions dated 2025 — re-verify.
Conceptual backing for the prompt-injection and least-privilege questions on the agent front.