Why AI Initiatives Fail (and the 10-20-70 Rule)
The failure is organizational, not technical
- Explain the BCG 10-20-70 rule and why winners spend the majority of their effort on people and process, not algorithms
- Diagnose the 'learning gap' (MIT NANDA) — that pilots stall on integration and adoption, not on model quality
- Recognize the five recurring failure modes that stall AI programs: thin pilots, weak C-suite engagement, no measurement, bolting AI onto legacy workflows, and reaching for agents too early
- Articulate why workflow redesign has the single biggest effect on EBIT impact from AI
- Explain why CEO oversight of AI governance is the attribute most correlated with bottom-line impact
- Ask the diagnostic questions that separate a stalled tool rollout from a value-creating transformation
Most enterprise AI initiatives stall not because the technology is weak but because the organization around it is. The evidence is blunt: a 2025 MIT NANDA study found ~95% of enterprise GenAI pilots delivered no measurable P&L return, and the root cause was a 'learning gap' in workflow and adoption, not model quality. This lesson gives leaders the diagnosis (the BCG 10-20-70 rule), the named failure modes, and the single highest-leverage fixes — workflow redesign and CEO ownership — so you can stop funding pilot purgatory and start capturing value.
- 1The headline: everyone has AI, almost nobody has AI value
- 2The 10-20-70 rule: where the value (and the failure) actually lives
- 3The 'learning gap': pilots stall on adoption, not model quality
- 4The five failure modes to recognize on sight
- 5The highest-leverage fix: redesign the workflow, don't bolt AI on
- 6The attribute that predicts success: CEO oversight of governance
- 7Putting it together: the diagnostic the board should run
The headline: everyone has AI, almost nobody has AI value
Start with the uncomfortable number. In MIT NANDA's The GenAI Divide: State of AI in Business 2025 (July 2025), roughly 95% of enterprise GenAI pilots delivered no measurable P&L return — only about 5% reached rapid value or production — despite an estimated $30–40B of investment. The pattern repeats across every major 2025 study, which is why it's a finding, not an outlier.
| Source (2025) | The gap it documents |
|---|---|
| MIT NANDA — GenAI Divide | ~95% of pilots show no measurable P&L return; only ~5% reach rapid value |
| McKinsey — State of AI | ~90% of orgs use AI, but only ~39% report any enterprise-level EBIT impact; only ~6% are high performers (5%+ EBIT) — the 'gen AI paradox' |
| BCG — Widening AI Value Gap | Only ~5% of firms are 'future-built' value creators; ~60% generate no material value despite investment |
| Deloitte — State of AI in the Enterprise | Share of firms abandoning most AI initiatives reportedly rose from ~17% (2024) to ~42% (2025) |
These are time-stamped findings from 2025 surveys, not timeless laws. The numbers will move; the shape of the gap — wide adoption, thin value — is the durable lesson. Re-verify each figure against the live source before you quote it.
The term practitioners use is 'pilot purgatory': an organization runs demo after demo, each one impressive, none of them changing a number on the P&L. The critical reframe for a leader is that this is not a technology failure. The models work. The failure is organizational — and that means it is yours to fix, not the vendor's.
Key insight
The reframe
'Our pilot didn't work' almost never means 'the model isn't good enough.' It means the work around the model — workflow, data access, adoption, measurement — was never done. That work is a leadership job, not an engineering one.
Watch out
Where leaders get it wrong
Treating AI as an IT or tooling project. Tooling projects end when the software is deployed. AI value only appears when the workflow and the people change — which is a transformation program, with all the executive ownership that implies.
The 10-20-70 rule: where the value (and the failure) actually lives
The single most useful framework for understanding why initiatives fail is BCG's 10-20-70 rule. It describes how effort, cost, and attention should be split in a successful AI transformation:
| Layer | Share of effort | What it is | Who owns it |
|---|---|---|---|
| Algorithms | ~10% | The model itself — choosing it, tuning it | Vendors and a small technical team |
| Technology & data | ~20% | Pipelines, integration, data quality, platforms | IT and data teams |
| People & process | ~70% | Adoption, workflow redesign, reskilling, incentives, change management | The executive |
The insight that changes behavior: winners invest in the 70%; stalled pilots invert the ratio, pouring attention into model selection and infrastructure while starving adoption and process redesign. The 70% is precisely the part a non-technical leader is best positioned to own — and the part that gets neglected because it's harder, slower, and less exciting than a demo.
The corroborating evidence is consistent: Deloitte found only ~37% of organizations invest significantly in change management, even though ~70% of the value (and the failure) sits in the people-and-process layer. You cannot buy your way out of the 70% with a better model.
Tip
The leadership move
When you review an AI proposal, ask to see the budget split. If more than ~30% is on models and tech and almost nothing is earmarked for workflow redesign, change management, and reskilling, you are looking at a pilot that will stall. Reallocate toward the 70% before you approve it.
Example
BCG: the value gap is real and widening
BCG's 2025 'Widening AI Value Gap' research found AI leaders out-earn laggards by roughly 1.7x revenue growth, 1.6x EBIT margin, and 3.6x three-year total shareholder return. Nearly 100% of 'future-built' firms had deeply engaged C-suites with real funding and leaders who personally use AI — the human-and-leadership layer, not a better algorithm.
The 'learning gap': pilots stall on adoption, not model quality
MIT NANDA gave the failure a precise name: the 'learning gap.' Pilots fail not because the underlying models are weak, but because the tools don't learn, adapt, or integrate into how work actually gets done. There is no memory of the organization, no feedback loop, no redesign of the surrounding process — so the tool sits beside the workflow instead of inside it, and the value never compounds.
McKinsey describes the same phenomenon from a different angle as the 'gen AI paradox': horizontal copilots and chatbots scale fast but deliver diffuse, unmeasurable gains, while the transformative vertical, function-specific use cases — where the real money is — stay roughly 90% stuck in pilots. The easy wins are shallow; the deep wins require integration work that most firms skip.
The practical implication is liberating: the bottleneck is rarely the next, better model. Waiting for a more capable model to rescue a stalled pilot is usually a category error. The fix is integration and adoption — wiring the tool into a redesigned workflow, giving it access to the right data, building the feedback loops, and getting people to actually use it as part of their job.
Key insight
The reframe
Stop asking 'is the model good enough yet?' and start asking 'have we changed the workflow and the data access so this tool can actually do work here?' The first question stalls you for years; the second one is answerable this quarter.
Watch out
Where leaders get it wrong
MIT also found that more than half of GenAI budgets go to visible front-office functions while better ROI often sits in less glamorous back-office automation — and that 90%+ of employees use personal LLMs for work because the sanctioned tools never got integrated. Spending follows visibility, not value.
The five failure modes to recognize on sight
Stalled programs fail in recognizable, repeatable ways. Learn to spot these five — each has a one-line tell and a one-line fix.
| # | Failure mode | The tell | The fix |
|---|---|---|---|
| 1 | Pilots spread too thin | Dozens of shallow experiments, none owned, none scaling | Concentrate on 2–3 high-value vertical use cases tied to a real number |
| 2 | Weak C-suite engagement | AI delegated to a lab or to IT; leaders don't use it themselves | CEO/C-suite ownership, a unified vision, real funding, and role-modeling |
| 3 | No measurement | Success reported as logins, seats, tokens — 'adoption theater' | A three-tier value framework: usage → workflow efficiency → P&L impact |
| 4 | Bolting AI onto legacy workflows | The tool sits beside the process; nothing was redesigned | Redesign the end-to-end workflow around the AI (see next section) |
| 5 | Reaching for agents too early | Autonomous multi-step agents before the basics are integrated | Sequence: prove integrated copilots first; add autonomy where data and rules are clear |
Numbers 1, 2, and 4 are the ones leaders most often cause personally. Number 3 is the one that hides the other four — without honest measurement, a program can look healthy on a dashboard while creating no value, exactly as it did at Klarna.
A note on the fifth mode, because it is the 2025–2026 trap. Gartner forecasts that over 40% of agentic AI projects will be canceled by the end of 2027 — driven by unclear value, cost, and weak risk controls. Agents add autonomy, and autonomy multiplies both value and risk. They belong after you have an integrated, measured, well-governed foundation, not as the opening move.
Tip
The leadership move
Run your current AI portfolio against this table in a single meeting. Count how many of the five failure modes you can see. Most struggling programs exhibit three or four at once — and that, not the model, is your roadmap.
Example
Klarna — when measurement hides the failure
In 2023 Klarna replaced ~700 customer-service roles with an AI assistant that handled roughly two-thirds of queries. Volume metrics looked excellent. But CSAT and quality quietly deteriorated, and by mid-2025 the company reversed course and re-hired humans for a blended model. CEO Siemiatkowski: 'We focused too much on efficiency and cost… the result was lower quality, and that's not sustainable.' Volume metrics masked a quality collapse — track quality, not just activity.
The highest-leverage fix: redesign the workflow, don't bolt AI on
If you change only one thing, change this. McKinsey's research finds that workflow redesign has the single biggest effect on the EBIT impact firms get from AI — and yet only about 21% of firms have fundamentally redesigned their workflows. The overwhelming majority 'bolt AI onto' legacy processes and then wonder why the returns are diffuse.
Re-verify the 21% figure live before teaching: McKinsey, The State of AI: How organizations are rewiring to capture value (2025). The finding — redesign beats bolt-on — is durable; the exact percentage will drift.
The distinction in plain terms:
| Bolt-on (most firms) | Redesign (the winners) |
|---|---|
| Drop the tool next to the existing process | Re-imagine the end-to-end process around what AI can now do |
| Same steps, one of them slightly faster | Steps removed, re-sequenced, or fully automated |
| Incremental, hard-to-measure gains | Step-change in cost, speed, or quality |
| 'We gave everyone a copilot' | 'We rebuilt how claims/underwriting/support actually runs' |
This is why McKinsey's four pillars for agentic value all point the same way: (1) pick 2–3 high-value vertical use cases; (2) embed AI into reimagined workflows — not legacy ones; (3) scale with infrastructure, cross-functional teams, and strong governance; (4) build a new skills mix. Pillar 2 is the whole ballgame. Redesigning a workflow is, by definition, a 70%-layer activity — which is exactly why it is the executive's job and exactly why it gets skipped.
Key insight
The reframe
'Are we rolling out a tool, or redesigning a workflow?' is the most diagnostic question you can ask about any AI initiative. Tool rollouts produce demos and pilot purgatory. Workflow redesigns produce EBIT.
Watch out
Where leaders get it wrong
Assuming the productivity study transfers. A task-level speedup (e.g. a developer finishing one task faster) does not automatically become organization-level throughput unless the surrounding workflow is redesigned to absorb the gain. Without redesign, the saved time leaks away.
The attribute that predicts success: CEO oversight of governance
Across roughly 25 organizational attributes McKinsey tested for correlation with bottom-line impact, one stood out: CEO oversight of AI governance is the attribute most strongly correlated with EBIT impact, especially at large firms. It is not a coincidence that the most-cited 'leaders get it wrong' signal is the inverse — accountability for AI is diffuse, with only a minority of firms saying the CEO or board owns AI governance.
Re-verify live before teaching: McKinsey, The State of AI (2025). The ranking is the durable point; the exact attributes and percentages refresh annually.
Why would governance ownership — not, say, technical talent or budget — be the strongest predictor? Because CEO oversight of governance is a proxy for genuine ownership of the whole transformation. A CEO who owns governance is, in practice, the kind of leader who:
- treats AI as a strategic transformation, not an IT line item;
- funds and protects the 70% (workflow redesign, change management, reskilling);
- role-models use — and BCG found high performers are far more likely to have senior leaders who personally use AI daily;
- sets and tracks real KPIs, so the program is measured on value, not activity;
- and engages the C-suite — deeply engaged C-suite teams are reported to be many times more likely to be top AI value creators.
Governance here is not a brake. Done well, it is the mechanism by which a leader steers the program toward value and away from the five failure modes. The opposite — delegating AI to a central lab and hoping it diffuses upward — is the single clearest predictor of a stalled program.
Tip
The leadership move
Put one accountable executive's name on the AI transformation — ideally with the CEO owning governance — and have them personally use the tools weekly. Visible, hands-on senior ownership is the cheapest, highest-correlation intervention available, and it costs no model budget.
Example
JPMorgan — scoped, owned, measured
JPMorgan reportedly runs 450+ AI use cases in production with benefits growing ~30–40% year over year. The contrast with pilot purgatory is instructive: scoped use cases, senior ownership, and KPIs anchored to real numbers — the disciplined opposite of sprinkling shallow experiments and hoping a lab sorts it out.
Putting it together: the diagnostic the board should run
The themes converge into a short diagnostic any leader can run on any AI initiative. If the answers trend right, you have a transformation; if they trend left, you have a pilot heading for purgatory.
| Ask this | Stall signal (left) | Value signal (right) |
|---|---|---|
| Where is the budget going? | 70%+ on models and tech | Majority on the 70% — people, process, workflow |
| What are we actually doing? | Rolling out a tool | Redesigning an end-to-end workflow |
| How many bets? | Dozens of thin pilots | 2–3 concentrated, high-value vertical use cases |
| Who owns it? | A lab or IT; leaders don't use it | A named executive; CEO owns governance; leaders role-model |
| How do we measure success? | Logins, seats, tokens | Three tiers: usage → workflow efficiency → P&L impact |
| Are we reaching for agents yet? | Autonomous agents before integration | Integrated copilots first; autonomy where data and rules are clear |
The through-line of this entire lesson: the failure is organizational, not technical, and that is good news — because organizational levers are exactly the ones an executive controls. You do not need to become technical to fix a stalled AI program. You need to fund the 70%, redesign the workflow, concentrate the bets, measure value honestly, and own it from the top.
Note
Carry this forward
The next lessons build directly on this diagnosis: the AI operating model (hub-and-spoke / Center of Excellence) is how you structure the 70%, and AI governance (NIST AI RMF, ISO/IEC 42001, EU AI Act) is how CEO ownership becomes an operating discipline rather than a slogan.
Try it: Run the failure-mode diagnostic on your own AI portfolio
Goal: turn this lesson into a one-page diagnosis of where your organization actually stands — a strategic exercise, no technology required. 1) Inventory. List every AI initiative currently running or proposed in your organization (include shadow AI you know about). For each, note its sponsor, its stated success metric, and whether it changed a real number on the P&L. 2) Score against the five failure modes. For each initiative, mark which of the five it exhibits: pilots spread too thin, weak C-suite engagement, no measurement, bolting AI onto legacy workflows, agents too early. Most struggling initiatives will show three or four. 3) Check the 10-20-70 split. For your two largest AI investments, estimate the rough budget split across algorithms / tech & data / people & process. Flag any where less than ~30% goes to the people-and-process 70%. 4) Apply the tool-vs-redesign test. For each initiative, write one sentence answering 'are we rolling out a tool, or redesigning a workflow?' Be honest — most will be tool rollouts. 5) Name the owner. For your single most important initiative, write down who owns it, whether the CEO has oversight of its governance, and whether that owner personally uses the tool. 6) Produce two recommendations. Based on the diagnostic, write (a) one initiative to kill or consolidate, and (b) one workflow to redesign end-to-end with a named executive owner and a P&L-linked KPI. Deliverable: a single page you could present to your board that answers 'why are our AI initiatives not yet creating value, and what are the two highest-leverage moves to change that?'
Key takeaways
- 1Most AI initiatives fail organizationally, not technically: ~95% of enterprise GenAI pilots showed no measurable P&L return in MIT NANDA's 2025 study — and the cause was integration and adoption, not weak models.
- 2BCG's 10-20-70 rule says ~10% of effort is algorithms, ~20% tech and data, and ~70% people and process — winners invest in the 70%, the layer the executive owns; stalled pilots invert the ratio.
- 3The 'learning gap' (MIT) and the 'gen AI paradox' (McKinsey) describe the same trap: shallow horizontal copilots scale but deliver diffuse value, while deep vertical use cases stay stuck in pilots for lack of workflow integration.
- 4Five recurring failure modes: spreading pilots too thin, weak C-suite engagement, measuring activity instead of value, bolting AI onto legacy workflows, and reaching for autonomous agents too early.
- 5Workflow redesign has the single biggest effect on AI's EBIT impact (McKinsey), yet only ~21% of firms have fundamentally redesigned workflows — the rest bolt AI on and get incremental, hard-to-measure gains.
- 6CEO oversight of AI governance is the attribute most correlated with bottom-line impact (McKinsey) — a proxy for genuine, hands-on, top-down ownership of the whole transformation.
Quiz
Lock in what you learned
Check your understanding
0 / 4 answered
1.A board member argues, 'Our AI pilots keep failing — we should wait for the next, more powerful model before trying again.' Based on the 2025 evidence, what is the best response?
2.BCG's 10-20-70 rule describes how successful AI transformations allocate effort. What does the '70' refer to, and why does it matter to an executive?
3.According to McKinsey's research, which organizational change has the single biggest effect on the EBIT impact a firm gets from AI — and what is the catch?
4.McKinsey tested ~25 organizational attributes for correlation with bottom-line AI impact. Which one ranked highest, and why is it a meaningful signal?
Go deeper
Hand-picked sources to keep learning
The source of the '~95% of pilots show no P&L return' finding and the 'learning gap' diagnosis. Re-verify the figures before quoting.
The 'gen AI paradox' and the finding that CEO oversight of AI governance is most correlated with EBIT impact. Volatile stats — re-verify live.
Source for workflow redesign having the single biggest effect on EBIT impact, and the ~21% redesigned figure.
The 10-20-70 rule and the people-and-process emphasis that define this lesson's core framework.
The widening value gap: leaders vs. laggards on revenue growth, EBIT margin, and TSR. Re-verify the multiples live.
The canonical 'volume metrics masked a quality collapse' case — why you measure quality, not just activity.