Agentic AI AcademyAgentic AI Academy

The Next Value Wave — and the Paradox

Horizontal copilots vs. vertical agentic workflows

Intermediate 13 minDecision-maker
What you'll be able to do
  • Explain the gen AI paradox — why near-universal adoption has produced little bottom-line impact
  • Distinguish shallow horizontal copilots from deep vertical agentic workflows, and say where each creates value
  • Describe the "agentic organization" fix: reengineering core processes around human + agent coworkers
  • Assess the realistic 2025–2026 state of agent adoption — early, concentrated, and unevenly proven
  • Articulate the services-as-software / digital-worker thesis and why it expands AI's addressable market
  • Ask the right boardroom questions before greenlighting an agentic bet
At a glance

Almost everyone has adopted AI, yet almost no one has captured material P&L impact — the "gen AI paradox." This lesson explains why: value has been trapped in shallow, horizontal copilots, while the real prize sits in deep, vertical agentic workflows that reengineer how core work actually gets done. You will leave able to tell copilot value from agent value, judge the realistic state of agent adoption, and recognize the "services-as-software" shift that aims AI at the labor budget, not just software seats.

  1. 1The gen AI paradox: everyone adopted, almost nobody profited
  2. 2Why the value leaked out: shallow horizontal copilots vs. deep vertical workflows
  3. 3The fix: the "agentic organization"
  4. 4The realistic state of agent adoption (2025–2026)
  5. 5Services as software: agents aim at the labor budget
  6. 6The right questions to ask before greenlighting an agentic bet

The gen AI paradox: everyone adopted, almost nobody profited

Start with the uncomfortable headline every board should sit with: near-universal adoption, negligible bottom-line impact.

McKinsey frames it as the "gen AI paradox": roughly 80% of companies use generative AI, yet roughly 80% report no material impact on the P&L (McKinsey, Seizing the Agentic AI Advantage, 2025). The broader State of AI survey puts adoption near-universal — ~88% of organizations use AI in at least one function — but only ~39% report enterprise-level EBIT impact, and most of those attribute under 5% of EBIT to it (McKinsey State of AI, Nov 2025).

This is not a one-firm finding. Independent research converges on the same gap:

Source (date)FindingImplication
McKinsey State of AI (Nov 2025)~88% adopt; ~39% report EBIT impact; most <5% of EBITAdoption ≠ value
MIT NANDA GenAI Divide (Jul 2025)~95% of enterprise GenAI pilots show no measurable P&L returnThe "learning gap," not model quality
BCG Widening AI Value Gap (Sep 2025)Only ~5% "future-built"; ~60% "laggards"A small minority capture value at scale

The diagnosis: the problem is almost never the model. MIT NANDA names the root cause the "learning gap" — flawed workflow integration and process fit, not technology quality. If a pilot isn't moving a number, look at the workflow, not the model.

(All figures are time-stamped survey findings that refresh roughly annually — treat them as a 2025–2026 snapshot and re-verify against the live source before quoting.)

Key insight

The reframe

"Everyone has AI; almost nobody has AI value yet." The differentiator between the 5% and the 60% is operating model and workflow redesign — not which model they bought. That makes the value problem a leadership problem, which is exactly the part executives own.

Watch out

Where leaders get it wrong

Reading near-universal adoption as success. Seat counts, logins, and "we've rolled out Copilot" are activity, not impact. The board question is not "are people using it?" but "what number on the P&L moved, and by how much?"

Why the value leaked out: shallow horizontal copilots vs. deep vertical workflows

If adoption is everywhere and value is nowhere, where did the value leak out? McKinsey's answer is the heart of this lesson: most spend went into shallow, horizontal copilots — broad assistants that make many people a little faster at many tasks — when the prize sits in deep, vertical agentic workflows that reengineer a specific core process end-to-end.

The distinction is the whole story:

Horizontal copilotVertical agentic workflow
ScopeBroad — everyone, every taskNarrow — one core process, end-to-end
What it changesIndividual task speedThe process itself
Value shapeDiffuse, shallow, hard to measureConcentrated, deep, P&L-visible
EffortBuy seats, light changeRedesign the workflow + governance
P&L impactRarely moves EBITWhere the real value lands

Copilots make existing tasks faster — incremental productivity that's genuinely real but spreads too thin to surface on the income statement. Agents can do a whole workflow end-to-end — a step change that requires more redesign and more governance, but lands where value concentrates.

This is why the same research that documents the paradox also documents the cure: McKinsey finds workflow redesign has the single biggest effect on whether AI moves EBIT — yet only about a fifth of firms have fundamentally redesigned a workflow; most simply bolt AI onto a legacy process and capture little.

Tip

The leadership move

Before funding the next AI initiative, ask which it is: "Are we buying a tool to make a task faster, or redesigning a workflow to change how the work is done?" The first rarely moves EBIT; the second is where value lives. Fund a few of the second, not a hundred of the first.

Example

Bolting on vs. reengineering

A horizontal rollout gives every employee a chat assistant and counts logins — broad, shallow, invisible on the P&L. A vertical play takes one process (say, claims intake or invoice reconciliation), redesigns it around human + agent steps with measurement built in, and reports a cost or cycle-time number the CFO can see. Same budget; only the second shows up in EBIT.

The fix: the "agentic organization"

The corrective McKinsey proposes is not "add more agents." It is to stop treating AI as a feature bolted onto today's org chart and start reinventing core processes around human + agent coworkers — what McKinsey calls "the agentic organization."

The logic chains together:

  1. Chat and copilots gave broad but shallow gains — useful, but they don't reengineer the process, so they don't move the P&L.
  2. Agents execute whole workflows, which is what makes them an operating-model change rather than a productivity feature.
  3. The value only materializes when the workflow is redesigned around the agent — not when an agent is dropped into an unchanged legacy process.

McKinsey's prescription for getting there (its "four pillars") is deliberately the opposite of "agents everywhere":

PillarWhat it means for a leader
FocusPick 2–3 high-value vertical use cases — not scores of shallow experiments
ReimagineEmbed agents into redesigned workflows, not legacy ones
ScaleBack winners with infrastructure, cross-functional teams, and strong governance
ReskillBuild a new human + agent skills mix

The recurring governance principle underneath all of it: automation for execution, humans for judgment. Agents run the steps; people own the decisions, the exceptions, and the accountability.

Key insight

The reframe

An agent dropped into an unchanged process is a faster version of a broken workflow. The "agentic organization" idea says the unit of transformation is the workflow, not the tool — you are redesigning how work flows between people and agents, not installing software.

Watch out

Where leaders get it wrong

Launching "agents everywhere" as a banner program. Spreading thin across dozens of disconnected pilots consumes budget without coordinated value. Winners concentrate on a few end-to-end workflows; stalled programs run scores of shallow ones.

The realistic state of agent adoption (2025–2026)

Before championing agents internally, calibrate to reality — agent adoption today is early, concentrated, and unevenly proven. The hype is running ahead of the reliability, and a clear-eyed leader is more credible than an evangelist.

What the time-stamped evidence shows:

Signal (source, date)Reading
~23% scaling agents in ≥1 function; ~39% experimenting (McKinsey State of AI, Nov 2025)Most firms are early — pilots, not production at scale
No single function runs agents at full scale in more than ~10% of firms (McKinsey, 2025)Adoption is thin even where it exists
Agents ~17% of AI value in 2025 → ~29% projected by 2028 (BCG, Sep 2025)A growing share of the value pie...
46% piloting agents, but only ~16% of those show tangible value yet (BCG, Sep 2025)...but a minority of pilots have proven out
>40% of agentic AI projects predicted canceled by end of 2027 (Gartner, Jun 2025)Cost, unclear value, and weak risk controls bite

Where agents actually work today is narrow and structured: the most common enterprise use cases are IT / service-desk triage and resolution and knowledge "deep research" agents that pull from internal sources (McKinsey, 2025). Production agents "tend to be simple and human-supervised" (Anthropic) — reliable for bounded, well-structured, lower-stakes tasks, not autonomous "set-and-forget" digital employees.

Two cautions for the boardroom. First, "agent washing" is rampant — Gartner estimates only ~130 of thousands of self-styled "agentic" vendors are genuine; the anti-hype test is autonomous reasoning + tool orchestration + persistent context. Second, data and integration readiness gates everything — about 8 in 10 firms cite data limitations as the roadblock to scaling agents.

Tip

The leadership move

Position agents as a next step, not a starting point. The prerequisites are strong data foundations, scaled AI capabilities, and clear governance. Greenlight a few high-value end-to-end workflows with workforce training — and budget for the realistic chance a pilot gets killed or reshaped, rather than promising the board autonomous digital employees by next quarter.

Example

The Klarna both-sides cautionary tale

Klarna's AI assistant handled ~two-thirds of support chats in month one (~2.3M conversations, ~700 FTE-equivalent, ~$40M projected 2024 profit uplift; Feb 2024, vendor-sourced). Then it partly reversed — by mid-2025 Klarna rehired humans after CSAT declined, the CEO admitting it had "focused too much on efficiency and cost." The lesson: volume metrics masked a quality collapse. Track quality (CSAT/NPS), not just throughput, and blend humans with agents on nuanced work.

Services as software: agents aim at the labor budget

There is a bigger reason agents are positioned as the next wave and not just a better copilot — and it's an economic one, not a technical one. The investor framing (a16z) is "services as software" / digital workers: the unit of sale shifts from software seats to work delivered.

The distinction reshapes the size of the prize:

Software-as-a-service (the old model)Services-as-software (the agent model)
What you buyA tool a human operatesAn outcome, delivered
ExampleA contract-review application, licensed per seatThe contract review itself, completed
Budget it targetsThe software lineThe labor / services line
Market sizeSeats × priceA far larger services/labor spend

Instead of licensing a tool that a person uses to do the work, an agent (internal or vendor) delivers the outcome the work used to produce — a reconciled invoice batch, a reviewed contract, a researched supplier comparison. This is why a16z frames agents as a far larger market than SaaS: they go after the services and labor budget, not just the software budget.

For an executive, this reframes two things at once:

  • The opportunity — AI stops being an IT cost line and becomes a lever on the largest cost in most businesses: people and outsourced services.
  • The risk and the responsibility — "digital workers" sit under human oversight, not instead of it. The augmentation narrative (work redesigned around human + agent, not wholesale replacement) is both the realistic near-term path and the defensible one.

Key insight

The reframe

Copilots sell speed; agents sell outcomes. Once the unit of value is "work delivered" rather than "a tool used," the relevant budget is the services and labor line — orders of magnitude larger than the software line. That is the real reason agents are framed as a step change, not an upgrade.

Watch out

Where leaders get it wrong

Hearing "digital worker" and modeling straight headcount replacement. The proven near-term pattern is augmentation under human oversight — and the Klarna reversal shows what happens when leaders optimize for cost-elimination and let quality slip. Frame agents as coworkers with guardrails, not as a layoff plan.

The right questions to ask before greenlighting an agentic bet

You don't need to evaluate the technology to govern it well — you need to ask the questions that separate a real, value-bearing agentic bet from a shallow copilot rollout or an "agent-washed" demo. Bring these to the table:

On the value thesis

  • Is this a tool rollout or a workflow redesign? (Only the second reliably moves EBIT.)
  • Which one core process does it reengineer end-to-end — and what P&L number do we expect to move?
  • Are we concentrating on 2–3 vertical use cases, or spreading thin across scores of pilots?

On readiness

  • Is our data clean, governed, and accessible enough to feed this? (8 in 10 firms stall here.)
  • Have we redesigned the workflow, or are we bolting an agent onto a legacy process?

On "is it even a real agent"

  • Does it show autonomous reasoning + tool orchestration + persistent context — or is it a rebranded chatbot? (The anti–agent-washing test.)

On governance and measurement

  • What is our human-in-the-loop policy for a system that acts, and can we halt or roll back?
  • Are we tracking quality (CSAT/NPS, error rates), not just volume? (The Klarna lesson.)
  • What is our plan if this pilot is one of the >40% Gartner expects to be canceled — kill it, or reshape it?

The throughline: champion agents where the workflow is redesigned, the data is ready, the value is concentrated, and a human owns the judgment. Everywhere else, you are funding activity, not impact.

Tip

The leadership move

Make "tool rollout or workflow redesign?" the first question in every AI funding conversation. It instantly sorts shallow horizontal copilots (fund sparingly) from deep vertical agentic bets (fund deliberately) — and signals to the organization that you understand where value actually comes from.

Try it: Strategic exercise: from copilot sprawl to a vertical agentic bet

Goal: practice the core executive judgment of this lesson — telling shallow copilot activity from a deep vertical bet that can actually move the P&L. No technology required; this is a one-page strategic worksheet you could take to a leadership meeting. 1) Inventory the sprawl. List every place your organization already uses AI (include shadow/unsanctioned use). For each, mark whether it is a horizontal copilot (broad, task-speed) or a vertical workflow (one core process, end-to-end). Most will be copilots — that's the point. 2) Diagnose the paradox locally. For your top 3 copilot uses, write the honest answer to: 'What number on the P&L has this moved?' If you can't name one, note why — almost always it's diffuse value, not a bad tool. 3) Pick one vertical candidate. Choose a single high-value core process (e.g., claims intake, invoice reconciliation, customer-journey support, supplier research) where an end-to-end agentic workflow could reengineer the work. State the one EBIT or cycle-time number you'd expect to move. 4) Pressure-test it with the right questions. Score your candidate against: tool rollout vs. workflow redesign? data clean/governed/accessible? a real agent (reasoning + tool orchestration + persistent context) or agent washing? human-in-the-loop and halt/rollback defined? measuring quality (CSAT/NPS, error rate), not just volume? 5) Apply the services-as-software lens. Note whether this bet targets the software budget (a tool) or the labor/services budget (an outcome delivered) — the latter signals a bigger prize and bigger governance need. 6) Write the one-paragraph recommendation: which 2–3 vertical use cases you'd fund deliberately, which copilot sprawl you'd stop funding, and the single metric by which you'll judge success in 90 days. Deliverable: a one-page opportunity map plus a five-question screen you can reuse on every future AI proposal.

Key takeaways

  1. 1The gen AI paradox (McKinsey, 2025): ~80% of firms use generative AI yet ~80% report no material P&L impact — value leaked into shallow horizontal copilots instead of deep vertical workflows. (~88% adopt; ~39% report EBIT impact, most under 5% — McKinsey State of AI, Nov 2025.)
  2. 2The fix is the "agentic organization": reengineer 2–3 core processes around human + agent coworkers — workflow redesign, not bolting agents onto legacy processes, is what moves EBIT.
  3. 3Agent adoption is early and concentrated: ~23% scaling, ~39% experimenting (McKinsey, 2025); agents ~17%→~29% of AI value (BCG, 2025); most common in IT/service-desk and "deep research" — and only ~16% of agent pilots show tangible value yet.
  4. 4Services as software / digital workers (a16z): agents target the labor and services budget, not just software seats — they deliver the outcome (e.g., a contract reviewed) rather than licensing a tool. That is why agents are framed as a step change.
  5. 5Stay clear-eyed: ~95% of GenAI pilots show no P&L return (MIT NANDA, 2025), >40% of agentic projects may be canceled by 2027 (Gartner), "agent washing" is rampant, and the Klarna reversal warns that volume metrics can hide a quality collapse.
  6. 6The governing question for any agentic bet: "Is this a tool rollout or a workflow redesign?" — with automation for execution and humans for judgment, plus a halt/rollback policy and quality (not just volume) metrics.

Quiz

Lock in what you learned

Check your understanding

0 / 4 answered

1.A board member says, "We've rolled AI out to 90% of staff — clearly our AI strategy is working." Which insight from the gen AI paradox should you raise?

2.Why does McKinsey argue that deep vertical agentic workflows, not horizontal copilots, are where AI value concentrates?

3.Your CTO proposes "deploying autonomous agents across every function this year." Based on the 2025–2026 evidence, what is the most accurate caution?

4.What does the "services as software" / digital-worker thesis (a16z) claim is strategically significant about agents?

Go deeper

Hand-picked sources to keep learning