The Moat Is the Data, Not the Model
Why models commoditize and data compounds
- Explain why foundation models are a 'strategic commodity' and what the collapsing cost of intelligence means for strategy
- Identify the four places durable AI advantage actually lives once the model is a commodity
- Describe the data flywheel and why proprietary data is the only AI moat that compounds
- Recognize that siloed, ungoverned data is a liability — 'AI agents cannot query what they cannot find' — and what the board owes it
- Make the case for model portability and multi-model architecture to avoid single-vendor lock-in
- Ask the right board-level questions to tell a commodity capability apart from a moat-building one
Foundation models are becoming a strategic commodity: the cost of GPT-3.5-level intelligence fell roughly 280x in about 18 months (Stanford HAI AI Index 2025), and open-weight models have nearly closed the quality gap. Durable advantage is therefore migrating up the stack — to proprietary data, workflow integration, customer relationships, and vertical specialization. This lesson makes the board-level case that governed, discoverable data is the one AI moat that compounds, and that you should architect for model portability rather than lock into any single vendor.
- 1The model is becoming a commodity
- 2The moat is what you build around the model
- 3Data is the only moat that compounds
- 4Ungoverned data is a liability, not a moat
- 5Architect for portability — avoid single-vendor lock-in
- 6The board's questions — commodity or moat?
The model is becoming a commodity
Start with the single most consequential shift in AI economics for a board: the model itself is commoditizing. Gartner has gone so far as to call foundation models a "strategic commodity" — meaning differentiation based on raw model performance alone is unlikely to last.
Three forces drive this:
- The cost of intelligence is collapsing. The inference cost to reach GPT-3.5-level quality fell from about $20 per million tokens (late 2022) to about $0.07 (late 2024) — a roughly 280x reduction in ~18 months (Stanford HAI AI Index 2025). Treat the exact figure as time-stamped and re-verify it before each board cycle — but the direction is the durable point.
- Quality is converging. Leading models share similar architectures and training methods; open-weight models have nearly closed the quality gap. The frontier still moves, but the gap between "the best" and "good enough for the job" keeps shrinking.
- Anyone can deploy one. Cloud APIs, fine-tuning frameworks, and open weights let any enterprise use a capable model without building one.
The strategic translation: the intelligence you are buying gets dramatically cheaper and more interchangeable every year. That is wonderful news for adoption — and a warning for anyone who thinks owning or betting on one model is the strategy.
Key insight
The reframe
If your AI advantage is 'we use the best model,' you have no advantage — your competitor can rent the same model next quarter for a fraction of today's price. A commodity, by definition, is something everyone can buy. The model is becoming exactly that.
Watch out
Where leaders get it wrong
Signing rigid multi-year, fixed-price model commitments — or assuming today's leading model is a moat. The price you are locked into is the price your competitor will undercut once intelligence gets cheaper. Don't anchor strategy to a depreciating asset.
The moat is what you build around the model
If the model is the commodity, where does durable advantage go? It migrates up the stack — away from raw intelligence and toward how that intelligence is applied, integrated, owned, and governed. As Marc Andreessen's firm a16z frames it: "The moat is not the model — it's what you build around it."
Four places durable advantage actually lives:
| Where advantage lives | Why a rival can't easily copy it | The leadership question |
|---|---|---|
| Proprietary data | They cannot buy, borrow, or replicate your unique operational data | Does this use case compound a dataset only we own? |
| Workflow integration | The value is in redesigned operating processes, not a bolted-on tool | Are we redesigning the workflow or just adding a chatbot? |
| Customer relationships & distribution | Trust and reach are earned over years, not licensed | Does this deepen a relationship only we have? |
| Vertical / domain specialization | Legal, medical, and industrial AI each need different data, workflows, and compliance | Do we own domain depth a generalist vendor lacks? |
Notice what's not on this list: the model. The model is the raw engine everyone can rent. The moat is the proprietary asset the engine is pointed at.
Tip
The leadership move
For every proposed AI use case, ask one sorting question: 'Is this capability available to all my competitors off the shelf — or does it compound a unique asset we own?' If it's off-the-shelf, buy the commodity cheaply and move on. If it compounds proprietary data, workflow, or relationships, that's where you invest to own it.
Example
JPMorgan: scoped to where it compounds
JPMorgan runs 450+ AI use cases in production with benefits growing ~30-40% year over year — not by chasing the best model, but by pointing AI at its proprietary data and workflows, anchored to KPIs. The advantage is the bank's data and processes; the model is interchangeable plumbing.
Data is the only moat that compounds
Of the four, one stands apart: proprietary data is the one thing competitors cannot buy, borrow, or replicate. And uniquely, it can compound — through what's called a data flywheel.
The flywheel works like this: when AI is architected so that every interaction generates feedback data — what the user accepted, corrected, or rejected; what outcome followed — that data feeds back into improving the product, which attracts more usage, which generates more data. The advantage doesn't just persist; it widens over time. A rival starting later faces a gap that grows every day you operate.
This is why data is the only AI moat that compounds. The model gets cheaper for everyone. Your customer relationships grow linearly. But a well-architected data flywheel grows the gap non-linearly — and it is the asset a competitor with a bigger checkbook still cannot purchase.
"The one thing competitors cannot buy, borrow, or replicate is your data." As models commoditize, proprietary data becomes the differentiator — and the board's strategic asset to protect.
The board-level reframe: turn your unique operational knowledge into a defensible, revenue-generating, board-reportable strategic asset. That changes data from an IT cost center into a P&L-relevant moat the board is accountable for stewarding.
Key insight
Why 'compounds' is the load-bearing word
Most assets depreciate or stay flat. A data flywheel appreciates: more use generates more feedback data generates a better product generates more use. That positive loop is the only AI advantage that gets stronger while the model layer gets cheaper and more equal for everyone.
Example
Klarna: feedback data is two-sided
Klarna's AI assistant handled ~2/3 of customer chats and cut resolution from ~11 minutes to under 2 — then the company rehired humans by mid-2025 after customer satisfaction fell ('we focused too much on efficiency... not sustainable' — CEO). The lesson for the flywheel: the feedback data you capture must include quality (CSAT/NPS), not just volume, or you optimize the wrong loop.
Ungoverned data is a liability, not a moat
Here is the trap. Leaders hear "data is the moat" and assume they already have one — we have decades of data! But raw data is not a moat. Data is only a moat if it is governed and usable.
Proprietary data becomes an asset only when it is discoverable, accessible under policy, and auditable. If it lives in disconnected silos — no unified catalog, no access controls, no lineage — it is not a dormant moat. It is an active liability: a compliance exposure, a security surface, and a blocker to the very AI you want to deploy.
The phrase to remember, because it makes the abstract concrete:
"AI agents cannot query what they cannot find."
An agent pointed at your business is only as good as the data it can actually reach under policy. Roughly 80% of the real AI work is data, governance, and workflow integration (MIT Sloan) — not the model. McKinsey calls the precondition "data productization": treating data products as managed, owned assets, which is what unlocks scaling agents at all.
| Data as a moat (governed) | Data as a liability (ungoverned) |
|---|---|
| Discoverable via a unified catalog | Trapped in disconnected silos |
| Accessible under clear policy | No access controls — all-or-nothing |
| Auditable lineage you can show a regulator | No record of where it came from |
| Feeds the flywheel | Cannot be queried by your own agents |
| A board-reportable strategic asset | A breach and compliance exposure |
Watch out
Where leaders get it wrong
Confusing 'we own a lot of data' with 'we have a data moat.' Volume in silos is not an advantage — it is unrealized liability. The bottleneck on most AI programs is not the model; it is that the data is not discoverable, governed, or usable by the systems meant to consume it.
Tip
The leadership move
The board's three verbs for data: govern it, protect it, activate it. Charter a data-readiness assessment before the flagship AI bet — ask whether your data is clean, catalogued, access-controlled, and accessible enough to feed the use case. If the honest answer is no, that foundation is the first investment, not the model.
Architect for portability — avoid single-vendor lock-in
If the model is a commodity whose price keeps falling and whose quality keeps converging, the strategic posture follows directly: don't lock the company into one vendor's model or price card. Architect for model portability — the ability to swap the underlying model with minimal disruption — and assume multi-model is the norm, not the exception.
The evidence backs this up. Multi-model is already standard enterprise practice (a16z reports ~37% of enterprises use 5+ models, differentiating by use case), and roughly three-quarters of enterprise use cases are bought, not built. A telling nuance leaders miss: the household-name consumer brand is not the enterprise-spend leader. Per Menlo Ventures' 2025 survey, enterprise LLM API spend was led by Anthropic (40%) over OpenAI (27%) and Google (21%) — a reversal from 2023. Treat any single market-share snapshot as volatile; the durable lesson is the rankings move, so keep your optionality.
What portability buys you:
- Pricing leverage — as intelligence gets cheaper, you capture the savings instead of being locked into yesterday's rate.
- Resilience — if a vendor fails, is breached, changes terms, or deprecates a model, you switch rather than scramble.
- Best-fit per use case — different models lead at coding, reasoning, or multimodal work; portability lets you route each task to the right engine.
The inverse mistake is equally common on the buying side: blanket per-seat licensing on day one — reports show 30-40% of company-wide Copilot/Gemini seats sit unused within 90 days. Start with high-value cohorts and expand on evidence.
Tip
The leadership move
Treat the model layer like cloud compute: a swappable, competitively-priced commodity you keep optionality over. Invest your scarce, hard-to-reverse capital in the moat layer — proprietary data, workflow redesign, governance — not in long, rigid commitments to a depreciating model price.
Watch out
Where leaders get it wrong
Assuming 'OpenAI = the market' because ChatGPT is the household name, then betting the architecture on one vendor. Enterprise spend tells a multi-vendor story, and the leaders shift year to year. Single-vendor lock-in trades a falling-price commodity for a rising strategic risk.
The board's questions — commodity or moat?
Bring this to the boardroom as a small set of durable questions that separate spend on commodities from investment in moats. They outlast any model ranking or price.
On every use case:
- Is this capability available to all competitors off the shelf (a commodity — buy it, move on), or does it compound a unique asset we own (a differentiator — invest to own it)?
- If this is a moat, which asset does it compound — proprietary data, workflow, customer relationship, or vertical depth?
On data:
- Is our data clean, governed, catalogued, and accessible enough to actually feed this use case? Could an agent find it?
- Are we treating our proprietary data as a board-reportable strategic asset — governed, protected, activated — or as an IT cost we ignore until it leaks?
- Does this design create a flywheel — does every interaction generate feedback data that compounds our advantage?
On vendors and architecture:
- Can we swap the underlying model without re-platforming? Are we architected for portability?
- Which critical capabilities are really someone else's model — and what happens if that vendor fails, is breached, or changes terms?
- Are we about to license seats for the whole workforce on day one, or start with high-value cohorts and expand on evidence?
The through-line: spend like a buyer on the commodity layer; invest like an owner on the moat layer.
Key insight
One distinction to carry into every AI decision
Buy the model; build the moat. The model is rented intelligence that gets cheaper for everyone. The moat is the proprietary data, workflow, and relationships the intelligence is pointed at — the only part a richer competitor still cannot purchase.
Try it: Commodity-vs-Moat Audit: separate what to buy from what to own
Goal: leave with a one-page board artifact that sorts your AI portfolio into 'buy the commodity' vs. 'build the moat,' plus a data-readiness verdict. No code — this is a strategic exercise. 1) List 5-8 current or proposed AI use cases. Pull them from any function (support, marketing, finance, ops, product). 2) Apply the sorting question to each: is this capability available off the shelf to all competitors (COMMODITY — buy it, move on), or does it compound a unique asset we own (MOAT — invest to own it)? For each 'moat' case, name which asset it compounds: proprietary data, workflow, customer relationship, or vertical depth. 3) Flywheel test. For your top 2 moat candidates, write one sentence on whether every interaction generates feedback data that compounds advantage — and whether you're capturing quality (CSAT/NPS), not just volume (the Klarna lesson). 4) Data-readiness verdict. For each moat case, rate the underlying data on three governed-data tests — discoverable (is there a catalog?), accessible under policy (are there access controls?), auditable (is there lineage?). Mark any silo where an agent 'cannot query what it cannot find.' 5) Portability check. Note where you are locked into a single model/vendor and what it would take to swap; flag any rigid long-term price commitment or blanket day-one seat licensing. 6) Write the board recommendation (half a page): which use cases to buy as commodities, which one moat to invest in first, the single biggest data-governance gap to fix before that bet, and one model-portability action. The deliverable is the artifact a board needs to spend like a buyer on the model layer and invest like an owner on the moat layer.
Key takeaways
- 1Foundation models are a 'strategic commodity' (Gartner): the cost of GPT-3.5-level intelligence fell ~280x in ~18 months (Stanford HAI AI Index 2025, time-stamped — re-verify), and quality has converged. Raw model performance is no longer a durable advantage.
- 2Durable advantage migrates up the stack to four places: proprietary data, workflow integration, customer relationships, and vertical specialization — 'the moat is not the model, it's what you build around it' (a16z).
- 3Proprietary data is the only AI moat that compounds: a data flywheel means every interaction generates feedback data that widens the gap non-linearly — an asset a richer competitor still cannot buy.
- 4Raw data is not a moat — governed, discoverable data is. Siloed, ungoverned data is a liability: 'AI agents cannot query what they cannot find.' The board's job is to govern, protect, and activate it.
- 5Architect for model portability and multi-model; avoid single-vendor lock-in. Intelligence keeps getting cheaper and the enterprise-spend leaders shift year to year — keep optionality and don't sign rigid long-term price commitments.
- 6The board's sorting question for every use case: is this an off-the-shelf commodity (buy it, move on) or does it compound a unique asset we own (invest to own it)? Buy the model; build the moat.
Quiz
Lock in what you learned
Check your understanding
0 / 4 answered
1.An executive argues, 'Our AI strategy is sound because we've contracted for the highest-performing foundation model on the market for the next three years.' What is the strongest objection a board member should raise?
2.Once foundation models commoditize, where does durable competitive advantage migrate?
3.Why is proprietary data described as the only AI moat that 'compounds'?
4.A company owns decades of operational data sitting in disconnected, ungoverned silos with no catalog or access controls. How should the board regard this for AI strategy?
Go deeper
Hand-picked sources to keep learning
The core 'moat moves up the stack' thesis — proprietary data, workflow, and data productization as durable advantage.
Source for the ~280x inference-cost decline (GPT-3.5-level: ~$20 → ~$0.07 per million tokens, 2022→2024). Volatile — re-verify each cycle.
The 'moat is not the model, it's what you build around it' framing and the value-migrates-up-the-stack thesis.
Enterprise LLM spend share (Anthropic 40% / OpenAI 27% / Google 21%) and the ~76%-bought-not-built figure. Volatile market snapshot.
Evidence that the bottleneck is the 'learning gap' — data, workflow, integration — not model quality; ~95% of pilots show no P&L return.
Board-level framing of proprietary, governed data as the compounding moat and of ungoverned data as a liability.