Enterprise AI adoption,
2023-2026.
Every claim sourced, plus agents & skills, plus why the studies disagree.
This is the report I keep current on the state of enterprise AI: what adoption actually looks like once you separate surveys from measurements, where the value lands and where the pilots die, what the tooling layer (MCP, Agent Skills) changed, and what people actually want from agents. Every number carries its source, and every source carries its incentive.
Four findings, up front.
- Adoption is near-universal. Value is rare. 88% of organizations report using AI (McKinsey, Nov 2025), up from 55% in 2023 (Stanford HAI). Yet MIT Project NANDA (Jul 2025) found only ~5% of integrated pilots extract real value; the rest show no measurable P&L impact.
- Value lands in three places: coding, customer-support deflection, back-office automation. Nearly everything else is still pilots. Anthropic overtook OpenAI at 40% of enterprise LLM spend (Menlo, Dec 2025).
- The tooling layer moved faster than the value did. MCP went from launch (Nov 2024) to 10,000+ public servers in ~12 months. Agent Skills went from a Claude feature (Oct 2025) to a cross-vendor standard in under six months. Both shipped without a security model: 36.8% of audited skills contain a flaw (Snyk ToxicSkills, Feb 2026).
- What people actually want from agents is not replacement. Anthropic's Economic Index Survey (Jun 2026, ~9,700 linked respondents) found the top hope is collaboration on meaningful work, second is automating drudgery. And the heaviest delegators are the most optimistic about their own jobs, not the least.
Who said it, and what
they sell.
Every source used, what it actually measures, and who paid for it. Read this before trusting any number below.
| # | Source | Date | Method / sample | Measures | Incentive flag |
|---|---|---|---|---|---|
| S1 | McKinsey, The State of AI | Nov 2025 | Survey, 1,993 respondents, 105 countries, fielded Jun-Jul 2025 | Self-reported adoption, EBIT impact | Sells AI transformation |
| S2 | Stanford HAI AI Index | 2025, 2026 | Meta-analysis of public + survey data | Adoption, capability, investment | Academic, lowest conflict |
| S3 | US Census BTOS | Ongoing, biweekly | Nationally representative, ~1.2M firms | Actual AI use in production | Government, no incentive |
| S4 | Ramp AI Index | Ongoing | Transaction data, 70,000+ firms | Paid AI vendor spend | Ramp sells corporate cards |
| S5 | Menlo Ventures, State of GenAI in the Enterprise | Dec 9, 2025 | Survey, 495 US decision-makers, fielded Nov 7-25, 2025 + bottoms-up market model | Spend, vendor share, build/buy, architecture | VC; invested in Anthropic, Supabase, Pinecone, Databricks, Lovable |
| S6 | MIT Project NANDA, The GenAI Divide | Jul 2025 | 300+ initiatives, 52 interviews, 153 surveys, ~6 months | Pilot-to-P&L conversion | Academic, preliminary, not peer-reviewed |
| S7 | METR, Measuring the Impact of Early-2025 AI on Developer Productivity | Jul 2025 | RCT, 16 experienced OSS devs, 246 real tasks | Measured dev speed | Non-profit eval org, low conflict |
| S8 | Gartner press releases | Jun 2025, 2026 | Analyst forecast + client data | Project cancellation, agent deployment | Sells research subscriptions |
| S9 | BCG, The Widening AI Value Gap | Sep 30, 2025 | 1,250 executives | Value capture tiers | Sells AI consulting |
| S10 | Deloitte, State of AI in the Enterprise | Jan 2026 | 3,235 leaders, 24 countries, fielded Aug-Sep 2025 | Agent maturity, governance, transformation depth | Sells AI consulting |
| S11 | KPMG Global AI Pulse | Mar 31, 2026 | 2,110 senior leaders, 20 countries | Agent scaling, budgets | Sells AI advisory |
| S12 | KPMG US AI Quarterly Pulse Q2 | Jun 24, 2026 | US senior leaders, quarterly panel | Agent deployment, cost visibility | Same |
| S13 | Anthropic Economic Index, Cadences | Jun 26, 2026 | Privacy-preserving telemetry + survey linked to usage, ~9,700 respondents | How and why people use AI | Vendor-published; own product data |
| S14 | Anthropic Economic Index, Jan 2026 report | Jan 2026 | Telemetry, Nov 2025 sample | Automation vs augmentation split | Same |
| S15 | Anthropic Economic Index, geography | 2025 | Telemetry, global | API vs consumer usage patterns | Same |
| S16 | Snyk Labs, ToxicSkills | Feb 5, 2026 | Scanned 3,984 skills from ClawHub + skills.sh | Agent-skill security | Snyk sells security scanning |
| S17 | VentureBeat, Agent Skills open standard | Dec 18, 2025 | Reporting + Anthropic PM interview | Skills standardization | Trade press |
| S18 | The New Stack, Agent Skills spec | Dec 2025 | Reporting | Skills timeline | Trade press |
| S19 | GitHub Octoverse 2025 | 2025 | Platform telemetry, 180M devs | Developer AI adoption | Microsoft-owned; sells Copilot |
| S20 | Klarna/OpenAI announcement + Bloomberg follow-up | Feb 2024 / May 2025 | Company self-report, then CEO reversal | Support automation outcomes | Both parties promotional at launch |
| S21 | Goldman Sachs Research | 2025-2026 | Macro modeling | AI capex, productivity | Sell-side research |
| S22 | arXiv: Emerging Threats of the Agent Skill Ecosystem | May 2026 | Threat taxonomy, real samples | Skills security | Preprint, not peer-reviewed |
| S23 | Agent Skills '26 / SkillsBench | 2026 | Benchmark across public skills | Skill quality effect on agent performance | Academic workshop |
| S24 | Ecosystem trackers (Termdock, Agentman, OSS Insight, SpecWeave) | 2026 | Secondary aggregation, self-published | Skill counts, catalog size | Low reliability, unverified blogs. Treat as directional only. |
Everyone adopted.
Few converted.
1.1 Headline trend
| Metric | 2023 | 2024 | 2025 | 2026 YTD | Source |
|---|---|---|---|---|---|
| Orgs using AI (≥1 function) | 55% | 72-78% | 88% | - | S1, S2 |
| Orgs scaling enterprise-wide | - | ~33% | ~33% | - | S1 |
| Report any EBIT impact | - | - | 39% | - | S1 |
| High performers (>5% EBIT) | - | - | 6% | - | S1 |
| Census BTOS (production use) | - | ~4.6% | 10% (old def) / 17.3% (new def) | 17-20% | S3 |
| Ramp (paid AI adoption) | - | - | 44.5% Aug → 46.6% Dec | >50% Mar 2026 | S4 |
| Enterprise GenAI spend | $1.7B | $11.5B | $37B | - | S5 |
| Enterprises scaling agents | - | - | 23% | 11% "AI leaders" | S1, S11 |
Read this carefully: the four adoption rows measure four different things. See §6.
1.2 Spend detail (all S5)
- Total 2025: $37B, 3.2x YoY. Applications $19B, infrastructure $18B.
- Applications = 6% of the entire software market within three years of ChatGPT's launch.
- Application split: horizontal $8.4B, departmental $7.3B, vertical $3.5B.
- Coding alone: $4.0B, up from $550M in 2024, the single largest category anywhere in the app layer.
- Infrastructure split: foundation model APIs $12.5B, training infra $4.0B, data/orchestration $1.5B.
- Excluded from these figures: chips, cloud inference (AWS/GCP/Azure), and AI features bolted into existing SaaS. So the real number is higher; this is the net-new AI market.
Menlo Ventures bottoms-up market model (S5). Excludes chips, cloud inference, and AI features inside existing SaaS, the net-new market only.
1.3 Macro context (S21)
- AI capex ≈ 0.8% of US GDP, below the 1.5%+ seen in past tech booms.
- Hyperscaler capex to exceed $500B in 2026.
- Net US GDP impact only ~0.1-0.3pp in 2026, because much of the hardware is imported.
- Goldman found no meaningful economy-wide AI/productivity relationship yet. Gains are localized to coders and support reps.
1.4 Build vs. buy: corrected
S5 reports: 47% built / 53% bought (2024) → 24% built / 76% bought (2025).
Do not read this as "enterprises stopped building." Four corrections:
- Menlo's own text notes continued strong investment in internal builds; roughly a third of AI budgets still go to them.
- The unit is use cases, not dollars or effort. Ten bought SaaS tools + three deep internal builds scores 77% "buy."
- Buying Cursor or Supabase counts as "buy", then you build with it. The categories overlap.
- The infra/app spend split is ~50/50 ($18B vs $19B), which is the closest proxy for buy-to-build vs buy-to-use. That undercuts the headline.
There is no company-size segmentation on build/buy in S5. If you need mid-market vs. large-enterprise, use S10 (Deloitte, segments by size and geography).
The failure numbers, and
the success numbers.
2.1 The failure numbers
| Finding | Figure | Source |
|---|---|---|
| Integrated pilots extracting real value | ~5% | S6 |
| Pilots with no measurable P&L impact | ~95% | S6 |
| Companies past proof-of-concept to real value | 26% | BCG, Oct 2024 |
| "Future-built" firms capturing full value | 5% | S9 |
| Laggards with little/no value | 60% | S9 |
| Agentic AI projects to be cancelled by end-2027 | >40% | S8 |
| Orgs that have deployed AI agents | 17% | S8 (2026 Hype Cycle) |
| Enterprises scaling agents enterprise-wide | 11% | S11 |
| Real-time visibility into AI running costs | 26% | S12 |
| Mature governance model for agentic AI | 21% | S10 |
2.2 The success numbers
| Finding | Figure | Source |
|---|---|---|
| AI deals converting to production | 47% (vs 25% SaaS) | S5 |
| Report AI delivering meaningful business outcomes | 64% | S11 |
| Report first-year ROI | 74% | Google Cloud ROI report, vendor |
| Avg cost savings, early adopters | 15.2% | S8 |
| Avg productivity improvement | 22.6% | S8 |
| Pilots blending internal + external expertise | 67% success | S6 |
| IT-only internal builds | 22% success | S6 |
| Future-built firms' revenue growth multiple | 1.7x | S9 |
2.3 Why projects fail
- Not model quality. S6's diagnosis: tools can't retain feedback, adapt to context, or improve over time, and organizations bolt AI onto legacy processes instead of redesigning them.
- S11's version of the same finding: the 11% who scale redesign the process first, then deploy agents into it. The 89% do the reverse.
- S1's blockers: data quality, workflow rigidity, operating-model inertia, measurement gaps.
- Budget misallocation (S6): over half of AI budgets went to sales and marketing, which produced low ROI, while back-office automation delivered the actual returns.
- Shadow AI (S6, S5): 90%+ of firms have employees using personal AI accounts for work. S5 estimates PLG + shadow adoption is close to 40% of application AI spend.
2.4 The coding counter-evidence (S7)
The most rigorous study in the whole corpus, and it cuts against the hype.
METR ran a randomized controlled trial: 16 experienced open-source developers, 246 real tasks in codebases they already knew, mostly Cursor Pro with Claude 3.5/3.7 Sonnet.
- Devs predicted a 24% speedup beforehand.
- After finishing, they estimated AI had made them 20% faster.
- Measured result: AI made them 19% slower.
Randomized controlled trial: 16 experienced open-source developers, 246 real tasks in codebases they already knew, early-2025 tooling. The 39-point gap between felt and measured speed, in the wrong direction, is the finding.
Caveats that matter: small sample, mature codebases with high context load, early-2025 tooling. It does not generalize to juniors, greenfield work, or prototyping, where gains of 27-90% are reported elsewhere. METR's Feb 2026 follow-up was judged an unreliable signal.
What to take from it: the perception gap is the finding. People are bad at estimating their own AI-assisted productivity, in a consistent direction. Any client self-report of "we're 30% faster" is unverified until measured.
Where the money goes,
and where it comes back.
| Domain | Spend / adoption | Reported value | Status | Source |
|---|---|---|---|---|
| Software eng / coding | $4.0B; 50% of devs daily (65% top-quartile); ~80% of new GitHub devs use Copilot in week 1; 90% of Fortune 100 | 15%+ velocity claimed; METR RCT measured −19% for experts | Proven at scale, claims inflated | S5, S19, S7 |
| Customer service | $630M departmental; largest agentic category | Intercom Fin ~76% avg resolution across 12,000 customers; Anthropic internal 50.8% resolution, 1,700+ hrs saved; cost/resolution $0.10-0.99 vs $6-20 human | Proven for tier-1; quality risk | S5, Fin.ai, Anthropic |
| Sales & marketing | Marketing $660M; 78% startup share in sales tools | 30% cut in external agency spend (one MIT case) | Mixed, high spend, S6 flagged low ROI | S5, S6 |
| Finance & accounting | 91% startup share (AI-first ERPs) | Back-office automation = highest ROI in S6 ($2-10M savings) | Emerging, best ROI/effort ratio | S5, S6 |
| Legal | $650M vertical market | Contract review, triage | Emerging | S5 |
| HR / recruiting | 5% of departmental spend | - | Experimental | S5 |
| Healthcare / life sci | $1.5B (43% of vertical AI); ambient scribes $600M (+2.4x YoY) | Scribes cut documentation time >50% | Proven in documentation | S5 |
| IT operations | $700M | Incident response, infra management | Emerging | S5 |
| Supply chain / mfg | - | Multi-objective optimization (cost vs time-to-market) | Experimental | S10 |
| Product / design | Design 7% of departmental spend | - | Emerging | S5 |
3.1 The Klarna arc: the single most instructive case (S20)
Feb 2024 (company announcement): the AI assistant handled 2.3M conversations, about two-thirds of Klarna's support chats, described as the equivalent work of 700 full-time agents. Resolution time fell from 11 minutes to under 2. Projected profit improvement: $40M for 2024.
May 2025 (CEO to Bloomberg): Siemiatkowski said the cuts had gone too far. AI-only support produced lower quality. Klarna began rehiring humans into a hybrid model.
The lesson: resolution rate is not resolution quality. Both facts are true and both were reported by the same company 15 months apart. Any support-automation business case should carry a quality floor, not just a deflection target.
The layer that moved
faster than the value.
4.1 Model Context Protocol
Timeline
- Nov 2024: Anthropic open-sources MCP.
- Mar 2025: OpenAI adopts (ChatGPT, Agents SDK, Responses API).
- 2025: Google DeepMind, Microsoft, AWS, Cloudflare follow.
- Dec 2025: donated to the Linux Foundation's Agentic AI Foundation (co-founded by Anthropic, Block, OpenAI; backed by Google, Microsoft, AWS, Cloudflare, Bloomberg).
Scale (Anthropic, Dec 2025)
- 10,000+ active public MCP servers.
- 97M+ monthly SDK downloads, up from ~100K at launch, roughly 970x in 18 months.
- Official registry lists 6,400+ servers.
Security: unresolved
- OWASP frames the risk as tool poisoning, sitting between LLM01 (prompt injection) and LLM05 (supply-chain).
- Two CVEs put it on the map: MCPoison (CVE-2025-54136), CurXecute (CVE-2025-54135).
- Apr 2025: researchers demonstrated indirect prompt injection via emails, docs, and web pages.
- Nov 2025: a WhatsApp MCP integration flaw allowed extraction of full message histories via poisoned tool descriptions.
- The structural flaw: tool descriptions are reviewed once at connect time, but tool responses reach the model context at runtime with no equivalent check.
Enterprise footprint: Microsoft Security Blog (Feb 10, 2026), citing the Microsoft Data Security Index 2026: over 80% of Fortune 500 run active AI agents, but only 47% of those organizations have implemented specific security controls over them. (A frequently repeated "28% of Fortune 500 have implemented MCP servers" figure traces to vendor estimates with no named primary survey, do not cite it.)
4.2 Agent Skills: the newer layer
Timeline (S17, S18)
- Oct 2025: Anthropic launches Agent Skills: folders containing instructions, scripts, and resources that teach an agent a repeatable workflow. A skill is a SKILL.md file with YAML frontmatter plus markdown body.
- Dec 2025: released as an open standard at agentskills.io with a reference SDK.
- Adopters named by Anthropic's PM: Microsoft (VS Code, GitHub), Cursor, Goose, Amp, OpenCode. OpenAI adopted a structurally identical architecture in ChatGPT and Codex CLI, same file naming, same metadata format, same directory layout.
- Feb 2026: enterprise controls: org-wide provisioning on Team/Enterprise plans, plus stock plug-ins for finance, legal, HR.
Why it spread so fast: it answers a specific question cheaply, how do you make an assistant consistently good at specialized work without fine-tuning a model. The barrier to publish is a markdown file and a GitHub account.
Does it work? (S23: SkillsBench)
- Curated skills raise agent pass rates by +16.2 points on average.
- 2-3 focused skills deliver +18.6 points; monolithic "everything in one doc" skills reduce performance by 2.9 points.
- Self-generated skills hurt performance on nearly a third of tasks.
- Average quality score across public skills: 6.2 out of 12. Benchmarks used only top-quartile skills (9+).
Security: worse than MCP (S16, S22)
- Snyk scanned 3,984 skills across ClawHub and skills.sh (Feb 5, 2026): 1,467 (36.8%) had at least one security flaw; 534 (13.4%) critical; 76 confirmed malicious payloads: credential theft, reverse shells, data exfiltration.
- Daily publishing rate went from under 50 in mid-Jan 2026 to over 500 by early Feb, 10x in weeks. Vetting capacity did not scale with it.
- Five named threat actors operated across multiple platforms. ClawHub was subsequently shut down.
- OWASP published an Agentic Skills Top 10 on Apr 27, 2026. AST01 is Malicious Skills, rated Critical.
- The mechanism: skills are not sandboxed plugins. They execute with the host agent's full privileges, filesystem, terminal, network, credentials.
Practical rule: treat a community skill exactly like an unaudited npm package with code-execution rights. Read the source. Curated internal libraries beat public catalogs on both quality and safety.
4.3 Agent architecture reality (S5)
- Only 16% of enterprise and 27% of startup deployments qualify as true agents, where the model plans, executes, observes, and adapts.
- The rest are fixed-sequence or routing workflows around a single model call. S5's phrasing: basic if-then logic around a model call.
- Customization techniques by frequency: prompt design first, then RAG. Fine-tuning, tool calling, context engineering, and RL remain niche.
- Copilots dwarf agents in spend: $7.2B (86%) vs $750M (10%).
4.4 LLM vendor share, enterprise API usage (S5)
| Vendor | 2023 | 2024 | 2025 |
|---|---|---|---|
| Anthropic | 12% | 24% | 40% |
| OpenAI | 50% | ~34% | 27% |
| 7% | - | 21% | |
| Open-weight (Llama et al.) | - | 19% | 11% |
Estimated dollars based on self-reported production API usage share (S5), see the methodology caveat below. Missing bars are years the source published no figure.
- Top three = 88% of enterprise LLM API usage.
- Coding specifically: Anthropic ~54% vs OpenAI 21%, up from 42% six months earlier, driven by Claude Code.
- Chinese models: ~1% of enterprise API usage, but rising fast among startups and indie devs via OpenRouter and vLLM (Qwen, DeepSeek, GLM). Airbnb uses Qwen for user-facing features; Cursor used it as the base for an internal model.
- Methodology caveat: these are estimated dollars based on self-reported production API usage share, weighted by application scale and triangulated with public financials. Not audited revenue. And Menlo is an Anthropic investor.
4.5 High-growth infra tools
| Company | ARR | Valuation | Note | Source |
|---|---|---|---|---|
| Supabase | ~$170M (May 2026), from ~$101M end-2025, $30M end-2024 | $10.5B post (Jun 2026 Series F, GIC) | Backend of the vibe-coding boom, Lovable and Bolt run on it; 4M+ devs | Sacra, press |
| Cursor / Anysphere | ~$2B (Feb 2026) → ~$4B (mid-2026); $100M Jan 2025 → $1B Nov 2025 | $29.3B (Nov 2025); SpaceX agreed to acquire ~$60B all-stock (Jun 2026) | Fastest B2B software ramp on record; 70% of Fortune 1,000 | Press |
| Lovable | $100M (mid-2025) → $400M+ (Feb 2026) | $6.6B (Dec 2025, CapitalG/Menlo) | A reported $12B round was in talks, not closed | Sacra, press |
| Replit | ~$10M → $100M in ~6 months (2025); ~$250M Oct 2025 | $9B (Mar 2026 Series D) | Press | |
| Vercel | $200M+ (mid-2025) | $9.3B post (Sep 2025 Series F) | Press | |
| n8n | ~$40M (Jul 2025) | $2.5B (Oct 2025 Series C, Accel; Nvidia participating) | Built on open-source community adoption before enterprise sales | Press |
| Zapier | ~$420M (Q1 2026) | $5B (2021 secondary, stale) | Zapier Agents GA May 2025; Zapier MCP; AI-task volume +760% | Sacra |
| Databricks | $5.4B run-rate (+65% YoY); AI products $1.4B annualized | $134B (Series L, Feb 2026) | A reported $165B+ round was in discussions | Company |
| Snowflake | Public | - | 9,100+ customers use AI products weekly; $100M AI run-rate hit ahead of plan | Company |
| Pinecone | - | $750M (2023, stale) | Reportedly weighing a sale | Press |
| Hugging Face | - | $4.5B (2023, stale) | Press | |
| Weaviate | - | $50M Series B (2023) | Later Series C unconfirmed | Press |
Signal in the stale rows: vector-DB valuations are 2+ years old with no fresh priced round. Incumbents added vector search and squeezed the category. S5 confirms the direction, incumbents hold 56% of AI infrastructure spend, because even AI-native app builders keep choosing Databricks, Snowflake, MongoDB, and Datadog.
The section on
purpose.
This is the section on purpose. Most reports measure spend and adoption. Very few measure why. The best data here comes from Anthropic's Economic Index (S13, S14, S15), vendor-published, but it is telemetry linked to a survey, not a self-report questionnaire, which makes it structurally harder to game than anything in §1.
5.1 The core frame: automation vs augmentation
Anthropic classifies every conversation into five interaction modes, grouped into two families (S13, S14):
Automation, the human delegates
- Directive: hand over a complete task, minimal back-and-forth. "Translate this document."
- Feedback loop: the human relays real-world outcomes back to the model. "Make this email more casual."
Augmentation, the human collaborates
- Task iteration: work through it together, human refining outputs.
- Learning: ask for explanation or understanding, not task completion.
- Validation: ask the model to check your own work.
5.2 The purpose trend, tracked over 18 months
| Period | Augmentation | Automation | Directive share | Source |
|---|---|---|---|---|
| Jan 2025 | 56% | 41% | 27% | S14 |
| Aug 2025 | - | overtook augmentation | 39% | S14 |
| Nov 2025 | 52% | 45% | 32% | S14 |
| Feb 2026 | slightly up | - | - | Mar 2026 report |
Read the shape, not the points. Directive use climbed hard through Aug 2025, then fell back 7pp. Anthropic's own reading: the August spike overstated how fast delegation was arriving, but the underlying direction is still toward automation. The Nov pullback is attributed partly to product changes, file creation, memory, and Skills: that encourage more collaborative, iterative work.
5.3 Enterprise purpose looks completely different from consumer purpose
This is the sharpest split in the data (S15):
- API / enterprise traffic: 77% automation patterns, mostly directive. Only 12% augmentation.
- Claude.ai / consumer: roughly an even split.
Anthropic Economic Index telemetry (S15). The concrete enterprise patterns (S14): email classification, invoice processing, calendar scheduling, back-office throughput, exactly where S6 found the ROI actually was.
Businesses do not buy AI to think alongside. They buy it to hand work over. S14 names the concrete enterprise patterns: email classification, invoice processing, calendar scheduling, back-office throughput, exactly where S6 found the ROI actually was.
5.4 What they're producing (S13, Apr-Jun 2026 sample)
93% of conversations produce an identifiable artifact. Top categories:
| Artifact | Share of all conversations | Work-related share |
|---|---|---|
| Explanations | 17% | - |
| Documents & reports | 15% | - |
| Guidance | 11% | 80%+ personal |
| Marketing content | - | 80% work |
| Blogs / articles | - | 81% work |
| Database queries | - | 82% work |
Flipped by purpose: work conversations most often produce documents and reports (20%), then explanations (9%), email drafts (7%), analyses and summaries (6%).
5.5 The autonomy gradient: and why the product matters more than the model
S13's most useful finding for anyone building agent workflows.
AI autonomy is rated 1-5. Across 26 of 31 output types, autonomy is higher in Claude Code than in chat. The average gap is 0.37 points.
The concrete example: producing a blog post. The median chat session takes 13 rounds of back-and-forth. The median Claude Code session doing the same job contains a single human prompt.
Two-thirds of the gap is the same tasks being executed with more delegation. One-third is a different mix of work.
And it isn't the model. Claude Code runs Opus far more often (54% vs 10% in chat). But comparing only Sonnet sessions, Claude Code still shows 0.26 points more autonomy. S13's conclusion: the product surface matters more than the underlying model.
5.6 More valuable work costs more compute (S13)
- Token consumption rises with the wage of the occupation the task maps to. Marketing managers earn ~2x what editors do; their conversations use ~2.5x the tokens.
- Building apps uses 3x the median conversation's tokens. A typical explanation uses about a fifth.
- Autonomy and token use rise together (r = 0.68).
- But the human doesn't disappear at the top end. In higher-wage work, Claude produces 1.34x more per turn and users engage 1.53x more turns, with extended thinking on more often. S13's read: these move together, which looks labor-augmenting rather than labor-displacing.
5.7 How people get into using them: the on-ramp
Four documented pathways, in order of how much of the market they explain:
1. Product-led, bottom-up (S5). 27% of AI application spend arrives through PLG, nearly 4x the 7% rate in traditional software. Counting shadow AI on personal cards, close to 40%. Cursor reached $200M revenue before hiring a single enterprise sales rep. n8n formalized contracts only after hundreds of employees were already using it. Lovable, OpenRouter, ElevenLabs, Gamma, Wispr Flow followed the same pattern.
2. Delegation as learning-by-doing (S13). People who delegate more report AI can do more of their work, and expect it to do more next year. Anthropic offers two readings: either delegation teaches you what AI can actually finish, or people who already believe it can do their job are the ones willing to hand it over. Both are plausible; the data can't separate them. Either way, usage drives belief, not the reverse.
3. Skills as the specialization step. Once a team is using an agent, skills are how they encode house rules without touching a model. S23 quantifies the payoff: +16.2 points pass rate from curated skills, +18.6 from 2-3 focused ones. Also the trap: self-generated skills hurt performance on nearly a third of tasks. So the on-ramp is curate, not generate.
4. Enterprise governance retrofit (S17, S10). Central provisioning of skills arrived Feb 2026, after the bottom-up adoption. Governance is being fitted to existing usage, not preceding it. Only 21% of firms have a mature agentic governance model (S10).
5.8 What people say they want: the direct answer to "what is the human purpose"
S13 asked ~9,700 linked respondents an open-ended question: what do you hope an AI-shaped economy looks like in ten years? Classified themes, top three:
| Rank | Theme | Share |
|---|---|---|
| 1 | Augmentation: collaborating on work that feels meaningful, careers still mattering, new industries and jobs | Over half |
| 2 | Automation of drudgery: offloading tedious work for more free time and meaning outside work | Just over half |
| 3 | Shared prosperity: that the economic gains are widely distributed | ~1/3 |
Not replacement. Not headcount reduction. The top two hopes are keep the meaningful part, remove the boring part.
Supporting evidence, same survey:
- Productivity gains reported: speed 86%, scope 82%, quality 69%. 27% report cost savings on services they'd otherwise buy.
- 68% report learning more with AI; 57% say AI made their skills more valuable.
- Over a third expect AI to handle most or nearly all of their work tasks within 12 months.
- Only 10% rate losing their own job as likely, slightly below the US annualized involuntary-separation rate (~13.4%). But they're far more worried for others: over a third put a junior colleague's job-loss probability above 60%.
5.9 The counterintuitive finding worth carrying into the next conversation
The heaviest delegators are the most optimistic, on all six dimensions measured (pay, job security, ability to find a new job, meaning, autonomy, human interaction). Largest effects on expected pay and job-finding ability.
And the common fear, that delegating means offloading thinking and eroding skill, does not show up in the data. Heavy delegators report learning at the same rate as everyone else, and are more likely to say their skills grew in market value.
Two honest caveats (S13's own):
- Selection can't be ruled out, enthusiasts may both delegate more and feel better. Though the effect survives controlling for account tenure.
- These are self-assessments. Skills can erode even while someone reports learning more. The data does not disprove skill erosion; it just doesn't find it.
5.10 Who these findings do not describe
S13 is explicit that the survey is not representative:
- Computer & mathematical occupations: ~30% of respondents vs 4% of US employment.
- Management: 23% of respondents vs 7% of employment.
- Transportation, food service, construction: heavily under-represented.
- Women are 12% of the linked sample, and use Claude measurably differently: 6.3pp lower Claude Code share, 7.3pp lower automation share, more iterative use, more active minutes in chat, even after controlling for occupation.
So §5 describes technical and managerial knowledge workers who already use AI heavily. Do not generalize it to a whole workforce.
Staring at an AI statistic you are not sure you trust? That is the conversation we like having.
Book a call →Four true numbers,
one reality.
The single most useful section for judging any AI statistic you're handed.
6.1 The adoption number is four different numbers
| Source | Figure | What it actually asks | Why it's high or low |
|---|---|---|---|
| McKinsey (S1) | 88% | "Does your org use AI in at least one function?" | Highest. One person in one department counts. Respondents are AI-interested leaders. Firm sells AI transformation. |
| Stanford (S2) | 78-88% | Aggregates survey data | Inherits survey bias, but neutral analysis |
| Ramp (S4) | >50% | Did the company pay an AI vendor? | Middle. Hard transaction data, but misses in-house builds and open-source, and skews to Ramp's customer base (US, tech-forward) |
| Census (S3) | 17-20% | "Did this firm use AI in producing goods or services?" | Lowest. Nationally representative across ~1.2M firms including small and non-tech. Best methodology, narrowest question. |
All four are true; they answer different questions. Self-reported adoption runs 2-4x higher than transaction- or production-based measures.
The reconciliation: nearly every large, tech-forward company has touched AI. Roughly half pay a vendor. Fewer than one in five have it in actual production of goods and services. All four are true. They're answering different questions.
6.2 MIT's 95% vs Menlo's 47%: a direct collision
| MIT NANDA (S6) | Menlo (S5) | |
|---|---|---|
| Headline | ~95% of pilots show no P&L impact | 47% of AI deals reach production (vs 25% SaaS) |
| Unit measured | Value delivered to the income statement | Procurement conversion to deployment |
| Sample | 300+ initiatives, 52 interviews, 153 surveys | 495 US decision-makers |
| Method | Qualitative + survey, ~6 months | Survey + bottoms-up market model |
| Published by | Academic, preliminary, not peer-reviewed | VC with portfolio stakes in the market measured |
| Direction of bias | Toward pessimism (interview-led, failure-salient) | Toward optimism (portfolio value, explicit rebuttal framing) |
These are not actually contradictory. A project can reach production (Menlo's bar) and still move no P&L line (MIT's bar). Menlo measures deployment. MIT measures value. Both can be right simultaneously, and probably are.
Menlo's Dec 2025 report explicitly frames itself against MIT's finding, which tells you the framing was chosen, not discovered.
Methodology critiques of MIT worth knowing: the sample is small and preliminary; the authors describe it as directionally accurate rather than definitive; critics have called it methodologically fragile. Use it as a directional truth about the pilot-to-value gap, not as a precise failure rate.
6.3 Agent deployment: 11%, 16%, 17%, 42%, 53%
Five credible sources, five very different numbers, all within twelve months.
| Figure | Source | Date | What it counts |
|---|---|---|---|
| 11% | KPMG Global (S11) | Mar 2026 | Scaling agents enterprise-wide with business outcomes ("AI leaders") |
| 16% | Menlo (S5) | Dec 2025 | Deployments that are architecturally true agents (plan-execute-observe-adapt) |
| 17% | Gartner (S8) | 2026 | Organizations that have deployed agents at all |
| 42% | KPMG US | Sep 2025 | Orgs that have deployed at least some agents |
| 53% | KPMG US (S12) | Jun 2026 | Orgs deploying agents, quarterly panel |
One measure (share of organizations), five definitions of "agent." Menlo's 16% is the strictest technical test; KPMG's 53% is the loosest; KPMG's own 11% is the strictest business test.
The spread is the definition of "agent," not disagreement about reality. Menlo's 16% is the strictest, it's a technical architecture test. KPMG's 53% is the loosest, any deployment counts. KPMG's own 11% "AI leaders" figure is the strictest business test.
Practical translation: roughly half of large firms have something they call an agent. About one in six is technically an agent. About one in nine gets enterprise-wide value from it.
Gartner also notes widespread "agent washing": of thousands of self-described agentic AI vendors, they estimate only about 130 are real.
6.4 Skills ecosystem counts: 3,984 vs 22,511 vs 47,150 vs 490,000
Four numbers circulating for "how many skills exist." All from 2026. They differ by two orders of magnitude.
| Count | Source | What it is | Reliability |
|---|---|---|---|
| 3,984 | Snyk ToxicSkills (S16) | Skills scanned from ClawHub + skills.sh, Feb 5, 2026 | High: primary security research, defined corpus |
| 22,511 | Secondary aggregation (S24) | Skills in a broader security audit | Medium, audit not independently verified |
| 47,150 | SkillsBench (S23) | Public skills analyzed for quality | Medium-high, academic benchmark |
| 490,000+ | Ecosystem blogs (S24) | Claimed total across three marketplaces, Mar 2026 | Low: self-published, no methodology |
Why the gap: these are corpora, not censuses. Snyk scanned what it could scan. SkillsBench analyzed what it could benchmark. The 490K figure counts everything ever published including abandoned duplicates. OSS Insight notes the long tail is vast and largely unused, thousands published, nobody installs them.
6.5 Survey vs measurement: the widest gap in the whole field
| Domain | Self-reported | Measured | Gap |
|---|---|---|---|
| Dev productivity | +20% (devs' own post-hoc estimate, S7) | −19% (RCT, S7) | 39 points, wrong direction |
| Org adoption | 88% (S1) | 17-20% (S3) | 4x |
| AI project ROI | 74% report first-year ROI (Google Cloud) | 39% report any EBIT impact (S1); ~5% real value (S6) | 2-15x |
| Support automation | 2/3 of chats automated (Klarna, Feb 2024) | Quality dropped, humans rehired (Klarna, May 2025) | Reversed within 15 months |
The pattern is consistent and directional: people overestimate AI's benefit to their own work, and organizations overestimate their own AI maturity. This isn't dishonesty, METR showed devs were wrong about their own measured performance in real time.
6.6 A checklist for the next AI statistic you're handed
Who fielded it and what do they sell?
Consultancies sell transformation. VCs hold portfolios. Vendors sell tools. All three publish real data with chosen framing.
Survey or measurement?
Self-report inflates 2-4x. Transaction, telemetry, and RCT data don't.
What's the unit?
Use cases, dollars, deployments, and organizations give wildly different answers to "how much AI is there."
What's the definition?
"Agent" swings a number from 11% to 53% with no change in underlying reality.
What's the sample frame?
495 US decision-makers ≠ 1.2M US firms ≠ 9,700 heavy Claude users.
Is the effect measured against a baseline?
Almost never. This is why the METR result was surprising.
Eight rules the data
actually supports.
Define the P&L metric and its baseline before building anything.
The pilot-to-value gap (§2) is the market's central failure. If you can't name the number that moves, run discovery, don't build. Measure the baseline, or you'll never distinguish the METR effect from real gain.
Start where value is documented: internal tooling/coding, tier-1 support deflection, back-office automation.
Avoid sales and marketing content as a first project, S6 found it absorbed over half of budgets and returned the least.
Consider co-building the third option, not a compromise.
S6's split is the strongest datapoint here: 67% success for internal+external blended teams vs 22% for IT-only builds. The build/buy binary in §1.4 erases exactly that model.
Redesign the process, then deploy the agent.
S11's diagnosis of the 89% is that they lay AI over existing workflows. BCG's 10/20/70 rule says the same: 10% algorithms, 20% tech and data, 70% people and process.
Package for delegation deliberately.
§5.5 is the actionable finding: same model, different surface, 13 turns vs 1. Decide upfront whether a workflow should be delegated or collaborative, and build the surface to match. Don't leave it to chance.
Curate a skills library; never install from public catalogs.
36.8% flaw rate, 13.4% critical, code execution with full host privileges (§4.2). But curated skills are worth +16.2 points on agent pass rates. The value is real and so is the risk, which makes curation work worth budgeting for, not overhead.
Treat MCP as the default integration layer and assume it's insecure.
Build for MCP compatibility. Put a gateway in front that validates tool schemas before they reach the model. Never connect an untrusted server to write access or sensitive data.
Start discovery with shadow AI, not procurement.
27-40% of AI app spend enters bottom-up (§5.7). Ask what people already run on personal accounts. That's the real adoption baseline and the fastest path to a working use case.
Thresholds that flip the recommendation
| Signal | Action |
|---|---|
| Measured gain <10% after 90 days | Wrong tool for that workflow, stop, don't tune |
| Resolution rate up, CSAT down | Cap automation, go hybrid (the Klarna line) |
| Company under ~$20M ARR | Almost never build custom; off-the-shelf pays back in 3-9 months |
| No named success metric | Discovery engagement, not a build |
| Agent handles a decision with no reversal path | Deterministic workflow instead |
| Skill sourced from a public catalog | Read the source or don't ship it |
What this report
can't claim.
- Surveys and hard data disagree by 2-4x and I've kept both rather than averaging them. §6.1 explains why. Any single adoption number in this document is incomplete without its method.
- Vendor-incentive flags are in §0 and repeated inline. Menlo (VC, Anthropic investor), Google Cloud, Anthropic, GitHub, Ramp, Deloitte, BCG, McKinsey, KPMG, Snyk all sell into this market. Least conflicted: US Census, Stanford HAI, METR. MIT NANDA is unconflicted but preliminary.
- MIT's 95% is preliminary and contested. Small qualitative sample, not peer-reviewed, described by its own authors as directional. Use it for the shape of the challenge, not as a failure rate.
- §5 rests heavily on Anthropic's own telemetry. It is the best data available on why people use agents, and telemetry beats self-report, but it covers Claude users only, skews technical and managerial, and Anthropic has an interest in the augmentation framing landing well. The automation/augmentation classification is Anthropic's own taxonomy, not an industry standard.
- Forward-looking figures are labelled, not asserted. Lovable's ~$12B and Databricks' ~$165B+ rounds were in talks, not closed. Gartner's 40% cancellation rate is a forecast.
- Private ARR figures are estimates (Sacra, press) unless company-disclosed. Pinecone, Hugging Face, and Zapier valuations are 2+ years stale.
- Skill counts above 50,000 come from self-published ecosystem blogs with no methodology. Flagged S24 throughout. Don't put them in a deck.
- Everything here is US- and English-weighted. Menlo's sample is US-only. Census is US-only. Stanford noted China and Europe posted the highest YoY adoption increases, but the granular data behind that is thinner than the US picture.
Deciding what to build
from numbers like these?
If you are trying to work out where AI would actually move a number in your business, and which statistics to ignore on the way, that is the conversation we like having.