// research report · updated august 2026

Enterprise AI adoption,
2023-2026.

Every claim sourced, plus agents & skills, plus why the studies disagree.

This is the report I keep current on the state of enterprise AI: what adoption actually looks like once you separate surveys from measurements, where the value lands and where the pilots die, what the tooling layer (MCP, Agent Skills) changed, and what people actually want from agents. Every number carries its source, and every source carries its incentive.

AI adoptionAI agentsAgent SkillsMCPAI ROIMethodology
// tl;dr

Four findings, up front.

  • Adoption is near-universal. Value is rare. 88% of organizations report using AI (McKinsey, Nov 2025), up from 55% in 2023 (Stanford HAI). Yet MIT Project NANDA (Jul 2025) found only ~5% of integrated pilots extract real value; the rest show no measurable P&L impact.
  • Value lands in three places: coding, customer-support deflection, back-office automation. Nearly everything else is still pilots. Anthropic overtook OpenAI at 40% of enterprise LLM spend (Menlo, Dec 2025).
  • The tooling layer moved faster than the value did. MCP went from launch (Nov 2024) to 10,000+ public servers in ~12 months. Agent Skills went from a Claude feature (Oct 2025) to a cross-vendor standard in under six months. Both shipped without a security model: 36.8% of audited skills contain a flaw (Snyk ToxicSkills, Feb 2026).
  • What people actually want from agents is not replacement. Anthropic's Economic Index Survey (Jun 2026, ~9,700 linked respondents) found the top hope is collaboration on meaningful work, second is automating drudgery. And the heaviest delegators are the most optimistic about their own jobs, not the least.
// §0 · source register

Who said it, and what
they sell.

Every source used, what it actually measures, and who paid for it. Read this before trusting any number below.

#SourceDateMethod / sampleMeasuresIncentive flag
S1McKinsey, The State of AINov 2025Survey, 1,993 respondents, 105 countries, fielded Jun-Jul 2025Self-reported adoption, EBIT impactSells AI transformation
S2Stanford HAI AI Index2025, 2026Meta-analysis of public + survey dataAdoption, capability, investmentAcademic, lowest conflict
S3US Census BTOSOngoing, biweeklyNationally representative, ~1.2M firmsActual AI use in productionGovernment, no incentive
S4Ramp AI IndexOngoingTransaction data, 70,000+ firmsPaid AI vendor spendRamp sells corporate cards
S5Menlo Ventures, State of GenAI in the EnterpriseDec 9, 2025Survey, 495 US decision-makers, fielded Nov 7-25, 2025 + bottoms-up market modelSpend, vendor share, build/buy, architectureVC; invested in Anthropic, Supabase, Pinecone, Databricks, Lovable
S6MIT Project NANDA, The GenAI DivideJul 2025300+ initiatives, 52 interviews, 153 surveys, ~6 monthsPilot-to-P&L conversionAcademic, preliminary, not peer-reviewed
S7METR, Measuring the Impact of Early-2025 AI on Developer ProductivityJul 2025RCT, 16 experienced OSS devs, 246 real tasksMeasured dev speedNon-profit eval org, low conflict
S8Gartner press releasesJun 2025, 2026Analyst forecast + client dataProject cancellation, agent deploymentSells research subscriptions
S9BCG, The Widening AI Value GapSep 30, 20251,250 executivesValue capture tiersSells AI consulting
S10Deloitte, State of AI in the EnterpriseJan 20263,235 leaders, 24 countries, fielded Aug-Sep 2025Agent maturity, governance, transformation depthSells AI consulting
S11KPMG Global AI PulseMar 31, 20262,110 senior leaders, 20 countriesAgent scaling, budgetsSells AI advisory
S12KPMG US AI Quarterly Pulse Q2Jun 24, 2026US senior leaders, quarterly panelAgent deployment, cost visibilitySame
S13Anthropic Economic Index, CadencesJun 26, 2026Privacy-preserving telemetry + survey linked to usage, ~9,700 respondentsHow and why people use AIVendor-published; own product data
S14Anthropic Economic Index, Jan 2026 reportJan 2026Telemetry, Nov 2025 sampleAutomation vs augmentation splitSame
S15Anthropic Economic Index, geography2025Telemetry, globalAPI vs consumer usage patternsSame
S16Snyk Labs, ToxicSkillsFeb 5, 2026Scanned 3,984 skills from ClawHub + skills.shAgent-skill securitySnyk sells security scanning
S17VentureBeat, Agent Skills open standardDec 18, 2025Reporting + Anthropic PM interviewSkills standardizationTrade press
S18The New Stack, Agent Skills specDec 2025ReportingSkills timelineTrade press
S19GitHub Octoverse 20252025Platform telemetry, 180M devsDeveloper AI adoptionMicrosoft-owned; sells Copilot
S20Klarna/OpenAI announcement + Bloomberg follow-upFeb 2024 / May 2025Company self-report, then CEO reversalSupport automation outcomesBoth parties promotional at launch
S21Goldman Sachs Research2025-2026Macro modelingAI capex, productivitySell-side research
S22arXiv: Emerging Threats of the Agent Skill EcosystemMay 2026Threat taxonomy, real samplesSkills securityPreprint, not peer-reviewed
S23Agent Skills '26 / SkillsBench2026Benchmark across public skillsSkill quality effect on agent performanceAcademic workshop
S24Ecosystem trackers (Termdock, Agentman, OSS Insight, SpecWeave)2026Secondary aggregation, self-publishedSkill counts, catalog sizeLow reliability, unverified blogs. Treat as directional only.
// §1 · the adoption curve

Everyone adopted.
Few converted.

1.1 Headline trend

Metric2023202420252026 YTDSource
Orgs using AI (≥1 function)55%72-78%88%-S1, S2
Orgs scaling enterprise-wide-~33%~33%-S1
Report any EBIT impact--39%-S1
High performers (>5% EBIT)--6%-S1
Census BTOS (production use)-~4.6%10% (old def) / 17.3% (new def)17-20%S3
Ramp (paid AI adoption)--44.5% Aug → 46.6% Dec>50% Mar 2026S4
Enterprise GenAI spend$1.7B$11.5B$37B-S5
Enterprises scaling agents--23%11% "AI leaders"S1, S11

Read this carefully: the four adoption rows measure four different things. See §6.

1.2 Spend detail (all S5)

  • Total 2025: $37B, 3.2x YoY. Applications $19B, infrastructure $18B.
  • Applications = 6% of the entire software market within three years of ChatGPT's launch.
  • Application split: horizontal $8.4B, departmental $7.3B, vertical $3.5B.
  • Coding alone: $4.0B, up from $550M in 2024, the single largest category anywhere in the app layer.
  • Infrastructure split: foundation model APIs $12.5B, training infra $4.0B, data/orchestration $1.5B.
  • Excluded from these figures: chips, cloud inference (AWS/GCP/Azure), and AI features bolted into existing SaaS. So the real number is higher; this is the net-new AI market.
// enterprise genai spend · net-new market (s5)
$1.7B to $37B in two years

Menlo Ventures bottoms-up market model (S5). Excludes chips, cloud inference, and AI features inside existing SaaS, the net-new market only.

1.3 Macro context (S21)

  • AI capex ≈ 0.8% of US GDP, below the 1.5%+ seen in past tech booms.
  • Hyperscaler capex to exceed $500B in 2026.
  • Net US GDP impact only ~0.1-0.3pp in 2026, because much of the hardware is imported.
  • Goldman found no meaningful economy-wide AI/productivity relationship yet. Gains are localized to coders and support reps.

1.4 Build vs. buy: corrected

S5 reports: 47% built / 53% bought (2024) → 24% built / 76% bought (2025).

Do not read this as "enterprises stopped building." Four corrections:

  • Menlo's own text notes continued strong investment in internal builds; roughly a third of AI budgets still go to them.
  • The unit is use cases, not dollars or effort. Ten bought SaaS tools + three deep internal builds scores 77% "buy."
  • Buying Cursor or Supabase counts as "buy", then you build with it. The categories overlap.
  • The infra/app spend split is ~50/50 ($18B vs $19B), which is the closest proxy for buy-to-build vs buy-to-use. That undercuts the headline.

There is no company-size segmentation on build/buy in S5. If you need mid-market vs. large-enterprise, use S10 (Deloitte, segments by size and geography).

// §2 · completion and success rates

The failure numbers, and
the success numbers.

2.1 The failure numbers

FindingFigureSource
Integrated pilots extracting real value~5%S6
Pilots with no measurable P&L impact~95%S6
Companies past proof-of-concept to real value26%BCG, Oct 2024
"Future-built" firms capturing full value5%S9
Laggards with little/no value60%S9
Agentic AI projects to be cancelled by end-2027>40%S8
Orgs that have deployed AI agents17%S8 (2026 Hype Cycle)
Enterprises scaling agents enterprise-wide11%S11
Real-time visibility into AI running costs26%S12
Mature governance model for agentic AI21%S10

2.2 The success numbers

FindingFigureSource
AI deals converting to production47% (vs 25% SaaS)S5
Report AI delivering meaningful business outcomes64%S11
Report first-year ROI74%Google Cloud ROI report, vendor
Avg cost savings, early adopters15.2%S8
Avg productivity improvement22.6%S8
Pilots blending internal + external expertise67% successS6
IT-only internal builds22% successS6
Future-built firms' revenue growth multiple1.7xS9

2.3 Why projects fail

  • Not model quality. S6's diagnosis: tools can't retain feedback, adapt to context, or improve over time, and organizations bolt AI onto legacy processes instead of redesigning them.
  • S11's version of the same finding: the 11% who scale redesign the process first, then deploy agents into it. The 89% do the reverse.
  • S1's blockers: data quality, workflow rigidity, operating-model inertia, measurement gaps.
  • Budget misallocation (S6): over half of AI budgets went to sales and marketing, which produced low ROI, while back-office automation delivered the actual returns.
  • Shadow AI (S6, S5): 90%+ of firms have employees using personal AI accounts for work. S5 estimates PLG + shadow adoption is close to 40% of application AI spend.

2.4 The coding counter-evidence (S7)

The most rigorous study in the whole corpus, and it cuts against the hype.

METR ran a randomized controlled trial: 16 experienced open-source developers, 246 real tasks in codebases they already knew, mostly Cursor Pro with Claude 3.5/3.7 Sonnet.

  • Devs predicted a 24% speedup beforehand.
  • After finishing, they estimated AI had made them 20% faster.
  • Measured result: AI made them 19% slower.
// metr rct · perception vs measurement (s7)
Devs felt 20% faster. They measured 19% slower.

Randomized controlled trial: 16 experienced open-source developers, 246 real tasks in codebases they already knew, early-2025 tooling. The 39-point gap between felt and measured speed, in the wrong direction, is the finding.

Caveats that matter: small sample, mature codebases with high context load, early-2025 tooling. It does not generalize to juniors, greenfield work, or prototyping, where gains of 27-90% are reported elsewhere. METR's Feb 2026 follow-up was judged an unreliable signal.

What to take from it: the perception gap is the finding. People are bad at estimating their own AI-assisted productivity, in a consistent direction. Any client self-report of "we're 30% faster" is unverified until measured.

// §3 · use cases by business domain

Where the money goes,
and where it comes back.

DomainSpend / adoptionReported valueStatusSource
Software eng / coding$4.0B; 50% of devs daily (65% top-quartile); ~80% of new GitHub devs use Copilot in week 1; 90% of Fortune 10015%+ velocity claimed; METR RCT measured −19% for expertsProven at scale, claims inflatedS5, S19, S7
Customer service$630M departmental; largest agentic categoryIntercom Fin ~76% avg resolution across 12,000 customers; Anthropic internal 50.8% resolution, 1,700+ hrs saved; cost/resolution $0.10-0.99 vs $6-20 humanProven for tier-1; quality riskS5, Fin.ai, Anthropic
Sales & marketingMarketing $660M; 78% startup share in sales tools30% cut in external agency spend (one MIT case)Mixed, high spend, S6 flagged low ROIS5, S6
Finance & accounting91% startup share (AI-first ERPs)Back-office automation = highest ROI in S6 ($2-10M savings)Emerging, best ROI/effort ratioS5, S6
Legal$650M vertical marketContract review, triageEmergingS5
HR / recruiting5% of departmental spend-ExperimentalS5
Healthcare / life sci$1.5B (43% of vertical AI); ambient scribes $600M (+2.4x YoY)Scribes cut documentation time >50%Proven in documentationS5
IT operations$700MIncident response, infra managementEmergingS5
Supply chain / mfg-Multi-objective optimization (cost vs time-to-market)ExperimentalS10
Product / designDesign 7% of departmental spend-EmergingS5

3.1 The Klarna arc: the single most instructive case (S20)

Feb 2024 (company announcement): the AI assistant handled 2.3M conversations, about two-thirds of Klarna's support chats, described as the equivalent work of 700 full-time agents. Resolution time fell from 11 minutes to under 2. Projected profit improvement: $40M for 2024.

May 2025 (CEO to Bloomberg): Siemiatkowski said the cuts had gone too far. AI-only support produced lower quality. Klarna began rehiring humans into a hybrid model.

The lesson: resolution rate is not resolution quality. Both facts are true and both were reported by the same company 15 months apart. Any support-automation business case should carry a quality floor, not just a deflection target.

// §4 · the tooling and infrastructure layer

The layer that moved
faster than the value.

4.1 Model Context Protocol

Timeline

  • Nov 2024: Anthropic open-sources MCP.
  • Mar 2025: OpenAI adopts (ChatGPT, Agents SDK, Responses API).
  • 2025: Google DeepMind, Microsoft, AWS, Cloudflare follow.
  • Dec 2025: donated to the Linux Foundation's Agentic AI Foundation (co-founded by Anthropic, Block, OpenAI; backed by Google, Microsoft, AWS, Cloudflare, Bloomberg).

Scale (Anthropic, Dec 2025)

  • 10,000+ active public MCP servers.
  • 97M+ monthly SDK downloads, up from ~100K at launch, roughly 970x in 18 months.
  • Official registry lists 6,400+ servers.

Security: unresolved

  • OWASP frames the risk as tool poisoning, sitting between LLM01 (prompt injection) and LLM05 (supply-chain).
  • Two CVEs put it on the map: MCPoison (CVE-2025-54136), CurXecute (CVE-2025-54135).
  • Apr 2025: researchers demonstrated indirect prompt injection via emails, docs, and web pages.
  • Nov 2025: a WhatsApp MCP integration flaw allowed extraction of full message histories via poisoned tool descriptions.
  • The structural flaw: tool descriptions are reviewed once at connect time, but tool responses reach the model context at runtime with no equivalent check.

Enterprise footprint: Microsoft Security Blog (Feb 10, 2026), citing the Microsoft Data Security Index 2026: over 80% of Fortune 500 run active AI agents, but only 47% of those organizations have implemented specific security controls over them. (A frequently repeated "28% of Fortune 500 have implemented MCP servers" figure traces to vendor estimates with no named primary survey, do not cite it.)

4.2 Agent Skills: the newer layer

Timeline (S17, S18)

  • Oct 2025: Anthropic launches Agent Skills: folders containing instructions, scripts, and resources that teach an agent a repeatable workflow. A skill is a SKILL.md file with YAML frontmatter plus markdown body.
  • Dec 2025: released as an open standard at agentskills.io with a reference SDK.
  • Adopters named by Anthropic's PM: Microsoft (VS Code, GitHub), Cursor, Goose, Amp, OpenCode. OpenAI adopted a structurally identical architecture in ChatGPT and Codex CLI, same file naming, same metadata format, same directory layout.
  • Feb 2026: enterprise controls: org-wide provisioning on Team/Enterprise plans, plus stock plug-ins for finance, legal, HR.

Why it spread so fast: it answers a specific question cheaply, how do you make an assistant consistently good at specialized work without fine-tuning a model. The barrier to publish is a markdown file and a GitHub account.

Does it work? (S23: SkillsBench)

  • Curated skills raise agent pass rates by +16.2 points on average.
  • 2-3 focused skills deliver +18.6 points; monolithic "everything in one doc" skills reduce performance by 2.9 points.
  • Self-generated skills hurt performance on nearly a third of tasks.
  • Average quality score across public skills: 6.2 out of 12. Benchmarks used only top-quartile skills (9+).

Security: worse than MCP (S16, S22)

  • Snyk scanned 3,984 skills across ClawHub and skills.sh (Feb 5, 2026): 1,467 (36.8%) had at least one security flaw; 534 (13.4%) critical; 76 confirmed malicious payloads: credential theft, reverse shells, data exfiltration.
  • Daily publishing rate went from under 50 in mid-Jan 2026 to over 500 by early Feb, 10x in weeks. Vetting capacity did not scale with it.
  • Five named threat actors operated across multiple platforms. ClawHub was subsequently shut down.
  • OWASP published an Agentic Skills Top 10 on Apr 27, 2026. AST01 is Malicious Skills, rated Critical.
  • The mechanism: skills are not sandboxed plugins. They execute with the host agent's full privileges, filesystem, terminal, network, credentials.
36.8%
of 3,984 scanned public skills had at least one security flaw (Snyk ToxicSkills, Feb 2026)
13.4%
had a critical flaw, 534 skills across the two catalogs scanned
76
confirmed malicious payloads: credential theft, reverse shells, data exfiltration

Practical rule: treat a community skill exactly like an unaudited npm package with code-execution rights. Read the source. Curated internal libraries beat public catalogs on both quality and safety.

4.3 Agent architecture reality (S5)

  • Only 16% of enterprise and 27% of startup deployments qualify as true agents, where the model plans, executes, observes, and adapts.
  • The rest are fixed-sequence or routing workflows around a single model call. S5's phrasing: basic if-then logic around a model call.
  • Customization techniques by frequency: prompt design first, then RAG. Fine-tuning, tool calling, context engineering, and RL remain niche.
  • Copilots dwarf agents in spend: $7.2B (86%) vs $750M (10%).

4.4 LLM vendor share, enterprise API usage (S5)

Vendor202320242025
Anthropic12%24%40%
OpenAI50%~34%27%
Google7%-21%
Open-weight (Llama et al.)-19%11%
// enterprise llm api usage share · 2023 → 2025 (s5)
Anthropic and OpenAI traded places in two years
2023 2024 2025

Estimated dollars based on self-reported production API usage share (S5), see the methodology caveat below. Missing bars are years the source published no figure.

  • Top three = 88% of enterprise LLM API usage.
  • Coding specifically: Anthropic ~54% vs OpenAI 21%, up from 42% six months earlier, driven by Claude Code.
  • Chinese models: ~1% of enterprise API usage, but rising fast among startups and indie devs via OpenRouter and vLLM (Qwen, DeepSeek, GLM). Airbnb uses Qwen for user-facing features; Cursor used it as the base for an internal model.
  • Methodology caveat: these are estimated dollars based on self-reported production API usage share, weighted by application scale and triangulated with public financials. Not audited revenue. And Menlo is an Anthropic investor.

4.5 High-growth infra tools

CompanyARRValuationNoteSource
Supabase~$170M (May 2026), from ~$101M end-2025, $30M end-2024$10.5B post (Jun 2026 Series F, GIC)Backend of the vibe-coding boom, Lovable and Bolt run on it; 4M+ devsSacra, press
Cursor / Anysphere~$2B (Feb 2026) → ~$4B (mid-2026); $100M Jan 2025 → $1B Nov 2025$29.3B (Nov 2025); SpaceX agreed to acquire ~$60B all-stock (Jun 2026)Fastest B2B software ramp on record; 70% of Fortune 1,000Press
Lovable$100M (mid-2025) → $400M+ (Feb 2026)$6.6B (Dec 2025, CapitalG/Menlo)A reported $12B round was in talks, not closedSacra, press
Replit~$10M → $100M in ~6 months (2025); ~$250M Oct 2025$9B (Mar 2026 Series D)Press
Vercel$200M+ (mid-2025)$9.3B post (Sep 2025 Series F)Press
n8n~$40M (Jul 2025)$2.5B (Oct 2025 Series C, Accel; Nvidia participating)Built on open-source community adoption before enterprise salesPress
Zapier~$420M (Q1 2026)$5B (2021 secondary, stale)Zapier Agents GA May 2025; Zapier MCP; AI-task volume +760%Sacra
Databricks$5.4B run-rate (+65% YoY); AI products $1.4B annualized$134B (Series L, Feb 2026)A reported $165B+ round was in discussionsCompany
SnowflakePublic-9,100+ customers use AI products weekly; $100M AI run-rate hit ahead of planCompany
Pinecone-$750M (2023, stale)Reportedly weighing a salePress
Hugging Face-$4.5B (2023, stale)Press
Weaviate-$50M Series B (2023)Later Series C unconfirmedPress

Signal in the stale rows: vector-DB valuations are 2+ years old with no fresh priced round. Incumbents added vector search and squeezed the category. S5 confirms the direction, incumbents hold 56% of AI infrastructure spend, because even AI-native app builders keep choosing Databricks, Snowflake, MongoDB, and Datadog.

// §5 · agents and skills: what humans actually want from them

The section on
purpose.

This is the section on purpose. Most reports measure spend and adoption. Very few measure why. The best data here comes from Anthropic's Economic Index (S13, S14, S15), vendor-published, but it is telemetry linked to a survey, not a self-report questionnaire, which makes it structurally harder to game than anything in §1.

5.1 The core frame: automation vs augmentation

Anthropic classifies every conversation into five interaction modes, grouped into two families (S13, S14):

Automation, the human delegates

  • Directive: hand over a complete task, minimal back-and-forth. "Translate this document."
  • Feedback loop: the human relays real-world outcomes back to the model. "Make this email more casual."

Augmentation, the human collaborates

  • Task iteration: work through it together, human refining outputs.
  • Learning: ask for explanation or understanding, not task completion.
  • Validation: ask the model to check your own work.

5.2 The purpose trend, tracked over 18 months

PeriodAugmentationAutomationDirective shareSource
Jan 202556%41%27%S14
Aug 2025-overtook augmentation39%S14
Nov 202552%45%32%S14
Feb 2026slightly up--Mar 2026 report

Read the shape, not the points. Directive use climbed hard through Aug 2025, then fell back 7pp. Anthropic's own reading: the August spike overstated how fast delegation was arriving, but the underlying direction is still toward automation. The Nov pullback is attributed partly to product changes, file creation, memory, and Skills: that encourage more collaborative, iterative work.

That's a notable finding: adding skills made usage less fully-delegated, not more. Skills gave people a reason to stay in the loop.

5.3 Enterprise purpose looks completely different from consumer purpose

This is the sharpest split in the data (S15):

  • API / enterprise traffic: 77% automation patterns, mostly directive. Only 12% augmentation.
  • Claude.ai / consumer: roughly an even split.
// automation share of traffic · api vs consumer (s15)
Businesses hand work over. Consumers think alongside.

Anthropic Economic Index telemetry (S15). The concrete enterprise patterns (S14): email classification, invoice processing, calendar scheduling, back-office throughput, exactly where S6 found the ROI actually was.

Businesses do not buy AI to think alongside. They buy it to hand work over. S14 names the concrete enterprise patterns: email classification, invoice processing, calendar scheduling, back-office throughput, exactly where S6 found the ROI actually was.

5.4 What they're producing (S13, Apr-Jun 2026 sample)

93% of conversations produce an identifiable artifact. Top categories:

ArtifactShare of all conversationsWork-related share
Explanations17%-
Documents & reports15%-
Guidance11%80%+ personal
Marketing content-80% work
Blogs / articles-81% work
Database queries-82% work

Flipped by purpose: work conversations most often produce documents and reports (20%), then explanations (9%), email drafts (7%), analyses and summaries (6%).

5.5 The autonomy gradient: and why the product matters more than the model

S13's most useful finding for anyone building agent workflows.

AI autonomy is rated 1-5. Across 26 of 31 output types, autonomy is higher in Claude Code than in chat. The average gap is 0.37 points.

The concrete example: producing a blog post. The median chat session takes 13 rounds of back-and-forth. The median Claude Code session doing the same job contains a single human prompt.

13 → 1
median human turns to produce a blog post: 13 rounds in chat, a single prompt in Claude Code, same job
26/31
output types where autonomy is higher in Claude Code than in chat; the average gap is 0.37 points
+0.26
the autonomy gap that remains comparing only Sonnet sessions, the surface, not the model, drives it

Two-thirds of the gap is the same tasks being executed with more delegation. One-third is a different mix of work.

And it isn't the model. Claude Code runs Opus far more often (54% vs 10% in chat). But comparing only Sonnet sessions, Claude Code still shows 0.26 points more autonomy. S13's conclusion: the product surface matters more than the underlying model.

Implication for engagements: how you package an agent determines how much people delegate to it, more than which model you pick. Same model, different surface, 13 turns vs 1.

5.6 More valuable work costs more compute (S13)

  • Token consumption rises with the wage of the occupation the task maps to. Marketing managers earn ~2x what editors do; their conversations use ~2.5x the tokens.
  • Building apps uses 3x the median conversation's tokens. A typical explanation uses about a fifth.
  • Autonomy and token use rise together (r = 0.68).
  • But the human doesn't disappear at the top end. In higher-wage work, Claude produces 1.34x more per turn and users engage 1.53x more turns, with extended thinking on more often. S13's read: these move together, which looks labor-augmenting rather than labor-displacing.

5.7 How people get into using them: the on-ramp

Four documented pathways, in order of how much of the market they explain:

1. Product-led, bottom-up (S5). 27% of AI application spend arrives through PLG, nearly 4x the 7% rate in traditional software. Counting shadow AI on personal cards, close to 40%. Cursor reached $200M revenue before hiring a single enterprise sales rep. n8n formalized contracts only after hundreds of employees were already using it. Lovable, OpenRouter, ElevenLabs, Gamma, Wispr Flow followed the same pattern.

What this means in practice: the tool is usually already inside the company before anyone signs anything. Discovery should start with "what are people already using on personal accounts," not "what should we buy."

2. Delegation as learning-by-doing (S13). People who delegate more report AI can do more of their work, and expect it to do more next year. Anthropic offers two readings: either delegation teaches you what AI can actually finish, or people who already believe it can do their job are the ones willing to hand it over. Both are plausible; the data can't separate them. Either way, usage drives belief, not the reverse.

3. Skills as the specialization step. Once a team is using an agent, skills are how they encode house rules without touching a model. S23 quantifies the payoff: +16.2 points pass rate from curated skills, +18.6 from 2-3 focused ones. Also the trap: self-generated skills hurt performance on nearly a third of tasks. So the on-ramp is curate, not generate.

4. Enterprise governance retrofit (S17, S10). Central provisioning of skills arrived Feb 2026, after the bottom-up adoption. Governance is being fitted to existing usage, not preceding it. Only 21% of firms have a mature agentic governance model (S10).

5.8 What people say they want: the direct answer to "what is the human purpose"

S13 asked ~9,700 linked respondents an open-ended question: what do you hope an AI-shaped economy looks like in ten years? Classified themes, top three:

RankThemeShare
1Augmentation: collaborating on work that feels meaningful, careers still mattering, new industries and jobsOver half
2Automation of drudgery: offloading tedious work for more free time and meaning outside workJust over half
3Shared prosperity: that the economic gains are widely distributed~1/3

Not replacement. Not headcount reduction. The top two hopes are keep the meaningful part, remove the boring part.

Supporting evidence, same survey:

  • Productivity gains reported: speed 86%, scope 82%, quality 69%. 27% report cost savings on services they'd otherwise buy.
  • 68% report learning more with AI; 57% say AI made their skills more valuable.
  • Over a third expect AI to handle most or nearly all of their work tasks within 12 months.
  • Only 10% rate losing their own job as likely, slightly below the US annualized involuntary-separation rate (~13.4%). But they're far more worried for others: over a third put a junior colleague's job-loss probability above 60%.

5.9 The counterintuitive finding worth carrying into the next conversation

The heaviest delegators are the most optimistic, on all six dimensions measured (pay, job security, ability to find a new job, meaning, autonomy, human interaction). Largest effects on expected pay and job-finding ability.

And the common fear, that delegating means offloading thinking and eroding skill, does not show up in the data. Heavy delegators report learning at the same rate as everyone else, and are more likely to say their skills grew in market value.

Two honest caveats (S13's own):

  • Selection can't be ruled out, enthusiasts may both delegate more and feel better. Though the effect survives controlling for account tenure.
  • These are self-assessments. Skills can erode even while someone reports learning more. The data does not disprove skill erosion; it just doesn't find it.

5.10 Who these findings do not describe

S13 is explicit that the survey is not representative:

  • Computer & mathematical occupations: ~30% of respondents vs 4% of US employment.
  • Management: 23% of respondents vs 7% of employment.
  • Transportation, food service, construction: heavily under-represented.
  • Women are 12% of the linked sample, and use Claude measurably differently: 6.3pp lower Claude Code share, 7.3pp lower automation share, more iterative use, more active minutes in chat, even after controlling for occupation.

So §5 describes technical and managerial knowledge workers who already use AI heavily. Do not generalize it to a whole workforce.

Staring at an AI statistic you are not sure you trust? That is the conversation we like having.

Book a call →
// §6 · why the investigations disagree

Four true numbers,
one reality.

The single most useful section for judging any AI statistic you're handed.

6.1 The adoption number is four different numbers

SourceFigureWhat it actually asksWhy it's high or low
McKinsey (S1)88%"Does your org use AI in at least one function?"Highest. One person in one department counts. Respondents are AI-interested leaders. Firm sells AI transformation.
Stanford (S2)78-88%Aggregates survey dataInherits survey bias, but neutral analysis
Ramp (S4)>50%Did the company pay an AI vendor?Middle. Hard transaction data, but misses in-house builds and open-source, and skews to Ramp's customer base (US, tech-forward)
Census (S3)17-20%"Did this firm use AI in producing goods or services?"Lowest. Nationally representative across ~1.2M firms including small and non-tech. Best methodology, narrowest question.
// "what share of companies use ai?" · four answers, 2025-26
Same question, four methodologies, a 4x spread

All four are true; they answer different questions. Self-reported adoption runs 2-4x higher than transaction- or production-based measures.

The reconciliation: nearly every large, tech-forward company has touched AI. Roughly half pay a vendor. Fewer than one in five have it in actual production of goods and services. All four are true. They're answering different questions.

Rule of thumb: self-reported adoption runs 2-4x higher than transaction or production-based measures.

6.2 MIT's 95% vs Menlo's 47%: a direct collision

MIT NANDA (S6)Menlo (S5)
Headline~95% of pilots show no P&L impact47% of AI deals reach production (vs 25% SaaS)
Unit measuredValue delivered to the income statementProcurement conversion to deployment
Sample300+ initiatives, 52 interviews, 153 surveys495 US decision-makers
MethodQualitative + survey, ~6 monthsSurvey + bottoms-up market model
Published byAcademic, preliminary, not peer-reviewedVC with portfolio stakes in the market measured
Direction of biasToward pessimism (interview-led, failure-salient)Toward optimism (portfolio value, explicit rebuttal framing)

These are not actually contradictory. A project can reach production (Menlo's bar) and still move no P&L line (MIT's bar). Menlo measures deployment. MIT measures value. Both can be right simultaneously, and probably are.

Menlo's Dec 2025 report explicitly frames itself against MIT's finding, which tells you the framing was chosen, not discovered.

Methodology critiques of MIT worth knowing: the sample is small and preliminary; the authors describe it as directionally accurate rather than definitive; critics have called it methodologically fragile. Use it as a directional truth about the pilot-to-value gap, not as a precise failure rate.

6.3 Agent deployment: 11%, 16%, 17%, 42%, 53%

Five credible sources, five very different numbers, all within twelve months.

FigureSourceDateWhat it counts
11%KPMG Global (S11)Mar 2026Scaling agents enterprise-wide with business outcomes ("AI leaders")
16%Menlo (S5)Dec 2025Deployments that are architecturally true agents (plan-execute-observe-adapt)
17%Gartner (S8)2026Organizations that have deployed agents at all
42%KPMG USSep 2025Orgs that have deployed at least some agents
53%KPMG US (S12)Jun 2026Orgs deploying agents, quarterly panel
// "have you deployed ai agents?" · five answers in twelve months
The spread is the definition, not the reality

One measure (share of organizations), five definitions of "agent." Menlo's 16% is the strictest technical test; KPMG's 53% is the loosest; KPMG's own 11% is the strictest business test.

The spread is the definition of "agent," not disagreement about reality. Menlo's 16% is the strictest, it's a technical architecture test. KPMG's 53% is the loosest, any deployment counts. KPMG's own 11% "AI leaders" figure is the strictest business test.

Practical translation: roughly half of large firms have something they call an agent. About one in six is technically an agent. About one in nine gets enterprise-wide value from it.

Gartner also notes widespread "agent washing": of thousands of self-described agentic AI vendors, they estimate only about 130 are real.

6.4 Skills ecosystem counts: 3,984 vs 22,511 vs 47,150 vs 490,000

Four numbers circulating for "how many skills exist." All from 2026. They differ by two orders of magnitude.

CountSourceWhat it isReliability
3,984Snyk ToxicSkills (S16)Skills scanned from ClawHub + skills.sh, Feb 5, 2026High: primary security research, defined corpus
22,511Secondary aggregation (S24)Skills in a broader security auditMedium, audit not independently verified
47,150SkillsBench (S23)Public skills analyzed for qualityMedium-high, academic benchmark
490,000+Ecosystem blogs (S24)Claimed total across three marketplaces, Mar 2026Low: self-published, no methodology

Why the gap: these are corpora, not censuses. Snyk scanned what it could scan. SkillsBench analyzed what it could benchmark. The 490K figure counts everything ever published including abandoned duplicates. OSS Insight notes the long tail is vast and largely unused, thousands published, nobody installs them.

Use the 3,984 and 47,150 figures. Don't cite 490,000.

6.5 Survey vs measurement: the widest gap in the whole field

DomainSelf-reportedMeasuredGap
Dev productivity+20% (devs' own post-hoc estimate, S7)−19% (RCT, S7)39 points, wrong direction
Org adoption88% (S1)17-20% (S3)4x
AI project ROI74% report first-year ROI (Google Cloud)39% report any EBIT impact (S1); ~5% real value (S6)2-15x
Support automation2/3 of chats automated (Klarna, Feb 2024)Quality dropped, humans rehired (Klarna, May 2025)Reversed within 15 months

The pattern is consistent and directional: people overestimate AI's benefit to their own work, and organizations overestimate their own AI maturity. This isn't dishonesty, METR showed devs were wrong about their own measured performance in real time.

6.6 A checklist for the next AI statistic you're handed

01

Who fielded it and what do they sell?

Consultancies sell transformation. VCs hold portfolios. Vendors sell tools. All three publish real data with chosen framing.

02

Survey or measurement?

Self-report inflates 2-4x. Transaction, telemetry, and RCT data don't.

03

What's the unit?

Use cases, dollars, deployments, and organizations give wildly different answers to "how much AI is there."

04

What's the definition?

"Agent" swings a number from 11% to 53% with no change in underlying reality.

05

What's the sample frame?

495 US decision-makers ≠ 1.2M US firms ≠ 9,700 heavy Claude users.

06

Is the effect measured against a baseline?

Almost never. This is why the METR result was surprising.

// §7 · what this means if you're buying or building

Eight rules the data
actually supports.

01

Define the P&L metric and its baseline before building anything.

The pilot-to-value gap (§2) is the market's central failure. If you can't name the number that moves, run discovery, don't build. Measure the baseline, or you'll never distinguish the METR effect from real gain.

02

Start where value is documented: internal tooling/coding, tier-1 support deflection, back-office automation.

Avoid sales and marketing content as a first project, S6 found it absorbed over half of budgets and returned the least.

03

Consider co-building the third option, not a compromise.

S6's split is the strongest datapoint here: 67% success for internal+external blended teams vs 22% for IT-only builds. The build/buy binary in §1.4 erases exactly that model.

04

Redesign the process, then deploy the agent.

S11's diagnosis of the 89% is that they lay AI over existing workflows. BCG's 10/20/70 rule says the same: 10% algorithms, 20% tech and data, 70% people and process.

05

Package for delegation deliberately.

§5.5 is the actionable finding: same model, different surface, 13 turns vs 1. Decide upfront whether a workflow should be delegated or collaborative, and build the surface to match. Don't leave it to chance.

06

Curate a skills library; never install from public catalogs.

36.8% flaw rate, 13.4% critical, code execution with full host privileges (§4.2). But curated skills are worth +16.2 points on agent pass rates. The value is real and so is the risk, which makes curation work worth budgeting for, not overhead.

07

Treat MCP as the default integration layer and assume it's insecure.

Build for MCP compatibility. Put a gateway in front that validates tool schemas before they reach the model. Never connect an untrusted server to write access or sensitive data.

08

Start discovery with shadow AI, not procurement.

27-40% of AI app spend enters bottom-up (§5.7). Ask what people already run on personal accounts. That's the real adoption baseline and the fastest path to a working use case.

Thresholds that flip the recommendation

SignalAction
Measured gain <10% after 90 daysWrong tool for that workflow, stop, don't tune
Resolution rate up, CSAT downCap automation, go hybrid (the Klarna line)
Company under ~$20M ARRAlmost never build custom; off-the-shelf pays back in 3-9 months
No named success metricDiscovery engagement, not a build
Agent handles a decision with no reversal pathDeterministic workflow instead
Skill sourced from a public catalogRead the source or don't ship it
// §8 · caveats

What this report
can't claim.

  • Surveys and hard data disagree by 2-4x and I've kept both rather than averaging them. §6.1 explains why. Any single adoption number in this document is incomplete without its method.
  • Vendor-incentive flags are in §0 and repeated inline. Menlo (VC, Anthropic investor), Google Cloud, Anthropic, GitHub, Ramp, Deloitte, BCG, McKinsey, KPMG, Snyk all sell into this market. Least conflicted: US Census, Stanford HAI, METR. MIT NANDA is unconflicted but preliminary.
  • MIT's 95% is preliminary and contested. Small qualitative sample, not peer-reviewed, described by its own authors as directional. Use it for the shape of the challenge, not as a failure rate.
  • §5 rests heavily on Anthropic's own telemetry. It is the best data available on why people use agents, and telemetry beats self-report, but it covers Claude users only, skews technical and managerial, and Anthropic has an interest in the augmentation framing landing well. The automation/augmentation classification is Anthropic's own taxonomy, not an industry standard.
  • Forward-looking figures are labelled, not asserted. Lovable's ~$12B and Databricks' ~$165B+ rounds were in talks, not closed. Gartner's 40% cancellation rate is a forecast.
  • Private ARR figures are estimates (Sacra, press) unless company-disclosed. Pinecone, Hugging Face, and Zapier valuations are 2+ years stale.
  • Skill counts above 50,000 come from self-published ecosystem blogs with no methodology. Flagged S24 throughout. Don't put them in a deck.
  • Everything here is US- and English-weighted. Menlo's sample is US-only. Census is US-only. Stanford noted China and Europe posted the highest YoY adoption increases, but the granular data behind that is thinner than the US picture.

Deciding what to build
from numbers like these?

If you are trying to work out where AI would actually move a number in your business, and which statistics to ignore on the way, that is the conversation we like having.

Book a call → See what we build