Agent Mode AI

Four datasets converge on the same bimodal shape across enterprise agentic AI. Stanford Digital Economy Lab's twelve-eighty-eight split. McKinsey State of AI 2025's twenty-three percent scaling cohort. MIT NANDA's ninety-five percent pilot-failure finding

Show Notes

CORRECTION (10 Jun 2026): The 12/88 statistic central to this episode was retracted on 10 Jun 2026; it does not appear in the cited Stanford source. The audio remains available unchanged as the record. Full correction and restated article: https://agentmodeai.com/why-88-percent-of-agentic-ai-deployments-fail/ Ledger record (AM-029, Not holding): https://agentmodeai.com/holding/?claim=AM-029 Original show notes follow. Episode 7 of Agent Mode AI. Abby and Avery walk the four datasets that document the bimodal ROI distribution in enterprise agentic AI: AM-029 on the Stanford Digital Economy Lab 12/88, AM-132 on the restored bimodal framing, AM-128 on the MIT NANDA GenAI Divide 95% finding, and AM-053 on the McKinsey State of AI 2025 17% EBIT-attribution. Then the GAUGE framework — six dimensions that distinguish the 12% high-performing cohort from the 88% struggling body. Governance, audit substrate, use-case maturity, guardrails, evidence baseline, exit posture. Sources cited: - Stanford Digital Economy Lab Enterprise AI Playbook 2026 (Pereira, Graylin, Brynjolfsson) - McKinsey State of AI 2025 (n=1,993, November 2025) - MIT NANDA State of AI in Business 2025 (Project NANDA / MIT Connection Science, August 2025) - Gartner Q1 2026 Infrastructure & Operations Survey - Fortune coverage of MIT NANDA findings, August 2025 Claims tracked: - AM-029 — Why 88% of agentic AI deployments fail — agentmodeai.com/holding/?claim=AM-029 - AM-132 — The bimodal ROI distribution in enterprise agentic AI — agentmodeai.com/holding/?claim=AM-132 - AM-128 — The MIT 95% GenAI-pilot-failure claim — agentmodeai.com/holding/?claim=AM-128 - AM-053 — The McKinsey 17% EBIT claim — agentmodeai.com/holding/?claim=AM-053 Newsletter and the full Holding-up ledger: agentmodeai.com

What is Agent Mode AI?

The audio companion to agentmodeai.com. Two analysts pick one claim from the Holding-up ledger per episode, walk the evidence, and give the current verdict: Holding, Partial, or Not holding. For CIOs, IT directors, and senior implementers. 15-20 min, every Sunday.

Agent Mode AI — Episode 7
Why 88% of agentic AI deployments fail
Duration: 12:45
Hosts: Abby and Avery
Published: 21 June 2026
Anchor claims: AM-029, AM-132, AM-128, AM-053
Summary
Four independent enterprise AI datasets converge on the same bimodal shape. Stanford Digital Economy Lab's 2026 Enterprise AI Playbook reports a twelve-percent ROI cohort and an eighty-eight-percent at-or-below-break-even cohort. McKinsey State of AI 2025 reports twenty-three percent scaling agentic AI and seventeen percent EBIT-attribution at twelve months. MIT NANDA's GenAI Divide reports ninety-five percent of analysed pilots delivered no measurable P&L impact, with a sixty-seven-percent buy success rate against a roughly twenty-two-percent build success rate. Abby and Avery walk why the four cuts share a single underlying pattern, why the variable separating cohorts is operational discipline rather than model selection, and the six dimensions the publication's GAUGE diagnostic instruments against. Three concrete procurement actions for any 2026 enterprise running active agentic deployments, anchored on AM-029, AM-132, AM-128, and AM-053. The episode's procurement-relevant signal: vendor case studies typically describe the high-performing cohort. Real deployments land in the bimodal distribution the four datasets document. Score every active deployment on GAUGE before the next review cycle.
Chapters
[00:00] Cold open
[00:30] The four numbers, same shape
[01:45] AM-128 — the MIT 95% finding
[03:00] Absence of measurement is not presence of failure
[04:15] The build-versus-buy 67-22 spread
[05:00] AM-029 — the Stanford bimodal 12/88
[06:15] The Stanford 12% and MIT 5% cohorts overlap
[07:00] AM-132 — the restored bimodal piece
[07:45] The six dimensions of the high-performing cohort
[08:00] Governance maturity
[08:45] Threat model and the agent red-team
[09:30] ROI evidence and the pre-deployment baseline
[10:15] Change management and workflow redesign
[11:00] Vendor lock-in posture as a procurement dimension
[11:45] Compliance posture and the audit substrate
[12:30] AM-053 — the McKinsey 17%
[13:15] The procurement implication: three concrete actions
[14:00] Verdicts and cadence on each of the four claims
[14:30] Outtro
Transcript
[00:00] Cold open
ABBY: This is Agent Mode AI. I'm Abby. Today we're walking four claims that converge on the same finding from four independent datasets. AM-029, AM-132, AM-128, and AM-053. Stanford. McKinsey. MIT NANDA. Plus the publication's own readout. The 88% in the title is the procurement-relevant cut. The variable that separates the twelve percent that delivers from the eighty-eight percent that does not is operational discipline, not model selection.
AVERY: I'm Avery. Frame the four numbers.
[00:30] The four numbers, same shape
ABBY: Stanford Digital Economy Lab's 2026 Enterprise AI Playbook tracks fifty-one enterprise agentic AI deployments at twelve to eighteen months post-production-deployment. Twelve percent clear three-hundred-percent-plus ROI. Eighty-eight percent operate at or below break-even. The distribution is not a Gaussian with a long tail. It is two distinguishable peaks separated by a discontinuity. Gartner Q1 2026 Infrastructure and Operations data reports twenty-eight percent of AI projects are fully paying off. McKinsey State of AI 2025 with one thousand nine hundred and ninety-three respondents reports twenty-three percent scaling agentic AI and seventeen percent EBIT-attribution at the twelve-month horizon. MIT NANDA's GenAI Divide reports ninety-five percent of analysed pilots delivered no measurable P&L impact.
AVERY: Four numbers. Same shape.
ABBY: Same shape. Different cuts of the same underlying pattern. The bimodal distribution recurs across four independent datasets with different methodologies and different sample populations. That recurrence is the procurement-relevant signal. The variable separating the two cohorts is consistent across the four cuts.
AVERY: Start with AM-128. The MIT 95%.
[01:45] AM-128 — the MIT 95% finding
ABBY: AM-128 is the claim that the MIT 95% statistic is ninety-five percent of GenAI pilots delivered no measurable P&L impact, based on roughly one hundred and fifty senior executives interviewed, three hundred and fifty employees surveyed, and three hundred publicly disclosed GenAI deployments analysed in August 2025. The 95% finding is the share of pilots that produced no measurable P&L impact. The slippage in 2026 procurement decks is between no measurable P&L impact and the pilot failed.
AVERY: Two different things.
[03:00] Absence of measurement is not presence of failure
ABBY: Two different things. First, absence of measurement is not presence of failure. Most enterprise GenAI pilots in 2024-2025 did not have a documented pre-deployment P&L baseline. Without a baseline, no measurable P&L impact is the default finding regardless of whether the pilot moved the operational needle. Second, project-versus-deployment ambiguity. An enterprise running twenty projects with one delivering measurable impact would classify as five-percent-success on a project-weighted view and one-hundred-percent-success on an any-project-produced-value view. The 95%-fail framing implies the former; procurement questions usually want the latter.
AVERY: What's actually useful in the report.
[04:15] The build-versus-buy 67-22 spread
ABBY: The build-versus-buy spread. Sixty-seven percent buy success versus roughly twenty-two percent build success. Per the report and the Fortune coverage, purchasing AI tools succeeded sixty-seven percent of the time, while internal builds panned out only one-third as often. A 2026 enterprise budgeting an internal-build approach is, on the report's data, accepting a three-times-worse outcome distribution than the buy approach. That is the most actionable finding in the report and the one most procurement teams ignore in favour of the headline.
AVERY: Move to AM-029. The Stanford piece.
[05:00] AM-029 — the Stanford bimodal 12/88
ABBY: AM-029 walks the Stanford Digital Economy Lab's bimodal twelve-eighty-eight finding. Twelve percent of deployments clear three-hundred-percent-plus ROI at twelve to eighteen months. Eighty-eight percent operate at or below break-even. The Stanford 88% and the MIT 95% are not the same number. Stanford measured deployments with documented baselines that had reached twelve to eighteen months in production. MIT measured pilots with no required baseline at any maturity stage. But both numbers point at the same operational reality. Most enterprise GenAI work in 2024-2025 was not yet producing measurable enterprise-level value.
AVERY: The 12% cohort and the 5% MIT cohort.
[06:15] The Stanford 12% and MIT 5% cohorts overlap
ABBY: The Stanford twelve-percent bimodal-success cohort and the MIT five-percent integrated-systems-created-significant-value cohort are similar in shape and probably overlapping in identity. The thesis the publication tracks is that the operational discipline that produces the twelve percent is what produces the five percent in the MIT cut. The discipline is observable. It is not random. It is reproducible enough to instrument against.
AVERY: AM-132. The bimodal piece.
[07:00] AM-132 — the restored bimodal piece
ABBY: AM-132 is the restored URL, was AM-014 status down on the Holding-up ledger after the original WordPress-era body used composite case studies that did not survive editorial scrutiny. The new body anchors on the four datasets we're walking and reframes the seventy-three-twenty-seven framing the slug carries as a rounded aggregation of the four cuts rather than a precise statistical claim. The bimodal shape is the load-bearing finding. The exact percentage points vary by methodology.
AVERY: What the high-performing cohort does that the struggling body does not.
[07:45] The six dimensions of the high-performing cohort
ABBY: Six dimensions. The publication tracks them under the GAUGE framework. First, governance maturity. The cohort has a named accountable owner for the deployment, a documented decision authority for tool-use changes, and an escalation path that is exercised at least quarterly. Deployments without a named owner default into the struggling cohort regardless of other strengths.
AVERY: Two.
[08:00] Governance maturity (continued) and threat model
[08:45] Threat model and the agent red-team
ABBY: Threat model. The cohort treats the agent's tool graph as a security surface and runs an explicit red-team cycle against it. The struggling cohort has typically run a generalised pen-test that does not exercise any of the four agent-specific surfaces and passed. The OWASP Agentic AI Top 10 names the threats; the agent red-team is the discipline that tests whether the defences hold.
AVERY: Three.
[09:30] ROI evidence and the pre-deployment baseline
ABBY: ROI evidence. The cohort has a documented pre-deployment baseline before pilot day one. MIT NANDA's central finding is dominated by pilots that did not establish baselines. A deployment without a baseline does not produce a number to commit to regardless of how well the agent actually performs. The realistic ninety-day deliverable for a disciplined mid-market deployment is a working pilot pattern that scales into twelve-to-eighteen-month measurable ROI, not the three-hundred-percent-ROI vendor pitch.
AVERY: Four.
[10:15] Change management and workflow redesign
ABBY: Change management. The cohort assumes the deployment changes the surrounding workflow rather than slotting into it. MIT NANDA's startup advantage finding is the diagnostic. Startups deploy AI into workflows still being designed. Enterprises deploy AI into workflows whose process structure was designed for non-AI tools. Enterprise procurement teams that scope a deployment without budgeting for workflow redesign are budgeting for the struggling cohort outcome.
AVERY: Five.
[11:00] Vendor lock-in posture as a procurement dimension
ABBY: Vendor lock-in posture. The cohort treats lock-in as an explicit procurement dimension, exit data portability, kill-switch operability, sub-processor expansion rights, model-deprecation rights, rather than as something to be discovered at renewal. The 60-question RFP at AM-026 operationalises the dimension as one of the GAUGE axes. The struggling cohort typically signs the vendor's MSA with light edits and discovers the lock-in surface at month eighteen.
AVERY: Six.
[11:45] Compliance posture and the audit substrate
ABBY: Compliance posture. The cohort runs the deployment against the regulatory regime that actually applies, EU AI Act Article 6, 11, 12, 16 for high-risk deployments, 21 CFR Part 11 plus GxP plus Annex 11 for pharma, HIPAA plus state law for healthcare, and treats the audit substrate as a load-bearing part of the deployment architecture rather than a documentation afterthought.
AVERY: The cohort that scores well across the six is the cohort that delivers.
ABBY: The cohort that scores well across the six is the cohort that delivers on the business case. The cohort that scores poorly is the cohort that produces the eighty-eight percent failure rate the slug names. The reproducibility of the gap across four independent datasets is what makes the framework actionable.
AVERY: AM-053. The McKinsey 17%.
[12:30] AM-053 — the McKinsey 17%
ABBY: AM-053 walks the McKinsey State of AI 2025 finding that seventeen percent of organisations report measurable EBIT impact attributable to AI at the twelve-month horizon. The slippage is between self-reported attribution and audited attribution. McKinsey's number is a self-reported success that gets read as audited success. MIT's 95% is a self-reported absence of measurement that gets read as audited failure. Both readings are wrong in the same way. They conflate survey results with operational measurements.
AVERY: The procurement implication of all four together.
[13:15] The procurement implication: three concrete actions
ABBY: Three concrete actions for a 2026 enterprise. First, score every active deployment on GAUGE before the next review cycle. Deployments scoring below twelve are in the struggling cohort regardless of how the project status is currently reported internally. Second, treat low-scoring deployments as a portfolio kill-or-fix decision, not a continuation default. The realistic move-from-eighty-eight-to-twelve horizon for a single deployment under sustained discipline is twelve months. Deployments where the team cannot commit to the discipline within one to two quarters are better killed than rescued. Third, anchor the next procurement against the cohort, not the average. Vendor case studies typically describe the high-performing cohort. The procurement team's deployment will land in the bimodal distribution that the four datasets document.
AVERY: Verdict update on each.
[14:00] Verdicts and cadence on each of the four claims
ABBY: AM-029 is Holding. Stanford's published methodology is intact. The deployment cohort is fixed at fifty-one. Cadence is sixty days. AM-132 is Holding. The four datasets are linked. The bimodal shape is documented across the four. Cadence is sixty days. AM-128 is Holding. The MIT NANDA report is published. The 95% framing and the build-versus-buy 67-22 spread are documented. Cadence is sixty days. AM-053 is Holding. The McKinsey 17% figure is published. The self-reported-versus-audited slippage is named. Cadence is ninety days.
AVERY: What would change any of them.
ABBY: For AM-029, a new Stanford DEL update with refreshed deployment cohort. For AM-132, vendor-side or analyst-side coordinated narrative shift on the bimodal framing. For AM-128, MIT NANDA published 2026 follow-up to the 2025 report. For AM-053, McKinsey State of AI 2026 mid-year refresh.
AVERY: Final word.
[14:30] Outtro
ABBY: The four claims, the four primary datasets, and the GAUGE diagnostic are linked at agentmodeai dot com slash holding. The Sunday brief ships every week with what moved on the ledger.
AVERY: Holding-up. See you next Sunday.
Claims tracked in this episode
AM-029 — Why 88% of agentic AI deployments fail
https://agentmodeai.com/holding/?claim=AM-029
AM-132 — Bimodal ROI distribution across four datasets
https://agentmodeai.com/holding/?claim=AM-132
AM-128 — MIT 95% GenAI-pilot-failure claim
https://agentmodeai.com/holding/?claim=AM-128
AM-053 — McKinsey 17% EBIT claim
https://agentmodeai.com/holding/?claim=AM-053
Primary sources
Stanford Digital Economy Lab — Enterprise AI Playbook 2026 — https://digitaleconomy.stanford.edu/app/uploads/2026/03/EnterpriseAIPlaybook_PereiraGraylinBrynjolfsson.pdf
McKinsey State of AI 2025 — https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
MIT NANDA — State of AI in Business 2025 Report — https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
Fortune coverage of MIT NANDA findings — https://fortune.com/2025/08/21/an-mit-report-that-95-of-ai-pilots-fail-spooked-investors-but-the-reason-why-those-pilots-failed-is-what-should-make-the-c-suite-anxious/
Frameworks referenced
GAUGE diagnostic — https://agentmodeai.com/gauge/
Subscribe
Newsletter and the full Holding-up ledger — https://agentmodeai.com
Editorial charter — https://agentmodeai.com/standards/
How it's written — https://agentmodeai.com/how-its-written/
Written by Claude · Curated and signed by Peter