KPMG's AI Trust Problem: The Agentic AI Report That Hallucinated Its Evidence
In October 2025, KPMG published a report with a confident title: Total Experience: Redefining Excellence in the Age of Agentic AI. In June 2026, the firm quietly took it down and said it was reviewing the circumstances of its publication. In between, a forensic review of the report's evidence found that much of it — the case studies and citations meant to prove that agentic AI was already remaking real companies — did not survive contact with reality.
The obvious way to tell this story is as irony: one of the world's largest professional-services firms, which happens to sell "trusted AI" advisory work, published a report about AI that appears to have been undone by AI. That reading is fair, and I'll come back to it. But the irony is the least interesting part. What actually went wrong here is not that a language model made things up — models do that; we've known it for years. What went wrong is that a document carrying a global firm's name reached publication with its factual claims unverified. That is not an AI failure. It's an institutional-control failure that AI happened to expose.
What the forensic review actually found
The most detailed public account of the report's problems comes from GPTZero, a company that builds AI-detection and content-forensics tools. On June 12, 2026, it published an analysis of the report's sourcing. It's worth being precise about what that analysis is and isn't: it's a careful, methodical forensic assessment of citations, not an omniscient verdict. Where I describe its conclusions below, treat them as GPTZero's findings, not settled fact.
GPTZero examined 45 citations in the report. Its breakdown matters, because the temptation is to flatten "flawed citations" into "invented sources," and the truth is more textured than that:
- 5 of the 45 accurately pointed to real, verifiable sources.
- 28 clustered around identifiable real sources but had paraphrased or altered titles and, in many cases, fabricated components — a real organization or study bent into a shape that no longer matched the original.
- 12 were too vague or too garbled to match to anything with confidence.
By GPTZero's count, 40 of the 45 citation titles were not exact matches to a real title, and at least 16 met its working definition of a hallucination. Stepping back from citations to the claims those citations were supposed to support, GPTZero assessed that roughly half of the citation-backed claims were fake or misattributed. Again: that is GPTZero's assessment, arrived at citation by citation. It's a strong, specific finding, and it's also the kind of finding a reasonable reader should hold as "carefully argued" rather than "proven beyond dispute."
The distinction between a fabricated source and a mangled one isn't pedantry. A wholly invented citation is a machine confabulating out of nothing. A real study with a rewritten title and an added, false detail is something subtler and, in a way, more dangerous: it looks like diligence. It has the texture of a real reference. It would pass a skim. That texture is exactly what a citation is supposed to earn honestly, and exactly what no one at the publishing end checked.
The companies that said "that's not us"
The sharpest evidence didn't come from a detector at all. It came from the companies the report held up as success stories, who read about their own AI deployments and said, in effect, that's not what happened.
These responses were first reported by the Financial Times; because that piece sits behind a paywall, I'm relying on the International Accounting Bulletin's summary for the specific wording, with the FT original linked for anyone who can reach it.
- UBS. The report described the bank integrating AI agents across investment advice, risk, and compliance on a platform co-developed with Microsoft. UBS called that account "factually incorrect."
- SBB, the Swiss rail operator, described as running AI journey-planning agents, said the claim was "not accurate."
- Transport for London, credited with a described congestion and multimodal-transport deployment, called the description "misleading."
- NHS Greater Manchester, cited for AI agents that predicted hospital readmissions, triaged patients, and automated referrals, said the claim "doesn't really align" with the release it was drawn from.
Read those together and a pattern emerges that's worse than random error. In each case there was a real seed — a real bank, a real rail operator, a real transit authority, a real health system, each doing some genuine work with technology. The report took the seed and grew a more impressive plant than the facts supported. That's the signature of generated text optimized to be persuasive rather than true.
The Emirates example is the cleanest illustration, so it's worth handling carefully. The report described "Sara" as a mobile chatbot capable of changing customers' bookings. Emirates did introduce a Sara in 2023 — but it was a physical robot check-in assistant, not a booking-altering mobile app, and GPTZero found no support for the claim that it could modify reservations. The name was real. The capability was not. If you didn't already know what Sara was, the sentence would read as perfectly ordinary corporate AI copy — which is the whole problem.
The number that contradicted KPMG's own number
One internal inconsistency deserves its own mention, because it didn't require any outside company to catch. Total Experience stated that 55% of CEOs rank AI as their top investment priority. In the same month, KPMG released its own 2025 Global CEO Outlook, which reported that 71% identify AI as a top investment priority.
I want to be fair about the gap. "Their top investment priority" and "a top investment priority" are not necessarily measuring the same question — the single-highest priority is a stricter bar than one priority among several — so the two figures aren't automatically in contradiction. But the report never made its underlying source clear, and that's the point. A reader can't reconcile 55% with 71% because the report doesn't give them the thread to pull. A firm's flagship AI report disagreeing with the firm's flagship CEO survey, in the same month, with no visible reconciliation, is the kind of thing a single accountable editor would have caught in an afternoon.
The real failure is the missing control, not the model
Here's the part I care most about. It would be easy, and lazy, to conclude "this is why you can't trust AI for research." That's the wrong lesson, and it lets the actual failure off the hook.
AI is genuinely good at accelerating research. It can draft, summarize, surface sources, and assemble a coherent narrative faster than any human team. What it cannot do — what it is structurally bad at — is guarantee that a specific sentence is true and that a specific citation points where it claims to. Confabulation isn't a bug that a better prompt eliminates; it's a property of systems that generate fluent text by prediction. Any organization using these tools at scale has to design around that property, not wish it away.
The design that was missing here is boring, and its absence is the whole story. A publication process that puts an institution's name on a document needs three controls that no model can provide for itself:
- Claim-level provenance. Every factual assertion should carry a traceable link back to a source a human has actually opened. Not a citation that exists — a citation that supports this specific claim. The 28 "close but altered" references are what you get when provenance is assumed instead of checked.
- Source verification. Someone has to confirm the source says what the draft says it says, and that the named organization, study, or figure is real and correctly quoted. This is the exact step that would have caught Sara, caught the UBS platform, caught the 55%-versus-71% gap.
- Accountable human sign-off. A named person who stakes their judgment on the document and can be asked, afterward, "did you verify this?" Diffusion of responsibility — where the model drafted it, an analyst assembled it, and no one owns the truth of it — is how a document with sixteen-plus hallucinated citations gets a global logo on the cover.
None of these are novel. They are the ordinary controls of serious publishing, the kind newsrooms and journals and, frankly, audit practices have used for a century. What AI changed is the volume and plausibility of unverified material flowing toward the top of that funnel. The tools got faster. The controls didn't scale to match. That mismatch is the failure.
Keeping the irony in proportion
To KPMG's credit, its public response named the right principle even as it acknowledged the miss. The firm said it takes accuracy and integrity seriously, that it had removed the report and was reviewing the circumstances of its publication, and that it expects people to follow its responsible-AI guidelines — including human oversight to validate content and verify independent sources. That last clause is the correct diagnosis. The controls existed on paper. They weren't applied to this document.
And yes, the irony is real: KPMG markets trusted-AI and responsible-AI advisory services, and here was a report about agentic AI apparently tripped up by the failure modes that practice is meant to prevent. It's fair to note that. It is not fair to leap from there to "the whole firm is fraudulent." A large organization can sell sound guidance and still ship a document that skipped its own process; the gap between the two is the lesson, not proof of bad faith.
KPMG isn't alone in learning this in public, either. In 2025, Deloitte Australia agreed to refund the final installment of an Australian government contract after revising a taxpayer-funded report that had contained AI-generated errors and fabricated references. Worth stating precisely: that was a refund of the last payment, not a full refund of the engagement — the distinction matters if you're keeping score honestly. Two of the Big Four, inside a year, tripped over the same wire. That's not a coincidence. It's an industry-wide control gap becoming visible one embarrassment at a time.
The uncomfortable takeaway is that the firms selling AI transformation are running the same unverified pipelines they'd flag in a client. The fix isn't to stop using AI for research; that ship has sailed and it's a good ship. The fix is to treat model output as a draft that has earned nothing until a human has verified its claims and put their name to the result. Provenance, verification, sign-off. Until those are as automatic as spellcheck, the next retracted report is already being generated — fluently, plausibly, and wrong.
Sources
- GPTZero, "Chasing the Hallucinations: KPMG's AI-Powered Attempt at Redefining Excellence" — the primary forensic citation review.
- Financial Times, original reporting (paywalled).
- International Accounting Bulletin, "KPMG drops AI report after false case studies exposed" — counterparty-response summary.
- KPMG 2025 Global CEO Outlook.
- The Register, summary and KPMG statement.
- The Guardian, on Deloitte Australia's partial refund.
- Business Traveller, background on Emirates' "Sara" robot assistant.
- KPMG Trusted AI services.