← back to the archiveCover illustration for “AI research needs evidence lineage”
POSTday 20·6 weeks ago·by Andy Padia

AI research needs evidence lineage

KPMG pulled an agentic-AI report after UBS, the NHS, Swiss Federal Railways, and TfL disputed its case studies. The lesson isn't "AI hallucinates" — it's that review without claim-to-source lineage is not a control.

KPMG has withdrawn a report on agentic AI — a report about the benefits of AI, which is the detail the internet will enjoy — after the organizations named in its case studies started disputing them. Per TechCrunch: UBS called claims about AI agents in its investment advice "factually incorrect". Swiss Federal Railways said the journey-planning agents KPMG described are not accurate. NHS Greater Manchester and Transport for London disputed their examples too. The report had been live since October 2025; outside researchers inferred hallucination patterns in the citations, and KPMG says it is investigating. Whether an LLM actually produced the errors is unconfirmed — and for the lesson worth taking, it doesn't matter.

Because here is the uncomfortable part: that report almost certainly passed review. A firm like KPMG does not publish a flagship thought-leadership piece without partners reading it. Competent, motivated humans read fluent, plausible, well-structured prose about named organizations — and approved it. The failure wasn't that nobody looked. It is that looking is not a control when the artifact gives the reviewer nothing to check.

A reviewer confronting the sentence "Swiss Federal Railways uses agents to plan and book journeys" has exactly two options: recognize it as false from personal knowledge, or find it plausible. Plausible is what hallucinations are. Without a link from the claim to its source — where it came from, the quotation in context, the date it was retrieved — the reviewer is not verifying; they are vibing with extra steps. Eight months of a false report on the website is what vibing signs off.

The pre-AI version of this control existed and worked: fact-checking against footnotes, painful and slow. AI-assisted research quietly dropped it, because generation became free while lineage stayed expensive — the model produces the claim but not the receipt. So organizations bolted the old review step onto the new pipeline and called it governance. Review inherited a job it cannot do.

The control that works is structural: quarantine consequential claims until lineage is attached. Every claim that names an organization, cites a number, or describes a deployment enters the draft flagged, and the flag clears only when four things are attached — the source, the quotation in its context, the retrieval date, and the named reviewer who opened that source and cleared it. No lineage, no publication; the sentence gets cut or rewritten as explicitly illustrative. This is boring, checkable, and delegable — the properties a control needs and "a partner read it" lacks.

At work I now put one question to every AI-assisted research deliverable before it goes near a client: pick any three material claims and show me the source behind each within two minutes. Teams with lineage pass instantly — the links are in the draft. Teams without it produce the distinctive silence of people realizing that what they reviewed was prose, not evidence. That silence, at a Big Four firm, is now a global news story with four blue-chip organizations issuing corrections.

Steal this threshold for your own shop: a claim is "material" if it names a real organization, states a figure, or would embarrass you in a correction. Material claims travel with receipts or they don't travel. Your review meeting stops asking "does this read well?" and starts asking "is every flag cleared?" — a question a junior person can answer accurately, which is precisely what makes it a control.

Hallucination is a model behavior; publishing one is a process failure — reviewers can only catch what the artifact lets them check, so ship the receipts with the prose.

#governance#hallucination#research#consulting#quality
← older drop
ARR velocity does not prove a product moat
newer drop →
Bounded agents beat executive titles

related drops

explore all 78 drops →
← back to the archiveday 59