I keep seeing the stories.
Deloitte produced a government report with incorrect and fabricated references. KPMG withdrew a report after investigators challenged many of its citations and case studies. PwC has now faced similar questions about thought-leadership reports published in the Middle East.
It would be very easy to laugh.
These are enormous professional services firms. They sell expertise, assurance, risk management and responsible AI. Clients pay serious money because the name on the document is meant to mean something.
But my honest reaction is different:
But for the grace of God go I.
I use AI constantly. It has helped me write, research, build and communicate in a way that was never available to me before. I know how powerful it is. I also know how quickly a sentence that sounds completely reasonable can arrive with a reference that looks completely reasonable.
That is the dangerous bit. The bad answer does not normally arrive wearing a red hat and shouting, "I made this up."
It arrives fluently.
So this is not a piece about how stupid somebody else was. It is about how any of us can make the same mistake, and why organisations need a stronger release process when AI-generated words are carrying evidence.
What has actually happened?
The public cases are not identical, and we should not pretend they are. Some have been formally corrected. Some remain under investigation. Some are recent reporting rather than a completed regulatory finding.
| Organisation | Publicly reported problem | Status as at 30 July 2026 |
|---|---|---|
| Deloitte Australia | An Australian government assurance review contained incorrect references, case citations and a fabricated quotation attributed to a Federal Court decision. | The report was corrected and replaced. The department published correspondence about the issue, and Deloitte agreed to a partial refund. |
| Deloitte Canada | Questions were raised about citations in a major Newfoundland and Labrador health workforce report. Deloitte said AI had supported a small number of research citations, not written the report. | Deloitte said it would revise citations and stood by the findings. The provincial professional accounting body opened an investigation in March 2026. That is not a final finding. |
| KPMG | GPTZero challenged citations and case studies in a report about customer experience and agentic AI. Organisations named in the report disputed claims attributed to them. | KPMG removed the report while it investigated. GPTZero reported that only five of 45 citations accurately pointed to the claimed source. That figure comes from the investigating vendor, not a regulator. |
| PwC Middle East | The Financial Times reported fake footnotes, misattributed claims and unverifiable material across four thought-leadership reports identified by GPTZero. | PwC told the FT that it was updating a limited number of supporting citations. The reporting is very recent and should not be described as a concluded regulatory finding. |
That last column matters. A proper evidence process must distinguish a correction from an allegation, an investigation from a judgement, and a vendor's analysis from a regulator's finding.
Otherwise, in writing an article about weak sourcing, I would simply be repeating the same failure.
The scandal is not that an AI hallucinated
Language models can produce false statements and false references. That is not a new discovery.
OpenAI's own research explains one reason: training and evaluation can reward a model for guessing instead of admitting uncertainty. The UK Government's AI Playbook is similarly direct. Generative AI output should not be trusted as factual without validation.
That does not make language models useless. It means we have to use the right tool for the right part of the work.
A language model is brilliant at helping us explore, structure, challenge, summarise and write. It can help locate possible evidence. It can help build a claim ledger. It can ask where our argument is weak.
It is not, on its own, a source of truth.
The deeper failure begins when an organisation takes probabilistic language generation, applies it to deterministic evidencing, and then wraps the result in institutional authority.
The model guessed. The organisation still signed its name.
Why do smart people miss this?
I do not think the answer is simply laziness or incompetence. There are several forces pushing in the same direction.
Fluency feels like evidence
We are used to awkward mistakes looking like mistakes. AI makes the opposite possible. A fabricated journal title can sound perfectly academic. A false quotation can sit naturally inside a well-written paragraph. A URL can look plausible enough that our eyes move past it.
The presentation quality rises faster than the evidence quality.
Verification is slower than generation
A model can draft 30 pages quickly. A person still has to open each source, find the relevant passage, check the date and scope, and confirm that the source supports the precise claim being made.
The more output we generate, the more verification debt we create.
Authority transfers too easily
Once text sits inside a branded report, readers stop seeing it as model output. It becomes a Deloitte claim, a KPMG case study or a PwC statistic.
The reputation of the institution lends credibility to the sentence, even when the evidence underneath has not earned it.
People are being asked to move faster
Organisations want more research, more proposals, more insight and more thought leadership. AI makes that volume possible. But if the review capacity does not grow with the production capacity, quality control becomes the bottleneck everyone is tempted to route around.
That is when "human in the loop" can mean little more than a tired person looking at a polished document five minutes before release.
How one bad reference becomes institutional fact
A false reference does not necessarily stay inside one report.
The report is published by a respected organisation. A journalist cites it. A university paper includes it. A search engine indexes it. A retrieval system gives it to another model. That model now sees several apparently independent places repeating the same claim.
The error has acquired a history.
This is why provenance, source authority and an audit trail matter. Search is not verification. Retrieval is not verification. Repetition is not verification.
The closer a piece of work gets to publication, advice, policy, client action or public authority, the more the evidence needs to move back towards the original source.
They already know what good governance looks like
There is an uncomfortable irony here.
Deloitte publishes guidance on trustworthy AI governance. KPMG has a Trusted AI framework. PwC promotes responsible AI controls. The Financial Reporting Council's 2026 guidance says that even when generative or agentic AI is used in audit, the human auditor remains accountable.
The answer is not another lovely framework sitting on a website.
The answer is to make the framework bite at the release gate.
Before a source-grounded document leaves the organisation, somebody should be able to answer:
- Does every cited source exist?
- Can the reviewer access it?
- Does it actually say what the document claims?
- Do the date, geography, population and scope match?
- Is this a primary source, a secondary report or somebody else's interpretation?
- Have allegations, estimates and uncertainty been labelled honestly?
- Who is the named human accepting responsibility for release?
Those checks are not anti-AI. They are how we make AI usable in serious work.
Use a red team before you publish
Here is a simple prompt readers can use in Codex, Claude or another capable research agent. Give it the draft and the reference list. Where possible, also give it access to the original source files or ask it to browse public sources.
Do not use the same conversation that wrote the article as your only review. A separate task, fresh context or different model is more likely to challenge the assumptions already embedded in the draft.
Prompt: Red Team My References
RED TEAM MY REFERENCES
Act as a sceptical evidence reviewer. Review the draft and every factual claim, quotation, number, case study and reference it contains. Your job is to find weaknesses, not to defend the draft.
For every material claim, produce a table with:
- claim as written
- cited source
- can the source be accessed?
- does the source exist?
- does the source actually support the precise claim?
- do the date, geography, population and scope match?
- is it a primary source, secondary source, allegation, estimate or opinion?
- status: verified / partially supported / unsupported / contradicted / inaccessible
- confidence
- correction required
Rules:
1. Open and inspect the source. Do not assume that a link, title or citation proves the claim.
2. Prefer the original official, regulatory, academic or first-party source. Independently locate it where possible.
3. Do not invent replacement references, quotations, page numbers, authors, dates or URLs.
4. If a source cannot be verified, say so plainly and recommend removing or softening the claim.
5. Distinguish an allegation or open investigation from a concluded finding.
6. Check whether one source is being used to support a broader claim than it can carry.
7. Flag circular sourcing, copied claims and sources that all trace back to the same unverified origin.
8. Quote only the shortest passage needed to show why a source does or does not support the claim.
9. End with: critical red flags, corrected wording, references to remove, references still needing human review, and an overall release recommendation: do not publish / publish after corrections / evidence checks passed.
This is a review-only task. Do not edit, publish, deploy, contact anyone or present the draft as approved. A human remains responsible for checking the evidence and deciding whether to release it.
The prompt is deliberately not clever. It is a checklist with teeth.
You can strengthen it further by dividing the work:
- One agent extracts the claims.
- A second checks the sources.
- A human reviews the disputed or consequential items.
- A final release check confirms that the corrected wording made it into the actual document.
That last step matters. A perfect review in a spreadsheet is useless if the wrong PDF is published.
A practical evidence workflow
My preferred workflow is becoming quite simple.
- Draft for meaning. Use AI to explore the argument and help express what you are trying to say.
- Extract the claims. Turn every important factual assertion into a claim ledger.
- Attach evidence. Link each claim to the strongest available original source.
- Check entailment. Confirm that the source supports the actual wording, not merely the general subject.
- Show uncertainty. Label estimates, limitations, contested claims and open investigations.
- Red-team separately. Ask another agent or person to find the failures.
- Sign off visibly. Record who accepted the evidence and the release decision.
- Keep corrections honest. Preserve a visible correction trail if something later proves wrong.
For repeated work, turn those checks into evals. Test whether references resolve. Test whether quoted text appears in the source. Test whether dates and named organisations match. Test whether every consequential claim has an owner and evidence.
Automation will not remove the need for judgement. It can make the boring checks much harder to forget.
But for the grace of God
I am not going to pretend that I will never publish an error.
I write a lot. I research quickly. I work with agents. The same speed that makes this possible also creates the risk.
The lesson from these cases is not that we should stop using AI. It is that we should stop pretending a fluent answer is a checked answer.
Professional authority is not evidence.
A citation is not evidence until somebody has opened it.
A human glance is not governance unless that human has the time, tools and responsibility to challenge the work.
So yes, hold Deloitte, KPMG and PwC to the standards their names promise. But also look at your own process.
Because the most useful response to somebody else's failure is not:
"How could they be so stupid?"
It is:
"What would stop the same thing happening here?"
Related reading
- The AI Screwed Up. Then What?
- You Cannot Vibe What You Do Not Know
- You Are The A In AI
- The Question Is Not Can We Trust AI
Sources and notes
- Australian Department of Employment and Workplace Relations: Targeted Compliance Framework assurance review
- Australian Department of Employment and Workplace Relations: Published correspondence about the review
- Associated Press: Deloitte to partially refund the Australian government over report errors
- The Independent: Deloitte responds to questions about the Newfoundland and Labrador healthcare report
- The Independent: Provincial accounting watchdog opens an investigation
- GPTZero: Investigation of KPMG's Total Experience report
- TechCrunch: KPMG removes the report while investigating
- Financial Times: PwC published reports on AI marred by AI hallucinations
- UK Financial Reporting Council: AI in audit guidance
- NIST AI 600-1: Generative Artificial Intelligence Profile
- OpenAI: Why language models hallucinate
- UK Government: Artificial Intelligence Playbook
- Deloitte: Trustworthy AI governance in practice
- KPMG: Trusted AI framework
- PwC: Responsible AI
