Skip to main content

10 July 2026

AI Governance

The Deloitte AI Hallucination Scandal: Lessons for Everyone

Deloitte was caught delivering a A$440k government report riddled with AI hallucination: fake citations and a made-up court quote. Here are the real lessons.

The Deloitte AI Hallucination Scandal: Lessons for Everyone, AI Governance, AI Risk analysis by Amjid Ali.

A A$440,000 report, delivered to a federal department, cited a book that was never written and quoted a judge who never said the words. The tool did not fail quietly. The people around it did.

In short, Deloitte Australia delivered a government assurance report containing AI-generated citations, references, and a court quote that were simply invented, then corrected the document and refunded part of its fee once a university researcher exposed the fabrications.

I have spent a lot of time helping organisations put generative AI into real workflows, and I keep coming back to one point: the technology is rarely the thing that fails you. Your process is. The Deloitte case is the cleanest public example of that I have seen, and it is worth walking through carefully, because the lessons apply to every business now using AI, not just to one of the world’s largest consultancies.

What actually happened with the Deloitte AI report?

Deloitte was paid roughly A$440,000 to review the IT system behind Australia’s welfare penalty framework, and the report it delivered contained multiple fabrications produced by generative AI.

In late 2024, Deloitte was engaged by the Department of Employment and Workplace Relations (DEWR) to conduct an independent assurance review of the Targeted Compliance Framework, the system that automates penalties for jobseekers who miss their mutual obligations. This is sensitive territory in Australia. It sits close to the shadow of Robodebt, the unlawful automated debt-recovery scheme that caused enormous harm. Getting an assurance review of welfare automation right actually matters.

The 237-page report was published on the department’s website in July 2025. It looked unremarkable until Dr Christopher Rudge, a law researcher at the University of Sydney, read it closely. As The Nightly reported, Rudge found citations to academic papers and experts that did not exist, including around ten references tied to a book titled “The Rule of Law and Administrative Justice in the Welfare State” that was never written.

Worse, the report misquoted a real Federal Court judgment. It attributed a fabricated quote to Justice Jennifer Davies and misspelled her surname as “Davis”, and referenced the case in a way that invented a passage of several lines that appears nowhere in the actual ruling. For a legal and assurance document, inventing judicial reasoning is about as serious as an error gets.

After the scrutiny began, Deloitte confirmed it had used a generative AI system, Microsoft’s Azure OpenAI GPT-4o, to help draft parts of the report. It issued a corrected version on 26 September 2025 with the false references removed and a disclosure of the AI tooling added, and it agreed to refund the final instalment of its fee. As ACS Information Age noted, this became one of the first public cases of a government clawing back money over undisclosed AI use in official work.

What was fabricated in the Deloitte report?

The report invented sources across three categories: non-existent academic references, a book that does not exist, and a misattributed, misquoted court judgment.

Here is the pattern, because the pattern is the point. These were not typos or sloppy paraphrases. They were confident, specific, professional-looking fabrications: an author here, a paper title there, a paragraph of judicial reasoning that reads exactly like real judicial reasoning. That is the signature of an AI hallucination. It does not look wrong. It looks right, which is precisely why a human skimming quickly would wave it through.

If you want the deeper implications for the profession, I wrote separately about how AI is reshaping the consulting industry, because this is not a one-firm story.

What is an AI hallucination, and why does it happen?

An AI hallucination is fluent, confident output that is factually false, and it happens because language models generate plausible text rather than retrieve verified facts.

A large language model is not a database. When you ask GPT-4o for a citation supporting a claim, it does not look up a library. It predicts the sequence of words most likely to follow your prompt, based on patterns in its training. If a real citation fits that pattern, great. If one does not exist, the model will happily manufacture something that has the shape of a citation: plausible author, plausible title, plausible year. It is not lying, because it has no concept of truth. It is completing a pattern.

This is why hallucinations are most dangerous in exactly the settings where they are hardest to catch: legal references, academic sources, statistics, quotes, case law. The output is specific and authoritative, and specificity reads as credibility. Understanding this is not optional for any leader deploying AI. If you have not built your team’s intuition for where models break, start with my piece on whether ChatGPT is safe for Australian business.

Is this an “AI is bad” story? No, and that framing misses the lesson

This was a governance and process failure, not a technology failure. The AI did exactly what these tools do. The controls that should have caught it were not there.

Let me be fair to Deloitte here, because the pile-on has been enthusiastic and not always useful. Generative AI is a legitimate, valuable tool for drafting, summarising, and accelerating research. Used well, it makes good consultants faster. The problem was never that Deloitte used AI. The problem is that AI-generated content reached a federal client without anyone verifying the citations it contained.

That is a verification gap, and verification is not a technology problem. It is a discipline problem. A junior analyst who invented a court quote would have been caught by a review process. The AI invented a court quote and the review process did not catch it, which tells you the process was either absent or not designed for this failure mode.

This is the part that should worry every organisation, not just consultancies. The Big Four sell AI governance frameworks to their clients. If a firm of that scale can let hallucinated content reach a government department, the honest question every business should ask is: what would stop the same thing happening to us? For most companies, the uncomfortable answer is nothing, because they have not built the controls yet. I have written about why so many AI pilots fail to deliver, and the root cause is almost always the same: enthusiasm for the tool, neglect of the operating discipline around it.

What controls would have caught this?

Every fabrication in the Deloitte report was catchable by a basic verification step. The failure was in not running those steps.

Here is the map of what went wrong against the control that would have stopped it. None of these controls are exotic. They are the sort of thing a well-run team should already do.

What went wrongThe control that catches it
Cited academic papers that do not existEvery citation checked against a primary source before delivery
Invented a book and attributed real-sounding authorsProvenance tracking: each claim traced to where it actually came from
Fabricated and misquoted a Federal Court judgmentLegal and factual claims verified against the original ruling
AI-generated draft reached a government client unreviewedA hard rule: no unreviewed AI output leaves the building
AI use was not disclosed in the original reportDisclosure of AI assistance as standard practice on deliverables

How do you stop AI hallucinations at work?

You accept that hallucinations cannot be eliminated, and you engineer a verification layer that assumes the AI will occasionally be confidently wrong.

Here is what I put in place with the teams I work with, and what I would recommend to any business, from a two-person firm to an enterprise.

Verify every factual claim against a primary source. Citations, quotes, statistics, names, dates, legal references. If AI produced it and it asserts a fact, a human confirms that fact before it goes anywhere. This is non-negotiable for anything client-facing.

Never send unreviewed AI output to a client, regulator, or government body. Treat the model’s output as a first draft from a fast but unreliable junior. It saves you time on structure and phrasing. It does not get to sign the document.

Disclose AI use where it matters. If AI materially shaped a deliverable, say so. Regulators and clients increasingly expect it, and the Australian government is already reviewing procurement rules to require disclosure of generative AI in commissioned work. Disclosure also changes behaviour: teams that know they must declare AI use tend to check it more carefully.

Build validation into the workflow, not as an afterthought. Quality gates before delivery. Named accountability for verification. A checklist that treats hallucination as an expected failure mode, not a surprise. This is exactly the kind of structure I lay out in my AI governance framework for Australian business.

The real lesson is about trust

AI does not remove professional accountability. It concentrates it. When you put a tool in front of your team that produces authoritative-sounding output at speed, you have not reduced the need for human judgement. You have raised the stakes on it. The faster the draft, the more disciplined the review has to be.

Deloitte’s reputation is built on trust, and trust is the one asset AI cannot generate for you. A client does not pay a consultancy for words on a page. Words are now nearly free. They pay for the assurance that those words are true, checked, and defensible. The moment an AI-drafted fabrication reaches a client unverified, the firm has sold the one thing it cannot outsource to a model.

The takeaway for everyone, not just consultants: the winners in this era will not be the organisations that adopt AI fastest. They will be the ones that pair AI’s speed with verification discipline that never blinks. Use the tool. Trust nothing it tells you until a human has checked it. That is not fear of AI. That is how you use it like a professional.

Amjid Ali is an AI and technology leader in Melbourne helping organisations adopt AI with governance that actually holds. To talk about putting real controls around AI in your business, get in touch.

Frequently asked.

What happened with the Deloitte AI report for the Australian government?
Deloitte delivered a roughly A$440,000 assurance report to the Department of Employment and Workplace Relations that contained fabricated academic citations, a non-existent book, and an invented quote attributed to a Federal Court judge. The errors came from generative AI. After they were exposed, Deloitte corrected the report, disclosed its use of Azure OpenAI, and agreed to a partial refund.
Who caught the errors in the Deloitte welfare compliance report?
Dr Christopher Rudge, a health and welfare law researcher at the University of Sydney, spotted the fabrications while reading the 237-page report. He noticed citations to papers and experts that did not exist, plus a misquoted and misspelled Federal Court judgment. His public analysis triggered the media scrutiny that led Deloitte to correct the document and refund part of its fee.
What is an AI hallucination and why does it happen?
An AI hallucination is when a large language model produces confident, fluent text that is factually false or entirely invented. It happens because these models predict plausible word sequences rather than retrieve verified facts. When asked for a citation or quote that does not exist in their training, they generate one that looks correct. The output reads authoritatively, which is exactly what makes it dangerous.
How do you stop AI hallucinations in business documents?
You cannot fully stop hallucinations, so you build controls around them. Verify every citation, quote, statistic, and factual claim against a primary source before it leaves your organisation. Never send unreviewed AI output to a client, regulator, or government body. Disclose AI use where it matters. Treat AI as a fast first draft, never as the final authority on facts.
What are the lessons from the Deloitte AI scandal for consultants?
The core lesson is that AI does not remove professional accountability, it concentrates it. A firm's reputation rests on verification discipline, not on the tool it uses. Consultants should adopt mandatory human fact-checking, provenance tracking for every claim, clear AI disclosure, and quality gates before delivery. The failure was a process failure, and process is exactly what clients pay a consultancy to have.

Picked by shared topic. The through-line is agentic AI shipped into production, not the pilot theatre.

Read another.