AI Lessons

A $440,000 Consulting Report Full of AI Sources That Never Existed

A $440,000 Consulting Report Full of AI Sources That Never Existed

A few weeks ago, one of the largest and most respected consulting firms in the world had to refund a government client because parts of the report it delivered were, in a real sense, fiction. Not wrong in judgment. Fiction: quotes from a court case that were never spoken, academic papers that don't exist, an expert citation attributed to a scholar who never wrote it.

The firm was Deloitte. The client was the Australian federal government. And the tool that produced the fabrications was a generative AI model doing exactly what these models do when no one checks their work.

If your business delivers expert work, reports, analyses, briefs, assessments, anything a client pays for because they trust your judgment, this is the most important AI story of the year for you. Because Deloitte didn't get burned by exotic technology. They got burned by the single most common way AI fails, and they had every resource in the world to catch it and didn't.

What Deloitte delivered

Late in 2024, Australia's Department of Employment and Workplace Relations hired Deloitte for an independent review, a contract worth about 440,000 Australian dollars. The subject was fitting in hindsight: the department wanted an assurance review of the IT system it used to automate penalties in the country's welfare program. An automated decision-making system, being audited.

Deloitte delivered a 237-page report, and the department published it in July 2025. For a while, nothing seemed wrong.

Then a researcher at the University of Sydney, Dr. Chris Rudge, actually started checking the references. What he found was that the report was, in his words, full of fabricated citations. It referenced research papers that did not exist. It attributed work to academics who had never produced it. Most striking, it included a quote from a real Australian federal court judgment, from paragraphs that don't exist in the actual ruling. The report had invented a passage and put it in a judge's mouth.

The confession in the footnotes

At first, Deloitte didn't say how the errors got there. When the corrected version was published, the explanation was in the methodology section: the work had used a generative AI model, Azure OpenAI's GPT-4o, as part of the process.

This is what makes the case so useful. The AI didn't break. It performed exactly as designed. Large language models generate plausible text, and a fake citation is plausible text. It has the shape of a real reference, the right author names, a believable title, a properly formatted quote. It is confident and well-formed and completely invented. The model was never promising it was true. It was only ever promising it was plausible. Someone treated plausible as true, and shipped it.

The detail that should stick with any professional-services firm: Deloitte sells responsible-AI advisory. This is a company that trains other organizations on how to deploy AI safely, and it still let fabricated content go out the door in a signed, six-figure government deliverable. Sophistication with the technology did not save them, because the failure wasn't technical. It was a missing step in their process: nobody verified the machine's output before it became the firm's word.

What it cost, and what it didn't

Deloitte refunded the final installment of its fee, roughly 97,000 Australian dollars. In the context of a firm that earns tens of billions a year, the money is a rounding error.

The real cost was reputational, and it's worth being precise about why. The damage wasn't "Deloitte used AI." Using AI to help produce a report is fine, arguably smart. The damage was that Deloitte used AI and didn't check it, and got caught by an outside academic reading the footnotes the firm apparently hadn't. For a business whose entire product is "trust our diligence," being publicly shown to have skipped the diligence is the injury. The refund is forgettable. The headline is not.

Why this is the failure most likely to hit your business

The chatbot cases and the robotaxi cases are dramatic, but they involve systems running live and unattended. This one is different, and closer to home, because it's about a human using AI as a tool to help produce work, exactly the way your team is probably starting to use it right now.

Someone on your staff, under deadline, asks an AI to help draft a client report, find supporting research, summarize a case, pull together citations. The output looks great. It's well-written, confident, specific. It's also, in some percentage of cases, partly invented, and the invented parts look identical to the real parts. That's the trap. A hallucination doesn't arrive flagged as a guess. It arrives in the same clean, authoritative prose as the truth.

How to use AI for expert work without becoming the next example

Treat AI output as a draft from a bright intern who lies confidently. You would never let a first-year's unchecked work go out under your firm's name with a client's name on it. AI output deserves exactly that posture: useful, fast, worth having, and never trusted without verification. The speed it gives you is real. The judgment still has to be yours.

Verify every specific, checkable claim. Names, quotes, statistics, citations, case references, dates. These are precisely what language models fabricate, and precisely what an outside reader can check. If the AI gives you a source, open the source. If it gives you a quote, find the quote. The rule is simple: anything falsifiable gets confirmed against the original before it ships.

Ground the AI in real material instead of asking it to recall. An AI asked to "find research on X" will invent research on X. An AI given your actual documents, your real sources, your verified data, and asked to work only from those, has far less room to fabricate. How you set up the task determines how much the machine has to make up to answer you.

Name an owner for the final word. Every deliverable that leaves your shop needs a specific person accountable for having checked it, not "the team," not "our process," a name. Deloitte's report went out because that accountability was diffuse enough that the machine's work reached the client without a human standing behind each claim.

The real lesson

We build AI into professional workflows for a living, including for firms whose product is expertise and whose reputation is everything. Done right, AI makes expert teams dramatically faster: it drafts, it researches, it structures, it handles the mechanical weight so your people can spend their time on judgment.

But it only works when someone owns the output. The Deloitte case isn't a warning against using AI for serious work. It's a warning against using it without the one step that makes it safe: a human who verifies before it becomes your name on the page. Deloitte had the talent, the tools, and the AI expertise to do that. They just didn't do it on this one. That's all it takes.

Andrew Lay

Written by

Andrew Lay

Andrew Lay is the founder and CEO of Hiero, a Michigan-based development studio that helps businesses use AI, automation, and custom software to improve how they operate. A business strategist specializing in AI, Andrew brings more than 20 years of experience building apps, digital products, and operational systems. His work focuses on the part of AI adoption most companies skip: identifying the right business problem, determining whether AI is actually the right solution, defining a defensible return, and putting the controls and feedback loops in place to protect that return after launch. Andrew is the author of the forthcoming book, Lessons from Bad AI Implementations and How to Guarantee ROI With AI, a practical field guide built from 34 verified failure cases and the Hiero implementation method. He also hosts the Hiero Exclusive podcast and speaks on AI strategy, entrepreneurship, and operational growth.

All posts by Andrew