Legal AI hallucinations were supposed to be a solved problem. The vendors who promised that have quietly rewritten their websites — and left the hardest question in the category unanswered.
In May of this year, Damien Charlotin published something more revealing than another sanctions statistic. He went back through the marketing pages of the major legal AI vendors and documented what they used to say versus what they say now.
LexisNexis once ran a blog post titled “How Lexis+ AI Delivers Hallucination-Free Linked Legal Citations.” The same post now reads “How Lexis+ AI Delivers Trustworthy Linked Legal Citations.” Thomson Reuters’ CaseText had stated plainly that “CoCounsel does not make up facts, or ‘hallucinate’.” That language softened. vLex removed a section headed “Safeguarding against hallucination” and replaced it with messaging about verification.
No press releases. No corrections. Just a series of quiet edits, made at different times, by different companies, in the same direction.
Charlotin’s phrase for it is better than anything I could write: the industry is “growing up, one quietly edited webpage at a time.”
I’ve been thinking about those edits for three months, because of what sits in the space they left behind.
What The Legal AI Hallucination Research Actually Found
The reason for the edits is not mysterious.
In 2024, a team from Stanford’s RegLab and Institute for Human-Centered AI — Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher Manning and Daniel Ho — ran the first systematic test of the purpose-built legal AI research tools the profession was being sold. Not consumer chatbots. The products marketed specifically to lawyers, with retrieval systems designed to prevent exactly this failure.
Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI hallucinated between 17% and 33% of the time.
The paper went on to be published in the Journal of Empirical Legal Studies in 2025, which matters: this is not a blog post someone can wave away. It is peer-reviewed literature, and its conclusion about the marketing was unambiguous.The providers’ claims are overstated.
To be fair to those tools, the same study found they beat general-purpose models substantially. Retrieval-augmented generation reduces legal AI hallucinations. It just doesn’t eliminate them the way “hallucination-free” implies, and a one-in-five error rate is a very different product from the one described on the landing page.
The edits followed. Which is, in its own way, a decent outcome — the claims were wrong and they were corrected.
But look at what the vendors reached for as a replacement. Trustworthy. Trust but verify. Verification-focused messaging. The industry retreated from a promise about the machine to something closer to an instruction for the user. The burden moved. It moved onto you.
What It Costs When Nobody Checks
Damien Charlotin also maintains the public database of court decisions in which legal AI hallucinations reached a filing. As of yesterday it held 1,952 cases worldwide.
776 of them involved lawyers.Twenty-eight involved judges.
The money has followed the volume. In the first quarter of this year alone, US courts imposed at least $145,000 in sanctions for AI-generated fake citations. In Whiting v. City of Athens, the Sixth Circuit fined two attorneys $30,000 each, plus fees and double costs — the highest federal appellate sanction tied to fabricated citations on record. Oregon courts accounted for roughly $109,700 in aggregate across several matters.
And these are only the ones that produced a written decision. The fabricated citation caught by a junior associate at eleven o’clock the night before filing does not appear in any database. Neither does the client advised on a regulation that says something slightly different from what the summary said. Those failures are invisible, and there is no reason to think they are rarer.
There’s an asymmetry worth sitting with here too. A Northwestern study cited in the sanctions reporting found that over 60% of federal judges are themselves using AI tools in their work. The bench is not standing outside this technology looking in. It is inside it, which is precisely why it recognises fabricated authority so quickly now, and why patience has run out.
The Strange Silence
So here is what I find genuinely odd about this market.
The single most valuable thing an AI legal product can offer right now is a credible answer to one question: can I rely on this? Not speed. Not coverage. Not a better interface. Reliability, evidenced.
That answer is worth more than any feature on any roadmap. And a significant part of the category is sitting on it, unspoken.
Try the exercise yourself. Open the product pages of the AI legal tools on your shortlist and read the top third — the part above the fold, where a company puts what it most wants you to know. Count how many use the words reviewed, checked, edited by, sourced by. They are all versions of the same word, and the word is verified.
The number is low. What you find instead is speed, scale, coverage, “AI-powered,” “real-time,” dashboards. All genuine benefits. None of them answer the question.
There are three reasons a company wouldn’t advertise that humans check its work — the one safeguard that actually catches legal AI hallucinations before they reach a filing.
- They don’t. This exists, and it is more common than the marketing suggests, but it’s not the interesting case.
- They do, and they’re faintly embarrassed about it. Human review reads as un-scalable. It complicates the story you tell investors. It is a payroll line in a deck that would rather show a margin curve bending the right way. So it goes in the FAQ, or nowhere.
- They do, and it never occurred to them to mention it. This is the most common of the three. It’s simply how the work gets done — of course a person reads the filings — so nobody thought to put it on the website.
Two and three are unforced errors. In a market where legal AI hallucinations have been documented in the flagship products a third of the time, and where the vendors have visibly retreated from their accuracy claims, the company willing to say a named human checked this before it reached you is holding the most valuable sentence available. And it’s whispering it.
Why We Say It Out Loud
I should be straight about my position here. I work for The Birdella Group, and we built Thelonious — a platform tracking AI legal and policy developments across more than 145 jurisdictions.
We put “human-verified” at the front of everything we say about it. Not as a caveat buried in the FAQ. As the headline.
That decision costs us.
It means we can’t ship as fast as a fully automated competitor. It puts a permanent payroll line where a rival has an efficiency. It quietly concedes that our product is not magic — that there is a person in the loop, and there is going to keep being a person in the loop.
Here is what it actually means in practice.
Our AI monitors continuously, across every jurisdiction we cover — litigation, regulation, policy guidance, industry developments. This is the part machines are genuinely better at than we are. No research team reads everything, everywhere, every day.
Then a legal or policy expert on our team checks what surfaced. Is it real, is it current, does it say what we think it says. The primary source travels with every entry, so you can check our work rather than take our word for it. And when something changes, it gets re-verified — not just re-scraped, which is how static trackers quietly rot.
Step two is where the cost sits.
It is also the step this market keeps trying to engineer away.
We think it’s the step you’re paying for.
The Question You’ll Be Asked
A general counsel stopped me mid-demo earlier this year. We were three screens into the Regulatory Records tracker when she raised a hand and asked one question.
“Who checked this?”
Not how fast. Not how many jurisdictions. Not what the AI does.
Who checked this.
I told her: our legal and policy team, every entry, before it appears — and here’s the primary source, one click away, so you can check us. She said something I’ve thought about ever since. “You’re the first person who’s had an answer.”
She wasn’t buying into an AI company. She was buying the ability to tell her board something true.
Every AI tool in legal will eventually face that question, from someone who is not in a mood to be impressed. A regulator, a client, a judge, an insurer, a partner whose name is on the filing.
The vendors who spent two years promising hallucination-free output have already found this out, and have edited their websites accordingly. The rest of the category still has a choice about how it answers.
If a tool can’t tell you who checked the output, that isn’t a small gap in the marketing.
That’s the gap where your professional risk lives.
Sources
- Magesh, Surani, Dahl, Suzgun, Manning & Ho, Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, Stanford RegLab & HAI (2024), published in the Journal of Empirical Legal Studies (2025)
- Damien Charlotin, Tracking hallucination marketing claims from legal tech vendors, LLRX, 31 May 2026
- Damien Charlotin, AI Hallucination Cases Database, updated 23 August 2026
- Rob Robinson, The AI Sanction Wave: $145K in Q1 Penalties, ComplexDiscovery, 6 April 2026
(Disclosure: The featured header image for this article was generated using artificial intelligence (Google Gemini).)