ICYMI
The AI Verification Failure Law Firms Don’t Want to Talk About
When Policies Aren’t Enough: Sullivan & Cromwell, Pinsent Masons, and the AI Verification Failure Nobody Wants to Talk About
The policies existed. The training requirements existed. The safeguards existed. The errors made it to court anyway.
That is the sentence that should unsettle every senior partner reading about Sullivan & Cromwell’s apology to Chief Judge Martin Glenn of the US Bankruptcy Court for the Southern District of New York. In a letter dated 18 April, restructuring co-head Andrew Dietderich confirmed that a motion filed nine days earlier contained fabricated case citations, misquoted provisions of the US Bankruptcy Code, and a range of errors so extensive that the correction required a three-page single-spaced attachment. The firm did not lack AI governance. It had comprehensive policies and training requirements specifically designed, in Dietderich’s own words, “to prevent exactly this situation.” They did not prevent it.
Within weeks, a parallel failure landed on this side of the Atlantic. ICC Judge Mark Mullen issued a public admonishment of Pinsent Masons after a junior solicitor, identified only as Lawyer A, used AI to research an insolvency matter and produced hallucinated statutory wording that was submitted to the High Court — not once, but twice. When the judge queried the original letter, the firm prepared an explanatory response. That response was also generated with AI assistance. It was also wrong. Pinsent Masons has referred itself and three of its solicitors to the SRA.
The Same AI. The Same Warning.
There is a detail in the Pinsent Masons judgment that deserves more attention than it has received. The AI tool used by Lawyer A did not silently produce the hallucinated content. According to the ruling reported by Legal Futures, the system repeatedly warned the lawyer to verify the material against authoritative legal sources before using it. The AI flagged its own unreliability. The warnings were not acted upon.
Judge Mullen’s conclusion was pointed: the junior solicitor had “almost entirely outsourced the thinking process” to the AI program. The supervising partner and senior associate had failed to check the work. The result was material that the judge described as “impossible to accept” — and his concern, on reading the firm’s attempted correction, was that a “cavalier attitude” was being taken to the accuracy of what was placed before the court.
This is not primarily a technology failure. It is a process failure in which the presence of AI created a false confidence that displaced the professional scrutiny that should have caught it.
The Policy Gap
The Sullivan & Cromwell situation illustrates a distinction that the legal sector is slowly being forced to confront. The firm’s own letter acknowledges that its policies on AI use “were not followed in connection with the preparation of the motion.” This matters. The failure was not a gap in the policy — it was a gap between the policy and what happened in practice. As we have argued before, fast is not the same as right. Comprehensive AI governance on paper does not constitute verification.
The distinction between AI-generated and verified is not semantic. It is the single most important operational question in legal AI deployment right now. A database of 330 instances of AI-hallucinated citations in US court filings — and growing — suggests that policies without process enforcement are largely decorative. Courts are running out of patience, and regulators are beginning to ask harder questions.
What Verified Actually Means
In the rush to adopt AI in legal practice, the word “verified” has been stretched to the point of meaninglessness. Firms say their AI tools have verification features. They say lawyers are responsible for checking outputs. They say training and policy frameworks are in place. Sullivan & Cromwell said all of these things — and 42 inaccuracies still reached a federal judge.
Human verification is not a feature. It is an architectural commitment to ensuring that a legal mind has actually engaged with the content before it reaches anyone who will rely on it. That means a person who understands not only whether a summary is technically accurate, but whether it captures the significance of what it describes — and who stands behind it professionally.
This is precisely the standard Thelonious was built to. Every piece of intelligence that reaches our subscribers — every case summary, every regulatory update, every enforcement alert — is reviewed by a legal mind before it arrives on their screen. Not reviewed by an AI. Not flagged for optional human spot-check. Reviewed. The primary source is linked. The analysis is checked. The confidence level meets the standard legal professionals require when briefing clients, advising boards, or making decisions that carry professional liability.
The cases that broke into the news this month are not outliers. They are the visible surface of a much larger problem. Lawyer Monthly’s analysis of the Pinsent Masons ruling noted that “verification failures are no longer confined to isolated procedural disputes or inexperienced users” — they are now appearing in complex, high-value matters at some of the world’s most sophisticated practices.
The question every legal team should now be asking of every AI tool they rely on is not “does it have policies?” It is: “Has a legal professional checked this before it reached me?” If the answer is anything other than yes, the professional risk sits with the team, not the tool.