The Answer Can Be Sourced, Confident, and No Longer the Law
Table of Contents
A legal AI can quote a real statute, link the official source, and still be wrong, because the provision it retrieved is no longer the law. This failure survives retrieval, and nothing on the surface shows it.
A legal AI can hand you a statute, quote it, link to the official source, and still be wrong. Not because it invented the section. Because the section it correctly retrieved is no longer the law. Regulators repealed it, or a later provision superseded it, or an amendment changed it in a way the quoted text does not show. The answer is sourced, fluent, and immediately incorrect, and nothing on the surface shows this flaw.
This is the failure that matters because it survives the obvious fix. Add retrieval, and the model stops inventing citations; it returns real ones. But a real citation to a provision that no longer governs is its own kind of wrong, and a harder one to catch, because everything about it looks right. The reader’s instinct to check is exactly what a clean, sourced answer disarms.
That is the shape of hallucination in legal texts. Not the wild fabrication you spot on sight. A right-looking answer resting on authority that has moved out from under it.
The Failure Has a Shape
People picture legal AI failure as the fabricated case, the citation to a decision that never happened. Those exist, and courts have sanctioned lawyers for filing them. But the more common failure is quieter. The rule is roughly right, and the citation is stale, or repealed, or renumbered, or attached to the wrong section. The answer reads like something a competent associate would write. That is precisely why it slips through.
It slips through because the better the model gets, the better the failure hides. Zhou and colleagues, writing in Nature in 2024, found that larger and more instruction-tuned models give an apparently sensible but wrong answer more often than older ones, including on the hard questions human reviewers overlook. The same models are tuned to never state, “I’m not sure.” They answer. A confident wrong citation, produced by a capable model, read by a busy professional, is close to the perfect conditions for an undetected error.
In Law, the Citation Is the Product
For most uses of these tools, a shaky source is an inconvenience. In legal work, it is a whole new game. An answer a lawyer cannot trace to authority in force is worse than no answer, because it arrives carrying the authority of confidence.
The numbers bear this out. Magesh and colleagues, writing in the Journal of Empirical Legal Studies in 2025, tested the leading commercial legal-research tools and found they still produced hallucinations between 17 and 33 percent of the time, and that vendors overstated their claims of “hallucination-free” results.
Dahl and colleagues, in the Journal of Legal Analysis of the year before, found that general-purpose models hallucinate on legal questions at least 58 percent of the time, and that they accept a user’s incorrect legal premise rather than correct it. The tool agrees with you, in citation form, whether or not you are right.
The Friendly Version, for Contrast
There is a gentle version worth separating out because it is the one a grounded tool already handles. When California reorganised its Public Records Act in 2023, the ten business response rule moved from Government Code section 6253 to section 7922.535 with no change of substance. An older model still cites the dead section 6253; a search-grounded one now returns 7922.535. The rule did not change, so even the wrong citation led to a correct answer. The reader who clicked through would find the live section and move on, mildly annoyed. This is the case people reach for, and it is the one retrieval solves.
On the surface, the ungentle case appears identical and is what a technical buyer should be concerned about. The cited provision did not just get renumbered; rather, legislators amended, repealed, or superseded it with a later section. The model, or even a correct retrieval over a stale corpus, returns the section as written and states it as the rule. It is no longer the rule. Same shape, actual harm, and no signal on the page telling the reader which version they are holding.
Take the federal personal exemption. Internal Revenue Code section 151 let a taxpayer deduct a fixed amount for themselves and each dependent, and that figure is quoted across decades of guides and forms a model trained on. The 2017 tax law set the amount to zero through 2025; the 2025 reconciliation law (the One Big Beautiful Bill Act) removed that expiry and kept the amount at zero with no end date. Section 151 is still in the code, so a model can retrieve it, quote it, and link to it. The exemption it describes is gone. Quote the old amount to a taxpayer and you have stated dead law, with a correct citation on it.
So the shape has contours worth naming. The right rule under a dead number. A repealed provision quoted as if it still governs. Or a confident answer to a question whose premise was already wrong. Different faces of the same failure: the model is reproducing the law as a pattern it absorbed, not as a body of authority it can check.
The Good News in the Shape
One can engineer against a shape. A truly random failure would be hopeless. This one is not. It clusters around specific, predictable seams: where the law moved, where it underwent rewriting, where the question assumed something untrue. Addressing those seams is possible, but not with the tool most teams initially choose.
A better prompt does not fix this. The behaviour is not a wording problem. The behaviour stems from the models’ construction and their programmed rewards, which is why it persists despite every clever instruction you provide. That is the next thing worth understanding: not that legal AI hallucinates, which everyone now knows, but why it hallucinates in this particular, citation-shaped way.
References
Expand (3 sources)
Dahl, M., Magesh, V., Suzgun, M., & Ho, D.E. (2024). Large legal fictions: Profiling legal hallucinations in large language models. Journal of Legal Analysis, 16(1), 64-93. https://doi.org/10.1093/jla/laae003
Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C.D., & Ho, D.E. (2025). Hallucination-free? Assessing the reliability of leading AI legal research tools. Journal of Empirical Legal Studies, 22, 1-27. https://doi.org/10.1111/jels.12413
Zhou, L., Schellaert, W., Martínez-Plumed, F., Moros-Daval, Y., Ferri, C., & Hernández-Orallo, J. (2024). Larger and more instructable language models become less reliable. Nature, 634, 61-68. https://doi.org/10.1038/s41586-024-07930-y