Retrieval-augmented generation was supposed to make this problem manageable. Ground a language model in an authoritative legal database and it stops inventing cases. Vendors said so directly; one marketed "hallucination-free legal citations."
In 2024, a Stanford and Yale team tested the claim. The answer was no, and it produced the number that has anchored law firm AI conversations ever since: legal research tools hallucinate between 17% and 33% of the time Stanford / JELS.
That number is real, carefully produced, and routinely misused. It describes a five-day test window in May 2024, uses a definition of hallucination most vendors do not, and measures systems that have since been rebuilt on newer models. Meanwhile a different number circulates: 0.2%. Also real. The two cannot be compared, and understanding why is worth more to a firm evaluating legal AI in 2026 than either figure.
What the Stanford Study Actually Found
The study, published in the Journal of Empirical Legal Studies in 2025, preregistered 202 questions and ran them May 23 to 27, 2024. Its most valuable contribution is not the percentage. It is the definition. Correctness asks whether a statement of law is true. Groundedness asks whether the cited source actually supports it.
Key factual propositions make valid references to relevant legal documents.
The right to same sex marriage is protected under the U.S. Constitution. Obergefell v. Hodges, 576 U.S. 644 (2015).
Key factual propositions are cited, but the source does not support the claim.
The right to same sex marriage is protected under the U.S. Constitution. Miranda v. Arizona, 384 U.S. 436 (1966).
Key factual propositions are not cited at all.
The right to same sex marriage is protected under the U.S. Constitution. No citation offered.
The three groundedness states, with the researchers' own illustrations. Only the middle one counts as a hallucination, and only the middle one looks like competent work.
A response is hallucinated if it is either incorrect or misgrounded: in the authors' words, "if a model makes a false statement or falsely asserts that a source supports a statement." They call misgrounding potentially more dangerous than outright fabrication, because catching it requires reading each cited source and comparing it to the proposition it supposedly supports.
A fabricated case fails the first check anyone runs. A misgrounded citation survives verification by any reviewer who confirms the case exists and stops there.
| System, as tested in May 2024 | Accurate | Incomplete | Note |
|---|---|---|---|
| Lexis+ AI | 65% of all queries | 18% | Highest accuracy of the tools tested |
| Westlaw AI-Assisted Research | 41% of all queries | 25% | Roughly one third of responses contained a hallucination |
| Ask Practical Law AI | 19% of all queries | 62% | Draws only from Practical Law, not primary law |
Two findings matter more than the headline range. Refusal flatters the metric. Ask Practical Law AI seldom hallucinated, and it also gave incomplete answers 62% of the time. A system that declines to answer cannot assert a falsehood, so any benchmark reporting hallucination without responsiveness rewards silence. Length drives the rate. Westlaw wrote the longest answers, 350 words on average, and longer answers contain more falsifiable propositions. When the unit of measurement is the response, thoroughness is penalized. One more finding cuts against comfortable readings: accuracy was highest on trick questions built around false premises, and lower on the categories representing real attorney work.
Why 17-33% Is Not a 2026 Product Ranking
The authors bound their own result. Testing was a five-day window more than two years ago, on models that have since been replaced, inside closed systems where, as they put it, it is not possible to say precisely where an error occurs. The 202-question set was deliberately hard: bar exam material, circuit splits, doctrine in motion. And it covered chat-style research only. CoCounsel was never evaluated, and the authors name contract review and memoranda as open benchmarking problems. A firm citing this study about contract review is citing something it does not measure.
None of that weakens the finding. It bounds it: strong evidence that RAG did not eliminate hallucination as of May 2024, and that misgrounding is real and hard to spot. Not a leaderboard.
Two Rates, Both Accurate, Not Comparable
Harvey's BigLaw Bench work reports its Assistant at roughly 0.2%, defining a hallucination as a factual claim that can be demonstrably disproven against a source of truth, excluding reasoning errors, and dividing hallucinated sentences by total sentences Harvey, vendor-created. Watch what the unit choice alone does to one unchanged answer.
Counted by response
Stanford / JELS
1 hallucinated response of 1
100%
Counted by sentence
Harvey BigLaw Bench
1 hallucinated sentence of 20
5%
The same 350-word answer. The same single error. Two published metrics, twenty times apart, before anyone disagrees about what counts as a hallucination.
| Stanford / JELS | Harvey BigLaw Bench | |
|---|---|---|
| Unit of measurement | The response | The sentence |
| Counts misgrounding | Yes, central to the definition | Not within the stated definition |
| Counts reasoning errors | Correctness failures count | Explicitly excluded, tracked separately |
| Task type | Open-ended legal research | Reasoning over multiple long documents |
| Who ran it | Independent academic, preregistered | Vendor internal, automated with human review |
| Versions and dates | Stated: queries run May 2024 | Published Oct 2024. No model versions, no test window |
Then add the definitions. Stanford's central case, the real citation that does not support the proposition, is not clearly captured by "demonstrably disproven." Reasoning errors, which Stanford's correctness prong counts, are excluded by design. Both teams documented their methods. The numbers simply answer different questions.
The test to apply before any comparison
Ask whether two hallucination percentages measure the same failure, on the same kind of task, with the same denominator, evaluated by the same kind of party. If any of those differ, the comparison is arithmetic, not evidence.
How Hard This Is, Even Done Well
This is not a criticism of Harvey, which publishes its definition, formula, sample sizes, and process. Most vendors quoting a hallucination figure publish none of that. It is what makes the example useful: this is the careful end of vendor benchmarking, and a firm still cannot safely use the number for comparison. Start with the published figures. Check the arithmetic.
Two of the three published pairs reconcile. The third does not. A rate of 1.9% implies roughly 1 in 53, and a ratio of 1 in 110 implies roughly 0.9%. The accompanying data table lists 1.9%, which suggests the ratio is the slip rather than the percentage.
Almost certainly a typo, and worth exactly one paragraph: it sits in the summary line of the most-quoted vendor hallucination figure in the legal industry, and it has circulated uncorrected because nobody divides. What the disclosure leaves unsettled matters more, and none of it is an error.
| What a firm would need | What the publication provides | Why it matters in 2026 |
|---|---|---|
| Which models were compared | The names "Claude," "ChatGPT" and "Gemini," with no version identifiers | The post is dated October 2024. Those names refer to model generations that have since been replaced more than once |
| When testing ran | Not stated. Publication date bounds it at October 7, 2024 | Roughly four months after the Stanford queries. Both figures firms treat as "current" are from 2024 |
| What counted as a failure | A factual claim demonstrably disproven against a source of truth. Reasoning errors tracked separately | Excludes two of the three failure modes the Stanford authors found most common |
| What is being compared | A product with retrieval and workflow scaffolding, against foundation models used directly | A reasonable thing to measure, and not a like-for-like model comparison. The scaffolding is the point of the product |
| Whether length was controlled | States the result holds despite longer, more detailed answers | Under a per-sentence denominator, longer answers add to numerator and denominator together, so length is largely neutralized already |
That is the ordinary condition of vendor benchmarking: a company measuring its own product, in good faith, on marketing's timeline rather than replication's.
If this is what the careful end of vendor benchmarking looks like, no firm should be outsourcing its reliability judgment to published numbers.
Hallucination Is One Failure Mode of Nine
Most of the ways legal AI fails involve inventing nothing. The taxonomy is Songbird's operational framework, informed by the failure types the Stanford team documented.
| Failure mode | What it looks like | Why review often misses it | Detectability | |
|---|---|---|---|---|
| 1 | Fabrication | Invented case, quote, statute, or procedural history | Rarely missed. Fails the first check anyone runs. | Easy to catch |
| 2 | Misgrounding | Real authority cited for a proposition it does not support | Survives any reviewer who confirms the case exists and stops. | Hard |
| 3 | Reasoning error | Sources real and fairly described, conclusion unsound | Requires substantive legal judgment, not verification. | Hard |
| 4 | Material omission | Controlling or adverse authority never surfaces | Leaves no artifact. Nothing in the output signals the absence. | Very hard |
| 5 | Currency error | Superseded statute, overruled precedent, stale guidance | Output looks correct against the source it actually used. | Moderate |
| 6 | Jurisdiction error | Legitimate authority that does not govern this matter | Citation checks pass. The case is real and good law elsewhere. | Moderate |
| 7 | False-premise acceptance | Model reasons from an incorrect assumption in the prompt | The answer is responsive, which reads as competence. | Hard |
| 8 | Retrieval failure | Generation is sound, retrieval never surfaced the right document | Invisible without knowing what should have been retrieved. | Very hard |
| 9 | Workflow execution error | Agent edits the wrong clause, document, or matter | Text-level review does not examine actions taken. | Hard |
A hypothetical makes rows 2 and 4 concrete. On a preemption question, the system returns five real cases, accurately summarized and correctly cited, and never surfaces the controlling sixth case that cuts the other way. Every citation check passes. Nothing in the output is false. The memo is still wrong, and no citation-focused review will catch it.
Reliability Is a Stack, Not a Model
Two products built on the same frontier model can differ substantially, because reliability is produced, or lost, at every layer between the model and the lawyer.
"It uses Claude" tells you about one layer and nothing about the other nine. A legal platform can outperform its base model by adding retrieval, authoritative data, and verification. It can also underperform it, with poor retrieval or constraints that suppress useful answers. Both happen.
"Claude vs. Legal AI" Is Converging, Asymmetrically
On May 12, 2026, Anthropic released Claude for Legal: twelve practice-area plugins under an Apache 2.0 license, plus more than twenty MCP connectors into systems firms already run Anthropic. The next day, LexisNexis announced it had integrated those plugins into Lexis+ with Protégé, inside the Lexis environment LexisNexis. The boundary is blurring, and it is blurring in one direction faster than the other.
Shipped. The authoritative corpus and the citator stay inside the vendor environment.
Shipped research connectors reach public dockets and secondary sources. The dominant paid corpora are not in the default set.
Read the repository, not the coverage. The shipped research connectors are CourtListener, Descrybe, Trellis, TopCounsel, Definely, and Solve Intelligence. Thomson Reuters sits under "Wanted connectors," and LexisNexis does not appear in the file at all CONNECTORS.md, whatever the trade press reports. Verify the current state yourself; this is the fastest-moving claim in this article. Out of the box, Claude's legal research runs on public dockets and secondary sources, which is genuinely useful and is not a licensed primary-law platform with citator coverage.
The question is no longer "Claude or Lexis." It is: for this task, which corpus is this system actually reading, and what verifies the link between its assertions and that corpus?
What Current Benchmarks Can and Cannot Tell You
| Benchmark | What it tests | Unit | What it can tell you | What it cannot |
|---|---|---|---|---|
| Stanford / JELS queries run May 2024 Independent | Open-ended legal research, 202 hard questions | Response | That RAG did not eliminate hallucination, and that misgrounding is common and hard to spot | Current product performance. Contract review. Drafting. CoCounsel, which was not tested |
| Vals VLAIR Feb 2025 Opt-in | Seven tasks including extraction, document Q&A, summarization, redlining, chronology | Task score vs lawyer baseline | Comparative task quality among participating vendors | Hallucination, which was not the measured outcome. Lexis+ AI withdrew, and vendors chose which tasks to enter |
| Vals legal research Oct 2025 Opt-in | 210 questions across nine research categories, scored on accuracy, authoritativeness, appropriateness | Weighted rubric | That tested tools scored 79% to 81% against a 71% lawyer baseline, and that general-purpose ChatGPT scored within one point of the best legal-specific tool | Broad market comparison. Thomson Reuters, LexisNexis and vLex all declined to participate. Hallucination was not measured |
| Harvey BigLaw Bench published Oct 2024 Vendor | Reasoning over multiple long documents | Sentence | Harvey's measured rate under Harvey's definition | Comparison to Stanford's figure. Reasoning errors, excluded by design. No model versions and no test window, so it cannot be dated or reproduced |
No benchmark here measures what most firms think it measures. Two of the four do not measure hallucination at all, the most rigorous one is two years old, and the lowest number is vendor-produced under a definition that excludes reasoning errors. The October 2025 Vals result is the strongest published evidence of convergence, with tested tools at 79% to 81% against a 71% lawyer baseline and a general-purpose model within a point of the best legal-specific tool Vals 2025. Its caveat is structural: Thomson Reuters, LexisNexis, and vLex all declined to participate VLAIR. A benchmark the incumbents skip describes the participants, not the market.
The Metric That Belongs in the Buying Decision
Verification burden is the time, expertise, and source access a qualified reviewer needs to determine that an output is safe to rely on. The framing is Songbird's; the observation is the Stanford authors', who noted that longer answers require proportionally more checking, proposition by proposition. Operational risk compounds three ways.
How often the system produces a material error on this kind of task.
How often the firm's review process fails to catch it. This is where tool design has the most leverage, and where demos reveal nothing.
What happens if the error survives into client work or a filing.
Procurement usually prices only the first. Picture two tools that save the same hour: Tool A links every proposition to its supporting passage and flags what it could not support, while Tool B writes better prose that must be reconstructed to check. Tool B can win the benchmark while Tool A is the lower-risk purchase. Slightly more accurate but substantially harder to verify is operationally worse.
Run Your Own Evaluation
Vendor benchmarks answer the vendor's question. Build 30 to 50 tasks from closed matters where the right answer is already known, across four tiers.
Summarization, extraction, basic drafting.
Research, multi-document synthesis, contract interpretation, chronology building.
Conflicting authority, jurisdiction-sensitive questions, recent developments, large document sets.
False premises, requests for authority that does not exist, ambiguous jurisdiction, superseded law, deliberately incomplete document sets.
Score each dimension separately; composites hide failure modes. The three starred below are the ones firms most often skip, and where the hardest-to-detect failures live.
| Dimension | What to test | What good looks like | Warning sign |
|---|---|---|---|
| Substantive correctness | Compare against the known outcome | Errors rare and legally minor | Fluent answers that are subtly wrong |
| Citation existence | Does every authority exist as cited | No fabrications | Any fabrication at all |
| Citation support★ | Read the source. Does it support the proposition | Propositions traceable to specific passages | You must read the full case to tell |
| Authority quality | Controlling versus persuasive versus secondary | Hierarchy respected and explained | Treats all sources as equivalent |
| Completeness★ | Seed a task with known adverse authority | Surfaces it, or flags a gap | Confident answer, missing authority |
| Currency | Include a superseded provision | Flags supersession | Silently relies on stale law |
| Jurisdiction | Ask a question governed by one state | Stays in jurisdiction or says it cannot | Cites real law from elsewhere |
| False-premise resistance | Embed a wrong assumption in the prompt | Corrects the premise | Reasons helpfully from the error |
| Uncertainty calibration | Ask something genuinely unsettled | Says so | Manufactures a clean answer |
| Reproducibility | Run the same task three times | Materially consistent | Different conclusions per run |
| Workflow correctness | For agents, check actions, not just text | Right document, right clause, logged | Correct text, wrong target |
| Provenance | Can you reconstruct what it saw | Sources, dates, retrieval visible | Conclusions without a trail |
| Reviewer effort★ | Time a qualified reviewer to sign-off | Minutes, with sources at hand | Reviewer rebuilds the research |
Set thresholds by workflow and consequence, not from published benchmarks, and record who tested, which product and model version, and when, because the answer expires. The recurring mistakes are consistent: treating a 2024 percentage as a current specification, comparing rates with different denominators, checking that citations exist but not that they support the proposition, testing only vendor-supplied examples, evaluating a model when the firm is buying a workflow, and never retesting after the product changes. ABA Formal Opinion 512 makes verification task-dependent and treats uncritical reliance as malpractice ABA FO 512, and that cuts both ways: it forecloses blanket bans justified by a stale statistic as firmly as it forecloses uncritical adoption. State rules vary Model Rules.
Procure for Verification, Not Infallibility
No system will be reliably right about the law. The standard is a system in which material errors are uncommon, the failure modes are named, the evidence is visible, citations can be checked against the propositions they support, and a lawyer can find the remaining errors in an amount of time the matter can absorb. Firms that evaluate this way can adopt aggressively, because they know what they are watching for. Firms that buy on a hallucination percentage have bought a number whose denominator they did not check.
Take the last AI-assisted research memo your firm produced, and ask the reviewing lawyer two questions.
Which propositions did you verify against the source text, rather than confirming the case exists?
How would you have known if the controlling case never appeared in the output at all?
If neither has a confident answer, the firm has an adoption process, not a reliability process.
What This Connects To
This article covers evaluation. The rest of the cluster covers what surrounds it. ai-tool-due-diligence covers the vendor review that precedes a pilot. how-to-run-a-legal-tech-pilot covers the structure the evaluation set belongs inside. governing-claude-legal-work covers the workflow governance a firm needs once a tool is approved, including task-specific review standards. privilege-confidentiality-ai covers the confidentiality analysis that governs what may enter the system in the first place. firm-ai-policy covers the policy structure. The Legal AI Matrix covers the product landscape.
This article is operational guidance and does not constitute legal advice or ethics guidance. ABA Model Rules and formal opinions are interpretive; state rules and ethics opinions vary, and jurisdictions may impose different requirements. Product capabilities and benchmark results described here reflect published sources as of August 2026 and change, in some cases quickly. Songbird Strategies is a legal technology consulting firm, not a law firm, and has not independently benchmarked any product named in this article. Firms should consult qualified ethics counsel on jurisdiction-specific questions. See Sources & Notes for every benchmark and vendor claim cited, including which are vendor-produced.