AI & Confidentiality

Legal AI Reliability in 2026: Why "Which Tool Hallucinates Least?" Is the Wrong Question

The most-quoted number in legal AI comes from a five-day test in May 2024, and it does not mean what most firms think it means. Two published hallucination rates can both be accurate and still be incomparable. What matters now is whether your firm can detect the errors that survive.

Retrieval-augmented generation was supposed to make this problem manageable. Ground a language model in an authoritative legal database and it stops inventing cases. Vendors said so directly; one marketed "hallucination-free legal citations."

In 2024, a Stanford and Yale team tested the claim. The answer was no, and it produced the number that has anchored law firm AI conversations ever since: legal research tools hallucinate between 17% and 33% of the time Stanford / JELS.

That number is real, carefully produced, and routinely misused. It describes a five-day test window in May 2024, uses a definition of hallucination most vendors do not, and measures systems that have since been rebuilt on newer models. Meanwhile a different number circulates: 0.2%. Also real. The two cannot be compared, and understanding why is worth more to a firm evaluating legal AI in 2026 than either figure.

What the Stanford Study Actually Found

The study, published in the Journal of Empirical Legal Studies in 2025, preregistered 202 questions and ran them May 23 to 27, 2024. Its most valuable contribution is not the percentage. It is the definition. Correctness asks whether a statement of law is true. Groundedness asks whether the cited source actually supports it.

Grounded

Key factual propositions make valid references to relevant legal documents.

The right to same sex marriage is protected under the U.S. Constitution. Obergefell v. Hodges, 576 U.S. 644 (2015).

Misgrounded

Key factual propositions are cited, but the source does not support the claim.

The right to same sex marriage is protected under the U.S. Constitution. Miranda v. Arizona, 384 U.S. 436 (1966).

Ungrounded

Key factual propositions are not cited at all.

The right to same sex marriage is protected under the U.S. Constitution. No citation offered.

The three groundedness states, with the researchers' own illustrations. Only the middle one counts as a hallucination, and only the middle one looks like competent work.

A response is hallucinated if it is either incorrect or misgrounded: in the authors' words, "if a model makes a false statement or falsely asserts that a source supports a statement." They call misgrounding potentially more dangerous than outright fabrication, because catching it requires reading each cited source and comparing it to the proposition it supposedly supports.

A fabricated case fails the first check anyone runs. A misgrounded citation survives verification by any reviewer who confirms the case exists and stops there.

System, as tested in May 2024 Accurate Incomplete Note
Lexis+ AI 65% of all queries 18% Highest accuracy of the tools tested
Westlaw AI-Assisted Research 41% of all queries 25% Roughly one third of responses contained a hallucination
Ask Practical Law AI 19% of all queries 62% Draws only from Practical Law, not primary law

Two findings matter more than the headline range. Refusal flatters the metric. Ask Practical Law AI seldom hallucinated, and it also gave incomplete answers 62% of the time. A system that declines to answer cannot assert a falsehood, so any benchmark reporting hallucination without responsiveness rewards silence. Length drives the rate. Westlaw wrote the longest answers, 350 words on average, and longer answers contain more falsifiable propositions. When the unit of measurement is the response, thoroughness is penalized. One more finding cuts against comfortable readings: accuracy was highest on trick questions built around false premises, and lower on the categories representing real attorney work.

Why 17-33% Is Not a 2026 Product Ranking

The authors bound their own result. Testing was a five-day window more than two years ago, on models that have since been replaced, inside closed systems where, as they put it, it is not possible to say precisely where an error occurs. The 202-question set was deliberately hard: bar exam material, circuit splits, doctrine in motion. And it covered chat-style research only. CoCounsel was never evaluated, and the authors name contract review and memoranda as open benchmarking problems. A firm citing this study about contract review is citing something it does not measure.

None of that weakens the finding. It bounds it: strong evidence that RAG did not eliminate hallucination as of May 2024, and that misgrounding is real and hard to spot. Not a leaderboard.

Two Rates, Both Accurate, Not Comparable

Harvey's BigLaw Bench work reports its Assistant at roughly 0.2%, defining a hallucination as a factual claim that can be demonstrably disproven against a source of truth, excluding reasoning errors, and dividing hallucinated sentences by total sentences Harvey, vendor-created. Watch what the unit choice alone does to one unchanged answer.

Counted by response

Stanford / JELS

1 hallucinated response of 1

100%

Counted by sentence

Harvey BigLaw Bench

1 hallucinated sentence of 20

5%

The same 350-word answer. The same single error. Two published metrics, twenty times apart, before anyone disagrees about what counts as a hallucination.

Stanford / JELS Harvey BigLaw Bench
Unit of measurement The response The sentence
Counts misgrounding Yes, central to the definition Not within the stated definition
Counts reasoning errors Correctness failures count Explicitly excluded, tracked separately
Task type Open-ended legal research Reasoning over multiple long documents
Who ran it Independent academic, preregistered Vendor internal, automated with human review
Versions and dates Stated: queries run May 2024 Published Oct 2024. No model versions, no test window

Then add the definitions. Stanford's central case, the real citation that does not support the proposition, is not clearly captured by "demonstrably disproven." Reasoning errors, which Stanford's correctness prong counts, are excluded by design. Both teams documented their methods. The numbers simply answer different questions.

The test to apply before any comparison

Ask whether two hallucination percentages measure the same failure, on the same kind of task, with the same denominator, evaluated by the same kind of party. If any of those differ, the comparison is arithmetic, not evidence.

How Hard This Is, Even Done Well

This is not a criticism of Harvey, which publishes its definition, formula, sample sizes, and process. Most vendors quoting a hallucination figure publish none of that. It is what makes the example useful: this is the careful end of vendor benchmarking, and a firm still cannot safely use the number for comparison. Start with the published figures. Check the arithmetic.

1 in 500 implies 0.20% published as 0.2% Consistent
1 in 150 implies 0.67% published as 0.7% Consistent
1 in 110 implies 0.91% published as 1.9% Cannot both be right

Two of the three published pairs reconcile. The third does not. A rate of 1.9% implies roughly 1 in 53, and a ratio of 1 in 110 implies roughly 0.9%. The accompanying data table lists 1.9%, which suggests the ratio is the slip rather than the percentage.

Almost certainly a typo, and worth exactly one paragraph: it sits in the summary line of the most-quoted vendor hallucination figure in the legal industry, and it has circulated uncorrected because nobody divides. What the disclosure leaves unsettled matters more, and none of it is an error.

What a firm would need What the publication provides Why it matters in 2026
Which models were compared The names "Claude," "ChatGPT" and "Gemini," with no version identifiers The post is dated October 2024. Those names refer to model generations that have since been replaced more than once
When testing ran Not stated. Publication date bounds it at October 7, 2024 Roughly four months after the Stanford queries. Both figures firms treat as "current" are from 2024
What counted as a failure A factual claim demonstrably disproven against a source of truth. Reasoning errors tracked separately Excludes two of the three failure modes the Stanford authors found most common
What is being compared A product with retrieval and workflow scaffolding, against foundation models used directly A reasonable thing to measure, and not a like-for-like model comparison. The scaffolding is the point of the product
Whether length was controlled States the result holds despite longer, more detailed answers Under a per-sentence denominator, longer answers add to numerator and denominator together, so length is largely neutralized already

That is the ordinary condition of vendor benchmarking: a company measuring its own product, in good faith, on marketing's timeline rather than replication's.

If this is what the careful end of vendor benchmarking looks like, no firm should be outsourcing its reliability judgment to published numbers.

Hallucination Is One Failure Mode of Nine

Most of the ways legal AI fails involve inventing nothing. The taxonomy is Songbird's operational framework, informed by the failure types the Stanford team documented.

Failure mode What it looks like Why review often misses it Detectability
1 Fabrication Invented case, quote, statute, or procedural history Rarely missed. Fails the first check anyone runs. Easy to catch
2 Misgrounding Real authority cited for a proposition it does not support Survives any reviewer who confirms the case exists and stops. Hard
3 Reasoning error Sources real and fairly described, conclusion unsound Requires substantive legal judgment, not verification. Hard
4 Material omission Controlling or adverse authority never surfaces Leaves no artifact. Nothing in the output signals the absence. Very hard
5 Currency error Superseded statute, overruled precedent, stale guidance Output looks correct against the source it actually used. Moderate
6 Jurisdiction error Legitimate authority that does not govern this matter Citation checks pass. The case is real and good law elsewhere. Moderate
7 False-premise acceptance Model reasons from an incorrect assumption in the prompt The answer is responsive, which reads as competence. Hard
8 Retrieval failure Generation is sound, retrieval never surfaced the right document Invisible without knowing what should have been retrieved. Very hard
9 Workflow execution error Agent edits the wrong clause, document, or matter Text-level review does not examine actions taken. Hard

A hypothetical makes rows 2 and 4 concrete. On a preemption question, the system returns five real cases, accurately summarized and correctly cited, and never surfaces the controlling sixth case that cuts the other way. Every citation check passes. Nothing in the output is false. The memo is still wrong, and no citation-focused review will catch it.

Reliability Is a Stack, Not a Model

Two products built on the same frontier model can differ substantially, because reliability is produced, or lost, at every layer between the model and the lawyer.

10 Frontier model Language understanding, synthesis, reasoning Which model and version? What happens when it changes?
09 Workflow design Instructions, task decomposition, guardrails Built for this task, or general-purpose with a legal label?
08 Source corpus What the system is able to see at all Primary law? Which jurisdictions? Secondary only? Our documents?
07 Retrieval Finding candidate sources How would we know a relevant authority was never retrieved?
06 Authority selection Ranking and choosing among candidates Does it distinguish controlling from persuasive?
05 Currency and treatment Validity and negative treatment Is citator data integrated, or is the model asserting currency?
04 Synthesis Turning sources into an answer Does it distinguish holding from dicta, court from litigant?
03 Attribution and verification Linking propositions to sources Can we check proposition to source, or only that the case exists?
02 Workflow controls For agentic systems, the actions taken What can it change? What requires approval?
01 Human review The lawyer What is the reviewer checking, and how long does it take?
Largely vendor-determined Shared Firm-determined

"It uses Claude" tells you about one layer and nothing about the other nine. A legal platform can outperform its base model by adding retrieval, authoritative data, and verification. It can also underperform it, with poor retrieval or constraints that suppress useful answers. Both happen.

"Claude vs. Legal AI" Is Converging, Asymmetrically

On May 12, 2026, Anthropic released Claude for Legal: twelve practice-area plugins under an Apache 2.0 license, plus more than twenty MCP connectors into systems firms already run Anthropic. The next day, LexisNexis announced it had integrated those plugins into Lexis+ with Protégé, inside the Lexis environment LexisNexis. The boundary is blurring, and it is blurring in one direction faster than the other.

Platform absorbs model
Claude legal plugins Lexis+ with Protégé

Shipped. The authoritative corpus and the citator stay inside the vendor environment.

Model reaches out
Claude connectors CourtListener · Trellis · Descrybe Westlaw · Practical Law (listed as wanted)

Shipped research connectors reach public dockets and secondary sources. The dominant paid corpora are not in the default set.

Read the repository, not the coverage. The shipped research connectors are CourtListener, Descrybe, Trellis, TopCounsel, Definely, and Solve Intelligence. Thomson Reuters sits under "Wanted connectors," and LexisNexis does not appear in the file at all CONNECTORS.md, whatever the trade press reports. Verify the current state yourself; this is the fastest-moving claim in this article. Out of the box, Claude's legal research runs on public dockets and secondary sources, which is genuinely useful and is not a licensed primary-law platform with citator coverage.

The question is no longer "Claude or Lexis." It is: for this task, which corpus is this system actually reading, and what verifies the link between its assertions and that corpus?

What Current Benchmarks Can and Cannot Tell You

Benchmark What it tests Unit What it can tell you What it cannot
Stanford / JELS
queries run May 2024
Independent
Open-ended legal research, 202 hard questions Response That RAG did not eliminate hallucination, and that misgrounding is common and hard to spot Current product performance. Contract review. Drafting. CoCounsel, which was not tested
Vals VLAIR
Feb 2025
Opt-in
Seven tasks including extraction, document Q&A, summarization, redlining, chronology Task score vs lawyer baseline Comparative task quality among participating vendors Hallucination, which was not the measured outcome. Lexis+ AI withdrew, and vendors chose which tasks to enter
Vals legal research
Oct 2025
Opt-in
210 questions across nine research categories, scored on accuracy, authoritativeness, appropriateness Weighted rubric That tested tools scored 79% to 81% against a 71% lawyer baseline, and that general-purpose ChatGPT scored within one point of the best legal-specific tool Broad market comparison. Thomson Reuters, LexisNexis and vLex all declined to participate. Hallucination was not measured
Harvey BigLaw Bench
published Oct 2024
Vendor
Reasoning over multiple long documents Sentence Harvey's measured rate under Harvey's definition Comparison to Stanford's figure. Reasoning errors, excluded by design. No model versions and no test window, so it cannot be dated or reproduced

No benchmark here measures what most firms think it measures. Two of the four do not measure hallucination at all, the most rigorous one is two years old, and the lowest number is vendor-produced under a definition that excludes reasoning errors. The October 2025 Vals result is the strongest published evidence of convergence, with tested tools at 79% to 81% against a 71% lawyer baseline and a general-purpose model within a point of the best legal-specific tool Vals 2025. Its caveat is structural: Thomson Reuters, LexisNexis, and vLex all declined to participate VLAIR. A benchmark the incumbents skip describes the participants, not the market.

The Metric That Belongs in the Buying Decision

Verification burden is the time, expertise, and source access a qualified reviewer needs to determine that an output is safe to rely on. The framing is Songbird's; the observation is the Stanford authors', who noted that longer answers require proportionally more checking, proposition by proposition. Operational risk compounds three ways.

01  Error probability

How often the system produces a material error on this kind of task.

02  Detection failure

How often the firm's review process fails to catch it. This is where tool design has the most leverage, and where demos reveal nothing.

03  Consequence

What happens if the error survives into client work or a filing.

Procurement usually prices only the first. Picture two tools that save the same hour: Tool A links every proposition to its supporting passage and flags what it could not support, while Tool B writes better prose that must be reconstructed to check. Tool B can win the benchmark while Tool A is the lower-risk purchase. Slightly more accurate but substantially harder to verify is operationally worse.

Run Your Own Evaluation

Vendor benchmarks answer the vendor's question. Build 30 to 50 tasks from closed matters where the right answer is already known, across four tiers.

Routine

Summarization, extraction, basic drafting.

Substantive

Research, multi-document synthesis, contract interpretation, chronology building.

Difficult

Conflicting authority, jurisdiction-sensitive questions, recent developments, large document sets.

Adversarial

False premises, requests for authority that does not exist, ambiguous jurisdiction, superseded law, deliberately incomplete document sets.

Score each dimension separately; composites hide failure modes. The three starred below are the ones firms most often skip, and where the hardest-to-detect failures live.

Dimension What to test What good looks like Warning sign
Substantive correctness Compare against the known outcome Errors rare and legally minor Fluent answers that are subtly wrong
Citation existence Does every authority exist as cited No fabrications Any fabrication at all
Citation support Read the source. Does it support the proposition Propositions traceable to specific passages You must read the full case to tell
Authority quality Controlling versus persuasive versus secondary Hierarchy respected and explained Treats all sources as equivalent
Completeness Seed a task with known adverse authority Surfaces it, or flags a gap Confident answer, missing authority
Currency Include a superseded provision Flags supersession Silently relies on stale law
Jurisdiction Ask a question governed by one state Stays in jurisdiction or says it cannot Cites real law from elsewhere
False-premise resistance Embed a wrong assumption in the prompt Corrects the premise Reasons helpfully from the error
Uncertainty calibration Ask something genuinely unsettled Says so Manufactures a clean answer
Reproducibility Run the same task three times Materially consistent Different conclusions per run
Workflow correctness For agents, check actions, not just text Right document, right clause, logged Correct text, wrong target
Provenance Can you reconstruct what it saw Sources, dates, retrieval visible Conclusions without a trail
Reviewer effort Time a qualified reviewer to sign-off Minutes, with sources at hand Reviewer rebuilds the research

Set thresholds by workflow and consequence, not from published benchmarks, and record who tested, which product and model version, and when, because the answer expires. The recurring mistakes are consistent: treating a 2024 percentage as a current specification, comparing rates with different denominators, checking that citations exist but not that they support the proposition, testing only vendor-supplied examples, evaluating a model when the firm is buying a workflow, and never retesting after the product changes. ABA Formal Opinion 512 makes verification task-dependent and treats uncritical reliance as malpractice ABA FO 512, and that cuts both ways: it forecloses blanket bans justified by a stale statistic as firmly as it forecloses uncritical adoption. State rules vary Model Rules.

Procure for Verification, Not Infallibility

No system will be reliably right about the law. The standard is a system in which material errors are uncommon, the failure modes are named, the evidence is visible, citations can be checked against the propositions they support, and a lawyer can find the remaining errors in an amount of time the matter can absorb. Firms that evaluate this way can adopt aggressively, because they know what they are watching for. Firms that buy on a hallucination percentage have bought a number whose denominator they did not check.

Take the last AI-assisted research memo your firm produced, and ask the reviewing lawyer two questions.

Which propositions did you verify against the source text, rather than confirming the case exists?

How would you have known if the controlling case never appeared in the output at all?

If neither has a confident answer, the firm has an adoption process, not a reliability process.

What This Connects To

This article covers evaluation. The rest of the cluster covers what surrounds it. ai-tool-due-diligence covers the vendor review that precedes a pilot. how-to-run-a-legal-tech-pilot covers the structure the evaluation set belongs inside. governing-claude-legal-work covers the workflow governance a firm needs once a tool is approved, including task-specific review standards. privilege-confidentiality-ai covers the confidentiality analysis that governs what may enter the system in the first place. firm-ai-policy covers the policy structure. The Legal AI Matrix covers the product landscape.

This article is operational guidance and does not constitute legal advice or ethics guidance. ABA Model Rules and formal opinions are interpretive; state rules and ethics opinions vary, and jurisdictions may impose different requirements. Product capabilities and benchmark results described here reflect published sources as of August 2026 and change, in some cases quickly. Songbird Strategies is a legal technology consulting firm, not a law firm, and has not independently benchmarked any product named in this article. Firms should consult qualified ethics counsel on jurisdiction-specific questions. See Sources & Notes for every benchmark and vendor claim cited, including which are vendor-produced.

Can Your Firm Explain Why Its AI Output Is Trustworthy?

If your AI evaluation consists of vendor demos, benchmark headlines and informal attorney testing, you do not yet have a reliability evaluation. A useful assessment runs the system against your firm's actual workflows, measures the failure modes that matter for that work, and defines how the output will be verified in practice and by whom.

See the Legal AI Matrix →
Book a Free Strategy Call

30 minutes. No sales pitch.