You don’t know from the vendor’s claim. You know from a 50-query acceptance test you design, run, and score yourself — because every published measurement of the grounded legal research tools has come in below what their marketing said. The Stanford RegLab study behind those numbers ran 202 preregistered queries against the two flagship platforms and found Lexis+ AI answered 65% accurately while hallucinating on 17% of queries, and Westlaw AI-Assisted Research answered 41% accurately while hallucinating on 33%. Both were sold at the time as hallucination-free. Both retrieve from the exact case-law corpora their marketing points to.
That is the whole problem in one sentence: retrieval grounding reduces fabrication, and does not eliminate hallucination. This page is the test protocol to run before you sign.
What “grounded” means, and what it does not mean
Grounded means every load-bearing proposition in the answer traces to a document the system actually retrieved and read, not to the model’s training memory. The Stanford typology splits responses three ways, and the middle category is where buyers get hurt:
- Grounded — the key factual propositions make valid references to relevant legal documents.
- Ungrounded — the key factual propositions carry no citation at all.
- Misgrounded — the propositions are cited, but the source does not support the claim, or the source is inapplicable to the jurisdiction asked about.
A response counts as a hallucination if it is incorrect or misgrounded. That second half is the one vendors leave out.
So here is what grounding is not. It is not a working hyperlink. A citation that resolves to a real case in a real database, with a live link, a correct reporter cite, and a real quotation, is still a hallucination if the case does not stand for the proposition the answer attached it to. Link-resolution is machine-checkable and every serious vendor now passes it. Proposition-support is not machine-checkable, and that is where the residual error lives. Any pilot that scores “did the link work?” is measuring the thing that was already fixed.
The five failure modes
| Mode | What happened | Who can catch it |
|---|---|---|
| Fabricated authority | The case or statute does not exist | Machine — citator lookup |
| Fabricated quotation | Real case, invented quoted language | Machine — string match against source |
| Misgrounding | Real case, real quote, does not support the proposition | Lawyer — read the pin cite |
| Jurisdictional misgrounding | Correct proposition, wrong circuit or state | Lawyer — check the court against the question |
| Stale authority | Real and on point, but overruled, superseded, or distinguished into irrelevance | Machine flag, lawyer judgment |
Modes 1 and 2 are the ones vendors mean when they say “hallucination-free.” Modes 3 through 5 are the ones that lose a motion. Budget your evaluation time accordingly: the cheap checks are already automated, so spend your lawyer hours on pin-cite reading.
The protocol: a 50-query acceptance test
Build the query set yourself, from your own matters. Never let the vendor supply it — a demo set is tuned to the retrieval corpus, and the whole point of the exercise is to find the edges the vendor did not tune for. Compose the 50 queries in these proportions, which mirror the Stanford dataset’s structure:
- 20 general research questions in the practice areas you actually staff.
- 15 jurisdiction- or date-specific questions, where the correct answer changes by circuit or by year. This is the category that separates a retriever with real metadata filters from one doing semantic similarity.
- 10 false-premise questions — ask about a doctrine that does not exist, or attribute a holding to a case that does not contain it. This is where sycophancy surfaces. A tool that constructs support for your false premise will do the same for a partner’s.
- 5 factual-recall questions with one verifiable answer (a holding, a date, a vote count).
Score every response on three binary axes — correct, grounded, complete — and apply these thresholds:
| Metric | Pass threshold | Why this number |
|---|---|---|
| Fabricated citations or quotes | 0 of 50 | Non-negotiable. One fabrication in 50 predicts a fabrication in a filing. |
| Misgrounding rate | ≤ 2 of 50 (4%) | Below the 17% floor the best measured tool posted, and low enough that a review SOP can absorb it. |
| False-premise pushback | ≥ 8 of 10 | The tool must reject the premise, not build support for it. |
| Incomplete or refused | ≤ 25% | Match the better of the two flagships rather than the worse. |
That last row is the calibration nobody publishes. Accuracy and refusal trade off, and vendors buy accuracy with silence. In the Stanford run, incomplete-response rates were 18% for Lexis+ AI, 25% for Westlaw AI-Assisted Research, and 62% for Ask Practical Law AI — the last of which also posted the lowest accuracy, 19%. Score accuracy without scoring refusal and you will rank a tool that declines to answer above a tool that answers.
Worked example: what the published numbers buy you
Normalize the three measured tools to a 50-query run and the purchase decision gets concrete.
| Tool | Usable answers | Hallucinated | Incomplete |
|---|---|---|---|
| Lexis+ AI | ~33 | ~9 | ~9 |
| Westlaw AI-Assisted Research | ~21 | ~17 | ~13 |
| Ask Practical Law AI | ~10 | ~9 | ~31 |
The middle column is your real cost. Seventeen hallucinated answers out of 50 does not mean 17 wasted queries — it means all 50 answers now require a lawyer to verify, because you cannot tell which 17 they are from the output. That is the economics of a sub-95% tool: verification cost scales with total volume, not with error rate.
Price it before the pilot ends. Have one lawyer verify 10 answers end-to-end against the pin cites and time it, then multiply by five. If that number exceeds the research hours the tool is supposed to save, the tool is negative-value at any license price, and you should say so in the renewal memo rather than a year later.
Five questions to put to the vendor
- Does each citation link to a document your retriever actually read, or to a search result generated after the fact?
- Is the citator run on every citation before the answer renders, or only when I click through?
- What is the corpus cutoff, and what does the system do when the answer requires authority newer than it?
- What is your refusal rate on queries you cannot ground, and will you show it to me?
- Can I see the retrieved documents alongside the answer?
A vendor who declines question 5 has removed your ability to audit misgrounding. Treat that as disqualifying for litigation work, whatever the demo looked like.
Pitfalls and guards
- Scoring link resolution instead of proposition support. Guard: the scorer reads the pin cite, not the link target.
- Testing on questions with famous answers. Guard: at least 15 of 50 queries jurisdiction- or date-specific.
- Letting the vendor supply the query set. Guard: queries come from closed matters your team already knows the answers to.
- Treating “grounded in Westlaw content” as an accuracy claim. Guard: both measured platforms hallucinated while grounded in exactly that content.
- Delegating verification to the tool’s own citation checker. Guard: separate the generator from the checker; a self-graded exam is not evidence.
- Skipping the false-premise queries because they feel unfair. Guard: 10 of 50, minimum — opposing counsel will not be fair either.
- Treating verification as an engineering problem. Guard: ABA Formal Opinion 512 (July 2024) puts independent verification of AI output on the lawyer under Model Rules 1.1 and 3.3. It is not delegable to a vendor’s SLA. Write it into the AI policy for legal teams and the contract review SOP rather than the procurement checklist.
The stakes are documented rather than theoretical. Damien Charlotin’s AI Hallucination Cases database recorded 1,668 court decisions worldwide involving AI-fabricated material as of 2 July 2026, 1,163 of them in the United States. A practicing lawyer was the responsible party in 653 of them. This is not a self-represented-litigant problem.
Where this framework breaks down
The 50-query protocol tests authority-citing research output. It does not transfer unchanged to three adjacent cases:
- Document-grounded work — contract review, extraction, diligence. There is no citator to run. Substitute exact-span provenance: the tool must highlight the source span in the source document, and you score whether the span supports the extraction. See contract data extraction for that scoring shape.
- Closed-corpus assistants over your own DMS or matter files. Same substitution, plus a permissions test — ask a question whose answer sits behind a privilege boundary and confirm the retrieval declines it.
- Post-purchase monitoring. This is a gate, not a regime. Vendors swap the underlying model without notice, and a model swap invalidates your measurement. Re-run the 50 queries at every announced model change and at every renewal.
Related
- AI policy for legal teams — where the verification duty gets written down
- Legal AI vs legaltech — the category boundary this evaluation sits inside
- Harvey, Thomson Reuters CoCounsel, LexisNexis Protégé — the platforms to run the protocol against
- Contract review SOP — the review layer that absorbs a non-zero misgrounding rate