Back to Blog

When a Real Source Still Produces a Wrong Legal AI Answer

A practical taxonomy for legal AI errors that survive a basic source check, including misquotation, mischaracterisation, transformation and incomplete coverage.

A practical taxonomy for legal AI errors that survive a basic source check, including misquotation, mischaracterisation, transformation and incomplete coverage.

A source can exist, open to the right page and still fail to support the sentence beside it. That failure is harder to catch than an invented case because the first check succeeds. The authority or document is real. The reviewer can find it. The error appears only when the reviewer compares the precise output with the source and the rest of the record. The familiar category of "hallucination" is too broad to explain what went wrong or which control would catch it.

Five different source failures

Mary uses five categories to separate these errors.

1. Existence failure

The cited authority, document or quotation does not exist. This is the conventional fabricated-case problem. An existence check should expose it. Legal research platforms, citators and the court record can confirm whether the authority is real. The control is necessary and comparatively straightforward.

2. Quotation failure

The source exists, but the words attributed to it are inaccurate. The system may remove a qualification, join separate passages, alter a number or produce a sentence that is not present. A reviewer who checks only the case name and citation will miss the problem. The control is direct quotation comparison against the source text.

3. Proposition failure

The quoted words are present, but they do not support the proposition for which they are cited. A judgment may recite a party's submission without adopting it. A witness statement may record an allegation. A contract clause may apply only after a condition has been satisfied. The source is genuine and the quotation may be exact, yet the legal or factual characterisation is wrong. The reviewer has to read the surrounding passage and understand the role the statement plays in the source.

4. Transformation failure

The system receives the correct information and changes it while producing the next output. Matthew Dahl and Eric Martínez use a striking example in Bye-bye, Bluebook? A lawyer had found and verified a source and supplied the chatbot with the correct title, year and link. Asked to format the citation, the model changed the title and authors. The error entered after research and verification. The source did not need to be found again. The model altered information already in context.

Transformation failures can occur outside citations. A summary may change which person attended a meeting. A chronology may convert an inferred date into a stated date. A draft may turn "reported by the claimant" into an unqualified fact. The control is a comparison between the structured source information and the generated work product, ideally supported by deterministic checks where the transformation follows explicit rules.

5. Coverage failure

The cited source supports the sentence, but the answer is incomplete because other material changes the conclusion. An email may support that notice was sent. A delivery failure message may show it was not received. The first citation is valid. The matter-level answer is still misleading if the second document never surfaced. Coverage failure is difficult because reviewing the displayed source cannot reveal material the system did not use.

The control requires visibility into the scope of the review and a test for material omissions.

Why one "grounded" label is not enough

Grounding usually means that the model has access to external material or attaches sources to its output. That can reduce fabrication and make review faster. It does not prove that the system quoted the source accurately, characterised it correctly, preserved the supplied details or considered the material needed for a complete answer. The word describes an architectural feature. It does not describe the result of every factual or legal check.

A product evaluation should therefore ask several separate questions:

Each question tests a different failure mode.

The Bluebook result is useful beyond citation formatting

The Bluebook article sets out the full study and its benchmark results. One architectural result is relevant here. The researchers separated extraction from rule application. The model classified and structured the source information, while deterministic software applied the citation rules. That reduced one class of error because the model was no longer responsible for every step. The design does not eliminate extraction failure. If the model puts the wrong author or source type into the structured fields, the formatter will apply the rule to bad inputs.

It does show why failure categories matter. A system can apply different controls at different points rather than asking the same model to generate and approve the final answer.

Review should follow the claim

The right check depends on the work product. For a direct quotation, compare the words and surrounding passage. For a legal proposition, confirm that the authority adopts the rule being stated and that later history or jurisdiction does not change its use. For a factual proposition, confirm the actor, date, event and evidentiary status. Then inspect whether another source conflicts with it. For a transformation such as a citation, calculation or form field, compare the output with the structured input and use ordinary software for rules that can be encoded reliably.

For a whole-matter answer, test material-fact recall and inspect documents the system did not use. "Human review" does not specify any of those steps. The product should make the appropriate review path visible.

A real source can make the error look safer

Fabricated authorities attract attention because the failure is obvious once the citation is checked. A mischaracterised real source can look like routine legal work. The case name is familiar. The document opens. The page contains related language. The reviewer may stop before reaching the precise discrepancy. That makes interface design relevant. A source link should open to the exact passage with enough context to assess the claim. The user should also be able to see whether the statement is an allegation, quotation, inference or reviewed fact.

The verification whitepaper sets out the broader controls. The vendor-responsibility article explains why the burden cannot be solved by telling lawyers to check harder.

How to test for these errors

A closed-matter test should deliberately include examples from each category. Include:

The purpose is not to trick the product. It is to discover which controls operate before the firm relies on the same workflow in a live matter. The source may be real. The review still has to ask what the system did with it and what remained outside the answer.

Related reading

Notes and sources