
Verification in Legal AI Is a Design Problem
Lawyers remain responsible. Legal AI should make its work easy to check before they sign.

The easiest legal AI demonstration to pass is one where the buyer does not know the answer.
A polished result can look impressive when nobody in the room knows which documents matter, which facts are missing or how long a proper review should take. A closed matter gives the firm something much more useful: a reference point.
The team already knows what happened. It knows the difficult documents, the unresolved conflicts and the mistakes a system is likely to make. That makes the matter suitable for a controlled test.
Use a closed, de-identified or internally approved matter that is representative of the work the firm wants the product to perform.
The matter should be large enough to expose cross-document problems, but not so large that the firm cannot build a credible reference record. It should contain at least one factual conflict, one important item that is easy to overlook, and a question the team can answer from its own prior work.
Avoid a matter supplied by the vendor. A vendor-selected example may be useful for learning the interface, but it does not give the buyer an independent basis for assessing the result.
The test should also be task-specific. "Understand the matter" is too broad. A stronger instruction might be:
The firm's reference record should be completed before the products are run.
It does not need to contain every fact in the matter. It needs a defensible list of the facts that are material to the chosen task, the sources supporting them, and the known conflicts or gaps.
Record the basis for materiality. A fact may be material because it changes an element of a claim, affects a limitation issue, contradicts a witness, alters the damages position or points to further investigation. Writing that down prevents the team from changing the answer key after seeing which product performed best.
The reference record is still a human work product. It may contain mistakes or omit something the product finds. If that happens, the team should review the proposed addition and update the reference record where justified. The goal is a fair evaluation, not defending the first human answer at all costs.
Give each product the same documents, the same task and the same amount of contextual information.
Record:
A product that requires different inputs may still be valuable. The difference should be part of the result rather than hidden in the setup.
Run the initial test without revealing the private reference record. Known contradictions and omissions should remain concealed until the output is assessed.
A useful test should measure more than whether the final prose reads well:
Those measures should be kept separate. A product can have high precision and poor recall. It may avoid unsupported claims by returning very little. Another product may surface more of the matter while creating more review work. One headline percentage will obscure that difference.
Most product demonstrations focus on what appeared in the answer. The known-matter test should also examine what did not.
Compare the product's factual record against the reference record and classify each missing item. Was the source unread? Was the fact buried in an attachment? Did the system find the document but fail to treat the passage as material? Did the instruction exclude it?
The remedy depends on the cause. A file-processing failure, an extraction failure and a materiality judgment are different product problems.
Ask the reviewers to inspect a sample of documents that the product did not use. A list of citations shows where an answer came from. It does not establish that the system considered the rest of the matter.
The first run tests extraction and output. The second run tests whether the product can carry reviewed work forward.
Correct several known errors or classifications. Then ask for a different work product that depends on the same facts, such as a witness brief, deposition questions or a summary of the evidence on one issue.
Record whether the product:
A system that requires the lawyer to restate every correction is providing generation, not persistent matter state.
Do not average unrelated activities into one vendor score.
A legal research tool should be tested on legal research. An eDiscovery platform should be tested on the relevant collection and review task. A factual-record product should be tested on whole-matter factual coverage and reviewability. A general drafting model should be tested on a controlled drafting task with defined inputs.
The same firm may choose different products for different jobs.
The Avoiding the Hype Trap article covers the broader vendor questions around security, support and implementation. The Legal Ops pilot guide covers ownership, users and rollout. The known-matter test supplies the evidentiary test inside that pilot.
A closed matter is not a perfect prediction of future performance.
The product may behave differently on another practice area, language, document quality or volume. The firm's reference record may reflect the lawyers' own interpretation of materiality. A single test cannot establish a universal accuracy rate.
It can establish something narrower and more useful: how the product behaved on a real task where the firm could identify both the surfaced errors and the missing work.
NIST's work on AI testing and evaluation repeatedly stresses that measurement depends on the context in which a system is used. The same principle applies here. The firm should test the product it intends to buy, on the work it intends to give it, with the review process it will actually use.