Back to Blog

The Real ROI of Legal AI: Measure Time to Usable Work

A practical method for measuring legal AI ROI across generation, review, correction, omissions, implementation and reuse.

A practical method for measuring legal AI ROI across generation, review, correction, omissions, implementation and reuse.

A draft produced in thirty seconds can still take two hours to verify. That makes "time to first output" a poor measure of legal AI return on investment. The useful question is how long it takes the team to reach work it is prepared to use. The difference can be large. A fast answer may create extra source review, correction and rework. A slower first pass may reduce the total effort if its scope and evidence are easier to inspect.

A better unit: time to usable work

Mary proposes measuring: Time to usable work = generation + source review + correction + omitted-material review + rework + implementation overhead - repeated work avoided. The formula is simple. Applying it requires the firm to define each component for the task being tested: Generation: preparing inputs, writing instructions and waiting for the output. Source review: checking material propositions against the underlying record. Correction: fixing factual, legal and formatting errors. Omitted-material review: determining what the product failed to surface. Rework: regenerating, redrafting or repairing downstream work affected by an earlier error. Implementation overhead: configuration, training, integration and support during the measured period. Repeated work avoided: time saved later because reviewed facts, corrections or workflow settings persist and can be reused. The last component is easy to miss. A factual record may take time to establish and then reduce repeated review across several later tasks. Measuring only the first output can understate that benefit.

Define the task narrowly

ROI cannot be measured credibly across "legal work" as one category. A chronology, a legal research memo, a bank analysis and a first draft have different inputs and review requirements. The same tool may create a strong return on one task and none on another. For each use case, record: the starting material. the expected output. who performs the work. the manual baseline. the review standard. the period over which reuse is measured. A baseline should be based on actual work where possible. Asking lawyers how long a task "usually takes" can produce inconsistent estimates. A closed matter or repeated workflow gives the firm a stronger comparison.

Measure review rather than assuming it

Many AI business cases treat human review as a constant that applies equally to every product. It is not. A source-linked factual output may take less time to review than a polished narrative with no visible scope. A product that exposes failed documents and missing material may make omitted-material review faster. Another product may generate a better first draft while requiring a lawyer to reconstruct the evidence behind it. Record review time separately from generation time. Also record who performs it. Ten minutes of partner review and ten minutes of paralegal checking have different costs and may reflect different risk. Morae's 2026 survey of 850 senior legal professionals found that 67% were concerned the cost of human verification could outweigh AI's productivity benefits, while 48% said AI outputs were materially fixed before use. Those figures do not prove that legal AI has poor ROI. They show why review has to appear in the calculation.

Include omissions

An output can be quick, accurate and incomplete. If the lawyer has to reopen the full matter to determine whether anything material was missed, much of the work remains. A factual workflow should record: material facts in the reference record. facts surfaced by the product. unsupported facts. sources the product did not use. expected documents or periods that were missing. time spent checking the negative space. This is where the known-matter test becomes useful. The firm can measure both the quality of the surfaced work and the effort required to establish what was absent.

Separate one-off setup from ongoing use

A pilot may contain costs that should not be assigned to every future matter. Initial integration, security review and workflow design may be substantial. The firm should record them, then decide over what period and number of matters they will be amortised. Training also changes over time. A first group of users may need close support. Later users may benefit from established templates and internal knowledge. Do not hide those costs. Do not assume they repeat forever. The Legal Ops pilot guide provides a structure for recording setup, user behaviour and support requirements during a controlled test.

Measure reuse across the matter

Legal AI ROI often compounds downstream. A reviewed factual record built during early case assessment can support a chronology, witness preparation, requests for admission, mediation and drafting. If each task begins from the approved record, the team avoids reconstructing the same people, dates and events. The saving should be measured where it occurs. For example, the first task may save only one hour after review. Four later tasks may each avoid another hour of factual reconstruction. The total return belongs to the workflow, not solely to the first artefact. The reverse is also true. An early factual error can contaminate several outputs and create later rework. A useful ROI model includes that cost rather than counting each fast generation as an independent success.

Do not convert every saved hour into revenue

Time saved can create several forms of value: more matters handled by the same team. faster response to clients or courts. lower cost on fixed-fee or contingency work. more time for strategy and client contact. reduced weekend or deadline work. improved consistency across volume matters. A saved hour is not automatically an additional billed hour. ABA Formal Opinion 512 also makes a separate professional point for US lawyers: an hourly lawyer may charge only for actual time spent, even where AI makes the work faster. The economic value may appear in capacity, margin, pricing or client retention rather than a direct increase in billed time. The firm should choose the value measure that matches its business model.

Use a simple scorecard

For each tested task, capture: Manual process total time. role mix. error and rework observed. usable output standard. AI-assisted process setup and generation time. source and omission review. corrections and regeneration. time to approval. later work avoided. Outcome net time difference. cost difference. quality or coverage difference. adoption and support burden. conditions under which the result is likely to repeat. The scorecard should keep quality beside time. A faster process that misses a material fact has not delivered a useful saving.

Published results should name the task

Claims such as "saves 80%" are hard to assess without the task, sample, baseline and review method. A stronger result says which firm performed which work, on what type of matter, and what was included in the measured time. The National Compensation Lawyers case study is one example. The published claim is attached to a 12-person personal injury firm and its document-review workflow rather than presented as a universal result for all legal work. The calculation can be less dramatic and more useful when every part of the workflow is visible.

Related reading

Notes and sources