Back to Blog

The Perfect AI Is Actually a Combination of Tools: A Litigation Guide

No single AI tool covers litigation end to end - legal research, discovery, factual analysis, work product, and trial each need a differently-tuned tool, and treating them as one job contributes to hallucinations. This article provides a practitioners guide to choosing the right tool for the task at hand.

No single AI tool covers litigation end to end - legal research, discovery, factual analysis, work product, and trial each need a differently-tuned tool, and treating them as one job contributes to hallucinations. This article provides a practitioners guide to choosing the right tool for the task at hand.

This week, we hosted a webinar 'AI in Litigation: The Promise, the Pitfalls, the Sanctions' with litigators across the US who are using AI in their daily practice. While we unpacked several important themes, the top question that came up: Which AI tool do I use for each task?

Practitioners inherently understand that there are different tools better suited to different jobs (contrary to what many vendors convey). For many firms, that lesson is learned the hard way, through failure to live up to promises or gaps and inaccuracies. As one litigator put it: “In my mind, the perfect AI tool right now is actually a combination of tools.” - Catherine Thompson, AI in Litigation webinar, July 2026

Making sure you have the right tool for the task is critical. Not only do we continue to see a growing number of sanctions for hallucinated outputs in court (over 1,600 cases and counting according to Damien Charlotin’s AI hallucination tracker), reports are now emerging that challenge the time saving thesis. A report from Glean's Work AI Institute (covered by CIO Dive), that show companies aren’t realizing the gains from AI, and one of the biggest culprits is shadow work. Of the ~11 hours saved per employee, they are spending 6.4 hours maintaining the AI to ensure accurate outputs - through prompt engineering and review work.

So why can’t one tool do it all?

As we unpacked in our previous blog ‘Getting started safely’, the tuning of models for different tasks is a key consideration for where they best apply. Tuning refers to temperature settings in AI which determine the degree of creativity in responses.

For tasks like legal research and factual analysis, purpose built tools with a lower default temperature setting are critical. For these tasks, a narrower scope but more predictable, accurate and transparent set of results is required.

For more creative tasks, like brainstorming arguments, alternative framing, and first pass drafts, a higher default temperature setting may be desirable, and this is where general AI tools are adding value.

Selecting the right tool for the task is especially important given the risks of cognitive surrender, a phenomenon where the authoritativeness of an AI answer, exacerbated under time pressure and cognitive overload, sees people accept an incorrect AI answer 80% of the time (Shaw and Nave, 2026). It’s why despite warnings we continue to see hallucinated outputs appearing in courts.

This is a prescient consideration in the complex litigation practice areas Mary specializes in working with. In the more transactional areas of the law - in-depth and rigorous factual grounding may be less relevant. This blog is our ode to those working in these areas!

Which tools should I be using across my stack?

So which tools should litigation firms be using across the practice? There are fundamentally three types of activities - practice management, case management and document management, where different tools may be integrated.

In practice management, we are seeing specialization by practice area. In Family law, Smokeball, Clio and MyCase are dominant across most firms. For Am Law firms, frequently the biggest challenge is billing so we are seeing specialised billing platforms dominate at a practice level e.g. Aderant, Intapp.

In document management it depends on the size of the firm, we see many SME firms well served by Microsoft or Google; and work with Netdocs and Imanage for larger firms.

Case management is where the vast array of developments have occurred in the AI landscape as it is where the highest complexity lives so this is where we are going to focus our time.

Diving into case management

Within case management, the jobs to be done are:

  1. Legal research
  2. Identifying relevant documents
  3. Granular factual analysis, issues, evidence and weakness assessment
  4. Work product development
  5. Verification workflows
  6. Brainstorming, alternative framings, developing templates
  7. Litigation trial tools

Legal research:

Considerations for a tool here are accuracy and exhaustiveness. Dominant players in this space are LexisNexis and Westlaw. These tools are dominant as they are purpose built and tuned specifically to handle the problem of legal research. Grounding here is critical - especially in light of the hallucinated cases we are seeing in courts. While these tools are grounded, Dahl et al 2024 found these tools experienced hallucination at rates between 17% and 33%, while still significantly down on general AI tools, this demonstrates having a purpose built tool is not enough to ensure accuracy.

“Whenever I’m using something that confidently tells me a case says something, I check it myself, even when I get it from CoCounsel - the Westlaw legal research tool.” - Catherine Thompson, AI in Litigation webinar, July 2026

Relevant documents:

The solution required here depends on document volume. For customers dealing with millions of records, modern eDiscovery is required to narrow down the evidence set from millions to hundreds of thousands. Relativity, Reveal, Everlaw, DISCO, and Nuix are some of the dominant players in this space - with purpose built capability to handle a wide context window. For family law, personal injury and torts where the window is hundreds of thousands as opposed to millions, a tool like Mary which drives factual analysis on a narrower context window may be all that is required. The key deciding factor is how large the typical document sets are.

Granular factual analysis and discovery:

Considerations for a tool here are accuracy and exhaustiveness. Meaningful, granular factual analysis in litigation (certainly the complex areas Mary specialises in) remains a task completed exclusively by people. This stage in litigation involves the painstaking determination of issues in the case, mapping the body of evidence accurately and exhaustively against these issues and building appropriate causation narratives and case strategies. Once the issues mapping and causation narratives are complete, these are used to drive discovery requests and responses, and artefacts ranging from deposition questions to interrogatories to legal advice. As more advanced workflows in litigation are built, the dependency on an accurate and exhaustive factual harness at the case layer grows - and the risk of not having this increases as lawyers are abstracted away from the detail.

Accurate and exhaustive factual analysis is a dependency for correct and just downstream outcomes. Work completed at this stage needs to be robustly reviewed by the right people, with the right level of friction at the right time.

While providers outside of the fact management space offer tools like AI queries and chronology building, due to the context window issue in eDiscovery; and the non-exhaustive approach taken in general AI and some general legal AI, these tools fall short on both exhaustiveness and accuracy. Limitations with existing toolsets are one of the primary reasons Mary now works with some of the largest law firms in the world. One litigator described the role of Mary in practice:

“The case involved 21 boxes of material. How do I synthesize that? AI did that for me. Not only did it allow me to do that timeline and create the history, but it also gave me arguments - what’s my argument in this case, and what’s the strength of the other side’s case. And it also told me what was missing. I’m talking 10 to 15 minutes was the time it took for that program to make those conclusions and provide me with that resource.” - Seth Goldstein, Law Offices of Seth L. Goldstein, AI in Litigation webinar, July 2026

Work product development:

Considerations for a tool here are the nature of the work product you need and size of the firm. This is one of the areas that has seen the greatest convergence across the landscape. As work product is required across the litigation lifecycle, most tools in the ecosystem have built this capability.

For the largest firms in the world, many use Harvey, Legora or an enterprise AI model for this support. Where the inputs are dependent on granular, factual analysis - these firms also use Mary. The right tool for the work product depends on the level of precision required and grounding in the evidence. The greater the required grounding, the more likely a tool like Mary will be best suited.

There are specific work products that have a higher accuracy bar and are much harder for general tools to produce. This might include financial analysis or the production of bank statement indices to substantiate patterns accurately. Here, purpose built tools like Mary can complement a general toolset.

For smaller firms with a higher dependency on factual outputs such as matrimonial, abuse or PI, and probate - a purpose built tool like Mary may be complete and sufficient.

Verification workflows:

While Mary was built for the dependency in litigation, the reason some of the biggest firms in the world are staying with Mary is because of Mary’s verification workflows, grounding and what we call productive friction. Irrespective of the tools you integrate across the stack, high quality litigation outputs depend on the right people engaging at the correct level of granularity - which necessitates a verification workflow.

“A tool that builds in telling me - hey, this is the answer to your question. These are the things I’m 100% sure of. These are the things I’m 80% sure of. These are the things I’m 50% sure of. These are the things that sound right, but I’m not sure of at all. This is where I got the information.” - Catherine Thompson, AI in Litigation webinar, July 2026

Brainstorming, alternative framings, developing templates:

Considerations for a tool here are flexibility and size of firm. General AI and legal AI are built specifically to expedite and enhance the ‘sparring partner’ experience core to the socratic method, and can be used exceptionally well in this space. Leading practitioners are sharpening arguments and positioning through meaningful tension and debate with these tools. Where the debate is grounded in the evidence, leading Mary users are also using the platform in this way - exposing gaps, weaknesses and contradictions in their own and the counter-parties case.

Caution needs to be applied when using general purpose tools. Free versions typically train on uploaded data, and we have seen court rulings that information uploaded into a free general purpose tool was not privileged.

Litigation trial tools:

To choose a tool here really means that trial is a core part of your business model. This is a burgeoning space with a raft of purpose built tools emerging to enhance trial outcomes. Judge and court intelligence from Trellis, Lex Machina, a number of mock trial tools, and courtroom presentation tools.

For the smallest practices - a tool like Mary combined with an appropriate legal research approach is more than sufficient to cover the end to end litigation requirements. Larger firms are increasingly adding tools from the breadth of the ecosystem. To drive meaningful gains from a broad array of tools in large firms is requiring deep consideration of future workflows and is seeing many vendors (like Mary) building MCP capabilities to easily integrate agnostic of the broader stack. Leaders in the space are building with Mary as the factual harness and verification layer across their workflows.

Conclusion

So there you have it - our map of the considerations and right tools for the job. The larger the firm, the broader the suite of tools we are seeing built out in the ecosystem, whereas smaller firms are choosing 2-3 key suppliers based on where their need is greatest. Where that is factual grounding, Mary is helping across a broader set of needs.