The newest independent benchmark found AI beating lawyers on legal research accuracy, general models included. It also left out the one task that gets lawyers sanctioned. What that means for where a firm should begin.
AI Answer Summary
In an independent benchmark published in October 2025, AI tools averaged about 80% accuracy on 200 legal research questions, against 71% for a control group of practising lawyers. A general-purpose model, ChatGPT, matched the legal-specific tools on accuracy.
The benchmark excluded generating formatted citations, which is the failure behind court sanctions for fabricated authority. Where the legal-specific tools did pull ahead was in the quality of the sources they cited, not in accuracy.
So legal research is no longer the weak category it was taken to be in 2024. It is still the wrong first project for a law firm, and the reasons have little to do with accuracy: what a single error costs, the verification that has to happen regardless, and the fact that research does not change how the firm's work flows.
Is AI Accurate Enough for Legal Research? What the Newest Benchmark Found
On the newest independent evidence, yes. AI beat the lawyers.
Vals AI published the study in October 2025 as a follow-up to its February benchmark, which had deliberately held legal research back for separate treatment. It set 200 US legal research questions across the kinds of problems a private practice sees, using a dataset built with Reed Smith, Fisher Phillips, McDermott Will & Emery, Ogletree Deakins, Paul Hastings and Paul Weiss. Three legal AI products, ChatGPT, and a group of practising lawyers answered them. A consortium of law firms and academics graded every response blind, against a published rubric.
| Participant | Accuracy |
|---|---|
| Lawyer control group | 71% |
| Alexi | 80% |
| Counsel Stack | 81% |
| Midpage | 79% |
| ChatGPT | 80% |
Grouped together, the legal tools and the general model both landed at 80%, nine points above the lawyers. Where the AI beat the lawyers on a particular question, LawSites reported it did so by an average of 31 points. And the speed difference is not close: answers in seconds or minutes, against roughly 23 minutes per question for a lawyer.
This needs saying directly, because much of what has been written about legal AI, including earlier pieces on this site, leaned on older evidence. Peer-reviewed testing from 2024 measured purpose-built research tools hallucinating on 17% to 33% of queries. That study still stands for the specific products it tested at the time. The category has moved since.
What the Legal Research Benchmark Did Not Test
The benchmark left out the thing that gets lawyers sanctioned.
Vals states that the study covers general legal research only. It excludes drafting pleadings and generating formatted citations. That exclusion is the hinge of this whole article. A research answer can be correct, and the citation attached to it can still be invented. Courts do not sanction wrong answers. They sanction invented authority.
Vals names the risk itself. The report explains that legal research attracts so much scrutiny because of the risk that outputs incorrectly identify or invent arguments, sources and citations.
Three further limits are worth knowing before anyone quotes the headline number:
- Every question went in as a single prompt with no follow-up, which is not how a lawyer works a research problem
- Thomson Reuters, LexisNexis and vLex, the three largest AI research platforms, did not take part, so the tools most mid-market firms already license are missing from the results
- Responses were collected in July 2025, and these products change on a timescale of months
Taken together, the study is strong evidence about answers and says nothing at all about citations.
Why Legal Research Is the Wrong Place for a Firm to Start
Accuracy was never the best argument against starting here, and now it is not an argument at all. Five better ones remain.
The Failure Mode Is Asymmetric
A wrong summary gets caught by the colleague who reads it. A fabricated citation goes to a court. The public database of decisions involving AI-fabricated material records invented case law as by far the largest category, and in Johnson v. Dunn the consequence for three lawyers at a large firm was disqualification from the case, referral to the state bar, and an order to send the opinion to every client they represented. No other category of legal work in this series carries that cost for a single mistake.
Verification Consumes Most of the Time Saved
The speed gap is real for a single question. What does not shrink is the duty to open every cited authority in a primary source before anything is filed, and that applies however fast the first draft arrived. Much of the time a tool saves sits in the reading, and the reading was never where the risk lay. This is our interpretation rather than a measured finding: no study we know of has timed a research task end to end, including verification.
It Does Not Compound
Research is one lawyer asking one question. It does not touch work the firm is already writing off, and it does not change what reaches anyone's desk. On hourly matters faster research produces a smaller invoice, because ABA Formal Opinion 512 requires billing the time spent. Of the five workflow categories covered in this series, research is the least firm-level and the hardest to measure.
Your Firm Probably Already Owns It
Thomson Reuters launched the next generation of CoCounsel Legal in August 2026, and LexisNexis replaced Lexis+ AI with Lexis+ with Protégé in February 2026. For a firm already paying for Westlaw or Lexis, research AI is increasingly a feature of a subscription it already has, rather than a project that needs scoping, budgeting and a champion.
It Is Weakest Where Mid-Market Firms Need It Most
Every system in the benchmark lost around eleven points of accuracy on questions that required combining the law of several states. Lawyers still won roughly a third of the question categories, specifically the ones demanding deep interpretive work: distinguishing precedents that look alike, and reconciling authorities that conflict. A firm practising across several jurisdictions spends a good deal of its research time in exactly that gap.
Legal AI or ChatGPT for Legal Research: Where Each Wins
On accuracy they tied. On the quality of their sources, the legal tools won.
| Measure | Legal AI products | ChatGPT | What it means |
|---|---|---|---|
| Accuracy | About 80% on average | 80% | No meaningful difference |
| Authoritativeness of sources | About 76% on average | 70% | Legal tools cite better sources, drawn from proprietary databases |
| Questions needing the most current information | Behind | Ahead | ChatGPT searches the open web by default |
Vals put it plainly: accuracy did not prove a defining factor in the study, and authoritativeness did. That is the right distinction for a lawyer to care about. A correct answer resting on a weak or unverifiable source still has to be re-researched before anyone can rely on it, so the real saving comes from a tool whose sources you can open and check.
The one reversal is worth noticing. On the question type that needed the most up-to-date information, ChatGPT beat the legal tools on sourcing, because it searches the open web by default while legal products tend to restrict themselves to curated databases. For fast-moving regulatory questions, that matters.
A correction for our own readers: our earlier article on AI for lawyers said legal-specific research tools had been measured as more accurate than general models. That rested on the 2024 evidence, and the newer benchmark contradicts it. We have corrected that article to match.
How to Use AI for Legal Research Without Getting Sanctioned
Treat every citation as unverified until you have opened it yourself.
Four practices cover most of the risk:
- Open every cited authority in a primary source before anything is filed, whichever tool produced it and whoever in the firm inserted it
- Prefer tools that link each proposition to a source you can click, since sourcing is where the tools differ
- Treat any multi-jurisdictional answer as a first draft, given the measured drop in accuracy
- Record who verified the authorities on each filing, so the firm can show it later
The last point sounds like administration until you see what its absence costs. After the show cause order in Johnson v. Dunn, the firm commissioned a review in which 28 attorneys at another firm checked more than 2,400 citations across 330 filings, because it had no other way to prove the problem went no further. Our article on law firm AI policy covers what a firm should record instead.
What to Automate First Instead of Legal Research
Start where errors are caught before they leave the building, and where the work was already being written off.
| Workflow | Why it beats research as a first project |
|---|---|
| Inbound document review and triage | The strongest benchmark results of any category, and courts have approved machine-assisted review since 2012 |
| Intake completeness and conflicts screening | Errors surface inside the firm, before a matter is opened |
| Billing narrative quality | Targets work already being written down, so it costs nothing in collected revenue |
All three share a property research lacks: a mistake gets caught by someone inside the firm before a client or a court ever sees it.
That property changes how a first project goes. A firm learning to use AI will get things wrong at the start, and the question is where those early errors land. In document triage a mislabelled file gets corrected by the associate who opens it. In intake, a missing document gets chased before the matter begins. In billing, a weak narrative gets rewritten before the invoice leaves. Each gives the firm room to learn without anyone outside it noticing, which research, by its nature, does not.
The sequencing rule set out in our guide to AI for mid-market law firms applies here too. Automate the work you were already writing off before paying to make billable work faster.
How Codebridge Works with Mid-Market Law Firms
We do not build legal research tools, which is part of why this article can say research is not where to start.
We build the three workflows in the table above: document review triage, intake and conflicts screening, and billing narrative cleanup. One workflow goes live in three weeks, wired into the systems the firm already runs, with the approval checkpoint designed in during the first conversation. Your firm owns the repository, the prompts and the configuration from day one.
The closest reference we can offer, labelled for what it is: Knowledge Cloud, built for a Big Four tax and legal practice, runs an expert review queue with an immutable audit log, so a senior practitioner approves each output before the firm acts on it. A research platform rather than a law firm system, and what it demonstrates is the review pattern.
Our founding team spent more than a decade at KPMG.
If you want to work out which workflow should come before research, book a 20-minute call.
Is AI good enough for legal research?
On accuracy, yes. In an October 2025 independent benchmark, AI tools averaged about 80% accuracy on 200 legal research questions, against 71% for practising lawyers. But the study excluded generating formatted citations, which is the failure behind sanctions, so every citation still needs checking.
Is ChatGPT as accurate as legal AI for research?
In the October 2025 Vals benchmark, yes on accuracy, with both at about 80%. Legal-specific tools scored higher on the authority of their sources, about 76% against 70% for ChatGPT, because they draw on proprietary legal databases. ChatGPT did better on questions needing the most current information.
Can I trust AI-generated legal citations?
Not without checking each one. The most recent independent research benchmark explicitly excluded generating formatted citations, so no independent study establishes how reliable they are. Open every cited authority in a primary source before filing, whichever tool produced it.
Why do lawyers get sanctioned for using AI in legal research?
Because AI tools can produce citations to cases that do not exist, and lawyers who file them without checking breach their duty to the court. In Johnson v. Dunn, three lawyers at a large firm were disqualified from a case and referred to their state bar after filing fabricated citations.
Do Westlaw and Lexis include AI research tools?
Yes. Thomson Reuters launched the next generation of CoCounsel Legal in August 2026, and LexisNexis replaced Lexis+ AI with Lexis+ with Protégé in February 2026. Neither took part in the October 2025 independent research benchmark.
What should a law firm automate before legal research?
Workflows where a mistake is caught inside the firm before a client or court sees it: inbound document review and triage, intake completeness and conflicts screening, and billing narrative quality. Billing narrative work also targets time that was already being written down.

Heading 1
Heading 2
Heading 3
Heading 4
Heading 5
Heading 6
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.
Block quote
Ordered list
- Item 1
- Item 2
- Item 3
Unordered list
- Item A
- Item B
- Item C
Bold text
Emphasis
Superscript
Subscript


























