10 min read
TL;DR Tap to expand the short version
- AI contract review tools are usually marketed with a single accuracy percentage. That number is a blend that hides clause by clause variance.
- CUAD, built in 2021 by UC Berkeley researchers, tested 510 real contracts across 41 clause categories. Its strongest model reached only 44% precision at 80% recall, with performance breaking down unevenly by clause type.
- A 2025 benchmark reran the same test on today’s frontier models and found the same pattern.
- Adoption of AI in legal and contract work is happening fast but verification has not kept pace. Most firms have not changed pricing or process to account for it.
- Avoiding AI isn’t the fix. Asking vendors for clause level performance, knowing which clause types are weakest, and routing those to a human on purpose is.
- A four question framework gets a non-lawyer buyer most of the way to a sound decision.
Table of Contents
If a business signs a few dozen contracts a year, instead of a few thousand, the person deciding whether an AI tool caught a risky clause is often the same person who has to live with the contract if it did not. Vendors selling AI contract review commonly advertise accuracy rates above 90%. That number is usually an average across every clause type in a contract. And averages hide a lot. Fortunately for us, we have the real numbers, beyond the vendor benchmarks.
For instance, the most rigorous public test of AI contract review there is comes from 2021, when Berkeley researchers, working with legal experts from The Atticus Project, built CUAD: a dataset of over 500 real commercial contracts, hand labeled by trained legal reviewers across 41 clause categories that lawyers specifically watch for. (Liability caps. IP ownership assignment. Exclusivity terms. Change of control triggers.)
The strongest model tested, DeBERTa xlarge, reached 44.0% precision at 80% recall. At that precision level, a reviewer leaning on the tool would still need to read through roughly one irrelevant flagged clause for every real issue the model correctly surfaced.
But that average number is not the real story either. The paper’s own category by category breakdown shows performance is nowhere near uniform. Governing Law, Document Name, and Parties, all predictable, boilerplate clauses, sit near the top of the ranking, close to the ceiling of what the model can achieve.
Covenant Not to Sue, IP Ownership Assignment, ROFR/ROFO/ROFN, and Most Favored Nation sit at the very bottom, close to zero. Those are not fringe clause types either. These terms decide who owns intellectual property, who has the right to buy first, and whether one party quietly locked in better terms than another.
But again, CUAD is from 2021, and AI has evolved a lot since then. Exponentially, some would say. A more recent benchmark reran the same category by category test on today’s frontier models and found the same pattern holds: strong on boilerplate, weak on rare or high stakes clauses. (More on that in the FAQ below, if you want the specifics.)
This gap carries more weight for a business just adopting these tools than for one with an established legal department. A business between $1M and $150M in revenue is often making this decision without in house counsel to check the tool’s output, without budget for independent testing, and with the same person both signing the contract and evaluating whether the review caught everything.
What the Evidence Supports
The data actually argues for adopting AI contract review, with one condition attached.
Fine tuning closes real ground on the categories that are lagging. One comparative study found that fine tuning a smaller model on legal specific data improved classification accuracy by up to 20.6 percentage points, reaching a best overall performance of 87.8%, compared to GPT-4’s zero shot performance of 67.2% on the same task. That gap, between a tuned system and an off the shelf general model, is the difference between “reliable enough to lean on” and “needs a second set of eyes.”
The condition is knowing where the strong performance is actually happening. Based on what CUAD shows about clause level results, the gains cluster in the categories these systems already handle well: standard terms, predictable structure, high volume, low complexity language. That is real time saved and real risk reduced on the majority of a typical contract. The lower performing categories do not disappear just because the average looks good. They still need a human reading them, and that person needs to know which ones to slow down for.
This points to a practical way of evaluating any AI contract review tool, whether you are comparing platforms or deciding if your current one is good enough. Ask for category level performance, not a single accuracy number. A vendor who can answer that question with real data has done the work. A vendor who only offers a single blended percentage has not, or does not want you to see the breakdown.
LegalOn’s own 2026 benchmark is a useful example of what that kind of testing looks like, though it is vendor commissioned and should be read with that in mind. The company tested its platform against 11 AI models across 3,282 head to head reviews and 21 precision critical guidelines.
It ranked first across all provision types tested, while general purpose models failed more often on specific clause identifications, thresholds, multi part requirements, cross references, and absence checks. Whatever you think of a vendor grading its own test, the structure of that test, breaking performance down by provision type rather than reporting one number, is the right question to be asking of any tool.
For a business this size, the real decision point is whether the tool being evaluated can show its work by clause type, and whether the review process still puts a human eye on the categories most likely to create liability.
Evaluating Contract-Review Accuracy
You do not need to become a procurement expert to make a sound decision here. Four questions get you most of the way.
1. Request Clause-Level Results
A single blended percentage can look strong while hiding weak performance on exactly the clauses most likely to create liability. Ask the vendor to break down accuracy or precision by clause category, not just overall. If they cannot, or will not, treat that as information in itself.
2. Identify Weak Clauses
Every tool has weak spots. The question is whether the vendor knows theirs and is willing to say so. A vendor who names their own weak categories, ownership clauses, liability terms, exclusivity language, is showing you real testing. A vendor who insists there are no weak spots is not.
3. Route Weak Clauses
This is the actual practical fix for the accuracy gap, and it doesn’t require rejecting AI review. It requires deciding in advance which clause types still get a second read. Build a short internal list, three to six clause types your contracts commonly include that fall into higher risk categories, and treat those as always human reviewed regardless of what the AI flags.
4. Test Your Contracts
Performance on a vendor’s curated sample contract does not tell you how the tool performs on your actual agreements, your actual counterparties, your actual clause language. Run a small batch of your own recent contracts through the tool and manually check the categories you identified in step three. This does not need to be a formal audit. It needs to happen at least once before full adoption.
When Vendors Show Their Work
This risk is already playing out at scale across a large number of businesses, and the data on how it’s going is mixed.
Clio’s 2026 survey of legal professionals found very high AI adoption among solo practitioners and small firms, 71% and 75% respectively. The gap shows up in what comes next. Only about a third of those firms report an associated revenue increase, 32% of solo practitioners and 31% of small firms, with many saying the impact on revenue is unclear or has not shown up yet.
The clearest sign of the gap: 86% of solo firms and 78% of small firms have not changed their pricing at all since they started using AI. Adoption moved fast. Verification and process did not move with it, and that is the same gap the CUAD data points to from a different angle.
Businesses are bringing these tools in quickly, which makes sense given the time savings on offer. Fewer are building the second step, the process for checking which clause categories still need a human, into how they actually use the tool day to day. For a business with a finance function, even an informal one, this is worth a five minute conversation. The AI tool is almost certainly saving review time. The open question is whether anyone has checked which clause types it is weakest on, and whether someone is still looking at those.
You do not need to overhaul how your business handles contracts to close this gap. You need one conversation and one short list. If you are currently using, or considering, an AI contract review tool, ask for its clause level performance breakdown before your next renewal or purchase decision. If the vendor can answer that clearly, you have real information to work with. If they cannot, that tells you something too. Bring the four question framework above to that conversation. It takes longer to read than it does to use.
FAQ
Trustworthiness is not a universal ‘yes’ or ‘no’ because accuracy varies significantly by clause type. While AI performs well on predictable clauses like governing law, it often struggles with complex or rare provisions, making it essential to evaluate performance on a category-by-category basis.
CUAD, or the Contract Understanding Atticus Dataset, is an academic benchmark containing over 500 expert-labeled commercial contracts. It serves as the industry standard for performing apples-to-apples comparisons of how different AI models handle various clause types.
While the original 2021 dataset is older, newer benchmarks like the 2025 ContractEval continue to use the same category-based testing methodology. These tests confirm that while models have improved, the pattern of uneven performance across different clause types remains consistent.
Yes, small businesses are arguably more exposed because they often lack the in-house legal teams that larger companies use as a safety net. Relying solely on a tool that underperforms on critical clauses can create significant risks for organizations without a human backstop.
You should request clause-level performance data rather than a general accuracy percentage. Specifically, ask the vendor to identify which clause categories the tool performs most poorly on to determine if their product aligns with your risk tolerance.
Spotted an error in this piece? We correct publicly, tell us via the contact page. Read our corrections policy.
