The Legal Stack
Independent LegalTech Analysis
← Research Briefings
Research BriefingNo. 100 · October 02, 2026 · 10 min read
Legal AI · Research Report

The Legal AI 'Governing Law Mismatch' Benchmarking Report 2026: How Often AI Contract Review Tools Apply the Wrong Jurisdictional Standard to Enforceability Questions — and Which Clause Categories Are Most Exposed

The Legal Stack Research Briefing | Benchmarking Series | Q1 2026

The Legal Stack Research Briefing | Benchmarking Series | Q1 2026


Executive Summary

AI contract review tools are failing at one of transactional law's most fundamental tasks: applying the correct jurisdictional standard to enforceability questions when the governing law of a contract diverges from the tool's embedded default assumptions. Across a structured benchmarking exercise covering five clause categories and five governing law regimes, general-purpose LLM wrappers misapplied jurisdictional standards in 61% of tested instances, while purpose-built legaltech platforms performed meaningfully better but still posted a 34% failure rate on cross-border and cross-state enforceability questions. The practical consequences range from commercially negligible to potentially catastrophic — particularly where AI-generated redline suggestions effectively import California's near-total ban on non-competes into a New York-governed services agreement, or where a tool flags an English law liquidated damages clause as a "penalty" using U.S. common law doctrine that English courts haven't recognized since Cavendish Square Holding BV v Makdessi [2015] UKSC 67.

This report is directed at general counsel, legal operations procurement leads, and transactional partners who are evaluating, renewing, or auditing AI contract review vendor relationships. Our core finding: governing law jurisdiction is not a feature these tools reliably implement — it is an assumption most of them quietly embed and rarely disclose.


Methodology

Our benchmark tested seven platforms across two categories: (1) four general-purpose LLM wrappers — products built on GPT-4o or Claude 3.5-class models with a contract review interface layered on top, including commercially marketed tools from vendors we designate GP-1 through GP-4; and (2) three purpose-built legaltech platforms with proprietary legal training and clause libraries, including products representative of the Ironclad AI, Harvey, and Spellbook tier of the market.

We generated 175 test contracts — 35 per governing law regime (New York, California, English law, Singapore law, and Delaware) — each containing identical substantive commercial terms but specifying a different governing law clause. Each contract was submitted to all seven platforms for AI-assisted review. We measured:

  • Whether risk flags on targeted clause types accurately reflected the applicable jurisdiction's enforceability standard
  • Whether suggested replacement language was jurisdictionally calibrated or generically drafted
  • Whether the platform's user-facing documentation disclosed any governing law assumption

Clause categories tested: non-compete and restrictive covenant, liquidated damages, indemnification caps, penalty clauses, and arbitration carve-outs.


Findings by Clause Category

Non-Compete and Restrictive Covenant Clauses: 68% Overall Mismatch Rate

This was the single highest-failure category and the one with the most commercially damaging error profile. The root problem is well-documented: California Business and Professions Code § 16600 renders most employee and many contractor non-competes void, while New York, post the FTC Rule litigation and the state's own 2023 legislative history, still enforces reasonably scoped covenants. Delaware, English law, and Singapore law each apply distinct "legitimate business interest" frameworks with different geographic and temporal reasonableness thresholds.

All four general-purpose LLM wrappers, when reviewing a New York-governed NDA with a standard 12-month, 50-mile non-compete, flagged the clause as "potentially unenforceable — California law may prohibit." This flag is legally incorrect for a New York-governed instrument. More troublingly, GP-2 and GP-3 generated suggested replacement language that effectively deleted the covenant entirely, citing California standards in the rationale field — an outcome that would materially disadvantage a New York employer who is entitled to enforce a reasonable covenant under BDO Seidman v. Hirshberg (1999).

For Singapore-governed agreements, two of three purpose-built platforms applied generic U.S. common law "blue-penciling" doctrine, rather than the Singapore Court of Appeal's distinct approach articulated in Man Financial (S) Pte Ltd v Wong Bark Chuan David [2008] 1 SLR(R) 663, which recognizes enforceability through a legitimacy-of-interest lens more permissive than California but more structured than New York.

Liquidated Damages: 57% Mismatch Rate

The English law failure here is acute. Following Cavendish v. Makdessi, English courts abandoned the historical penalty doctrine in favor of a legitimate interest test: a clause is not a penalty merely because it exceeds a pre-estimate of loss, provided it serves a legitimate commercial purpose. Every general-purpose LLM wrapper reviewed flagged liquidated damages clauses in English law-governed agreements using language directly drawn from U.S. Restatement (Second) of Contracts § 356 — the "reasonable estimate at time of contracting" standard — which is not the applicable test in England and Wales. GP-1's output stated that the clause "constitutes an unenforceable penalty under applicable law" — a legally incorrect statement for an English law instrument post-2015.

Purpose-built platforms performed better (39% mismatch), but two of three still failed to distinguish between the pre-Cavendish and post-Cavendish frameworks when the clause involved a primary obligation structured as a price adjustment, the precise scenario Makdessi was designed to address.

Indemnification Caps: 41% Mismatch Rate

Indemnification cap failures were primarily a Delaware versus New York distinction problem. Delaware courts applying their contractualist philosophy under ABRY Partners V, L.P. v. F&W Acquisition LLC (Del. Ch. 2006) will uphold sophisticated parties' bargained indemnification limitations with minimal judicial second-guessing. New York applies a more interventionist unconscionability analysis in consumer-adjacent contexts. Three platforms incorrectly imported the New York standard into Delaware-governed M&A ancillary agreements, suggesting cap language modifications that would be unnecessary and commercially unusual in a Delaware-governed deal context.

Penalty Clauses: 53% Mismatch Rate (Cross-Border)

The California-to-English-law delta drove most of the penalty clause errors. California Civil Code § 1671 creates a rebuttable presumption that liquidated damages clauses are valid in commercial contracts — closer to the post-Makdessi English position than most U.S. states — yet platforms reviewing California-governed agreements consistently applied a more skeptical New York-style analysis. The inverse error also occurred: English law contracts were reviewed with California § 1671 standards rather than the Makdessi legitimate interest test.

Arbitration Carve-Outs: 29% Mismatch Rate (Lowest Failure Category)

Arbitration clause analysis showed the strongest cross-jurisdictional performance, likely because the Federal Arbitration Act provides a strong federal baseline that several platforms have heavily trained on, and because Singapore's International Arbitration Act and the English Arbitration Act 1996 produce broadly compatible outcomes in commercial contexts. The failures that did occur concentrated in emergency relief carve-out language, where Singapore courts' approach to anti-suit injunctions differs meaningfully from New York commercial division practice.


Platform Disclosure Failures

Perhaps the most significant procurement-relevant finding: six of seven platforms provided no user-accessible documentation disclosing any governing law assumption. One purpose-built platform (representative of the Harvey tier) disclosed in its API technical documentation — not its product UI — that "risk assessments are calibrated to U.S. commercial law defaults unless jurisdiction-specific modules are activated." No platform clearly surfaced that disclosure at the contract review workflow level where a practitioner would encounter it.

This is a material gap. A legal operations team deploying a tool organization-wide across a contract portfolio spanning English, Singapore, and New York governing law has no reasonable way, from product documentation alone, to know that the tool is applying a single baseline standard to fundamentally different legal regimes.


Procurement Implications and Recommendations

For GCs and legal ops procurement leads, the governing law mismatch problem should be a vendor evaluation line item, not an assumption. Specific due diligence recommendations:

  1. Require vendors to produce a jurisdiction coverage matrix as a contractual deliverable — specifying which governing law regimes the platform has validated output against, and at what clause granularity.

  2. Run your own benchmark before renewal using the clause categories identified here. Submit identical substantive contracts with differing governing law clauses and compare flag logic and suggested language output. This takes approximately two hours with a junior associate and produces directly actionable procurement intelligence.

  3. Negotiate SLA language around jurisdictional accuracy for your primary governing law regimes. If 60% of your portfolio is English law, the tool's English law enforceability standards should be a specified performance dimension.

  4. Treat general-purpose LLM wrappers as higher-risk for cross-border portfolios. The performance gap between purpose-built and wrapper-tier products on jurisdiction-specific questions was consistent across all five governing law regimes tested, even where the gap narrowed for domestic U.S. use cases.

  5. Require disclosure of training data jurisdictional composition where vendors will provide it. Tools trained predominantly on U.S. federal district court and New York commercial division decisions will systematically underperform on English, Singapore, and civil law-adjacent questions.

The governing law mismatch problem is not a theoretical risk. It is a systematic and measurable failure mode embedded in tools that legal teams are already deploying at scale across global contract portfolios. The benchmarking data suggests the risk is addressable through disciplined procurement and vendor accountability — but only if buyers know to ask.


Methodology notes and full platform-level data available to Legal Stack subscribers. Platform names anonymized per vendor engagement protocols; individual platform briefings available under NDA for procurement purposes.

Filed under Legal AI → · The Legal Stack accepts no vendor funding for its research.

More Research

View all →
No. 099
10 min
The Legal AI 'Playbook Divergence' Benchmarking Report 2026: How Much Do Internal Legal Department AI Playbooks Actually Differ From Outside Counsel AI Policies — and Where the Conflicts Are Creating Compliance Exposure
10 min
No. 098
10 min
The Legal AI Lateral Partner Portability Report 2026: How AI Tool Proficiency, Client Data Portability, and Firm AI Policy Conflicts Are Reshaping Lateral Partner Movement — and What Recruiting Firms Are Not Telling Either Side
10 min
No. 097
10 min
The Legal AI 'Change of Control' Clause Audit Report 2026: How AI Contract Review Tools Perform on the Specific Clause Category That Determines Whether Your Entire Agreement Survives an Acquisition
10 min
© 2026 The Legal Stack — Independent LegalTech Analysis