Can Lawyers Trust AI Translations? A 10-Model Case Study Exposes the Risk

A translated indemnity clause comes back from an AI translator. The grammar is clean. The formatting has survived. The clause reads exactly the way a clause is supposed to read. It says the counterparty will guarantee the other side against losses. The original said it will indemnify. In common law those are different obligations, one primary and one secondary, with different triggers and different remedies. Nothing in the output flags the change.
So can lawyers trust AI translation? Not on a single model’s output. In a controlled test of ten AI models, one indemnification clause produced the correct term in every Spanish rendering, split three ways in French, and returned only one legally precise version out of ten in Japanese. Fluency stayed constant across all three. Legal accuracy did not.

For Indian practitioners this stopped being hypothetical some time ago. Cross-border commercial work, foreign patent filings, institutional arbitration and enforcement of foreign awards all travel through translated documents. A patent agent working out how to translate machine learning terms into Japanese before a priority filing is not making a linguistic choice. They are fixing the scope of a claim, in a jurisdiction where the filed text governs.
The profession has responded by reaching for automation. Machine translation now sits inside contract review workflows, due diligence platforms and court registries. The Supreme Court of India’s own Vidhik Anuvaad Software, better known as SUVAS, has translated judgments into sixteen regional languages, a figure confirmed by the Chief Justice of India in 2026. The technology is already embedded in legal practice. The verification habit is not.
Why a fluent translation is more dangerous than a broken one
Machine translation used to fail loudly. A garbled sentence, a mangled idiom, a word order that collapsed. Those failures were annoying and harmless, because you could see them.
The research team behind the ten-model test sorted its results into two categories, and the distinction is the single most useful thing in the study for a practising lawyer. Visible failures are the ones you catch. In the same test run, one model translated “burning the midnight oil” into Vietnamese as an expression meaning burning torches. Anyone reviewing that output would send it back.

Silent failures are the ones you do not catch. The output is grammatical, plausible, contextually smooth and legally wrong. There is no marker, no confidence warning, no formatting artefact. It reads like a clause because it is a clause. It simply commits your client to something different.
Only one of those two categories creates professional exposure. A visible failure gets returned. A silent failure gets signed.
What the 10-model case study actually tested
The methodology matters, so here it is in full. Eight test sentences were constructed to probe five specific pressure points: idiom, legal terminology, medical terminology, register and formality, and deliberate ambiguity. Each sentence was run in two rounds. Round one used five production AI models. Round two added five premium-tier models, among them DeepSeek V4-Pro, Claude Opus 4-7 and Gemini 3.5-Flash, for ten in total.
The first run covered English to Japanese and English to Vietnamese. The identical methodology was then repeated on English to Spanish and English to French. Same sentences, same rounds, same scoring, different language pairs. That repetition is what makes the results readable, because it isolates the variable that turns out to matter most.
One clause, ten models, three different legal obligations
The legal test sentence was an indemnification clause. Here is what came back.
In Spanish, all ten models across both rounds converged on the correct term. No divergence, no drift, nothing to review.
In French, the ten outputs split three ways. One of them, produced by Gemini 3.5-Flash in round two, rendered the obligation as a guarantee rather than an indemnity. A guarantee in common law is a secondary obligation that bites only on the principal’s default. An indemnity is a primary obligation that stands on its own. A reviewer without French would have no way of knowing that the exposure had changed shape.
In Japanese, every round-one model chose a compensation term rather than the legally precise formulation covering indemnity and exemption from liability. Only DeepSeek V4-Pro, added in round two, produced the precise term. One correct output in ten.


Infographic 1: the same indemnification clause produced three different risk profiles across three target languages. Source: 10-model AI translation benchmark, 2026.
Read that as a risk map rather than a scoreboard. The same clause, the same models and the same prompt produced a clean result in one language, a contested result in another and a near-total miss in a third. Risk in AI legal translation is not distributed by model. It is distributed by language pair, and most reviewers never see which pair they are exposed on.
Does adding more AI models fix the problem?
Partly, and the exceptions are the interesting part.
On formality, expanding the pool changed nothing. For a client-facing business message that called for the formal register, every single model in both rounds, in both Spanish and French, defaulted to the informal form. Ten models, zero flags. In a legal notice or a demand letter, register is not decoration. It is the difference between correspondence that reads as a formal communication and correspondence that does not.
On Japanese formality, the pool size did help. A message addressed to a chief executive drew plain-form casual Japanese from one round-two model. The consensus correctly excluded that output, but the consensus itself only moved from casual to formal once the model pool was widened. With five models, the wrong register would have carried the majority.
And on one test, consensus was simply wrong. A medical contraindication sentence referring to hepatic impairment split the round-two Vietnamese outputs six to four, between a phrase meaning liver failure and the clinically correct phrase meaning impaired liver function. The majority chose the less precise term. The split did not resolve as the pool grew. Two models from the same family consistently held the correct reading and were outvoted.
That result belongs in any honest account of this method. Comparing models surfaces disagreement and narrows the error surface. It does not confer certainty, and anyone selling it as certainty is overselling it.
Price does not rescue the position either. On a Japanese binding-arbitration clause, a model costing roughly one sixty-second of another per unit of output scored higher on the automated quality metric, while the expensive model produced the more legally appropriate register. The cheaper output won the score. The pricier output was the one a lawyer would have wanted. Cost is not a proxy for legal fitness.
Is choosing the best AI model a defence?
This is where the aggregate data is more useful than the case study. Across 78,324 translations spanning 3,682 language pairs in a single 28-day window, the gap between the best-performing and worst-performing engine on identical source text was 0.17 points on a ten-point scale. The gap between the best-performing and worst-performing language pair was 1.66 points, close to ten times larger.

In the same window, 8.50 per cent of scored outputs fell below the quality threshold, up from 5.96 per cent in the preceding 28 days as volume roughly tripled. Every engine in the panel declined against the prior period. Scale did not improve quality. It diluted it.

Infographic 2: engine choice moves quality far less than language pair does. Source: 10-point quality scoring across 78,324 translations, 28-day window, 2026.
Average model agreement across the ten highest-volume language pairs ran from 83.6 per cent on Japanese to English up to 91.9 per cent on English to Hindi. Put in review terms, between roughly eight and sixteen per cent of content is material on which the models do not land on the same wording. That is the fraction a single-model workflow never shows you.
The counterintuitive part is that agreement does not track how well resourced a language is. English to Spanish, one of the most heavily resourced pairs in existence, sat near the bottom of the agreement ranking. English to Hindi sat at the top. Agreement tracks ambiguity in the source text, not the volume of training data behind the target language, which is why a densely drafted English clause can be riskier to translate than a plainly drafted one regardless of the target.
None of this is a defect unique to any one provider. Research synthesised from Intento’s State of Translation Automation and the WMT24 shared task places single-model hallucination rates on translation tasks between 10 and 18 per cent. Forrester Research reported in 2025 that knowledge workers spend an average of 4.3 hours a week verifying AI output. The verification burden has not disappeared. It has moved onto the desk of whoever signs off.
What legal status does an AI translation actually have in India?
An AI translation has no independent legal status. Under Article 348(1)(a) of the Constitution, English remains the authoritative language for proceedings and judgments of the Supreme Court and the High Courts. Translated versions, including those produced through SUVAS, are issued for information and accessibility. The English text governs.
That distinction does real work. A regional-language judgment on the eSCR portal is a reading aid for a litigant, not the operative text a court construes. Treating it as the latter is an error of category, not of translation.
For documents a party actually files, the position is stricter. Indian courts generally require that a document in a language other than the language of the court be accompanied by a translation certified as correct, with a named individual, usually the advocate on record or an authorised translator, taking responsibility for that certification. The applicable rule varies between the Supreme Court Rules and individual High Court rules, so check the forum. What does not vary is the principle: certification attaches to a person, never to software. If the machine gets it wrong, the certifier owns it.
A verification protocol for lawyers using AI translation
The evidence does not support banning these systems, and no realistic practice could. It supports putting a structure around them. That structure is what separates raw model output from AI legal translation tools built for regulated work, where a professional reviewer signs off inside the same workflow and the output can be certified for filing. Seven steps that map directly onto the findings above:
- Classify the document before you translate it. Operative text that will be filed, executed or served sits in a different tier from background reading. Only the second tier tolerates an unverified output.
- Build the term list before translating, not after. Indemnify, guarantee, warrant, represent, undertake, best endeavours, reasonable endeavours, shall and may all carry defined weight. Fix the target-language equivalent in advance and hold every output to it.
- Run operative clauses through more than one model and compare the outputs. Disagreement is the signal. Where the models converge, your review time is better spent elsewhere. Where they split, you have found the clause that needs a human.
- Treat convergence as a floor, not as proof. The Vietnamese medical split shows a majority can be wrong together, and two dissenting models can be right.
- Check register as a separate pass from meaning. Every model in the test missed formality. Nothing in a meaning-level review will catch it.
- Keep a named human certifier in the chain for anything filed, executed or served. This is a professional responsibility question before it is a technology question.
- Log the terms the models disagreed on. Over a few matters that log becomes a firm glossary, and the glossary is what stops the same silent failure recurring on the next deal.
Frequently asked questions
Can AI translations be filed in Indian courts?
Not on their own. Indian courts generally require a translation certified as correct by a named person, typically the advocate on record or an authorised translator. AI output can form the first draft, but certification attaches to an individual who takes responsibility for accuracy. Always check the specific rule of the forum.
What is the difference between a certified translation and an AI translation?
A certified translation carries a signed declaration of accuracy from an identifiable person or agency, which creates accountability. An AI translation carries no declaration and no accountable party. The distinction is not about quality but about who answers for the document if a term turns out to be wrong.
Which legal terms does AI most often get wrong?
Terms whose closest everyday equivalent is legally weaker. In the ten-model test, indemnify drifted to compensate in Japanese and to guarantee in French. Similar risk attaches to warrant against represent, shall against may, and best endeavours against reasonable endeavours, where the natural translation loses the legal weight.
Does a more expensive AI model reduce legal translation risk?
Not reliably. On a Japanese arbitration clause, a model costing roughly one sixty-second of another per unit of output scored higher on automated quality metrics, while the expensive model produced the more legally appropriate register. Price tracks general capability, not fitness for a specific legal register in a specific language.
How many AI models should be compared for a legal document?
The test data shows five is enough to surface most disagreement and ten resolves cases five cannot, such as the Japanese formality miss. More important than the count is that the comparison happens at all, and that a human reviews every clause where the models split.
The question worth asking
The instinct when adopting AI translation is to ask which model is best. The data says that is close to the wrong question. The spread between the strongest and weakest engine on identical text was 0.17 points. The spread between language pairs was 1.66. Picking a favourite model buys almost nothing.
The question that does the work is narrower. On this clause, in this language pair, do the models agree, and if they do not, which one is right? A single model cannot answer that, because a single model has nothing to disagree with. It returns one fluent answer and gives you no reason to doubt it. That is precisely the condition under which a silent failure gets signed.
Cross-model comparison is not a productivity feature. For anyone whose signature ends up on a translated instrument, it is the closest thing available to a second reader.
Attention all law students and lawyers!
Are you tired of missing out on internship, job opportunities and law notes?
Well, fear no more! With 2+ lakhs students already on board, you don't want to be left behind. Be a part of the biggest legal community around!
Join our WhatsApp Groups (Click Here) and Telegram Channel (Click Here) and get instant notifications.




