Most teams running AI translation in production evaluate it the same way: someone who speaks the language reads a few pages and says it “looks fine.” Then a product name ships translated three different ways, or a warranty clause drifts in meaning, and the team discovers that “looks fine” was never a measurement.
Evaluating AI translation quality means scoring output against a defined error framework — categories, severities, and a scoring rule agreed in advance — so that quality becomes a number you can compare across batches, languages, and engines. Without the framework, review is opinion. With it, review is data.
The anatomy of a working error framework
The industry baseline is built on MQM/DQF categories. A practical working set:
| Category | What it catches | Typical severity range |
|---|---|---|
| Accuracy | Mistranslation, omission, addition, untranslated text | Minor → critical |
| Fluency | Grammar, spelling, awkward phrasing | Minor → major |
| Terminology | Glossary violations, inconsistent product terms | Major (often critical in regulated content) |
| Style | Tone, register, brand voice mismatch | Minor → major |
| Locale conventions | Dates, units, formats, market-specific requirements | Minor → major |
Two things turn the table into a system:
Severity weighting. A missing decimal in a dosage is not the same defect as a clumsy sentence. Each error gets a severity (minor / major / critical), each severity a weight, and each batch a score per thousand words. That is what makes quality comparable over time.
A pass threshold agreed in advance. A batch either clears the threshold or triggers rework. Without a threshold, scores are trivia; with one, they drive decisions.
Your framework or a standard one?
If your team already runs its own error typology — many e-commerce and enterprise teams do — a review partner should adopt it as-is: reviewers briefed on your categories and severities, results reported in your format, feeding your existing dashboards. The framework matters less than the discipline of using one consistently. What breaks comparability is switching frameworks mid-stream, not choosing the “wrong” one.
How much do you actually need to review?
Full review of all AI output usually defeats the economics that made AI translation attractive. A working setup matches depth to risk:
- High-risk content (regulated claims, safety text, legal clauses): full post-editing, every word.
- Revenue-facing content (product pages, campaigns): full or light post-editing depending on market exposure.
- High-volume, low-risk content (support macros, internal docs): sampled LQA — score a statistically meaningful slice per batch and monitor the trend.
The sampling cadence is the quality system. A team scoring 10% of volume every week against a stable framework knows more about its AI than a team that reviewed everything once, three months ago.
Which setup fits your situation
- If AI output ships to customers and errors carry legal or brand risk → full post-editing to publishable quality (ISO 18587 defines the discipline).
- If output is internal or low-exposure but must be understandable → light post-editing.
- If you mainly need to know whether the engine is good enough, and where it fails → LQA-only monitoring on a sampling cadence, no editing.
- If you run multiple engines or an internal AI solution → the same framework applied across all of them is the only way to compare fairly.
One more piece completes the loop: corrections must feed back. Recurring terminology errors go into the enforced glossary; recurring style misses go into the prompt or engine configuration. Evaluation that never changes anything upstream is bookkeeping, not quality management.
This is exactly the layer our human review of AI translation service provides — post-editing and LQA against an MQM/DQF-based framework or yours, at sampling depths matched to risk. The same evaluation discipline, applied to models rather than content, is part of our language data work for AI teams. And for where AI output actually fails, see what AI translation still gets wrong.