Translation Quality Metrics Explained: BLEU, COMET & Human Review

What each kind of translation-quality number actually measures, run on the text a PO file holds rather than on news benchmarks. Opens with the result that frames it: a correct paraphrase scored BLEU 9.5 while a sentence with its negation dropped scored 40.9, because overlap metrics rank surface similarity, not harm. Then each metric executed with sacrebleu 2.6.0 on eleven hand-built English to German cases: the default sentence-level BLEU gives an identical one-word translation 0.0 (sacrebleu warns, and effective_order=True fixes it), a synonym, a wrong word and a wrong word sense are indistinguishable at 0.0, a renamed placeholder scores 70.7, and the settings that silently move a score — case, the tokenizer (Chinese goes from 0.0 to 38.0), the number of references, and corpus versus sentence level. chrF gives partial credit and TER can reach 150. COMET is explained and its size, license and gating verified from public model metadata (the reference-based checkpoint is Apache-2.0 and ungated, the reference-free one is gated and non-commercial), but no COMET scores are quoted because it was not run. The centre is a fault-injection study on 1,746 real human German translations: a dropped placeholder still scores chrF 90 or above in 73% of cases, a changed number in 91%, a removed negation in 35%, while scrambling the word order is punished hardest. Plus why one UI string in four is too short for any 4-gram metric, a paired bootstrap showing that on 100 segments a one-point difference is noise and twice as many damaged segments can score higher, and the arithmetic of human review: MQM weighting, how many strings a reviewer must read (a clean sample of 100 still allows a 3.0% error rate), and Cohen kappa.

Back to the PO-File blog