Skip to content
Locale Lab

Chapter 20·Shipping translations·~12 min

Scoring an LQA round (MQM)

The delivery came back. Someone has to say whether it ships, and 'it reads fine to me' is not an answer a vendor can act on. MQM turns a review into a number.

MQMLQAqualityreview

The problem

Translation quality arguments go in circles when the only vocabulary is good and bad. A reviewer says the German feels off, the vendor says it is a preference, and nobody can point at a number. MQM (Multidimensional Quality Metrics) is the standard answer: a fixed error typology, four severity levels with penalty weights, and an arithmetic that normalizes penalties per 1000 words into a score a threshold can be set against. The same framework runs a vendor scorecard, an MT evaluation, and a release gate.

How it works

A typology turns opinions into data

Without a shared vocabulary, a review round is a list of complaints: the reviewer writes "this sounds wrong", the vendor replies "that is preference", and the program manager has nothing to put in a scorecard. An error typology makes the reviewer choose from a closed set of categories, so two reviewers on two rounds produce comparable data.

MQM, Multidimensional Quality Metrics, is the framework most of the industry uses. Its current top level has seven dimensions: Terminology, Accuracy, Linguistic conventions, Style, Locale conventions, Audience appropriateness, and Design and markup. Older MQM material calls Linguistic conventions "Fluency", so both names circulate. Each dimension nests further. A real scorecard usually ships a subset with the child types the product needs.

The dimension routes the fix. Terminology goes to whoever owns the termbase. Locale conventions is usually an engineering bug: code sets the date format, and the translator cannot change it. Design and markup belongs to whoever built the string container, and Accuracy returns to the translator. File a wrong-date bug under Accuracy and it reaches the wrong team.

Severity is the number that moves

Each logged error carries a severity, and severity is where the weight lives. The MQM default multipliers are neutral 0, minor 1, major 5, and critical 25. The steps are steep: one critical error costs the same as twenty-five minor ones, because one mistranslated dosage or one price promise the business cannot honor can outrank a page of comma slips.

Because the gap is steep, nearly every arbitration is a severity argument. The dispute is rarely about whether an error exists: both sides agree the word is wrong. They disagree whether the reader is slowed (minor) or misled (major). A team that defines severity with examples from its own product before the first round avoids most of these fights. The weights are also configurable: the numbers below come from MQM's own worked examples, and harmonized DQF-MQM weights differ.

One rule sits outside the arithmetic: a critical error count above zero fails the delivery, whatever the score says. A sample can score well overall and still contain the one mistranslated dosage that cannot ship.

neutral   0    logged, but costs the reader nothing
minor     1    noticeable, meaning survives
major     5    the reader is misled, confused, or blocked
critical  25   safety, legal, financial, or brand damage
MQM default severity multipliers.

The arithmetic, and why small samples lie

Scoring sums the severity weights of every logged error into the Absolute Penalty Total, then divides by the number of evaluated words and multiplies by 1000 to get the Normed Penalty Total: penalty points per 1000 words. Subtracting that from 100 gives the raw quality score.

The raw pass threshold sits near 99, not 90, because raw scores bunch up just under 100. A raw 95 is a bad result. Executives find that unintuitive, so the MQM Council recommends a calibrated model that stretches the narrow passing band onto a friendlier range. A threshold like 90 then becomes meaningful. Always say which model produced a number.

Normalization also punishes short samples. One major error in a 40-word sample is 5 penalty points over 40 words, which norms to 125 per 1000 words and drives the raw score below zero. The model is behaving correctly, which is why MQM and ISO 5060 both treat sampling as part of the method. When a representative sample is not available, the honest report is penalty points and error counts instead of a raw score.

APT = sum of severity weights
NPT = (APT / evaluated words) x 1000
raw score = 100 - NPT

1000 words, 2 minor + 1 major:
  APT = 2(1) + 1(5) = 7
  NPT = (7 / 1000) x 1000 = 7
  raw score = 93        -> FAIL against a threshold of 99
Raw MQM scoring, with the standard constants.

The demo

Try these, in order

Each step reproduces one specific failure in the demo below.

  1. 1
    Score the round without touching anything: leave every segment on No error and press Score my review. All 8 seeded errors go unfound, yet the review scores a perfect 100 and a PASS. A review that logs nothing always passes, so the demo tracks misses separately from the score.
  2. 2
    Flag segment 1 (the 3/4/26 date) and try each dimension in turn before checking the key. Accuracy is tempting because the reader gets the wrong date. The answer is Locale conventions: the German short date is 04.03.26, and code sets the format. That routes the bug to engineering, because the translator cannot fix it.
  3. 3
    Set segment 6 (the dropped 'on orders over $50' condition) to Accuracy and step its severity from minor to critical, watching the penalty total. The Absolute Penalty Total jumps from 1 to 25 on a single dropdown. One severity call outweighs every other error in the sample.
  4. 4
    Flag one of the clean segments, 3 or 9, with any dimension. The demo marks it as a false positive. The reviewer's own score drops even though the translation was fine. On a real vendor scorecard, over-flagging creates a dispute.
  5. 5
    Score the full round correctly and read the normalization line in the arithmetic panel. 48 penalty points over 39 evaluated words norms to over 1200 points per 1000 words, so the raw score clamps at 0. The sample is far too small for a stable raw score.
  6. 6
    In the Weight lab under the score panels, drag all three severity sliders to 0. The raw score climbs to 100 and the verdict still reads FAIL. The critical-error override counts errors rather than penalty points, so no weight configuration can pass a delivery that contains one.

Your review brief

Ten segments of a checkout flow, English into German, 39 source words. Classify each segment with an MQM dimension and a severity, or leave it as no error. Some segments are clean, and flagging one counts as a false positive: over-flagging wastes arbitration time and distorts the vendor scorecard.

Approved termbase

  • cart→WarenkorbApproved. Do not use Einkaufswagen or Korb.
  • sign in→anmeldenApproved. Do not use einloggen.
  1. 01Order confirmation banner

    Source (en-US)

    Your order will arrive on 3/4/26.

    Target (de-DE)

    Ihre Bestellung kommt am 3/4/26 an.

  2. 02Cart totals row

    Source (en-US)

    Total: $1,234.56

    Target (de-DE)

    Gesamt: $1,234.56

  3. 03Primary button, 20 character limit

    Source (en-US)

    Add to cart

    Target (de-DE)

    In den Warenkorb

  4. 04Empty state

    Source (en-US)

    Your cart is empty.

    Target (de-DE)

    Ihr Einkaufswagen ist leer.

  5. 05Section heading

    Source (en-US)

    Order summary

    Target (de-DE)

    Order summary

  6. 06Promotional line under the totals

    Source (en-US)

    Free shipping on orders over $50.

    Target (de-DE)

    Kostenloser Versand.

  7. 07Auth gate

    Source (en-US)

    Sign in to continue.

    Target (de-DE)

    Anmelden um fortzufahren.

  8. 08Promo badge

    Source (en-US)

    Save 20% today

    Target (de-DE)

    Sparen Sie 20% heute

  9. 09Payment error toast

    Source (en-US)

    We could not process your payment.

    Target (de-DE)

    Wir konnten Ihre Zahlung nicht bearbeiten.

  10. 10Banner with inline markup

    Source (en-US)

    <b>Limited time</b> offer

    Target (de-DE)

    <b>Angebot für kurze Zeit offer

The short version

A review round is only useful when it produces a number someone can act on: an error typology, a severity, a penalty, a threshold.

What to do about it

  • Argue about severity, not about existence. One critical error carries 25 penalty points, the same as 25 minor ones. Nearly every arbitration is really a severity dispute, so define severity with examples from the product itself before the first round.
  • Score a representative sample. Penalties are normed per 1000 words, so a 40-word sample multiplies every error by 25 and produces a wild score. When a bigger sample is impossible, report penalty points and error counts rather than a raw score.
  • Name the scoring model on the scorecard. Raw and calibrated models put the pass line in different places, so a bare score means nothing without it.
  • Publish the typology and the termbase before the round. A reviewer who invents categories mid-review produces a scorecard nobody can compare to the last one.

Where this comes up

Who it concerns

Localization PMVendor managerReviewersQA

Moments

  • ·Standing up a vendor scorecard or an SLA
  • ·Arbitrating a disputed review round
  • ·Deciding whether a delivery is shippable
  • ·Comparing MT engines or post-edit tiers with numbers

Field note

The most expensive review rounds leave no clean segments in the sample. A reviewer who flags every stylistic preference produces a scorecard that fails a delivery that was fine. The vendor disputes it, and two weeks disappear into arbitration. Good reviewer onboarding measures false positives as carefully as misses.

ISO 5060:2024, Evaluation of translation output ↗

Terms in this chapter

Where to read more

Related chapters