All lessons

Lesson 20·Unit 6 · Shipping translations·How to measure translation quality·Lv 201·~12 min

Scoring an LQA round (MQM)

The delivery came back. Someone has to say whether it ships, and 'it reads fine to me' is not an answer a vendor can act on. MQM turns a review into a number: an error typology, a severity, a penalty, a threshold.

MQMLQAqualityreview

By the end

  • ·File an error under the right MQM dimension, and hold that call when the vendor argues.
  • ·Say why one severity decision outweighs a dozen small findings.
  • ·Compute a raw MQM score, and know when your sample is too small to trust one.

The problem

Translation quality arguments go in circles when the only vocabulary is good and bad. A reviewer says the German feels off, the vendor says it is a preference, and nobody can point at a number. MQM (Multidimensional Quality Metrics) is the standard answer: a fixed error typology, four severity levels with penalty weights, and an arithmetic that normalizes penalties per 1000 words into a score you can put a threshold on. The same framework runs a vendor scorecard, an MT evaluation, and a release gate.

How it works

A typology turns opinions into data

Without a shared vocabulary, a review round is a list of complaints. The reviewer writes "this sounds wrong" and the vendor replies "that is preference". The program manager has nothing to put in a scorecard. An error typology makes the reviewer choose from a closed set of categories. Two reviewers on two rounds then produce comparable data.

MQM, Multidimensional Quality Metrics, is the framework most of the industry uses. Its current top level has seven dimensions: Terminology, Accuracy, Linguistic conventions, Style, Locale conventions, Audience appropriateness, and Design and markup. Older MQM material calls Linguistic conventions "Fluency", so you will meet both names. Each dimension nests further. A real scorecard usually ships a subset with the child types the product needs.

The dimension routes the fix. Terminology goes to whoever owns the termbase. Locale conventions is usually an engineering bug: code sets the date format, and the translator cannot change it. Design and markup belongs to whoever built the string container, and Accuracy returns to the translator. File a wrong-date bug under Accuracy and it reaches the wrong team.

How it works

Severity is the number that moves

Each logged error carries a severity, and severity is where the weight lives. The MQM default multipliers are neutral 0, minor 1, major 5, and critical 25. The steps are steep by design: one critical error costs the same as twenty-five minor ones. That shape is correct, because one mistranslated dosage or one price promise you cannot honor can outrank a page of comma slips.

Because the gap is steep, nearly every arbitration is a severity argument. The dispute is rarely about whether an error exists. Both sides agree the word is wrong. They disagree whether the reader is slowed (minor) or misled (major). Define severity with examples from your own product before the first round. The weights are also configurable: the numbers below come from MQM's own worked examples, and harmonized DQF-MQM weights differ.

Critical severity carries one extra rule outside the arithmetic. A critical error count above zero fails the delivery, whatever the score says. That is deliberate: a sample can score well overall and still contain the one mistranslated dosage you cannot ship.

neutral   0    logged, but costs the reader nothing
minor     1    noticeable, meaning survives
major     5    the reader is misled, confused, or blocked
critical  25   safety, legal, financial, or brand damage
MQM default severity multipliers.

How it works

The arithmetic, and why small samples lie

Scoring is simple. Sum the severity weights of every logged error to get the Absolute Penalty Total. Divide by the number of words you evaluated, then multiply by 1000. That gives the Normed Penalty Total: penalty points per 1000 words. Subtract that from 100 and you have the raw quality score.

Two things surprise people. First, the raw pass threshold sits near 99, not 90, because raw scores bunch up just under 100. A raw 95 is a bad result. Executives find that unintuitive, so the MQM Council recommends a calibrated model that stretches the narrow passing band onto a friendlier range. A threshold like 90 then becomes meaningful. Always say which model produced a number.

Second, normalization punishes short samples. One major error in a 40-word sample is 5 penalty points over 40 words. That norms to 125 per 1000 words and drives the raw score below zero. The model works as designed here, and this is why MQM and ISO 5060 both treat sampling as part of the method. When you cannot evaluate a representative sample, report penalty points and error counts instead of a raw score.

APT = sum of severity weights
NPT = (APT / evaluated words) x 1000
raw score = 100 - NPT

1000 words, 2 minor + 1 major:
  APT = 2(1) + 1(5) = 7
  NPT = (7 / 1000) x 1000 = 7
  raw score = 93        -> FAIL against a threshold of 99
Raw MQM scoring, with the standard constants.

See it yourself

Try these, in order

Each step triggers a specific failure you should recognize on sight.

  1. 1
    Score the round without touching anything: leave every segment on No error and press Score my review. You miss all 8 seeded errors, but your own flags produce a perfect 100 and a PASS. A review that logs nothing always passes. That is why the demo tracks misses separately from the score.
  2. 2
    Flag segment 1 (the 3/4/26 date) and try each dimension in turn before checking the key. Accuracy is tempting because the reader gets the wrong date. The answer is Locale conventions: the German short date is 04.03.26, and code sets the format. That routes the bug to engineering, because the translator cannot fix it.
  3. 3
    Set segment 6 (the dropped 'on orders over $50' condition) to Accuracy and step its severity from minor to critical, watching the penalty total. The Absolute Penalty Total jumps from 1 to 25 on a single dropdown. One severity call outweighs every other error in the sample. That is why most arbitration is about severity.
  4. 4
    Flag one of the clean segments, 3 or 9, with any dimension. The demo marks it as a false positive. Your own score drops even though the translation was fine. Over-flagging creates penalties, and on a real vendor scorecard it creates a dispute.
  5. 5
    Score the full round correctly and read the normalization line in the arithmetic panel. 48 penalty points over 39 evaluated words norms to over 1200 points per 1000 words, so the raw score clamps at 0. The sample is far too small for a stable raw score. This is the sampling caveat in practice.
  6. 6
    In the Weight lab under the score panels, drag all three severity sliders to 0. The raw score climbs to 100 and the verdict still reads FAIL. The critical-error override counts errors, not penalty points, so no weight configuration can pass a delivery that contains one.

Your review brief

Ten segments of a checkout flow, English into German, 39 source words. Classify each segment with an MQM dimension and a severity, or leave it as no error. Some segments are clean: flagging them costs you. A reviewer who over-flags wastes arbitration time and distorts the vendor scorecard.

Approved termbase

  • cartWarenkorbApproved. Do not use Einkaufswagen or Korb.
  • sign inanmeldenApproved. Do not use einloggen.
  1. 01Order confirmation banner

    Source (en-US)

    Your order will arrive on 3/4/26.

    Target (de-DE)

    Ihre Bestellung kommt am 3/4/26 an.

  2. 02Cart totals row

    Source (en-US)

    Total: $1,234.56

    Target (de-DE)

    Gesamt: $1,234.56

  3. 03Primary button, 20 character limit

    Source (en-US)

    Add to cart

    Target (de-DE)

    In den Warenkorb

  4. 04Empty state

    Source (en-US)

    Your cart is empty.

    Target (de-DE)

    Ihr Einkaufswagen ist leer.

  5. 05Section heading

    Source (en-US)

    Order summary

    Target (de-DE)

    Order summary

  6. 06Promotional line under the totals

    Source (en-US)

    Free shipping on orders over $50.

    Target (de-DE)

    Kostenloser Versand.

  7. 07Auth gate

    Source (en-US)

    Sign in to continue.

    Target (de-DE)

    Anmelden um fortzufahren.

  8. 08Promo badge

    Source (en-US)

    Save 20% today

    Target (de-DE)

    Sparen Sie 20% heute

  9. 09Payment error toast

    Source (en-US)

    We could not process your payment.

    Target (de-DE)

    Wir konnten Ihre Zahlung nicht bearbeiten.

  10. 10Banner with inline markup

    Source (en-US)

    <b>Limited time</b> offer

    Target (de-DE)

    <b>Angebot für kurze Zeit offer

If you remember one thing

A review round is only useful when it produces a number someone can act on: an error typology, a severity, a penalty, a threshold.

What to do about it

  • Argue about severity, not about existence. One critical error carries 25 penalty points, the same as 25 minor ones. Nearly every arbitration is really a severity dispute, so define severity with examples from your own product before the first round.
  • Score a representative sample. Penalties are normed per 1000 words, so a 40-word sample multiplies every error by 25 and produces a wild score. If you cannot sample enough, report penalty points and error counts rather than a raw score.
  • Remember the raw pass threshold sits near 99, not 90, because raw scores cluster just under 100. If a threshold like 90 is more legible to your stakeholders, use a calibrated model and say which one you used.
  • Publish the typology and the termbase before the round. A reviewer who invents categories mid-review produces a scorecard nobody can compare to the last one.

Use this with

Stakeholders

Localization PMVendor managerReviewersQA

Moments

  • ·Standing up a vendor scorecard or an SLA
  • ·Arbitrating a disputed review round
  • ·Deciding whether a delivery is shippable
  • ·Comparing MT engines or post-edit tiers with numbers

Field note

The most expensive review rounds leave no clean segments in the sample. A reviewer who flags every stylistic preference produces a scorecard that fails a delivery that was fine. The vendor disputes it, and two weeks disappear into arbitration. Good reviewer onboarding measures false positives as carefully as misses. That is why this exercise seeds clean segments and counts your over-flags.

ISO 5060:2024, Evaluation of translation output

Quick check

3 questions · pass at 2+

  1. Question 1/3

    Under the MQM default weights, one critical error carries the same penalty as how many minor errors?

  2. Question 2/3

    A German target leaves the US date format 3/4/26 in place. Which MQM dimension is the right file, and why does it matter?

  3. Question 3/3

    A review logs 1 major error in a 40-word sample. What does the raw MQM arithmetic report, and what should you conclude?

Words you'll hear

Where to read more

Related lessons