Skip to content

Text expansion in 55 languages

16/100
15px180px
Translating into 55 languages…

Where these numbers come from

The framing here follows the W3C i18n article Text size in translation, which is also where the IBM allowance quoted above comes from.

A string changes size for two reasons, and they do not agree. The linguistic one is character count. The typographical one is how much room those characters take, which depends on the script: Han characters set about twice as wide as an average Latin one, so Chinese sheds three quarters of the characters here and only half the width. The disagreement runs the other way too: Indonesian drops a character on this string and still sets marginally wider, and Thai comes out shorter on both counts while setting about a third taller, which is the axis a width budget cannot see.

The translations are machine translation, from Google Translate’s public endpoint. Machine and human output run to broadly similar lengths in a given language, so a row is a fair guide to the size a translator will hand back, and no guide at all to quality.

The sample strings are the exception: they are answered without a request. Their 55 rows apiece were translated once, reviewed, and checked into the source. 29 of those stored rows are hand corrections, nearly all of them the button label: “Save” with no context around it is translated as economize into Spanish, Danish and Chinese and as rescue into Dutch, Polish, Korean and Vietnamese, and Greek came back as the preposition “except”. A custom string still goes to the live endpoint, uncorrected.

Widths come from canvas.measureText in this browser, after the chosen face has loaded. No face here covers every script, so rows in Arabic, CJK, Thai and the Indic scripts are measured in the page’s fallback stack, which is what a product shipping no bundled font also gets. Ink height is actualBoundingBoxAscent plus actualBoundingBoxDescent: the painted box, not the line box.

Box fit lays each translation out the way a control does, and the line breaking is the hard part: Thai, Lao, Khmer, Burmese, Japanese and Chinese write no spaces, so splitting on them reports the whole string as one unbreakable word. Intl.Segmenter supplies the break candidates instead, which carries ICU’s dictionary segmenters for those scripts, and the truncation lands on a grapheme cluster so a Devanagari or Tamil row cannot end in a dotted circle. The percentage cut is width, not characters: the whole string minus the part shown, because measuring the hidden tail on its own would change the Arabic joining forms at the seam.

A short string grows by a larger fraction than a long one, so a button label is the honest test. The field guide chapter on text expansion covers why, and what to reserve.

Related: Field guide: the text expansion chapter · Budget estimator · price the round that produces these strings