Skip to content
Locale Lab

Chapter 02·Three visible failures·~6 min

Text expansion stress test

Translated copy is rarely the same length as the source. German and Russian run longer, Finnish builds compounds that don't fit in a button, and the shorter the string, the worse the swell. Designs built for English break first at the smallest breakpoint.

layoutswellQA

The problem

The default English string says "Sign in". The German translation is "Anmelden" (fine), the Finnish is "Kirjaudu sisään" (twice as long), and the Russian is "Войти в систему" (also longer, plus Cyrillic widths). A button that fits one breaks the others, and the design system's paddings, icon spacings, and badge widths break with it.

How it works

Why translations do not fit

Translations change length for structural reasons, not careless ones. German fuses ideas into compounds where English uses short phrases. Russian and Finnish carry grammar inside longer word forms, and French spends characters on articles and spacing. The shorter the source string, the larger the proportional growth, because a four-letter label has nowhere to hide the grammar another language needs. That is the basis of the IBM guidance W3C republishes: up to 10 source characters, plan for 200–300% growth.

The effect also runs the other way: Chinese and Japanese often say an English sentence in a fraction of the width, so the same component must look right both fuller and emptier than the mockup. The translator controls none of this: the layout decides whether length becomes a bug. Fixed widths truncate. Flexible containers with sensible minimums, wrapping room, and real-string testing absorb the change.

The counter in the editor is not the limit in the box

Sooner or later someone hands a translator a character budget: keep this button under 16. That budget is only useful if everyone agrees what a character is, and three common answers disagree: str.length in JavaScript counts UTF-16 code units, spreading the string counts code points, and Intl.Segmenter at grapheme granularity counts what a reader would call characters. All three match on plain ASCII, which is how the disagreement survives review and reaches production.

They diverge on the text localization deals in. Take a thumbs-up with a skin tone: one grapheme, two code points, four UTF-16 units. A budget enforced with length charges that emoji four times, rejecting a string that fits on screen. Or take an accented letter typed in decomposed form. It counts one higher than the same letter precomposed, so the same word passes or fails depending on which keyboard produced it.

The larger problem is that none of the three is what the UI enforces. A button truncates at a pixel width, and pixel width has no fixed relationship to any character count: six Japanese characters can render wider than eleven Latin ones. When the layout is the real constraint, measure the layout, and give translators the box rather than a number.

The demo

Try these, in order

Each step reproduces one specific failure in the demo below.

  1. 1
    Keep the “CTA button” scenario and drag the container width down until German (“In den Warenkorb”) breaks. Note the width where en-US still fits comfortably. That gap is where the bug appears. The % badges show each locale's swell against English.
  2. 2
    Find Chinese (加入购物车) and Japanese in the grid. CJK often contracts: the badge shows negative expansion.
  3. 3
    Switch the scenario to “Notification card”. Longer source strings swell by a smaller percentage: the 200–300% figure is a short-string problem. Compare the badges against the button scenario.
  4. 4
    Pick the narrowest width where every locale still renders acceptably. That is the component's minimum content width to hand the design system: a measured number instead of a guess.
en-US+0%

Add to cart

de-DE+45%

In den Warenkorb

fr-FR+55%

Ajouter au panier

es-ES+55%

Añadir al carrito

pt-BR+91%

Adicionar ao carrinho

it-IT+82%

Aggiungi al carrello

ru-RU+64%

Добавить в корзину

pl-PL+45%

Dodaj do koszyka

fi-FI+55%

Lisää ostoskoriin

zh-CN-55%

加入购物车

ja-JP-45%

カートに追加

ar-SA+18%

أضف إلى السلة

One string, three different sizes

Code points, UTF-8 bytes, and rendered pixels disagree: Mandarin usually has the fewest characters yet renders among the widest.

LocaleTitlecharsUTF-8 bytespxvs. en-US
en-USAdd to cart1111—0%
de-DEIn den Warenkorb1616—+45%
fr-FRAjouter au panier1717—+55%
es-ESAñadir al carrito1718—+55%
pt-BRAdicionar ao carrinho2121—+91%
it-ITAggiungi al carrello2020—+82%
ru-RUДобавить в корзину1834—+64%
pl-PLDodaj do koszyka1616—+45%
fi-FILisää ostoskoriin1719—+55%
zh-CN加入购物车515—-55%
ja-JPカートに追加618—-45%
ar-SAأضف إلى السلة1324—+18%

Once measured, the "vs. en-US" column compares rendered pixel width, the only one of the three that decides whether a translation overflows its box.

Writing to a limit

A translator receives a character budget and a box. The counter in the editor, the number in the spec, and the width the UI enforces are three different measurements. They disagree on non-ASCII text, which is where the limits bind.

Graphemes

16

what the reader sees

Code points

16

[...str].length

UTF-16 units

16

.length, what most tools count

Pixels

0

box is 120px

In the real container

In den Warenkorb

Inside both limits. Try the Japanese preset: six characters, and watch the pixel counter.

The three counts diverge only on non-ASCII text. An emoji with a skin tone is one grapheme, two code points, and four UTF-16 units. A budget expressed in .length charges it four times.

The "30% rule" does not hold

Localization training material repeats the line "translations expand by ~30%." For short UI strings (1–3 words) the real spread is −50% (CJK) to +70% (Portuguese, Russian, German). Size boxes in pixels, not characters. Give icon-only labels a generous min-width and allow subtitles to wrap to two lines. Avoid right-edge absolute positioning for CTAs: it lands on the wrong side in Arabic and Hebrew.

The short version

Budget +200–300% width for short strings and design layouts that flex. If the German fits, most things fit.

What to do about it

  • Budget width by source length. The W3C planning table runs from 200–300% growth for strings up to 10 characters down to about 30% for paragraph-length text. Verify by pseudo-localizing the design files (Figma plugins exist) or running the live UI through the pseudoloc chapter's transform.
  • Avoid fixed-width text containers. Use min-width + flexible max-width so buttons grow with their content instead of truncating.
  • Where truncation is unavoidable (status badges, avatars), pair the clipped label with a way to read the full string. Icon-only buttons avoid expansion entirely, but an icon's meaning does not always travel; confirm it reads the same way in each shipping market.
  • Build a "longest string" test that runs in CI against the real translation memory; the source-language bound is a poor substitute.

Where this comes up

Who it concerns

DesignEngineeringProduct

Moments

  • ·Design review for any UI change
  • ·Component library audit
  • ·Pre-launch QA for German / Russian / Finnish

Field note

IBM's classic expansion table, republished by W3C, is the industry rule of thumb: the shorter the English source, the worse the swell. Buttons, tabs, and badges are where teams hardcode widths, and also where translations are longest relative to source.

W3C: Text size in translation ↗

Terms in this chapter

Where to read more

Related chapters