Skip to content
Locale Lab

Chapter 01·Getting oriented·~6 min

Orientation: i18n, l10n, and the standards stack

One order confirmation (a count, a date, a price) rendered in four locales by the browser's own Intl APIs. Four layers did the work: Unicode supplied the characters, CLDR the locale data, the ICU/Intl engine applied it, and a BCP 47 tag named what was asked for.

i18n vs l10nCLDRICUBCP 47

The problem

Most teams defer this with one sentence: “we’ll localize later.” By the time later arrives, the codebase has hardcoded the assumptions English allowed: dates concatenated field by field, an s appended when count !== 1, layouts that only work left-to-right, prices formatted with toFixed(2). None of these are translation problems, so translators can’t fix them; they are design problems. The other cost of skipping orientation is vocabulary: a team that can’t distinguish i18n work from l10n work can’t scope, staff, or schedule either one.

How it works

Four words, two jobs

The field compresses its own vocabulary into numeronyms: take the first letter, count the letters in between, keep the last. "Internationalization" is i, then 18 letters, then n: i18n. "Localization" is l10n, "globalization" is g11n, "translation" is t9n. The abbreviations look interchangeable, so teams use the words interchangeably. They name different jobs.

Internationalization is engineering, done once: the product is designed so a team can adapt it later. No hardcoded date order, no count !== 1 plural logic, no layout that only works left to right. Localization is production, done per market: someone writes the German strings, chooses the Egyptian date conventions, and adapts imagery and legal text. Translation is the part of localization that converts text, so t9n is the narrowest of the four terms. Globalization (g11n; GILT also circulates) is the business umbrella term for all of it.

The split makes the work schedulable. i18n is a fixed cost that lives in the codebase. l10n recurs with every added market. A team that mixes them budgets neither: it over-scopes translation ("we need engineers for the French launch") or under-scopes engineering ("the translators will handle the dates"). Translators cannot fix a hardcoded date order, and that bug was cheap to fix at design time.

i18n = i + 18 letters + n   internationalization
l10n = l + 10 letters + n   localization
g11n = g + 11 letters + n   globalization
t9n  = t +  9 letters + n   translation
The numeronym rule: first letter + letter count + last letter.

The stack: who owns what

Every formatted string in the demo above passed through four layers, each with a different owner. At the bottom, the Unicode Consortium defines the characters: every character gets a code point and properties, including the bidirectional class that makes Arabic run right to left. The same consortium maintains CLDR, the Common Locale Data Repository. That is the data layer: month names per language, preferred digit shapes, plural rules, currency symbol placement, separators. Unicode owns what the characters are, and CLDR records how each locale uses them.

Data needs an engine. ICU (International Components for Unicode) is the C/C++ and Java library that applies CLDR data, and it is the industry's reference implementation. In a browser that machinery surfaces as Intl.*, an API that ECMA-402 standardizes. The locale data ships inside the JavaScript engine. That is why new Intl.PluralRules('ru') knows Russian's four cardinal shapes without a download: the rules shipped inside the browser, the same way they ship inside a phone's OS.

The top layer does no formatting at all. BCP 47 (an IETF standard, currently RFC 5646 for the tags and RFC 4647 for matching them) defines the identifier passed in: pt-BR is the language subtag pt plus the region subtag BR. It is the lookup key for every other layer, and many bugs live there: ask for the wrong tag, and three working layers return the wrong answer. Keeping the name of the request separate from the data and the engine is what lets one codebase serve any locale through a change to a single string.

"ar-EG"                              BCP 47: names the request
  plural rules, digits, month names  CLDR: the data for that locale
  Intl.NumberFormat / PluralRules    ICU / Intl (ECMA-402): applies it
  ٣ عناصر                            Unicode: the characters on screen
One request, top to bottom through the stack.

The demo

Try these, in order

Each step reproduces one specific failure in the demo below.

  1. 1
    Start on the ar-EG tab (selected when the page loads) and read the whole sentence. The digits are Arabic-Indic (٣, ٢٠٢٦), the month is أغسطس, and the currency amount reads right to left. Nobody authored any of it by hand: the count, date, and price all came out of the stack below.
  2. 2
    Under “Who did what”, select CLDR. The count, date, and price light up while the surrounding words dim. The dimmed words are translation: a person wrote them (l10n). The highlighted parts are locale data.
  3. 3
    Under Count, click 3, then 11. In “How the count resolved”, ar-EG moves from category few to many. The Arabic noun changes form (عناصر → عنصرًا), because 11 takes a singular-form noun in Arabic. No application code branched. A CLDR rule decided.
  4. 4
    Switch to the ja-JP tab and click every Count button, 1 through 100. The category never leaves other: Japanese has one cardinal category, so a single arm covers every count. The number of arms is a property of the language rather than of the UI.
  5. 5
    Select BCP 47. The highlight leaves the sentence entirely and lands on the active tab. The tag formats nothing: ja is a language subtag, JP a region subtag, and together they are the lookup key for the other three layers.
  6. 6
    Select Unicode, then flip between the en-US and ar-EG tabs. The static text highlights, and the direction note under the sentence changes: the Arabic characters carry a right-to-left bidirectional property. That is why ar-EG needs dir="rtl" and en-US does not.

One sentence, four locales

Same order: a count, a ship date, a price in USD.

طلبك المكوّن من ٣ عناصر يُشحن في ٥ أغسطس ٢٠٢٦ مقابل ‏٤٢٫٥٠ US$.

Arabic (Egypt) · dir="rtl"

Count

How the count resolved

Intl.PluralRules("ar-EG").select(3) → "few"

zero# عنصر
oneعنصر واحد
twoعنصرين
few# عناصر
many# عنصرًا
other# عنصر

ar-EG has 6 cardinal categories (a zero category also exists: 0 resolves to it). The arms are translations. Which arm renders is a CLDR rule, applied by the engine.

Who did what · select a layer

Four layers cooperated to render the sentence above. Select one to highlight its share of the work.

Inspect any BCP 47 tag

A tag starts with a language subtag, then optional script, region, variants, and a -u- extension. Type one to see it canonicalized, split, and expanded through CLDR likely-subtags.

Language

zhChinese

Script

HantTraditional

Region

HKHong Kong SAR China

Maximize (add likely subtags)

zh-Hant-HK

Minimize (drop the redundant)

zh-HK

Maximize adds the script and region CLDR considers most likely (so en becomes en-Latn-US). Minimize strips whatever is implied. Negotiation uses these to match zh-HK against a zh-Hant bundle.

The short version

Four layers render every localized string: Unicode supplies the characters, CLDR the locale data, the ICU/Intl engine applies it, and a BCP 47 tag names the request. Route every locale bug to its layer.

What to do about it

  • Learn the four layers before touching the APIs. Unicode (characters and their properties), CLDR (the locale data), ICU / Intl (the engine that applies it), BCP 47 (the tag that names the request). Every later chapter is one of these layers up close.
  • Route every bug to its layer. A wrong month name is data (CLDR); a crash or missing API is engine (Intl); text running the wrong direction is characters (Unicode); the wrong language entirely is usually the tag (BCP 47) or how it was negotiated.
  • Split the work by the two words: i18n is engineering that makes one codebase capable of rendering any locale; l10n is producing the content for one market. Budget them separately; one is a fixed cost, the other recurs per market.
  • Default to the platform. The browser already ships this entire stack; most of the failures in this guide come from reimplementing some layer of it by hand.

Where this comes up

Who it concerns

New hiresProductEngineering

Moments

  • ·Onboarding anyone who will touch international work
  • ·Scoping a first-market launch: splitting i18n from l10n line items
  • ·Triaging a locale bug to the layer that owns it

Field note

The W3C internationalization FAQ exists because teams keep asking whether i18n and l10n are the same thing. It is blunt about the cost of deferring the engineering half: a retrofit of a single-culture product for a global market is “much more difficult and time-consuming” than designing for it, a point it illustrates with the Y2K effort to undo two-digit year fields.

W3C: Localization vs. Internationalization ↗

Terms in this chapter

Where to read more

Related chapters