Chapter 06·The Intl toolbox·~6 min
Sorting: the alphabet is locale data
Polish ł sorts between l and m. Swedish ä sorts after z. German has two collations (dictionary vs phonebook), Spanish used to treat 'ch' as a single letter. Icelandic phone books sort by first name.
The problem
Array.sort() sorts by UTF-16 code-unit order, which files every accented letter by its encoding accident. A user list in a German UI shows Ö after Z. A Polish UI shows ł with the other accented Latin letters instead of between L and M. A mixed Hebrew-and-Latin list clumps by script, the basic Latin range sorting wholesale before the Hebrew block. The list is sorted, just not in an order any reader can predict.
How it works
Sorting is an algorithm with opinions, not a byte order
The Unicode Collation Algorithm compares strings in passes: base letters first, then accents, then case. That is why Élan files under E instead of after zebra: É differs from E only at the second level, and the second level only matters when the first is a tie. A code-point sort has no levels. It files every word that starts with an accented letter by its encoding accident.
Each language then tailors the default table to its own alphabet: Swedish appends å ä ö after z as letters 27–29, and Polish gives ł its own slot between l and m. German maintains two standardized orders, dictionary and phonebook, the second expanding ä to ae so Müller files with Mueller. Tailorings are selected in the locale tag itself: de-DE-u-co-phonebk is a complete, valid answer to "which alphabetical order?".
So any list a user reads as sorted goes through Intl.Collator with the user's locale. The same constructor's numeric: true fixes file10 sorting before file2 for free. Plain sort() is code-point order, which is wrong somewhere for nearly every language, including English the moment an accent or an uppercase letter appears.
Matching is a different tailoring than ordering
Sorting decides where a name files. Search decides whether two spellings count as the same name. CLDR ships separate rules for each job, and new Intl.Collator(locale, { usage: 'search' }) loads the second set. Pair it with sensitivity: base ignores accents and case, accent counts accents but not case, case counts case but not accents, variant counts both. The match test is compare(query, candidate) === 0: the collator compares whole strings, so a filter or find-as-you-type runs it against each candidate in turn.
The folding rules are locale opinions. To an English searcher, ü is a decorated u, so muller finds Müller. German search folds it to ue instead: muller misses and mueller hits, because that is the equivalence German spelling uses. Spanish refuses to fold ñ at all, even at base sensitivity, which is why nunez does not find Núñez under es. Stripping accents with a regex hardcodes one language's opinion into every market.
const de = new Intl.Collator("de", { usage: "search", sensitivity: "base" });
const en = new Intl.Collator("en", { usage: "search", sensitivity: "base" });
en.compare("muller", "Müller") // 0: match (ü folds to u)
de.compare("muller", "Müller") // ≠ 0: no match (ü ≡ ue in German)
de.compare("mueller", "Müller") // 0: matchThe demo
Try these, in order
Each step reproduces one specific failure in the demo below.
- 1On the default Polish list, scan the highlighted cells across the en and pl columns. Łukasiewicz and Łapińska move: Polish sorts ł as its own letter between l and m. English UCA folds it near l. Same list, different positions.
- 2Switch to “German names · dictionary vs phonebook ordering”. Müller vs Mueller swap between the de and de-DE-u-co-phonebk columns: the collation variant is selected in the BCP-47 tag itself (-u-co-phonebk).
- 3Switch to the Swedish list. Ångström, Älg, Östergren sink to the end under sv: å/ä/ö are letters 27–29, after z. The Danish column orders them differently again. Scandinavian is not one collation.
- 4Load “Case sensitivity” and compare the three English columns. Default, uppercase-first (-u-kf-upper), and sensitivity:'base' each order the apricot/Apricot and blueberry/Blueberry pairs differently: the default puts lowercase first, kf-upper puts capitals first, and base calls the pair equal and leaves it in input order. Case handling is one more tailorable collation decision.
- 5Type a few entries into “Your list · comma-separated”, and include at least one accented word, e.g. Zorro, Öberg, chávez, Lukas, łódź. Two or more entries replace the sample list and every column re-sorts on each keystroke. The accented entries land differently per column: the highlighting marks each row where a locale disagrees with the English default.
- 6In “Sort is not search”, keep the defaults (Search locale English, Sensitivity base, Search query muller), then switch Search locale to German (de). Under English, Müller matches (ü folds to u). Under German it does not: German search folding maps ü to ue, so the six-letter query no longer lines up. Now type mueller into Search query: German matches both Müller and Mueller, while English matches only Mueller.
Each column is the same list, sorted by Intl.Collator with the matching locale tag. Read across a row to see where a name lands in each sort. Or type two or more entries above to sort a custom list.
| # | en English (UCA default) | pl Polish |
|---|---|---|
| 01 | Łaba | Lewandowski |
| 02 | Łapińska | Lewiński |
| 03 | Lewandowski | Lipiński |
| 04 | Lewiński | Lis |
| 05 | Lipiński | Łaba |
| 06 | Lis | Łapińska |
| 07 | Łukasiewicz | Łukasiewicz |
| 08 | Maciejewski | Maciejewski |
| 09 | Małachowski | Małachowski |
| 10 | Modliński | Modliński |
Why this happens
CLDR ships a per-locale collation table: the canonical order of letters, the rules for accents, the rules for case. The Unicode Collation Algorithm (UCA) is the default, and locales override it. Intl.Collator reads the active locale, applies the overrides, and returns a comparator function for .sort().
- Polish ł is a separate letter that sorts between l and m. ASCII-sort puts it after z.
- German phonebook (DIN 5007-2) treats ä as if it were
ae, so Müller interleaves with names spelled Mueller. - Swedish sorts å, ä, ö after z as the last three letters of the alphabet. Apps that sort ä with a look broken to Swedish users.
- Icelandic sorts people by their given name because patronymics change per generation. The national phone book does this. CRM software that filters by "last name" breaks here.
- Case sensitivity is locale data too. For most apps the correct default is
sensitivity: "base"with a secondary tiebreaker.
The fix
Wherever a codebase calls .sort() on a user-visible list, replace the implicit string comparator with new Intl.Collator(locale).compare. That fixes sorting everywhere CLDR ships a tailoring, and turns the remaining edge cases into a visible bug a team can escalate to a linguist.
Ordering a list and deciding whether two strings mean the same name are different jobs. CLDR ships separate tailorings for them. usage: "search" loads the matching rules. sensitivity sets which differences count. A candidate matches when compare(query, name) === 0. The API compares whole strings, so a filter runs it against each entry in turn.
new Intl.Collator("en", { usage: "search", sensitivity: "base" })
- Müllermatch
- Muellerno match
- Möllerno match
- Núñezno match
- Åströmno match
- Célineno match
The folding rules are locale data too
Try it: under English with base, muller matches Müller. Switch the search locale to German and the match disappears, because German search collation folds ü toward ue, not toward u. Type mueller and German matches Müller while English does not. Each locale defines which spellings its users consider the same name, so a product that normalizes search with one hardcoded rule is wrong somewhere.
usage: "search" and an explicit sensitivity. Do not reuse the sort collator or strip accents with a regex. base is the forgiving default users expect from a search box.The short version
Alphabetical order is locale data. Use Intl.Collator for anything a user reads as a sorted list. Code-point sort is wrong in most languages, sometimes in two ways in one language.
What to do about it
- Always use
Intl.Collatorto compare strings for display. It uses CLDR collation data and respects the user's locale. - For multi-language sorting (one list, many locales), pick a "neutral" collator (
"und"or the UI's locale) and applysensitivity: "base"to fold accents together rather than guessing per language. - For numeric strings ("file2", "file10"), pass
{ numeric: true }. Otherwise file10 sorts before file2, every time. - On the database side, set the column collation to match the display locale (e.g. PostgreSQL
COLLATE "de-DE-x-icu"). Application-layer sort is cheap; database sort needs explicit configuration.
Where this comes up
Who it concerns
Moments
- ·Designing any sortable user list, table, or directory
- ·Database schema review
- ·Launching into a locale with non-trivial alphabet ordering
See it in a market: Sweden: where ä sorts after z →
Field note
Germany standardized two different alphabetical orders: DIN 5007-1 (dictionary: ä sorts with a) and DIN 5007-2 (phonebook: ä sorts as ae, so Müller files with Mueller). Neither is a bug: they serve different lookup tasks. A plain .sort() implements neither: it files ä after z, by code point.