Skip to content
Locale Lab

Chapter 14·Strings, names, and input·~8 min

Unicode pitfalls: strings that lie about what they are

'ID'.toLocaleLowerCase('tr') returns 'ıd', and a default-locale lowercase on a Turkish host is how Spring CVE-2024-38820 happened. 'family emoji'.length is 11, not 1.

Unicodeencodingnormalizationcase folding

The problem

JavaScript strings are sequences of UTF-16 code units, not characters. "café" is four units or five depending on whether the é is precomposed (NFC) or decomposed (NFD). Lowercase conversion is locale-dependent: Turkish has a dotless i. Equality comparison without normalization fails for usernames typed on different platforms, and a uniqueness check that skips it lets two accounts look identical.

How it works

Three different answers to the length of a string

A string has three lengths, and JavaScript defaults to the least useful one. There are code points (what Unicode assigns), UTF-16 code units (what .length counts, with surrogate pairs counting double), and grapheme clusters (what a human calls one character). The family emoji is one grapheme built from seven code points occupying eleven code units. Truncate by .length and a "character limit" can slice between a mother and her children. Intl.Segmenter with grapheme granularity is the only count that matches what users see.

Equality is layered the same way. Composed café (é as one code point) and decomposed café (e plus a combining accent) render identically yet compare unequal. Normalizing at the system's boundaries, usually to NFC, is what collapses the two spellings into one comparable form. Skip that step and two "identical" usernames coexist while lookups quietly miss.

Even lowercasing needs a locale. Turkish distinguishes dotted and dotless i, so 'ID'.toLocaleLowerCase('tr') is ıd, not id. That divergence caused a real Spring CVE, where a blocklist compared post-lowercase on the host's locale. The rules: normalize at the boundary, compare with the right tool (a pinned locale for security checks, Intl.Collator for humans), and never slice UI text by UTF-16 index.

The demo

Try these, in order

Each step reproduces one specific failure in the demo below.

  1. 1
    On the Turkish I tab, type ID (plain ASCII) into the input. en lowercases to “id”, while tr gives “ıd”, dotless. The equality badge flips false. That divergence is the CVE mechanism: a blocklist that compares post-lowercase misses on Turkish hosts.
  2. 2
    Open the NFC vs NFD tab and read the code-point dumps. Two visually identical cafés: 4 code points vs 5. The === badge says false until .normalize('NFC') makes them comparable. Filenames from older macOS volumes (HFS+) arrive NFD, and typed input arrives NFC.
  3. 3
    Still on that tab, scroll to “Normalize your own string”. Read the three rows for the seeded text, then replace it with any accented word or a filename pasted from Finder. The seed is café spelled the decomposed way. “As typed” shows .length = 5 and five code points ending in U+0301. The NFC row shows the same word at 4, and the input === input.normalize('NFC') badge reads false. Whatever the pasted text is, the badge reports whether that byte sequence would survive an NFC boundary unchanged.
  4. 4
    On the Emoji length tab, pick the family emoji 👨‍👩‍👧‍👦. .length says 11, code points say 7, graphemes say 1. Any character-limit, truncation, or cursor logic keyed to .length cuts through the middle of a person.
  5. 5
    Pick the Scotland flag 🏴󠁧󠁢󠁳󠁣󠁴󠁿. A black flag followed by invisible “tag characters” spelling gbsct: even longer in code units. One visible glyph, fourteen code units.

The same string, lowercased in English and in Turkish. Try İD, STRAİT, or the literal word ID.

toLocaleLowerCase("en-US")

i̇d

toLocaleLowerCase("tr-TR")

id

en === tr ?

false

Security implication

A blocked-list filter rejects fields named "id" or "userid", written with bare toLocaleLowerCase(), no locale argument. That call follows the host locale: on a Turkish user's device it lowercases ID to ıd (dotless ı), which the filter no longer matches. (The default sample İD shows the reverse: on a Turkish device it lowercases to plain id and collides with a name it should not match.)

Locale-independent toLowerCase() blocks "İD": no · Bare toLocaleLowerCase() on a Turkish device blocks it: yes

This is the shape of CVE-2024-38820 in Spring, and similar bugs in GitHub auth and Active Directory.

The fix: Never make security or routing decisions with bare toLocaleLowerCase(). Pin the locale with toLocaleLowerCase("en-US"), or use plain toLowerCase(), which is locale-independent (root-locale case mapping) by spec.

The short version

A string is not what it looks like. Normalize before comparing, segment by graphemes before counting or slicing, and pin the locale for case operations.

What to do about it

  • Normalize every user-supplied string on input with .normalize("NFC") before storing or comparing.
  • Keep security comparisons locale-free. JavaScript's .toLowerCase() already applies locale-independent mappings; Java-family runtimes need an explicit Locale.ROOT, because their default follows the host. Reserve .toLocaleLowerCase("tr") and friends for text shown to that locale's readers.
  • For string length that matches user perception, use [...str].length (counts code points) or Intl.Segmenter with granularity: "grapheme" (counts user-visible characters, including emoji ZWJ sequences).
  • Run confusable-skeleton checks (UTS #39) on usernames and handles to prevent homograph spoofing (e.g., Cyrillic а vs Latin a).

Where this comes up

Who it concerns

EngineeringSecurityQA

Moments

  • ·Designing username / handle uniqueness rules
  • ·Pre-merge code review for any string comparison
  • ·Security audit

Field note

Spring Framework's CVE-2024-38820 is the Turkish-i bug with a CVSS score. Case-insensitive field matching used a default-locale lowercase, so on certain locales a disallowed field name stopped matching its blocklist entry. A one-argument fix (Locale.ROOT / a pinned locale) prevents the whole class.

Spring: CVE-2024-38820 ↗

Terms in this chapter

Where to read more

Related chapters