Skip to content
Locale Lab

Chapter 15·Strings, names, and input·~7 min

Character encoding & mojibake

A string is code points; a file is bytes. The encoding is the map between them, and there is more than one map.

UTF-8UTF-16mojibakecharset

The problem

UTF-8 and UTF-16 are two ways to write the same Unicode code points as bytes, and nothing inside a byte stream says which one was used. When the writer says UTF-8 and the reader assumes Windows-1252, every non-ASCII character is misread: café is stored correctly and displayed as café. The pattern has a name, mojibake, and it accounts for most "why are there weird characters" bug reports.

How it works

A string is code points, a file is bytes

Text lives in a program as a sequence of Unicode code points: abstract numbers like U+0041 for A or U+20AC for the euro sign. A file, a network packet, or a database column holds bytes, not code points. An encoding is the map between the two. Unicode defines more than one: UTF-8 and UTF-16 store the same code points with different byte layouts.

UTF-8 is variable-width: an ASCII character is one byte. A is a single 0x41, byte-for-byte identical to 1960s ASCII. Accented Latin letters take two bytes, most other scripts three, and code points above U+FFFF (most emoji) four. UTF-16 uses a two-byte code unit for everything in the common range, so it spends two bytes even on plain ASCII. That overhead on ASCII-heavy text is a main reason UTF-8 became the web's default.

"A"    U+0041    UTF-8: 41             UTF-16: 0041
"€"    U+20AC    UTF-8: E2 82 AC       UTF-16: 20AC
"😀"   U+1F600   UTF-8: F0 9F 98 80    UTF-16: D83D DE00  (surrogate pair)
One code point, two byte layouts

Mojibake is a label bug

Nothing inside a stream of bytes announces which encoding produced it. Something must tell the reader: an HTTP charset, a <meta charset> tag, a database setting, or a convention. When the writer used UTF-8 and the reader assumes an 8-bit legacy encoding like Windows-1252, the reader misreads every multi-byte character one byte at a time.

That is how café becomes café. Its UTF-8 bytes are 63 61 66 C3 A9. Read as Windows-1252, C3 is à and A9 is ©, so the two-byte é splits into two Western-European glyphs. The bytes are intact, only mislabeled, so the repair belongs on the read side: declare the right encoding rather than hand-editing the garbled text.

The demo

Try these, in order

Each step reproduces one specific failure in the demo below.

  1. 1
    On the Bytes tab, clear the field and type Hello. UTF-8 bytes = 5, UTF-16 bytes = 10. Pure ASCII costs twice as much in UTF-16, because every code unit is two bytes.
  2. 2
    Click the Aé€😀 preset and read the table rows for é, €, and 😀. UTF-8 uses 2, 3, then 4 bytes as the code point climbs. The emoji row carries the tag 'surrogate pair': two UTF-16 units for one code point.
  3. 3
    Still on Bytes, note the emoji's UTF-16 column shows two units (D83D DE00). That pair is why "😀".length is 2 in JavaScript. .length counts UTF-16 code units, and this emoji occupies two.
  4. 4
    Switch to the Mojibake tab and leave the default text. Step 3 shows café and €, while the line below confirms the same bytes decode correctly as UTF-8. The bytes are fine. Only the read-side label was wrong.
  5. 5
    On the Mojibake tab, change 'decode as' from Windows-1252 to ISO-8859-1 (Latin-1). The € changes shape. Windows-1252 turns its 0x82 byte into a low quote mark (‚); ISO-8859-1 maps 0x80–0x9F to invisible C1 control characters, shown here as ⟨82⟩. The two tables disagree only in that range.
  6. 6
    On the Mojibake tab, replace the text with plain ASCII like Receipt total. No corruption appears. ASCII bytes are identical across UTF-8 and Latin-1, so only the non-ASCII characters turn to mojibake.

Graphemes

4

what a user sees

Code points

4

[...str].length

UTF-8 bytes

10

on the wire / on disk

UTF-16 bytes

10

JS string in memory

CharCode pointUTF-8 bytesUTF-16 units
AU+004141(1B)0041
éU+00E9C3 A9(2B)00E9
€U+20ACE2 82 AC(3B)20AC
😀U+1F600F0 9F 98 80(4B)D83D DE00surrogate pair
A plain-ASCII character is 1 byte in UTF-8 but 2 bytes in UTF-16. This is why UTF-8 became the web's default. Anything above U+FFFF (most emoji) is a single code point that UTF-16 must store as a surrogate pair of two units, which is why "😀".length is 2 in JavaScript.

The short version

A string holds code points, a file holds bytes, and the encoding is the map between them. Mojibake is what happens when the reader uses a different map than the writer.

What to do about it

  • Use UTF-8 everywhere: files, HTTP responses, database columns and connections (utf8mb4 in MySQL), message queues.
  • Declare the encoding, don't guess it. Send Content-Type: …; charset=utf-8, put <meta charset="utf-8"> first in <head>, and open files with an explicit UTF-8 encoding rather than the platform default.
  • Never "repair" mojibake by hand-editing the garbled glyphs. The bytes are usually fine; fix the label on the read side, or the mismatch moves downstream.
  • In JavaScript, "😀".length is 2: UTF-16 stores astral code points as surrogate pairs, so byte and length budgets computed on .length undercount real storage and miscount user-visible characters.

Where this comes up

Who it concerns

EngineeringData / infraQA

Moments

  • ·Debugging 'weird characters' in exports, emails, or CSV imports
  • ·Setting database and connection charset before a launch
  • ·Reviewing file / HTTP encoding in a code review

Field note

The most common production version of this is a database rather than a file. A MySQL column declared `utf8` (a three-byte subset) or a connection that negotiates `latin1` stores correct UTF-8 text as mojibake. The damage is invisible until someone reads it back. Set columns and the connection to `utf8mb4` before the first insert. A migration after the fact means decoding the corruption byte by byte.

Terms in this chapter

Where to read more

Related chapters