Chapter 15·Strings, names, and input·~7 min
Character encoding & mojibake
A string is code points; a file is bytes. The encoding is the map between them, and there is more than one map.
The problem
UTF-8 and UTF-16 are two ways to write the same Unicode code points as bytes, and nothing inside a byte stream says which one was used. When the writer says UTF-8 and the reader assumes Windows-1252, every non-ASCII character is misread: café is stored correctly and displayed as café. The pattern has a name, mojibake, and it accounts for most "why are there weird characters" bug reports.
How it works
A string is code points, a file is bytes
Text lives in a program as a sequence of Unicode code points: abstract numbers like U+0041 for A or U+20AC for the euro sign. A file, a network packet, or a database column holds bytes, not code points. An encoding is the map between the two. Unicode defines more than one: UTF-8 and UTF-16 store the same code points with different byte layouts.
UTF-8 is variable-width: an ASCII character is one byte. A is a single 0x41, byte-for-byte identical to 1960s ASCII. Accented Latin letters take two bytes, most other scripts three, and code points above U+FFFF (most emoji) four. UTF-16 uses a two-byte code unit for everything in the common range, so it spends two bytes even on plain ASCII. That overhead on ASCII-heavy text is a main reason UTF-8 became the web's default.
"A" U+0041 UTF-8: 41 UTF-16: 0041
"€" U+20AC UTF-8: E2 82 AC UTF-16: 20AC
"😀" U+1F600 UTF-8: F0 9F 98 80 UTF-16: D83D DE00 (surrogate pair)Mojibake is a label bug
Nothing inside a stream of bytes announces which encoding produced it. Something must tell the reader: an HTTP charset, a <meta charset> tag, a database setting, or a convention. When the writer used UTF-8 and the reader assumes an 8-bit legacy encoding like Windows-1252, the reader misreads every multi-byte character one byte at a time.
That is how café becomes café. Its UTF-8 bytes are 63 61 66 C3 A9. Read as Windows-1252, C3 is à and A9 is ©, so the two-byte é splits into two Western-European glyphs. The bytes are intact, only mislabeled, so the repair belongs on the read side: declare the right encoding rather than hand-editing the garbled text.
The demo
Try these, in order
Each step reproduces one specific failure in the demo below.
- 1On the Bytes tab, clear the field and type Hello. UTF-8 bytes = 5, UTF-16 bytes = 10. Pure ASCII costs twice as much in UTF-16, because every code unit is two bytes.
- 2Click the Aé€😀 preset and read the table rows for é, €, and 😀. UTF-8 uses 2, 3, then 4 bytes as the code point climbs. The emoji row carries the tag 'surrogate pair': two UTF-16 units for one code point.
- 3Still on Bytes, note the emoji's UTF-16 column shows two units (D83D DE00). That pair is why "😀".length is 2 in JavaScript. .length counts UTF-16 code units, and this emoji occupies two.
- 4Switch to the Mojibake tab and leave the default text. Step 3 shows café and €, while the line below confirms the same bytes decode correctly as UTF-8. The bytes are fine. Only the read-side label was wrong.
- 5On the Mojibake tab, change 'decode as' from Windows-1252 to ISO-8859-1 (Latin-1). The € changes shape. Windows-1252 turns its 0x82 byte into a low quote mark (‚); ISO-8859-1 maps 0x80–0x9F to invisible C1 control characters, shown here as ⟨82⟩. The two tables disagree only in that range.
- 6On the Mojibake tab, replace the text with plain ASCII like Receipt total. No corruption appears. ASCII bytes are identical across UTF-8 and Latin-1, so only the non-ASCII characters turn to mojibake.
Graphemes
4
what a user sees
Code points
4
[...str].length
UTF-8 bytes
10
on the wire / on disk
UTF-16 bytes
10
JS string in memory
| Char | Code point | UTF-8 bytes | UTF-16 units |
|---|---|---|---|
| A | U+0041 | 41(1B) | 0041 |
| é | U+00E9 | C3 A9(2B) | 00E9 |
| € | U+20AC | E2 82 AC(3B) | 20AC |
| 😀 | U+1F600 | F0 9F 98 80(4B) | D83D DE00surrogate pair |
"😀".length is 2 in JavaScript.The short version
A string holds code points, a file holds bytes, and the encoding is the map between them. Mojibake is what happens when the reader uses a different map than the writer.
What to do about it
- Use UTF-8 everywhere: files, HTTP responses, database columns and connections (
utf8mb4in MySQL), message queues. - Declare the encoding, don't guess it. Send
Content-Type: …; charset=utf-8, put<meta charset="utf-8">first in<head>, and open files with an explicit UTF-8 encoding rather than the platform default. - Never "repair" mojibake by hand-editing the garbled glyphs. The bytes are usually fine; fix the label on the read side, or the mismatch moves downstream.
- In JavaScript,
"😀".lengthis 2: UTF-16 stores astral code points as surrogate pairs, so byte and length budgets computed on.lengthundercount real storage and miscount user-visible characters.
Where this comes up
Who it concerns
Moments
- ·Debugging 'weird characters' in exports, emails, or CSV imports
- ·Setting database and connection charset before a launch
- ·Reviewing file / HTTP encoding in a code review
Field note
The most common production version of this is a database rather than a file. A MySQL column declared `utf8` (a three-byte subset) or a connection that negotiates `latin1` stores correct UTF-8 text as mojibake. The damage is invisible until someone reads it back. Set columns and the connection to `utf8mb4` before the first insert. A migration after the fact means decoding the corruption byte by byte.