Chapter 16·Strings, names, and input·~7 min
Parsing what users type: numbers, dates & digits
Intl.NumberFormat turns 1234.56 into 1.234,56 for a German user, and nothing in ECMA-402 turns it back. '01/02/03' names three different dates in three markets.
The problem
A checkout has a free-text amount field. A customer in Munich types 1.234,56; the backend runs parseFloat and books 1.234: three orders of magnitude short, with no error raised anywhere. A customer in Cairo types ١٬٢٣٤٫٥٦ and gets NaN. Formatting bugs at least show up on screen; parsing bugs corrupt stored data.
How it works
Formatting has an inverse, and nobody ships it
Every Intl API formats. Intl.NumberFormat has format, formatToParts, formatRange, and nothing that reads a string back. ECMA-402 ships no number-parsing API at all, so the moment a user types a localized value into a text field, the code is past the edge of the standard library.
JavaScript's own readers accept one grammar. Number() accepts the spec's numeric grammar (ASCII digits, one . as the decimal point, no grouping characters) and it is all-or-nothing: Number('1.234,56') is NaN. parseFloat() is the dangerous one. It reads the longest valid prefix and discards the rest. The German 1.234,56 becomes 1.234: a plausible number, three orders of magnitude off, and no error anywhere to catch.
The committee left parsing out for a reason: formatting is a function, and parsing is a guess. 1.234 means one thousand two hundred thirty-four to a German reader and a little over one to an American one. No API can decide between them unless the caller names the locale, so the committee shipped the half that has a right answer. The other half falls to application code, and Intl still does the hard part. formatToParts on a known number labels every character of the formatted output. That is how a parser derives a locale's group and decimal separators at runtime instead of hardcoding a table that CLDR eventually moves out from under it.
const parts = new Intl.NumberFormat("de-DE").formatToParts(1234567.89);
// [{type:"integer",value:"1"},{type:"group",value:"."}, …]
const group = parts.find(p => p.type === "group")?.value; // "."
const decimal = parts.find(p => p.type === "decimal")?.value; // ","Display format vs wire format
Keep two representations and never confuse them. The wire format is canonical and locale-free: ISO 8601 for dates (2003-02-01), plain JSON numbers or integer minor units for money. The display format is whatever the user's locale renders. It exists only at the edges: a formatter produces it on the way out, and the parser removes it immediately on the way in.
Dates show why the wire format matters most. 01/02/03 is January 2, 2003 to a month-first reader, February 1, 2003 to a day-first reader, and February 3, 2001 under year-first conventions: every field is a valid day, month, and year, so nothing fails. A number parsed with the wrong locale tends to produce an implausible value that stands out. A date parsed with the wrong locale produces a plausible wrong day.
The digits themselves are locale data too, and CLDR calls them numbering systems. ar-EG resolves to arab (٠١٢٣٤٥٦٧٨٩) and fa-IR to arabext, while a Japanese IME produces full-width forms even though ja-JP defaults to latn. Folding them is arithmetic, not a lookup table. Every decimal digit block stores its ten digits in order starting at zero, so subtracting the block's first code point gives the digit value. Subtract U+0660, U+06F0, or U+FF10 from any digit in those blocks and the result is ASCII.
The demo
Try these, in order
Each step reproduces one specific failure in the demo below.
- 1Leave "Paste as a user from" on "de-DE · Germany" and read the results table. Number(input) is NaN, but parseFloat(input) returns 1.234, flagged "kept only the leading digits": the amount shrank by three orders of magnitude with no error. The parser that fails loudly is the safe one here.
- 2Switch to "fr-FR · France" and check the Group separator card under "A parser that asks the locale first". The separator is U+202F, a narrow no-break space: it looks like a space but matches neither a typed space nor a hardcoded ",". The demo derived it from formatToParts at runtime instead of a lookup table.
- 3Switch to "ar-EG · Egypt". All three standard parsers fail: Arabic-Indic digits (U+0660–U+0669) are outside JavaScript's number grammar. In "The parse, step by step", step 1 folds each digit by subtracting the block start from its code point, and step 4 reads 1234.56.
- 4Switch to "ja-JP · full-width digits". Even the digits an IME produces (U+FF10–U+FF19) defeat parseFloat. The Numbering system card reads fullwide because the demo's tag carries u-nu-fullwide. Plain ja-JP defaults to Latin digits: the full-width forms come from the keyboard rather than the locale.
- 5Back on "de-DE · Germany", clear "Amount, as typed" and type 1,2,3. Press "Refill the example" when done. Step 3 turns both commas into decimal points, so step 4 returns NaN: the hand-built parser rejects malformed input instead of guessing. Deriving separators does not make bad input good.
- 6Scroll to "The same string, three dates" and compare the "Reads 01/02/03 as" column. January 2, 2003. February 1, 2003. February 3, 2001. Three readers, three days, and no error anywhere, because every field is a valid day, month, and year. This is why dates cross the wire as ISO 8601.
What the standard parsers do with it
Pick a market and the field fills with 1234.56 (123456.78 for India) as a user there would type it.
Dot groups thousands, comma marks the decimal: the exact mirror of en-US.
| Parser | Result | What happened |
|---|---|---|
| Number(input) | NaN | Accepts the whole string or nothing. Only U+002E is a decimal point. Grouping characters are not in the grammar. |
| parseFloat(input) | 1.234 ← kept only the leading digits | Reads until the first character it cannot use, then stops: no error, no NaN. |
| <input type="number"> | — | What this browser keeps when code assigns the string to a number input. Engines differ. |
A parser that asks the locale first
Intl will not parse, but it will tell the caller the rules. Format a known number with formatToParts() and read the separators out of the parts.
Deriving separators…
The same string, three dates
Numbers at least fail loudly. Short dates fail silently: every field of 01/02/03 is a valid day, month, and year. Each row reads its field order out of formatToParts() on the known date 2003-02-01.
01/02/03
| Locale | Field order | Formats 2003-02-01 as | Reads 01/02/03 as |
|---|---|---|---|
| Deriving field orders… | |||
Never parse a user-typed short date without knowing the locale. In the UI, a date picker or three labeled fields removes the ambiguity. Between systems, send ISO 8601 (2003-02-01), which reads the same everywhere.
The short version
Intl formats and never parses; the inverse falls to application code. Derive each locale's separators from formatToParts, fold digit blocks by code-point arithmetic, and keep the wire format (ISO 8601, plain numbers) locale-free.
What to do about it
- Derive, don't hardcode.
Intl.NumberFormat(locale).formatToParts()on a known number reveals the locale's group and decimal separators. Parse with those, never withreplace(",", ""). - Fold non-ASCII digits before validating: Arabic-Indic (U+0660–U+0669), Extended Arabic-Indic (U+06F0–U+06F9), and full-width (U+FF10–U+FF19) digits all map to ASCII by subtracting the block start from the code point.
- Keep locale strings at the edge. On the wire, send ISO 8601 dates and plain JSON numbers; render locale formats only at display time, parse them only at input time.
- Never parse a user-typed short date.
01/02/03is ambiguous by construction: use a date picker, separate labeled fields, or require the ISO order. - For identifiers (card numbers, postcodes), prefer
<input type="text" inputmode="numeric">overtype="number"; number inputs discard or round what users type without saying so.
Where this comes up
Who it concerns
Moments
- ·Reviewing any free-text amount, quantity, or date field
- ·Designing form validation for a multi-market launch
- ·API contract review: deciding what format crosses the wire
Use the tool: CLDR explorer · what users will type back →
Field note
In February 2020 the GOV.UK Design System dropped <input type="number">. Research showed browsers discarding letters users typed, rounding numbers of 16+ digits, and converting large values to exponential notation. NVDA also announced the field as an unlabeled spin button. Their replacement for numeric identifiers is <input type="text" inputmode="numeric">: the mobile number keypad without the destructive reinterpretation. Any control that reinterprets what the user typed is a parser, and it needs the same scrutiny as a hand-written one.