Skip to content
Locale Lab

Chapter 16·Strings, names, and input·~7 min

Parsing what users type: numbers, dates & digits

Intl.NumberFormat turns 1234.56 into 1.234,56 for a German user, and nothing in ECMA-402 turns it back. '01/02/03' names three different dates in three markets.

formatToPartsnumbering systemsparsing

The problem

A checkout has a free-text amount field. A customer in Munich types 1.234,56; the backend runs parseFloat and books 1.234: three orders of magnitude short, with no error raised anywhere. A customer in Cairo types ١٬٢٣٤٫٥٦ and gets NaN. Formatting bugs at least show up on screen; parsing bugs corrupt stored data.

How it works

Formatting has an inverse, and nobody ships it

Every Intl API formats. Intl.NumberFormat has format, formatToParts, formatRange, and nothing that reads a string back. ECMA-402 ships no number-parsing API at all, so the moment a user types a localized value into a text field, the code is past the edge of the standard library.

JavaScript's own readers accept one grammar. Number() accepts the spec's numeric grammar (ASCII digits, one . as the decimal point, no grouping characters) and it is all-or-nothing: Number('1.234,56') is NaN. parseFloat() is the dangerous one. It reads the longest valid prefix and discards the rest. The German 1.234,56 becomes 1.234: a plausible number, three orders of magnitude off, and no error anywhere to catch.

The committee left parsing out for a reason: formatting is a function, and parsing is a guess. 1.234 means one thousand two hundred thirty-four to a German reader and a little over one to an American one. No API can decide between them unless the caller names the locale, so the committee shipped the half that has a right answer. The other half falls to application code, and Intl still does the hard part. formatToParts on a known number labels every character of the formatted output. That is how a parser derives a locale's group and decimal separators at runtime instead of hardcoding a table that CLDR eventually moves out from under it.

const parts = new Intl.NumberFormat("de-DE").formatToParts(1234567.89);
// [{type:"integer",value:"1"},{type:"group",value:"."}, …]
const group   = parts.find(p => p.type === "group")?.value;   // "."
const decimal = parts.find(p => p.type === "decimal")?.value; // ","
The derivation the demo runs live: format a known number, read the separators out of the labeled parts.

Display format vs wire format

Keep two representations and never confuse them. The wire format is canonical and locale-free: ISO 8601 for dates (2003-02-01), plain JSON numbers or integer minor units for money. The display format is whatever the user's locale renders. It exists only at the edges: a formatter produces it on the way out, and the parser removes it immediately on the way in.

Dates show why the wire format matters most. 01/02/03 is January 2, 2003 to a month-first reader, February 1, 2003 to a day-first reader, and February 3, 2001 under year-first conventions: every field is a valid day, month, and year, so nothing fails. A number parsed with the wrong locale tends to produce an implausible value that stands out. A date parsed with the wrong locale produces a plausible wrong day.

The digits themselves are locale data too, and CLDR calls them numbering systems. ar-EG resolves to arab (٠١٢٣٤٥٦٧٨٩) and fa-IR to arabext, while a Japanese IME produces full-width forms even though ja-JP defaults to latn. Folding them is arithmetic, not a lookup table. Every decimal digit block stores its ten digits in order starting at zero, so subtracting the block's first code point gives the digit value. Subtract U+0660, U+06F0, or U+FF10 from any digit in those blocks and the result is ASCII.

The demo

Try these, in order

Each step reproduces one specific failure in the demo below.

  1. 1
    Leave "Paste as a user from" on "de-DE · Germany" and read the results table. Number(input) is NaN, but parseFloat(input) returns 1.234, flagged "kept only the leading digits": the amount shrank by three orders of magnitude with no error. The parser that fails loudly is the safe one here.
  2. 2
    Switch to "fr-FR · France" and check the Group separator card under "A parser that asks the locale first". The separator is U+202F, a narrow no-break space: it looks like a space but matches neither a typed space nor a hardcoded ",". The demo derived it from formatToParts at runtime instead of a lookup table.
  3. 3
    Switch to "ar-EG · Egypt". All three standard parsers fail: Arabic-Indic digits (U+0660–U+0669) are outside JavaScript's number grammar. In "The parse, step by step", step 1 folds each digit by subtracting the block start from its code point, and step 4 reads 1234.56.
  4. 4
    Switch to "ja-JP · full-width digits". Even the digits an IME produces (U+FF10–U+FF19) defeat parseFloat. The Numbering system card reads fullwide because the demo's tag carries u-nu-fullwide. Plain ja-JP defaults to Latin digits: the full-width forms come from the keyboard rather than the locale.
  5. 5
    Back on "de-DE · Germany", clear "Amount, as typed" and type 1,2,3. Press "Refill the example" when done. Step 3 turns both commas into decimal points, so step 4 returns NaN: the hand-built parser rejects malformed input instead of guessing. Deriving separators does not make bad input good.
  6. 6
    Scroll to "The same string, three dates" and compare the "Reads 01/02/03 as" column. January 2, 2003. February 1, 2003. February 3, 2001. Three readers, three days, and no error anywhere, because every field is a valid day, month, and year. This is why dates cross the wire as ISO 8601.

What the standard parsers do with it

Pick a market and the field fills with 1234.56 (123456.78 for India) as a user there would type it.

Dot groups thousands, comma marks the decimal: the exact mirror of en-US.

ParserResultWhat happened
Number(input)NaNAccepts the whole string or nothing. Only U+002E is a decimal point. Grouping characters are not in the grammar.
parseFloat(input)1.234 ← kept only the leading digitsReads until the first character it cannot use, then stops: no error, no NaN.
<input type="number">—What this browser keeps when code assigns the string to a number input. Engines differ.

A parser that asks the locale first

Intl will not parse, but it will tell the caller the rules. Format a known number with formatToParts() and read the separators out of the parts.

Deriving separators…

The same string, three dates

Numbers at least fail loudly. Short dates fail silently: every field of 01/02/03 is a valid day, month, and year. Each row reads its field order out of formatToParts() on the known date 2003-02-01.

01/02/03

LocaleField orderFormats 2003-02-01 asReads 01/02/03 as
Deriving field orders…

Never parse a user-typed short date without knowing the locale. In the UI, a date picker or three labeled fields removes the ambiguity. Between systems, send ISO 8601 (2003-02-01), which reads the same everywhere.

The short version

Intl formats and never parses; the inverse falls to application code. Derive each locale's separators from formatToParts, fold digit blocks by code-point arithmetic, and keep the wire format (ISO 8601, plain numbers) locale-free.

What to do about it

  • Derive, don't hardcode. Intl.NumberFormat(locale).formatToParts() on a known number reveals the locale's group and decimal separators. Parse with those, never with replace(",", "").
  • Fold non-ASCII digits before validating: Arabic-Indic (U+0660–U+0669), Extended Arabic-Indic (U+06F0–U+06F9), and full-width (U+FF10–U+FF19) digits all map to ASCII by subtracting the block start from the code point.
  • Keep locale strings at the edge. On the wire, send ISO 8601 dates and plain JSON numbers; render locale formats only at display time, parse them only at input time.
  • Never parse a user-typed short date. 01/02/03 is ambiguous by construction: use a date picker, separate labeled fields, or require the ISO order.
  • For identifiers (card numbers, postcodes), prefer <input type="text" inputmode="numeric"> over type="number"; number inputs discard or round what users type without saying so.

Where this comes up

Who it concerns

EngineeringQAProduct

Moments

  • ·Reviewing any free-text amount, quantity, or date field
  • ·Designing form validation for a multi-market launch
  • ·API contract review: deciding what format crosses the wire

Field note

In February 2020 the GOV.UK Design System dropped <input type="number">. Research showed browsers discarding letters users typed, rounding numbers of 16+ digits, and converting large values to exponential notation. NVDA also announced the field as an unlabeled spin button. Their replacement for numeric identifiers is <input type="text" inputmode="numeric">: the mobile number keypad without the destructive reinterpretation. Any control that reinterprets what the user typed is a parser, and it needs the same scrutiny as a hand-written one.

GOV.UK: why we changed the input type for numbers ↗

Terms in this chapter

Where to read more

Related chapters