Skip to content

Where every Atlas number comes from

Every dataset with its edition and refresh cadence, then the model constants and what the numbers leave out.

Sources

The Atlas fetches two datasets live and caches them on disk; the rest are curated tables in the repo. When a fetch fails, the Atlas serves the last good cache.

DatasetVintageRefreshWhat it feeds
Unicode Consortium
cldr-json main branch (latest published CLDR)Fetched, 7-day cachePopulation, literacy, long-run GDP, and per-territory L1+L2 language shares: the base layer of every Atlas view. src/lib/cldr.ts loads it from the cldr-json main branch. The disk cache expires after 7 days, with stale-cache fallback on network error.
World Bank Open Data
Most recent reported year per country (varies by country, year stored per record)Fetched, 30-day cacheIndicator IT.NET.USER.ZS. Converts population into online population: the addressable-market number the reach calculator and planner quote. Both the share and the derived online population render in the country page's market summary. src/lib/worldbank.ts loads it with mrnev=1 (most recent non-empty value). 30-day disk cache.
World Bank Open Data
Most recent reported year per country (varies by country)Fetched, 30-day cacheIndicator NY.GDP.PCAP.CD. Renders in the country page's market summary beside CLDR's long-run GDP estimate (different vintages, never blended), in the compare table, and as the GDP-per-capita metric in the explorer, coverage map, and opportunity matrix. A CLDR-derived figure fills in where the World Bank has none. src/lib/worldbank.ts loads it. 30-day disk cache.
World Bank Open Data
Most recent reported year per country (varies by country)Fetched, 30-day cacheIndicator IT.CEL.SETS.P2. Renders in the country page's market summary as subscriptions per 100 people. Values above 100 mean multi-SIM, not full coverage. src/lib/worldbank.ts loads it. 30-day disk cache.
World Bank Global Findex
Most recent survey wave per country (the Findex runs every ~3 years, wave year stored per record)Fetched, 30-day cacheIndicator FX.OWN.TOTL.ZS: adults with an account at a financial institution or a mobile-money provider. A payment-readiness signal: it bounds the share of a market able to complete a card or mobile-money checkout. It renders in the country page's market summary and becomes a briefing risk when it is low. The API rejects mrnev for this sparse indicator, so src/lib/worldbank.ts pulls the wave range and keeps the newest non-empty value per country. 30-day disk cache. Never enters planner scoring.
Ethnologue (SIL International)
2023 edition, rounded published estimatesStatic, updated by handDe-duplicated global speaker totals for the major languages: the cross-border 'how many people speak this' figure that CLDR's per-territory shares cannot give. Curated in src/lib/language-reach-data.ts, which is where the current list lives. Good for ordering, not citation.
W3Techs
Content-language survey, 2024Static, updated by handShare of surveyed websites per content language: the supply side paired against Ethnologue's audience side. Curated in src/lib/language-reach-data.ts.
EF Education First (Signum International AG)
2025 edition (released November 2025, EF SET test data from 2024). 116 of 123 published countries, cross-checked against two sourcesStatic, updated by handEnglish-proficiency score and band per country, curated in src/lib/english-proficiency.ts. Display-only context that never enters planner scoring. The curation drops countries whose score differs between the official PDF and the Wikipedia table of the same edition.
CSA Research
2020 survey of 8,709 consumers across 29 countriesStatic, updated by handThe why-localize citation: 76% of respondents preferred products with information in their own language, and 40% said they will not buy in other languages. Background reading: no Atlas number derives from it.
Curated (industry-typical defaults)
2025 defaults, every figure editable in the UIStatic, updated by handThe five-tier rate table in src/lib/budget-estimator.ts: raw MT through premium transcreation. A starting point for a budget conversation, not a vendor quote. The launch-cost model in src/lib/launch-costs.ts prices from the same table.
Natural Earth
1:110m admin-0 countries (177 features), bundled as a static fileStatic, updated by handpublic/ne_110m_admin_0_countries.geojson: the polygons behind both world maps. ISO codes resolve via resolveIsoA2() in src/lib/geo.ts because Natural Earth marks France, Norway, and Kosovo as "-99" and Taiwan as "CN-TW". Geometry only: the Atlas reads no statistics from it.

The EF EPI table ships 116 of the 123 countries in the 2025 edition. It includes a country only when the official report PDF and an independent transcription of the same edition agree on its score, and omits the seven where they disagreed.

The launch score

The launch planner scores each candidate market as a weighted blend of three components, each normalized to 0–100: incremental reach, GDP, and engineering ease. The weights come from the selected product profile:

ProfileReachGDPEaseRationale
B2C55%15%30%Weights people reached highest: consumer products monetize breadth.
B2B20%55%25%Weights GDP highest: deal size tracks purchasing power more than headcount.
SaaS35%40%25%Splits revenue and reach: self-serve products need both an audience and buyers in it.
Balanced34%33%33%A flat prior for when the product profile is undecided.

Reach (0–100)

Net-new speakers gained worldwide when the market's primary language is added, computed under the overlap model in the next section. A market already half-covered earns only half credit. The raw speaker count maps to the score linearly: 200,000,000 net-new speakers scores 100, and anything above saturates. The 200M anchor is a choice that makes the largest realistic language additions fill the scale; no market fact says 200M rather than 250M. The "% of world" figure beside it divides by a fixed 8.0 billion world-population constant.

GDP (0–100)

The country's CLDR long-run nominal GDP, normalized so the largest economy in the dataset scores 100. The result is a relative ordering; it does not forecast revenue. Scores are comparable within a run but not across datasets, since the normalizing maximum comes from the country list in play.

Engineering ease (0–100)

A checklist score for the candidate language: start at 70, then apply each adder that fires. The adder sizes are judgment, chosen to order markets by integration effort; no project data measured them. The model clamps the result to 0–100.

CheckAdderWhy it moves the score
Script already shipped+12Fonts, fallback stacks, and rendering QA for the script already exist.
Left-to-right layout+8No mirroring work. The counterpart below is the expensive case.
Right-to-left layout−10Mirrored layout, bidirectional text handling (UAX #9), and icon review.
No custom webfont needed+6OS-bundled fonts cover the glyphs. No font licensing or payload cost.
Custom webfont needed−4Font selection, subsetting, and load-performance work.
2 or fewer plural categories+4Every count-bearing string needs at most two message variants.
5 or more plural categories−6Arabic-class plural systems multiply message variants and test cases.

Plural-category counts are CLDR plural rules, checked against Intl.PluralRules. The model attributes a candidate language only to countries where it holds at least a 30% L1+L2 share, so Spanish surfaces against Mexico and Spain rather than the US.

The overlap model

CLDR reports each language's share of a territory as L1 plus L2 speakers, so a bilingual person counts once per language and the shares of one country routinely sum past 100%. Every Atlas aggregate handles this the same way: within each country, the model sums the selected-language shares and caps them at 100% of the population. Incremental reach applies the cap before and after, so adding a minority language to a saturated bilingual market contributes only the uncovered remainder.

The cap bounds the error without resolving it. Within the capped total, the model may still credit a trilingual speaker to whichever of their languages was selected first: the data gives share totals and never says which individuals overlap.

How speaker shares are printed

Every share of a territory's population is printed to one decimal place, rounded up, by a single formatter shared across the Atlas. Russian is spoken by 0.24% of the United States, 820,000 speakers, which whole-number rounding printed as 0%. Rounding up is safe here because CLDR's territory data is read at a 0.1% floor, so no row rounds up from zero, and the largest possible overstatement is 0.09 points on a figure that is already a survey estimate.

How Chinese is counted

Simplified and Traditional Chinese are separate languages in every reach and coverage figure. They are different written standards with different character sets and terminology, and CLDR's territory data already keeps their audiences apart: mainland China and Singapore carry a zh row, while Taiwan, Hong Kong, and Macau carry zh_Hant, along with Traditional-reading communities in a dozen more territories. Shipping one never earns credit for the other, so a product with only Simplified Chinese shows Taiwan as uncovered. When a source names a Chinese interface without saying which form (the App Store search API does this), the model counts it as Simplified, the majority case, and the page says so where it happens.

CLDR also records spoken Sinitic varieties as their own rows: Cantonese, Wu, Xiang, Hakka, Min Nan, Gan. These stay out of coverage maths entirely, because their speakers read the written standard of their territory, which the standard Chinese row already counts. The rows remain visible on country and language pages as demographic facts.

The cost model

Costs start from per-word rates times word counts, using this site's curated 2025 planning defaults. Three adjustments then move each estimate toward a real quote: a per-language-pair rate index, a translation-memory leverage discount, and a separate one-time engineering-setup line. One LQA round prices at 20% of the tier's per-word rate, charged on every word regardless of leverage.

Tier$/wordWords/dayAppropriate for
Raw machine translation$0.0001250,000Internal tooling, search indexes, low-stakes content where 'gist' is enough.
MT + light post-edit$0.045,000Support content, docs, FAQ: content read once and discarded.
MT + full post-edit$0.083,500Marketing pages, onboarding flows, anything the user reads carefully.
Standard human translation$0.152,500UI copy, legal text, brand-facing surfaces, contracts.
Premium / creative$0.251,500Hero copy, taglines, ad campaigns: anywhere the literal translation reads as flat.

Per-language-pair rate index

A word into Japanese, Icelandic, or a Nordic language costs more than the same word into Spanish. The cause is translator scarcity and source-market cost of living, not script difficulty. The estimator multiplies each locale's per-word rate by an index, with a mainstream pair at 1.0, the default for anything not listed. These are planning indices rather than sourced rates.

is1.8×fi1.55×sv1.5×da1.5×nb1.5×no1.5×ja1.4×ko1.3×he1.3×ar1.25×zh1.2×de1.2×nl1.2×el1.15×fr1.1×ru1.1×uk1.1×cs1.1×pl1.05×

Translation-memory leverage

Words already in the translation memory bill below the new-word rate: repetitions and exact matches pay 25% of it, fuzzy 75–99% matches 60%, and new words full price. A first launch has no memory to draw on (0% leverage); a mature product reuses a large share of prior wording. The estimate applies leverage to translation and to the schedule, never to LQA.

Engineering setup

The work a per-word rate never captures, priced once as its own line so an RTL or new-script launch is not quoted like a Latin one. Planning figures, derived from the locale set:

i18n framework$6,000One-time: externalize strings, wire ICU, pseudoloc CI gate, TMS hookup. Skipped if already internationalized.
RTL enablement$4,000One-time, charged once if any target locale is RTL: bidi isolation, logical-property CSS, icon mirroring.
Webfont (per script)$1,200Per distinct non-Latin script that needs a bundled font: subsetting and glyph QA.
Per-locale integration$500Per locale: file wiring and an in-context sign-off pass.

The launch planner prices three launch depths by assigning each content bucket (UI strings, docs, support content) one of the tiers above. Even the cheapest depth keeps a light human review pass on the UI strings.

DepthWhat shipsBuckets pricedLQA rounds
Machine-translation pilotShips UI strings only, machine-translated with a light human review pass. Docs and support stay in the source language.ui · MT + light post-edit0
UI at publication qualityShips UI strings edited to publication quality with one round of LQA (a human quality-review pass). Docs and support stay in the source language.ui · MT + full post-edit1
Full launchShips UI at human-translation quality, docs at full post-edit, and support content at light post-edit, with one round of LQA (a human quality-review pass).ui · Standard human translation; docs · MT + full post-edit; support · MT + light post-edit1

The figures above price a one-time launch; ongoing upkeep is estimated separately. The model assumes 40% of the source word corpus is retranslated each year (new features, edited strings, refreshed docs) and prices that through the same depth machinery: same tier rates, LQA rounds, and PM overhead. The churn rate is a planning assumption; replace it with a measured figure when one exists.

The briefing's derived numbers

The executive briefing combines the models above into a few numbers of its own, each with a fixed convention.

Budget ask

The decision block's budget ask auto-fills from the recommended market's launch cost (the first row of the cost appendix) plus that locale's one-time engineering setup, which the appendix itemizes. Setup credits what the portfolio already has: a second RTL language does not pay for RTL enablement again, an already-shipped script does not re-buy its webfont, and a portfolio with more than one language is assumed to have an i18n framework in place. The ask leaves out the all-markets total, because the decision on the table approves a single launch; the appendix prints that total on its own labeled row. A typed budget-ask override always replaces the auto-filled figure, and when the cost appendix is off or unpriced the drafted ask carries no dollar amount.

Proposed sequence

The ranking table scores every candidate against the same fixed portfolio, so its reach numbers must never be summed: two overlapping languages would double-count their shared markets. The briefing's Now / Next / Later tiles are the one surface that orders launches in time, and each tile's net-new count is recomputed as if the earlier tiles had already shipped. Those three numbers do sum cleanly.

Payback scenario

The payback card combines four terms of different provenance, and its footer labels each one ("headroom: estimate · ARPU: your data · capture: your assumption · upkeep: 40% annual string churn, an assumption"). Online headroom, the estimated term, is the anchor market's online population minus the reported MAU there. ARPU comes from the uploaded data, as that market's revenue divided by its MAU. Capture and churn are the assumptions. From there the arithmetic is short: users = headroom × capture, monthly revenue = users × ARPU, and months to recover = launch cost ÷ monthly revenue net of upkeep, rounded up. The card renders only when every term exists: a priced cost appendix plus an uploaded row for the anchor market carrying both revenue and an online-population estimate. When a term is missing, the briefing omits the card and never prints a zero. The card is a scenario, not a forecast.

Custom weights

The launch planner's Custom profile exposes three sliders. On every change, they renormalize to integers that sum to 100. The briefing accepts the same override as a URL parameter: weights=R-G-E, three non-negative integers, normalized by their sum. For example, weights=55-15-30 weighs reach 55%, GDP 15%, and ease 30%. The briefing ignores a malformed value and applies the selected profile preset. While the override is active, the memo header shows "custom weights" and the methodology footer prints the exact split and its source parameter. Picking any profile button clears the override.

Snapshot deltas

When the diff compares two uploads of the same product, per-market penetration deltas divide both snapshots' MAU by the same online-population estimate (the current one), so the delta isolates MAU movement rather than drift between two vintages of that estimate. Markets present in only one snapshot appear as added or removed, never as ±100% changes, and the totals cover only markets present in both. The diff can report what moved but not why.

What the numbers are not

  • Double-counting is capped but not resolved. L1+L2 shares still overlap inside the 100% cap, so the model can credit a multilingual person to the wrong selected language.
  • Two GDP figures coexist. CLDR's long-run GDP estimate drives the launch score and the nominal-GDP metric. World Bank GDP per capita (current US$, most recent year) appears on country pages and drives the per-capita lenses. CLDR long-run GDP ÷ population fills in where the World Bank has no figure. The two are different vintages and can disagree. Any single number comes from one source, never an average of the two.
  • No revenue estimates. The Atlas sizes audiences and costs, never revenue. The only usage numbers it shows are uploaded ones (Your data), and they stay display-only, never entering scoring.
  • The ease model has a coverage edge. Its metadata table covers 80 language codes. Outside it, the model defaults to the flattering shape (Latin script, LTR, 2 plural categories, no webfont). The planner therefore flags those rows as "No engineering metadata".
  • EF EPI is context, not an input. The index samples self-selected online test takers, skewed young and toward people already learning English. It never enters planner scoring.