Newly Minted data and methodology
Where Newly Minted’s words come from: Google Books Ngram data since 1800, and how each word’s take-off year is measured.
| What | How |
|---|---|
| Source | Google Books Ngram Viewer, English corpus (2022), years 1800 to 2022. |
| License | CC BY 3.0, credited to the Google Books Ngram Viewer. |
| What is measured | How common a word is in English-language books each year, as a share of all the words or phrases of the same length. All capitalizations are combined and spelling variants are summed. |
| Smoothing | A 5-year centered moving average. |
| Peak year | The year of a word’s highest smoothed use; the earliest year wins a tie. |
| Take-off year | The last year before the peak in which the smoothed curve rises up through a tenth of the peak. |
| Pool rules | Take-off from 1820 to 2019, little use in the decades before it, a peak of at least 20 uses per billion words, and 14 characters or fewer. |
| Pool size | 1,105 words in eight categories, with between 123 and 153 words in each. |
| Daily sets | Five words: exactly one from 1990 or later, at least one from 1929 or earlier, at least three decades, at most two per category, and no word repeated within 150 days. |
| Hint window | A 20-year span that always contains the take-off year, never centered on it. |
Where the words come from
Every year in Newly Minted is measured from the Google Books Ngram Viewer, a free tool that counts how often words and phrases appear in a large collection of digitized printed books. Newly Minted uses its English corpus (2022), which gives a figure for every year from 1800 through 2022: how common the term was that year, measured against all the words or phrases of the same length printed in it, with every capitalization combined, so “Sputnik” and “sputnik” count as one. The figures are gathered once, ahead of time, and stored with the game.
Credit and license
The Ngram data is published under the Creative Commons Attribution 3.0 license (CC BY 3.0) and credited here to the Google Books Ngram Viewer, at books.google.com/ngrams. Changes have been made: the yearly figures are summed across spelling variants, smoothed, and reduced to a take-off year, a peak year and a small curve for each word. Newly Minted is not affiliated with Google.
Smoothing
A word’s share in any single year can jump about for reasons that have nothing to do with language, such as a single widely reprinted book or a batch of poor scans. So the game smooths each word’s series with a five-year centered average: each year’s value becomes the mean of that year, the two years before it and the two years after it. At either end of the data, where fewer years exist, only the years that do exist are averaged. The peak year is the year of the highest smoothed value, and the earliest year wins a tie.
The take-off rule
The take-off year is the last year before the peak in which a word’s smoothed use rises up through a tenth of its peak. The game draws a line at a tenth of the peak’s height, walks forward through the years to the peak, and notes each year the curve climbs from below the line to at or above it. The last of those years is the answer, so an early blip that fell back again does not count: a word takes off when it rises and stays risen. A tenth is low enough to catch the moment a word starts to be noticed and high enough to ignore the faint background of early scans.
Two consequences follow, and the strategy guide shows how to use both. A word that arrives suddenly dates about two years early, because the five-year average looks two years ahead: “sputnik” took off in 1955, for a satellite launched in 1957. And a word that kept growing for a century dates late, because a tenth of a very high, very late peak is reached late: “electron” took off in 1936 and “protein” in 1907, long after chemists and physicists began writing about them.
Which words make the pool
A word joins the game only if it passes every rule. Its take-off must fall between 1820 and 2019: scans of earlier books are noisy, and a word that took off after 2019 may still be climbing when the data ends in 2022, so its real peak cannot be seen yet. Its use in the decades before take-off must be close to nothing: the average smoothed use from 40 years to 10 years before the take-off must be at most 3 percent of the peak, with at least 10 of those years inside the data, which filters out old words with new meanings. Its peak must reach at least 20 uses per billion words, because below that a curve is mostly noise, and its display form must fit in 14 characters.
A hand review comes last. Words whose take-off the rule gets wrong anyway are dropped, for example where an older meaning sets the year: “startup” and “broadband” are left out for that reason. The pool that remains has 1,105 words, of which 845 took off before 1990 (318 of them in 1929 or earlier) and 260 in 1990 or later, spread across tech, science, slang, culture, politics, business, food and health.
Spelling variants and hyphens
Some words are measured under more than one spelling, and the yearly figures for those spellings are added together, as with “email” and “e-mail”. There is a catch. The Ngram Viewer counts a hyphenated spelling as three tokens, the letter “e”, a hyphen and the word “mail”, so a hyphenated form is queried as written and measured as a three-token phrase. Its share is worked out against a slightly different total than a one-word form’s, so adding the two mixes denominators and the sum is approximate.
Why the year is not a dictionary’s first known use
A dictionary’s first known use is the earliest written record of a word that its editors have found, which can be a single sentence in a letter or a newspaper years before anyone else used the word. Newly Minted dates something different: the year a word’s use in printed books began a lasting climb. The two can be far apart. Books trail speech and, more recently, the web, so a word is often talked about for years before it is common in print, and the game’s year usually lands later than a dictionary’s. The smoothing can work the other way for a word that arrives in a single burst, dating it up to about two years early.
Limits of the data
The Ngram Viewer counts printed books that have been scanned, which is not everything ever printed and not everything ever said. Early scans contain more errors, the collection reflects what libraries and publishers supplied, and a word used mostly in speech, songs or the web can be under-counted for years. Newly Minted is a game about the trail words leave in print, not a scholarly dictionary, and its years should be read that way.
Updates and fairness
The five words chosen for a published puzzle never change once that puzzle is live. The word list and each word’s years are built once from the Viewer’s data and stored with the game, so the data stops at 2022. To report a data problem, email hello@guessday.com with the word and the puzzle number.
Attribution and independence
Newly Minted is not affiliated with Google. Google Books and the Ngram Viewer are Google’s, and this site uses no Google logos. The data describes printed books, not any individual person, and nothing on this site is linguistic, historical or scholarly guidance.