Free resource

Word frequency lists

The 5,000 most common words in nine languages, ranked by how often they occur in real speech. Every word comes with its IPA transcription, part of speech, an English meaning and an example sentence: on the page for the top 1,000, and in CSV and Anki downloads for all {listSize}.

Choose a language

What every list contains

  • Rank and frequency. The word's position and its raw count in the source corpus. A verb is listed once, under its infinitive, with the counts of all its conjugated forms added up.
  • IPA. A phonemic or phonetic transcription, so you can pronounce a word you have only read.
  • Type. Noun, verb, adjective, particle and so on, with a filter on every list to show only the types you want. For an inflected form other than a verb, the dictionary form it belongs to.
  • Meaning. A short English gloss of the word's primary sense.
  • Example sentence. A real sentence using the word, with an English translation.
  • Readings. Pinyin for Chinese, kana and rōmaji for Japanese, a romanization for Russian.

How the lists are built

The ranking comes from FrequencyWords (OpenSubtitles 2018), a word-count of film and television subtitles, the closest large corpus there is to everyday spoken language. Japanese uses the Leipzig Corpora Collection instead, because subtitle tokenizers split Japanese into fragments rather than words.

Each candidate word is then checked against Wiktionary: tokens that are not words of the language (names, typos, foreign words, tokenizer fragments) are dropped, and the survivors take their pronunciation, part of speech, meaning and example from the dictionary entry. Where Wiktionary has no usage example, one is drawn from Tatoeba, the open collection of translated sentences.

Because the lists are built from spoken-language corpora, inflected forms ("est", "está", "was") appear as their own entries, exactly as a learner meets them, each one tagged with the dictionary form it belongs to.

Using the lists in your classroom or blog

Every list is free to copy, print, embed and adapt. The frequency data and dictionary content are open-licensed (the exact licence is stated on each list's page); the lists themselves are released under the same terms. If you reuse one, link back to its page so your readers can get the latest version and the downloads.