Chinese language statistics, measured
Almost every number written about Chinese online is a round figure repeated from another blog. These are not. Each one is counted, here, from the 3,329 Chinese sentences this site publishes and its dictionary index of 4,995 words — with the method written out below, and every row linked to the page that teaches what it counts.
of the words in these sentences are covered by the whole of HSK 1–6
characters read 80% of the running text; 66 read half of it
of Chinese words are written with exactly two characters, not one
syllables in the whole of Mandarin, ignoring tone — see the pinyin chart
What does each HSK level actually buy you?
The question behind "is HSK 3 enough?" is really what share of the words in front of me will I know. Here it is counted: every word in 3,329 sentences was found by this site's own segmenter and graded against the official HSK lists. HSK 1–3 reaches 68.3%, HSK 4 reaches 77.7%, and the last two levels together add barely 1.9% — each of their words is one you meet rarely and need exactly.
| Level | Words added | Words known | This level | Covered so far | What it buys you |
|---|---|---|---|---|---|
| HSK 1 | 150 | 150 | 44.2% |
44.2%
|
Greetings, numbers, family and the grammar words that hold a sentence together. Not yet a text you can read, but the frame every later sentence hangs on. |
| HSK 2 | 151 | 301 | 12.9% |
57.1%
|
Everyday errands — shopping, time, travel, feeling ill. Short written messages start to make sense with a dictionary open beside you. |
| HSK 3 | 297 | 598 | 11.2% |
68.3%
|
The level where reading begins. Graded readers, simple signage and children's material become mostly comprehensible, and looking a word up stops interrupting every line. |
| HSK 4 | 598 | 1,196 | 9.3% |
77.7%
|
Opinions, comparisons and reasons. Conversation stops being transactional, and news headlines are readable even when the article is not. |
| HSK 5 | 1,300 | 2,496 | 1.5% |
79.2%
|
Abstract and written registers: argument, culture, work. This is where the words start being ones you meet rarely but need exactly. |
| HSK 6 | 2,499 | 4,995 | 0.4% |
79.6%
|
The formal and literary tail — idioms, set phrases and the vocabulary of essays and editorials. Each word here appears seldom, which is why the curve flattens. |
| Beyond HSK | — | — | 20.4% | Real words the dictionary knows but the HSK lists do not grade — names, places, and everyday vocabulary the syllabus simply leaves out. They are counted in the denominator above, not discarded, because dropping them is how coverage percentages get inflated. | |
Read this as a floor, not a ceiling. The corpus is learner-facing Chinese — graded readers, example sentences, travel phrases, grammar examples — so it is a fair description of the material a learner meets and an understatement of how much vocabulary a newspaper demands. Test yourself against it with the vocabulary size test or the HSK text analyzer, which reads your own text with the same segmenter that produced this table.
How many characters do you actually need?
Far fewer than the number people quote, because character frequency is brutally top-heavy. All 3,329 sentences here — 27,922 characters of running text — use only 1,261 different characters between them, and the first handful do most of the work.
The jump from 482 to 688 characters buys five percentage points; the first 66 bought fifty. That shape is the argument for learning characters in frequency order and for not worrying about the long tail until much later.
The commonest characters in this corpus
| # | Character | Pinyin | Meaning | Times | Running total |
|---|---|---|---|---|---|
| 1 | 我 | wǒ | I (me - myself) | 1,353 | 4.8% |
| 2 | 的 | de | Of (aim - the) | 822 | 7.8% |
| 3 | 了 | le | (past tense indicator); finish (end) | 736 | 10.4% |
| 4 | 一 | yī | One (single) | 576 | 12.5% |
| 5 | 这 | zhè | This (these); This (these) | 563 | 14.5% |
| 6 | 很 | hěn | Very | 531 | 16.4% |
| 7 | 他 | tā | He (him - another - other) | 493 | 18.2% |
| 8 | 个 | gè | ([measure word]) | 466 | 19.8% |
| 9 | 是 | shì | To be (is - are - am - yes) | 377 | 21.2% |
| 10 | 有 | yǒu | To have (has - there is - there are) | 350 | 22.4% |
| 11 | 天 | tiān | Day (heaven - sky) | 344 | 23.7% |
| 12 | 在 | zài | Located (in - at - exist) | 319 | 24.8% |
| 13 | 不 | bù | No (not) | 317 | 26.0% |
| 14 | 们 | men | ([plural indicator]) | 272 | 26.9% |
| 15 | 上 | shàng | Above (on - go up - upper - first) | 249 | 27.8% |
| 16 | 你 | nǐ | You | 234 | 28.7% |
| 17 | 要 | yāo | Want (need - request - must - demand) | 203 | 29.4% |
| 18 | 去 | qù | Go (leave - remove - send - past) | 201 | 30.1% |
| 19 | 请 | qǐng | Please (ask - invite - request) | 197 | 30.8% |
| 20 | 好 | hǎo | Good (well); Like | 187 | 31.5% |
| 21 | 子 | zǐ | Son (child - person - midnight) | 177 | 32.1% |
| 22 | 点 | diǎn | O'clock (dot - point - mw for suggestions) | 177 | 32.7% |
| 23 | 来 | lái | Come (arrive) | 175 | 33.4% |
| 24 | 下 | xià | Below (under - down - fall - next - lower - finish) | 169 | 34.0% |
| 25 | 家 | jiā | Family (home - domestic) | 159 | 34.5% |
| 26 | 里 | lǐ | In (inside - inner - neighbourhood - Chinese mile [half a km]) | 154 | 35.1% |
| 27 | 多 | duō | Many (more - how) | 153 | 35.6% |
| 28 | 学 | xué | Learn (study - imitate - science - ology) | 153 | 36.2% |
| 29 | 说 | shuō | Said (speak - say - talk - explain) | 151 | 36.7% |
| 30 | 时 | shí | When (time - hour - season - o'clock) | 148 | 37.3% |
| 31 | 看 | kàn | Look at (see - read - visit); Look after | 139 | 37.8% |
| 32 | 吗 | mǎ | ([question indicator]) | 128 | 38.2% |
| 33 | 到 | dào | To (arrive - go to) | 126 | 38.7% |
| 34 | 大 | dà | Big (oldest) | 126 | 39.1% |
| 35 | 儿 | ér | Son | 123 | 39.6% |
| 36 | 今 | jīn | Present (modern - today) | 116 | 40.0% |
| 37 | 以 | yǐ | Use (by - for - within - because - in order to) | 113 | 40.4% |
| 38 | 她 | tā | She | 113 | 40.8% |
| 39 | 人 | rén | Person (people) | 112 | 41.2% |
| 40 | 吃 | chī | Eat | 111 | 41.6% |
A frequency list from a corpus of learner Chinese is not the same as one from newspapers, and the difference is visible in the top rows: first- and second-person pronouns rank far higher here than they would in reported speech. That is a property of the material, stated rather than hidden. Practise writing any of them on the stroke-order animator or a 田字格 practice sheet.
Is a Chinese word one character?
No — and this is the misconception that costs beginners the most time. A character is closer to a syllable or a word root than to a word. Counted across the 4,995 headwords in this site's dictionary index:
| Written with | Words | Share of the index |
|---|---|---|
| 1 character | 708 |
14.2%
|
| 2 characters | 4,034 |
80.8%
|
| 3 characters | 132 |
2.6%
|
| 4 characters | 121 |
2.4%
|
Two characters is the overwhelming default. This is why looking a sentence up character by character produces nonsense, and why the pinyin annotator and the translators find word boundaries first — the same step that produced the HSK table above.
Which tone will you say most?
Tone drills usually give the four tones equal time. They do not appear equally. Counted over the 27,922 character occurrences in this corpus whose reading the dictionary supplies:
| Tone | What the pitch does | Share | Commonest character with it |
|---|---|---|---|
| First tone | high and level |
23.3%
|
一 yī |
| Second tone | rises |
14.1%
|
来 lái |
| Third tone | dips, then rises |
24.9%
|
我 wǒ |
| Fourth tone | falls sharply |
30.7%
|
这 zhè |
| Neutral tone | short, light, unstressed |
7.0%
|
的 de |
The fourth tone is the one to get right first. The neutral tone looks negligible at 7.0%, but it lands on some of the most frequent characters in the language, so you will say it constantly — and it is the one most courses skip. Hear your own pitch against the target on the tone pitch mirror, drill recognition on the tone trainer, or see every syllable and tone in the pinyin chart.
Which radicals repay learning first?
The ones that turn up in the most characters. Counted across the 3,616 simplified and traditional characters this site indexes, using the same component data the character finder searches:
| # | Form | Pinyin | Meaning | Characters | Share |
|---|---|---|---|---|---|
| 1 | 口 | kǒu | Mouth; opening | 313 | 8.7% |
| 2 | 氵 | shuǐ | Water (three drops) | 209 | 5.8% |
| 3 | 人 亻 | rén | Person | 206 | 5.7% |
| 4 | 扌 | shǒu | Hand (side form) | 205 | 5.7% |
| 5 | 木 | mù | Tree; wood | 178 | 4.9% |
| 6 | 曰 | yuē | To say; speak | 107 | 3.0% |
| 7 | 日 | rì | Sun; day | 99 | 2.7% |
| 8 | 刀 刂 | dāo | Knife | 96 | 2.7% |
| 9 | 辶 | chuò | Walk; movement (走之) | 94 | 2.6% |
| 10 | 火 灬 | huǒ | Fire | 93 | 2.6% |
| 11 | 土 | tǔ | Earth; soil | 92 | 2.5% |
| 12 | 月 | yuè | Moon; also 肉 flesh | 88 | 2.4% |
| 13 | 言 | yán | Speech; words | 88 | 2.4% |
| 14 | 宀 | mián | Roof (of a house) | 87 | 2.4% |
| 15 | 亠 | tóu | Lid; top | 83 | 2.3% |
| 16 | 心 | xīn | Heart; mind | 80 | 2.2% |
| 17 | 糸 | sī | Silk; thread | 76 | 2.1% |
| 18 | 艹 | cǎo | Grass; plant | 75 | 2.1% |
| 19 | 忄 | xīn | Heart (side form) | 68 | 1.9% |
| 20 | 貝 | bèi | Shell; money; wealth | 68 | 1.9% |
| 21 | 攵 攴 | pū | Tap; rap; script (folded… | 67 | 1.9% |
| 22 | 阝 | fù | Mound; hill (left ear) | 64 | 1.8% |
| 23 | 女 | nǚ | Woman; female | 61 | 1.7% |
| 24 | 小 ⺌ | xiǎo | Small; little | 59 | 1.6% |
| 25 | 田 | tián | Field; farmland | 55 | 1.5% |
The honest limits of this one. Nothing in this repository states which character contains which part — it is derived from stroke data by splitting each character at its radical and identifying each half by shape, and a part that cannot be named confidently is dropped rather than guessed at. So forms that deform heavily inside a character are undercounted, and two forms that are written almost identically — 日 and 曰 are the clearest pair — are separated by shape alone and can take each other's characters. Counting is also per written form: 氵 and 水 are one radical with one page but two shapes, so the busier shape takes the row rather than the two being added together. Read the most common radicals for the hand-written version of this list, and the radicals hub for all 100 of them.
Why Chinese has so many homophones
Because the sound system is tiny. Mandarin has 410 distinct syllables ignoring tone — built from 24 initials and 34 finals — and 1,183 syllable-and-tone combinations actually in use, out of the 1,640 that four tones would allow. English has thousands of syllables. Every Chinese character is exactly one of those 410, which is the whole explanation.
In this corpus the effect is measurable: its 1,261 distinct characters share only 723 different sounds between them — an average of 1.7 characters per sound, and far worse than that at the crowded end:
| Sound | Characters | Some of them |
|---|---|---|
| shì | 13 | 是事市试室视示士 |
| jì | 11 | 记计绩寄际纪继技 |
| fù | 8 | 付复附富负父傅副 |
| jí | 8 | 急及圾极级即籍集 |
| yì | 8 | 意议易谊忆艺译义 |
| jìng | 7 | 静净境镜竟敬竞 |
| jù | 7 | 句拒据距具剧聚 |
| shí | 7 | 时十实识拾石食 |
| jī | 6 | 机鸡奇积激基 |
| jiàn | 6 | 件见健建键荐 |
| lì | 6 | 力历利例励厉 |
| lù | 6 | 路律绿录虑率 |
And that is only the characters this corpus happens to use — a full dictionary is worse. It is also the practical reason tones are not optional: dropping the tone multiplies the number of things a syllable could be by four or five. The interactive pinyin chart has all 410 syllables with audio, and Bopomofo writes the same inventory a different way.
How this was measured — and what it does not say
A statistic without its method is a rumour, so here is the whole of it.
The corpus
3,329 distinct Chinese sentences, all of them published on this site and all of them readable in full: 2,266 hand-written dictionary examples, 469 lines from the graded readers, 464 travel phrases and 130 grammar examples. The word-length and radical tables are measured against different sets, each named beside its table: the dictionary index of 4,995 headwords, and the 3,616 characters the site's component data covers.
🔒 What the corpus is not
It is not a balanced corpus of native writing. It is Chinese written by people, for learners, at a spread of levels. That makes these figures a fair description of the material a learner actually meets — which is what the questions above are really asking — and an understatement of the vocabulary a newspaper, a novel or a contract demands. Anyone quoting a coverage percentage without saying what it was measured over is quoting nothing in particular; that is the mistake this page is written to avoid, not to repeat.
What a "token" is
Chinese is written without spaces, so "what share of the words" has no answer until something decides where the words are. Every sentence was segmented by greedy longest-match against the same in-memory dictionary the translators use — the identical step the HSK text analyzer performs on text you paste, so you can run the method yourself. That produced 18,996 word tokens (2,285 distinct words), an average of 5.7 words per sentence. Punctuation, Latin letters and digits are not tokens and are not counted anywhere on this page.
What is in the denominator
Everything. Words graded beyond HSK 6 and words the dictionary cannot identify are counted in the total rather than dropped from it. Removing them would raise every coverage figure on this page by several points and would be the single easiest way to make it wrong.
Tones and readings
Counted per character, not per word: every Chinese character is exactly one syllable, so a character-by-character pass gives one tone per occurrence with nothing to guess at. Each character's reading comes from the same dictionary as everything else, and the tone is read off the diacritic. Two honest caveats: a character with more than one reading is counted at its first, most common one, and a syllable that goes neutral only in context (the 了 of a finished action, a repeated 看看) is counted at its citation tone. A character the dictionary cannot read is left out of the tone table rather than defaulted to neutral.
When
Measured against the corpus as it stood on . Nothing here is stored: every figure is recomputed from the corpus each time the site restarts, so the page tracks the material rather than freezing at whatever it read the day it was written. If you quote a number from here, please quote the date with it and link back — the figures move as the corpus grows.
Questions people ask
How much of Chinese does HSK 3 cover?
Measured over the 3,329 Chinese sentences published on this site, HSK levels 1 to 3 together account for 68.3% of the words used — about 598 words of vocabulary. HSK 4 takes that to 77.7%, and all six levels together reach 79.6%. Those figures describe learner-facing Chinese — graded readers, example sentences, travel phrases and grammar examples — not a newspaper, which sits well above this. The last two levels move the number very little because the words they add are ones you meet rarely and need exactly.
How many Chinese characters do I need to know?
To read this corpus: 66 distinct characters cover half of the running text, 277 cover 80%, 482 cover 90% and 688 cover 95%. The whole corpus of 3,329 sentences uses only 1,261 different characters in total. The curve is steep at the start and very flat afterwards, which is why the first few hundred characters are worth far more per character than the next few thousand.
Is a Chinese word one character?
No, and that is the single most common misconception about the language. Of the 4,995 words in this site's dictionary index, 80.8% are written with exactly two characters and only 14.2% with one. Characters are closer to syllables, or to word roots, than to words.
Which Chinese tone is the most common?
The fourth tone — it accounts for 30.7% of the character occurrences in this corpus, against 14.1% for the second tone. The neutral tone is the rarest at 7.0%, but it lands on some of the most frequent characters in the language, so you say it constantly.
Which Chinese radicals are worth learning first?
The ones that appear in the most characters. Counted across the 3,616 characters this site indexes, 口 (mouth; opening) leads, followed by 氵, 人, 扌, 木. Learning twenty or thirty of these is the change that makes characters stop looking random.
How many syllables does Mandarin have?
410 distinct syllables ignoring tone, built from 24 initials and 34 finals, and 1,183 syllable-and-tone combinations in use out of the 1,640 that four tones would allow. English has thousands of syllables. That is why homophones are everywhere: the 1,261 characters in this corpus share only 723 different sounds between them, and shì alone is written 13 different ways.
What is this measured against?
3,329 Chinese sentences published on this site — 2,266 hand-written dictionary examples, 469 lines from the graded readers, 464 travel phrases and 130 grammar examples — plus the official HSK word lists and this site's dictionary index of 4,995 words. It is learner-facing Chinese written by people, not a balanced corpus of native writing, so it is a fair description of the material a learner actually meets and an understatement of how hard unsimplified Chinese is.
Can I quote these numbers?
Yes, with a link back to this page and the date they were measured, which is shown at the top. Every figure here is recomputed from the corpus each time the site restarts, so it tracks the corpus rather than freezing at whatever it read the day it was written.
Act on these numbers
- HSK word lists — all six levels, free, with printable PDFs
- Chinese Vocabulary Size Test — find out where on the coverage curve you already are
- HSK Text Analyzer — run the segmentation above on your own text
- Chinese Dictionary — the entry behind every character in the frequency table
- Radicals Hub — one page per component, in the order the table ranks them
- Example Sentence Bank — the corpus itself, searchable
- Tone Pitch Mirror — practise the tone you will say most