Chinese language statistics, measured
Almost every number written about Chinese online is a round figure repeated from another blog. These are not. Each one is counted, here, from the 10,884 Chinese sentences this site publishes and its dictionary index of 4,995 words — with the method written out below, and every row linked to the page that teaches what it counts.
of the words in these sentences are covered by the whole of HSK 1–6
characters read 80% of the running text; 100 read half of it
of Chinese words are written with exactly two characters, not one
syllables in the whole of Mandarin, ignoring tone — see the pinyin chart
What does each HSK level actually buy you?
The question behind "is HSK 3 enough?" is really what share of the words in front of me will I know. Here it is counted: every word in 10,884 sentences was found by this site's own segmenter and graded against the official HSK lists. HSK 1–3 reaches 53.6%, HSK 4 reaches 60.9%, and the last two levels together add barely 17.2% — each of their words is one you meet rarely and need exactly.
| Level | Words added | Words known | This level | Covered so far | What it buys you |
|---|---|---|---|---|---|
| HSK 1 | 150 | 150 | 34.5% |
34.5%
|
Greetings, numbers, family and the grammar words that hold a sentence together. Not yet a text you can read, but the frame every later sentence hangs on. |
| HSK 2 | 151 | 301 | 9.9% |
44.4%
|
Everyday errands — shopping, time, travel, feeling ill. Short written messages start to make sense with a dictionary open beside you. |
| HSK 3 | 297 | 598 | 9.2% |
53.6%
|
The level where reading begins. Graded readers, simple signage and children's material become mostly comprehensible, and looking a word up stops interrupting every line. |
| HSK 4 | 598 | 1,196 | 7.3% |
60.9%
|
Opinions, comparisons and reasons. Conversation stops being transactional, and news headlines are readable even when the article is not. |
| HSK 5 | 1,300 | 2,496 | 8.4% |
69.2%
|
Abstract and written registers: argument, culture, work. This is where the words start being ones you meet rarely but need exactly. |
| HSK 6 | 2,499 | 4,995 | 8.8% |
78.1%
|
The formal and literary tail — idioms, set phrases and the vocabulary of essays and editorials. Each word here appears seldom, which is why the curve flattens. |
| Beyond HSK | — | — | 21.9% | Real words the dictionary knows but the HSK lists do not grade — names, places, and everyday vocabulary the syllabus simply leaves out. They are counted in the denominator above, not discarded, because dropping them is how coverage percentages get inflated. | |
| Unidentified | — | — | 0.0% | Chinese the dictionary has no entry for at all. | |
Read this as a floor, not a ceiling. The corpus is learner-facing Chinese — graded readers, example sentences, travel phrases, grammar examples — so it is a fair description of the material a learner meets and an understatement of how much vocabulary a newspaper demands. Test yourself against it with the vocabulary size test or the HSK text analyzer, which reads your own text with the same segmenter that produced this table.
How many characters do you actually need?
Far fewer than the number people quote, because character frequency is brutally top-heavy. All 10,884 sentences here — 101,567 characters of running text — use only 2,734 different characters between them, and the first handful do most of the work.
The jump from 818 to 1,234 characters buys five percentage points; the first 100 bought fifty. That shape is the argument for learning characters in frequency order and for not worrying about the long tail until much later.
The commonest characters in this corpus
| # | Character | Pinyin | Meaning | Times | Running total |
|---|---|---|---|---|---|
| 1 | 的 | de | Of (aim - the) | 3,483 | 3.4% |
| 2 | 了 | le | (past tense indicator); finish (end) | 2,835 | 6.2% |
| 3 | 他 | tā | He (him - another - other) | 2,835 | 9.0% |
| 4 | 这 | zhè | This (these); This (these) | 2,602 | 11.6% |
| 5 | 一 | yī | One (single) | 1,876 | 13.4% |
| 6 | 我 | wǒ | I (me - myself) | 1,869 | 15.3% |
| 7 | 很 | hěn | Very | 1,446 | 16.7% |
| 8 | 个 | gè | ([measure word]) | 1,358 | 18.0% |
| 9 | 是 | shì | To be (is - are - am - yes) | 1,160 | 19.2% |
| 10 | 不 | bù | No (not) | 1,126 | 20.3% |
| 11 | 有 | yǒu | To have (has - there is - there are) | 1,045 | 21.3% |
| 12 | 在 | zài | Located (in - at - exist) | 1,017 | 22.3% |
| 13 | 要 | yāo | Want (need - request - must - demand) | 754 | 23.0% |
| 14 | 人 | rén | Person (people) | 736 | 23.8% |
| 15 | 上 | shàng | Above (on - go up - upper - first) | 735 | 24.5% |
| 16 | 们 | men | ([plural indicator]) | 715 | 25.2% |
| 17 | 天 | tiān | Day (heaven - sky) | 679 | 25.9% |
| 18 | 家 | jiā | Family (home - domestic) | 577 | 26.4% |
| 19 | 子 | zǐ | Son (child - person - midnight) | 568 | 27.0% |
| 20 | 大 | dà | Big (oldest) | 534 | 27.5% |
| 21 | 里 | lǐ | In (inside - inner - neighbourhood - Chinese mile [half a km]) | 521 | 28.0% |
| 22 | 请 | qǐng | Please (ask - invite - request) | 483 | 28.5% |
| 23 | 来 | lái | Come (arrive) | 480 | 29.0% |
| 24 | 学 | xué | Learn (study - imitate - science - ology) | 477 | 29.4% |
| 25 | 好 | hǎo | Good (well); Like | 472 | 29.9% |
| 26 | 得 | dé | Get; Have to (need) | 460 | 30.4% |
| 27 | 时 | shí | When (time - hour - season - o'clock) | 450 | 30.8% |
| 28 | 下 | xià | Below (under - down - fall - next - lower - finish) | 438 | 31.2% |
| 29 | 事 | shì | Matter (thing - accident - trouble - responsibility - job) | 428 | 31.7% |
| 30 | 到 | dào | To (arrive - go to) | 420 | 32.1% |
| 31 | 年 | nián | Year (age - old) | 402 | 32.5% |
| 32 | 地 | dì | Earth (ground - location); ~ly | 383 | 32.8% |
| 33 | 生 | shēng | Life (give birth - be born - grow) | 380 | 33.2% |
| 34 | 多 | duō | Many (more - how) | 378 | 33.6% |
| 35 | 点 | diǎn | O'clock (dot - point - mw for suggestions) | 373 | 34.0% |
| 36 | 说 | shuō | Said (speak - say - talk - explain) | 357 | 34.3% |
| 37 | 着 | zhāo | Move (tactic); Touch (affected by - ignited) | 354 | 34.7% |
| 38 | 去 | qù | Go (leave - remove - send - past) | 352 | 35.0% |
| 39 | 会 | huì | Meeting (understand - able to - can - will - shall) | 350 | 35.4% |
| 40 | 你 | nǐ | You | 341 | 35.7% |
A frequency list from a corpus of learner Chinese is not the same as one from newspapers, and the difference is visible in the top rows: first- and second-person pronouns rank far higher here than they would in reported speech. That is a property of the material, stated rather than hidden. Practise writing any of them on the stroke-order animator or a 田字格 practice sheet.
Is a Chinese word one character?
No — and this is the misconception that costs beginners the most time. A character is closer to a syllable or a word root than to a word. Counted across the 4,995 headwords in this site's dictionary index:
| Written with | Words | Share of the index |
|---|---|---|
| 1 character | 708 |
14.2%
|
| 2 characters | 4,034 |
80.8%
|
| 3 characters | 132 |
2.6%
|
| 4 characters | 121 |
2.4%
|
Two characters is the overwhelming default. This is why looking a sentence up character by character produces nonsense, and why the pinyin annotator and the translators find word boundaries first — the same step that produced the HSK table above.
Which tone will you say most?
Tone drills usually give the four tones equal time. They do not appear equally. Counted over the 101,519 character occurrences in this corpus whose reading the dictionary supplies:
| Tone | What the pitch does | Share | Commonest character with it |
|---|---|---|---|
| First tone | high and level |
23.8%
|
他 tā |
| Second tone | rises |
15.9%
|
人 rén |
| Third tone | dips, then rises |
19.8%
|
我 wǒ |
| Fourth tone | falls sharply |
33.3%
|
这 zhè |
| Neutral tone | short, light, unstressed |
7.1%
|
的 de |
The fourth tone is the one to get right first. The neutral tone looks negligible at 7.1%, but it lands on some of the most frequent characters in the language, so you will say it constantly — and it is the one most courses skip. Hear your own pitch against the target on the tone pitch mirror, drill recognition on the tone trainer, or see every syllable and tone in the pinyin chart.
Which radicals repay learning first?
The ones that turn up in the most characters. Counted across the 3,684 simplified and traditional characters this site indexes, using the same component data the character finder searches:
| # | Form | Pinyin | Meaning | Characters | Share |
|---|---|---|---|---|---|
| 1 | 口 | kǒu | Mouth; opening | 324 | 8.8% |
| 2 | 人 亻 | rén | Person | 215 | 5.8% |
| 3 | 氵 | shuǐ | Water (three drops) | 211 | 5.7% |
| 4 | 扌 | shǒu | Hand (side form) | 208 | 5.6% |
| 5 | 木 | mù | Tree; wood | 181 | 4.9% |
| 6 | 曰 | yuē | To say; speak | 107 | 2.9% |
| 7 | 日 | rì | Sun; day | 99 | 2.7% |
| 8 | 刀 刂 | dāo | Knife | 98 | 2.7% |
| 9 | 辶 | chuò | Walk; movement (走之) | 96 | 2.6% |
| 10 | 火 灬 | huǒ | Fire | 94 | 2.6% |
| 11 | 土 | tǔ | Earth; soil | 93 | 2.5% |
| 12 | 月 | yuè | Moon; also 肉 flesh | 92 | 2.5% |
| 13 | 言 | yán | Speech; words | 92 | 2.5% |
| 14 | 宀 | mián | Roof (of a house) | 89 | 2.4% |
| 15 | 亠 | tóu | Lid; top | 86 | 2.3% |
| 16 | 心 | xīn | Heart; mind | 84 | 2.3% |
| 17 | 糸 | sī | Silk; thread | 79 | 2.1% |
| 18 | 艹 | cǎo | Grass; plant | 76 | 2.1% |
| 19 | 貝 | bèi | Shell; money; wealth | 71 | 1.9% |
| 20 | 忄 | xīn | Heart (side form) | 69 | 1.9% |
| 21 | 攵 攴 | pū | Tap; rap; script (folded… | 67 | 1.8% |
| 22 | 阝 | fù | Mound; hill (left ear) | 65 | 1.8% |
| 23 | 女 | nǚ | Woman; female | 62 | 1.7% |
| 24 | 小 ⺌ | xiǎo | Small; little | 59 | 1.6% |
| 25 | 田 | tián | Field; farmland | 55 | 1.5% |
The honest limits of this one. Nothing in this repository states which character contains which part — it is derived from stroke data by splitting each character at its radical and identifying each half by shape, and a part that cannot be named confidently is dropped rather than guessed at. So forms that deform heavily inside a character are undercounted, and two forms that are written almost identically — 日 and 曰 are the clearest pair — are separated by shape alone and can take each other's characters. Counting is also per written form: 氵 and 水 are one radical with one page but two shapes, so the busier shape takes the row rather than the two being added together. Read the most common radicals for the hand-written version of this list, and the radicals hub for all 100 of them.
Why Chinese has so many homophones
Because the sound system is tiny. Mandarin has 410 distinct syllables ignoring tone — built from 24 initials and 34 finals — and 1,183 syllable-and-tone combinations actually in use, out of the 1,640 that four tones would allow. English has thousands of syllables. Every Chinese character is exactly one of those 410, which is the whole explanation.
In this corpus the effect is measurable: its 2,711 distinct characters share only 1,025 different sounds between them — an average of 2.6 characters per sound, and far worse than that at the crowded end:
| Sound | Characters | Some of them |
|---|---|---|
| shì | 18 | 是事市试视式示世 |
| yì | 18 | 意议易艺义益谊异 |
| yù | 17 | 育预遇域誉喻玉愈 |
| jì | 16 | 记计绩技纪际济迹 |
| xī | 15 | 息西吸希惜悉夕析 |
| fù | 14 | 复父负付富附傅副 |
| zhì | 14 | 制质致志治置至智 |
| bì | 13 | 必毕币弊避闭壁臂 |
| jí | 13 | 及急极级集即圾疾 |
| jiàn | 12 | 件见建健践鉴键箭 |
| jiāo | 12 | 教交焦跤娇浇蕉椒 |
| bó | 11 | 博勃搏薄泊膊脖驳 |
And that is only the characters this corpus happens to use — a full dictionary is worse. It is also the practical reason tones are not optional: dropping the tone multiplies the number of things a syllable could be by four or five. The interactive pinyin chart has all 410 syllables with audio, and Bopomofo writes the same inventory a different way.
How this was measured — and what it does not say
A statistic without its method is a rumour, so here is the whole of it.
The corpus
10,884 distinct Chinese sentences, all of them published on this site and all of them readable in full: 9,825 hand-written dictionary examples, 469 lines from the graded readers, 460 travel phrases and 130 grammar examples. The word-length and radical tables are measured against different sets, each named beside its table: the dictionary index of 4,995 headwords, and the 3,684 characters the site's component data covers.
🔒 What the corpus is not
It is not a balanced corpus of native writing. It is Chinese written by people, for learners, at a spread of levels. That makes these figures a fair description of the material a learner actually meets — which is what the questions above are really asking — and an understatement of the vocabulary a newspaper, a novel or a contract demands. Anyone quoting a coverage percentage without saying what it was measured over is quoting nothing in particular; that is the mistake this page is written to avoid, not to repeat.
What a "token" is
Chinese is written without spaces, so "what share of the words" has no answer until something decides where the words are. Every sentence was segmented by greedy longest-match against the same in-memory dictionary the translators use — the identical step the HSK text analyzer performs on text you paste, so you can run the method yourself. That produced 66,619 word tokens (7,499 distinct words), an average of 6.1 words per sentence. Punctuation, Latin letters and digits are not tokens and are not counted anywhere on this page.
What is in the denominator
Everything. Words graded beyond HSK 6 and words the dictionary cannot identify are counted in the total rather than dropped from it. Removing them would raise every coverage figure on this page by several points and would be the single easiest way to make it wrong.
Tones and readings
Counted per character, not per word: every Chinese character is exactly one syllable, so a character-by-character pass gives one tone per occurrence with nothing to guess at. Each character's reading comes from the same dictionary as everything else, and the tone is read off the diacritic. Two honest caveats: a character with more than one reading is counted at its first, most common one, and a syllable that goes neutral only in context (the 了 of a finished action, a repeated 看看) is counted at its citation tone. A character the dictionary cannot read is left out of the tone table rather than defaulted to neutral.
When
Measured against the corpus as it stood on . Nothing here is stored: every figure is recomputed from the corpus each time the site restarts, so the page tracks the material rather than freezing at whatever it read the day it was written. If you quote a number from here, please quote the date with it and link back — the figures move as the corpus grows.
Questions people ask
How much of Chinese does HSK 3 cover?
Measured over the 10,884 Chinese sentences published on this site, HSK levels 1 to 3 together account for 53.6% of the words used — about 598 words of vocabulary. HSK 4 takes that to 60.9%, and all six levels together reach 78.1%. Those figures describe learner-facing Chinese — graded readers, example sentences, travel phrases and grammar examples — not a newspaper, which sits well above this. The last two levels move the number very little because the words they add are ones you meet rarely and need exactly.
How many Chinese characters do I need to know?
To read this corpus: 100 distinct characters cover half of the running text, 459 cover 80%, 818 cover 90% and 1,234 cover 95%. The whole corpus of 10,884 sentences uses only 2,734 different characters in total. The curve is steep at the start and very flat afterwards, which is why the first few hundred characters are worth far more per character than the next few thousand.
Is a Chinese word one character?
No, and that is the single most common misconception about the language. Of the 4,995 words in this site's dictionary index, 80.8% are written with exactly two characters and only 14.2% with one. Characters are closer to syllables, or to word roots, than to words.
Which Chinese tone is the most common?
The fourth tone — it accounts for 33.3% of the character occurrences in this corpus, against 15.9% for the second tone. The neutral tone is the rarest at 7.1%, but it lands on some of the most frequent characters in the language, so you say it constantly.
Which Chinese radicals are worth learning first?
The ones that appear in the most characters. Counted across the 3,684 characters this site indexes, 口 (mouth; opening) leads, followed by 人, 氵, 扌, 木. Learning twenty or thirty of these is the change that makes characters stop looking random.
How many syllables does Mandarin have?
410 distinct syllables ignoring tone, built from 24 initials and 34 finals, and 1,183 syllable-and-tone combinations in use out of the 1,640 that four tones would allow. English has thousands of syllables. That is why homophones are everywhere: the 2,711 characters in this corpus share only 1,025 different sounds between them, and shì alone is written 18 different ways.
What is this measured against?
10,884 Chinese sentences published on this site — 9,825 hand-written dictionary examples, 469 lines from the graded readers, 460 travel phrases and 130 grammar examples — plus the official HSK word lists and this site's dictionary index of 4,995 words. It is learner-facing Chinese written by people, not a balanced corpus of native writing, so it is a fair description of the material a learner actually meets and an understatement of how hard unsimplified Chinese is.
Can I quote these numbers?
Yes, with a link back to this page and the date they were measured, which is shown at the top. Every figure here is recomputed from the corpus each time the site restarts, so it tracks the corpus rather than freezing at whatever it read the day it was written.
Act on these numbers
- HSK word lists — all six levels, free, with printable PDFs
- Chinese Vocabulary Size Test — find out where on the coverage curve you already are
- HSK Text Analyzer — run the segmentation above on your own text
- Chinese Dictionary — the entry behind every character in the frequency table
- Radicals Hub — one page per component, in the order the table ranks them
- Example Sentence Bank — the corpus itself, searchable
- Tone Pitch Mirror — practise the tone you will say most