Chinese language statistics, measured

Almost every number written about Chinese online is a round figure repeated from another blog. These are not. Each one is counted, here, from the 10,884 Chinese sentences this site publishes and its dictionary index of 4,995 words — with the method written out below, and every row linked to the page that teaches what it counts.

Measured
78.1%

of the words in these sentences are covered by the whole of HSK 1–6

459

characters read 80% of the running text; 100 read half of it

80.8%

of Chinese words are written with exactly two characters, not one

410

syllables in the whole of Mandarin, ignoring tone — see the pinyin chart

What does each HSK level actually buy you?

The question behind "is HSK 3 enough?" is really what share of the words in front of me will I know. Here it is counted: every word in 10,884 sentences was found by this site's own segmenter and graded against the official HSK lists. HSK 1–3 reaches 53.6%, HSK 4 reaches 60.9%, and the last two levels together add barely 17.2% — each of their words is one you meet rarely and need exactly.

Share of word tokens in 66,619 words of this site's Chinese. Every token counts, including the ones beyond HSK.
Level Words added Words known This level Covered so far What it buys you
HSK 1 150 150 34.5%
34.5%
Greetings, numbers, family and the grammar words that hold a sentence together. Not yet a text you can read, but the frame every later sentence hangs on.
HSK 2 151 301 9.9%
44.4%
Everyday errands — shopping, time, travel, feeling ill. Short written messages start to make sense with a dictionary open beside you.
HSK 3 297 598 9.2%
53.6%
The level where reading begins. Graded readers, simple signage and children's material become mostly comprehensible, and looking a word up stops interrupting every line.
HSK 4 598 1,196 7.3%
60.9%
Opinions, comparisons and reasons. Conversation stops being transactional, and news headlines are readable even when the article is not.
HSK 5 1,300 2,496 8.4%
69.2%
Abstract and written registers: argument, culture, work. This is where the words start being ones you meet rarely but need exactly.
HSK 6 2,499 4,995 8.8%
78.1%
The formal and literary tail — idioms, set phrases and the vocabulary of essays and editorials. Each word here appears seldom, which is why the curve flattens.
Beyond HSK 21.9% Real words the dictionary knows but the HSK lists do not grade — names, places, and everyday vocabulary the syllabus simply leaves out. They are counted in the denominator above, not discarded, because dropping them is how coverage percentages get inflated.
Unidentified 0.0% Chinese the dictionary has no entry for at all.

Read this as a floor, not a ceiling. The corpus is learner-facing Chinese — graded readers, example sentences, travel phrases, grammar examples — so it is a fair description of the material a learner meets and an understatement of how much vocabulary a newspaper demands. Test yourself against it with the vocabulary size test or the HSK text analyzer, which reads your own text with the same segmenter that produced this table.

How many characters do you actually need?

Far fewer than the number people quote, because character frequency is brutally top-heavy. All 10,884 sentences here — 101,567 characters of running text — use only 2,734 different characters between them, and the first handful do most of the work.

100
characters to read 50% of the text
459
characters to read 80% of the text
818
characters to read 90% of the text
1,234
characters to read 95% of the text

The jump from 818 to 1,234 characters buys five percentage points; the first 100 bought fifty. That shape is the argument for learning characters in frequency order and for not worrying about the long tail until much later.

The commonest characters in this corpus

Counted over every Chinese character in the corpus. Tap any character for its full dictionary entry.
# Character Pinyin Meaning Times Running total
1 de Of (aim - the) 3,483 3.4%
2 le (past tense indicator); finish (end) 2,835 6.2%
3 He (him - another - other) 2,835 9.0%
4 zhè This (these); This (these) 2,602 11.6%
5 One (single) 1,876 13.4%
6 I (me - myself) 1,869 15.3%
7 hěn Very 1,446 16.7%
8 ([measure word]) 1,358 18.0%
9 shì To be (is - are - am - yes) 1,160 19.2%
10 No (not) 1,126 20.3%
11 yǒu To have (has - there is - there are) 1,045 21.3%
12 zài Located (in - at - exist) 1,017 22.3%
13 yāo Want (need - request - must - demand) 754 23.0%
14 rén Person (people) 736 23.8%
15 shàng Above (on - go up - upper - first) 735 24.5%
16 men ([plural indicator]) 715 25.2%
17 tiān Day (heaven - sky) 679 25.9%
18 jiā Family (home - domestic) 577 26.4%
19 Son (child - person - midnight) 568 27.0%
20 Big (oldest) 534 27.5%
21 In (inside - inner - neighbourhood - Chinese mile [half a km]) 521 28.0%
22 qǐng Please (ask - invite - request) 483 28.5%
23 lái Come (arrive) 480 29.0%
24 xué Learn (study - imitate - science - ology) 477 29.4%
25 hǎo Good (well); Like 472 29.9%
26 Get; Have to (need) 460 30.4%
27 shí When (time - hour - season - o'clock) 450 30.8%
28 xià Below (under - down - fall - next - lower - finish) 438 31.2%
29 shì Matter (thing - accident - trouble - responsibility - job) 428 31.7%
30 dào To (arrive - go to) 420 32.1%
31 nián Year (age - old) 402 32.5%
32 Earth (ground - location); ~ly 383 32.8%
33 shēng Life (give birth - be born - grow) 380 33.2%
34 duō Many (more - how) 378 33.6%
35 diǎn O'clock (dot - point - mw for suggestions) 373 34.0%
36 shuō Said (speak - say - talk - explain) 357 34.3%
37 zhāo Move (tactic); Touch (affected by - ignited) 354 34.7%
38 Go (leave - remove - send - past) 352 35.0%
39 huì Meeting (understand - able to - can - will - shall) 350 35.4%
40 You 341 35.7%

A frequency list from a corpus of learner Chinese is not the same as one from newspapers, and the difference is visible in the top rows: first- and second-person pronouns rank far higher here than they would in reported speech. That is a property of the material, stated rather than hidden. Practise writing any of them on the stroke-order animator or a 田字格 practice sheet.

Is a Chinese word one character?

No — and this is the misconception that costs beginners the most time. A character is closer to a syllable or a word root than to a word. Counted across the 4,995 headwords in this site's dictionary index:

Written with Words Share of the index
1 character 708
14.2%
2 characters 4,034
80.8%
3 characters 132
2.6%
4 characters 121
2.4%

Two characters is the overwhelming default. This is why looking a sentence up character by character produces nonsense, and why the pinyin annotator and the translators find word boundaries first — the same step that produced the HSK table above.

Which tone will you say most?

Tone drills usually give the four tones equal time. They do not appear equally. Counted over the 101,519 character occurrences in this corpus whose reading the dictionary supplies:

Tone What the pitch does Share Commonest character with it
First tone high and level
23.8%
Second tone rises
15.9%
rén
Third tone dips, then rises
19.8%
Fourth tone falls sharply
33.3%
zhè
Neutral tone short, light, unstressed
7.1%
de

The fourth tone is the one to get right first. The neutral tone looks negligible at 7.1%, but it lands on some of the most frequent characters in the language, so you will say it constantly — and it is the one most courses skip. Hear your own pitch against the target on the tone pitch mirror, drill recognition on the tone trainer, or see every syllable and tone in the pinyin chart.

Which radicals repay learning first?

The ones that turn up in the most characters. Counted across the 3,684 simplified and traditional characters this site indexes, using the same component data the character finder searches:

Counted per written form. Every row links to that radical's own page.
# Form Pinyin Meaning Characters Share
1 kǒu Mouth; opening 324 8.8%
2 人 亻 rén Person 215 5.8%
3 shuǐ Water (three drops) 211 5.7%
4 shǒu Hand (side form) 208 5.6%
5 Tree; wood 181 4.9%
6 yuē To say; speak 107 2.9%
7 Sun; day 99 2.7%
8 刀 刂 dāo Knife 98 2.7%
9 chuò Walk; movement (走之) 96 2.6%
10 火 灬 huǒ Fire 94 2.6%
11 Earth; soil 93 2.5%
12 yuè Moon; also 肉 flesh 92 2.5%
13 yán Speech; words 92 2.5%
14 mián Roof (of a house) 89 2.4%
15 tóu Lid; top 86 2.3%
16 xīn Heart; mind 84 2.3%
17 Silk; thread 79 2.1%
18 cǎo Grass; plant 76 2.1%
19 bèi Shell; money; wealth 71 1.9%
20 xīn Heart (side form) 69 1.9%
21 攵 攴 Tap; rap; script (folded… 67 1.8%
22 Mound; hill (left ear) 65 1.8%
23 Woman; female 62 1.7%
24 小 ⺌ xiǎo Small; little 59 1.6%
25 tián Field; farmland 55 1.5%

The honest limits of this one. Nothing in this repository states which character contains which part — it is derived from stroke data by splitting each character at its radical and identifying each half by shape, and a part that cannot be named confidently is dropped rather than guessed at. So forms that deform heavily inside a character are undercounted, and two forms that are written almost identically — 日 and 曰 are the clearest pair — are separated by shape alone and can take each other's characters. Counting is also per written form: 氵 and 水 are one radical with one page but two shapes, so the busier shape takes the row rather than the two being added together. Read the most common radicals for the hand-written version of this list, and the radicals hub for all 100 of them.

Why Chinese has so many homophones

Because the sound system is tiny. Mandarin has 410 distinct syllables ignoring tone — built from 24 initials and 34 finals — and 1,183 syllable-and-tone combinations actually in use, out of the 1,640 that four tones would allow. English has thousands of syllables. Every Chinese character is exactly one of those 410, which is the whole explanation.

In this corpus the effect is measurable: its 2,711 distinct characters share only 1,025 different sounds between them — an average of 2.6 characters per sound, and far worse than that at the crowded end:

Tone counts as part of the sound: shī, shí, shǐ and shì are four different syllables here, not one.
Sound Characters Some of them
shì 18 是事市试视式示世
18 意议易艺义益谊异
17 育预遇域誉喻玉愈
16 记计绩技纪际济迹
15 息西吸希惜悉夕析
14 复父负付富附傅副
zhì 14 制质致志治置至智
13 必毕币弊避闭壁臂
13 及急极级集即圾疾
jiàn 12 件见建健践鉴键箭
jiāo 12 教交焦跤娇浇蕉椒
11 博勃搏薄泊膊脖驳

And that is only the characters this corpus happens to use — a full dictionary is worse. It is also the practical reason tones are not optional: dropping the tone multiplies the number of things a syllable could be by four or five. The interactive pinyin chart has all 410 syllables with audio, and Bopomofo writes the same inventory a different way.

How this was measured — and what it does not say

A statistic without its method is a rumour, so here is the whole of it.

The corpus

10,884 distinct Chinese sentences, all of them published on this site and all of them readable in full: 9,825 hand-written dictionary examples, 469 lines from the graded readers, 460 travel phrases and 130 grammar examples. The word-length and radical tables are measured against different sets, each named beside its table: the dictionary index of 4,995 headwords, and the 3,684 characters the site's component data covers.

🔒 What the corpus is not

It is not a balanced corpus of native writing. It is Chinese written by people, for learners, at a spread of levels. That makes these figures a fair description of the material a learner actually meets — which is what the questions above are really asking — and an understatement of the vocabulary a newspaper, a novel or a contract demands. Anyone quoting a coverage percentage without saying what it was measured over is quoting nothing in particular; that is the mistake this page is written to avoid, not to repeat.

What a "token" is

Chinese is written without spaces, so "what share of the words" has no answer until something decides where the words are. Every sentence was segmented by greedy longest-match against the same in-memory dictionary the translators use — the identical step the HSK text analyzer performs on text you paste, so you can run the method yourself. That produced 66,619 word tokens (7,499 distinct words), an average of 6.1 words per sentence. Punctuation, Latin letters and digits are not tokens and are not counted anywhere on this page.

What is in the denominator

Everything. Words graded beyond HSK 6 and words the dictionary cannot identify are counted in the total rather than dropped from it. Removing them would raise every coverage figure on this page by several points and would be the single easiest way to make it wrong.

Tones and readings

Counted per character, not per word: every Chinese character is exactly one syllable, so a character-by-character pass gives one tone per occurrence with nothing to guess at. Each character's reading comes from the same dictionary as everything else, and the tone is read off the diacritic. Two honest caveats: a character with more than one reading is counted at its first, most common one, and a syllable that goes neutral only in context (the 了 of a finished action, a repeated 看看) is counted at its citation tone. A character the dictionary cannot read is left out of the tone table rather than defaulted to neutral.

When

Measured against the corpus as it stood on . Nothing here is stored: every figure is recomputed from the corpus each time the site restarts, so the page tracks the material rather than freezing at whatever it read the day it was written. If you quote a number from here, please quote the date with it and link back — the figures move as the corpus grows.

Questions people ask

How much of Chinese does HSK 3 cover?

Measured over the 10,884 Chinese sentences published on this site, HSK levels 1 to 3 together account for 53.6% of the words used — about 598 words of vocabulary. HSK 4 takes that to 60.9%, and all six levels together reach 78.1%. Those figures describe learner-facing Chinese — graded readers, example sentences, travel phrases and grammar examples — not a newspaper, which sits well above this. The last two levels move the number very little because the words they add are ones you meet rarely and need exactly.

How many Chinese characters do I need to know?

To read this corpus: 100 distinct characters cover half of the running text, 459 cover 80%, 818 cover 90% and 1,234 cover 95%. The whole corpus of 10,884 sentences uses only 2,734 different characters in total. The curve is steep at the start and very flat afterwards, which is why the first few hundred characters are worth far more per character than the next few thousand.

Is a Chinese word one character?

No, and that is the single most common misconception about the language. Of the 4,995 words in this site's dictionary index, 80.8% are written with exactly two characters and only 14.2% with one. Characters are closer to syllables, or to word roots, than to words.

Which Chinese tone is the most common?

The fourth tone — it accounts for 33.3% of the character occurrences in this corpus, against 15.9% for the second tone. The neutral tone is the rarest at 7.1%, but it lands on some of the most frequent characters in the language, so you say it constantly.

Which Chinese radicals are worth learning first?

The ones that appear in the most characters. Counted across the 3,684 characters this site indexes, 口 (mouth; opening) leads, followed by 人, 氵, 扌, 木. Learning twenty or thirty of these is the change that makes characters stop looking random.

How many syllables does Mandarin have?

410 distinct syllables ignoring tone, built from 24 initials and 34 finals, and 1,183 syllable-and-tone combinations in use out of the 1,640 that four tones would allow. English has thousands of syllables. That is why homophones are everywhere: the 2,711 characters in this corpus share only 1,025 different sounds between them, and shì alone is written 18 different ways.

What is this measured against?

10,884 Chinese sentences published on this site — 9,825 hand-written dictionary examples, 469 lines from the graded readers, 460 travel phrases and 130 grammar examples — plus the official HSK word lists and this site's dictionary index of 4,995 words. It is learner-facing Chinese written by people, not a balanced corpus of native writing, so it is a fair description of the material a learner actually meets and an understatement of how hard unsimplified Chinese is.

Can I quote these numbers?

Yes, with a link back to this page and the date they were measured, which is shown at the top. Every figure here is recomputed from the corpus each time the site restarts, so it tracks the corpus rather than freezing at whatever it read the day it was written.

Act on these numbers