Shadow a Chinese sentence, then hear both back

Play a real sentence, say it over the top, and the page draws your pitch and your timing against the model's — then plays them back to back, which is where the difference is always obvious. Five levels, from four-character beginner lines to full-speed HSK 6. Nothing you record leaves this page: it is measured in your browser, never uploaded, and never stored.

Free — no account

Pick a level

Beginner — up to 6 characters, HSK 1 words. One short breath. The whole sentence fits inside a single phrase, so you can copy the melody without having to plan where to breathe. 80 sentences.

Why one syllable at a time is not enough

Most learners can produce all four tones correctly when they say them on their own. The same learner then says a ten-character sentence and the third tones stop dipping, the fourth tones stop falling all the way, and an English sentence melody — the one that rises towards the end of a question and falls at the end of a statement — gets laid straight over the top of the Chinese one. None of that is audible to you while you are speaking.

Shadowing is the fix that language teachers have used for decades: say the sentence with, or immediately after, a native model, and compare. What this page adds is the comparison — the model and your take drawn on the same time axis and played back to back — because the gap is usually one specific word, and until you can see or hear exactly which one, repeating the whole sentence is guesswork.

The half the apps leave out

The 2026 round-ups of Chinese apps disagree about nearly everything except one finding: no single app covers the whole job. And the criticism raised most often about the big beginner courses is that they are thin on tones and speaking. A lesson you pass by tapping the right tile never has to know what you sound like, and a streak can run for a year without anyone — including you — hearing you say a whole sentence.

This page is that missing half at sentence length, and it is free for a structural reason rather than a generous one: your pitch is measured by your own browser, with no recognition service behind it, so there is nothing here for anyone to bill you for. No account, no sign-in, no cap on takes. Which course or app to put beside it is a separate question, and the comparison of fifteen ways to learn Chinese answers it by saying what each one is best at and what each one is missing — this site included.

What the three scores mean

Score What is measured What a low one usually means
Melody The shape of your pitch across the whole sentence, against the model's. In Mandarin that shape is mostly the tones. The tones flattened out once the sentence got long, or an English question or list intonation was laid over the top.
Rhythm Where the loud and quiet moments fell — the stresses and the pauses — relative to the length of the sentence. Your syllables are lined up with the model's first, so landing one slightly early or late costs you a little; landing them in a different order or missing a pause costs you a lot. The sentence was grouped differently: usually an even, character-by-character delivery where the model runs words together and then pauses.
Pace How long you took compared with the model at the speed you last played it. Both listen buttons are slowed, and the score follows whichever one you used, so practising at the slower speed never costs you pace. Reading rather than speaking. It is the easiest of the three to fix and the one that most changes how fluent you sound.

Both pitch curves are drawn in semitones relative to your own average in that take, never in hertz, so a deep voice and a high voice score the same when the shape is right. Time is normalised before the melody and rhythm are compared, which is why being slower costs you pace and not the other two — a sentence said correctly but slowly has one thing to fix, and being told everything was wrong would be both untrue and useless. Each of the three is worth a fixed share of the total — melody the most, pace the least — and every take shows its own arithmetic under the bars, so the number is never a verdict you have to take on trust. Rhythm is marked the most kindly of the three, because the model is a speech synthesiser: its loudness is more even than any person's and it never breathes, so scoring you on how close your delivery is to that would be asking you to sound like a machine. Tones are the opposite — they are the language, so melody is held to the model. Your syllables are then lined up with the model's, within a limit of about a sixth of a second, before the melody is read: no one speaks to the millisecond, and a syllable said a fraction late is a timing difference, so it is charged to rhythm once rather than to both. The lining-up itself is decided by loudness alone — never by the pitch it is about to score, which could only manufacture the match it was supposed to be measuring.

“Slower than the model” means slower than the playback you heard, not slower than the clip. Both listen buttons re-time the model in your browser, and pace is measured against whichever one you played last, so copying the slow speed accurately scores the same as copying the faster one accurately. That is deliberate: a page that offers a slower playback and then marks you down for matching it would be penalising you for doing what it asked, and the bar would look identical either way. Only pace uses the speed — melody and rhythm are compared over normalised time, so re-timing cannot move them.

One sentence from each level

A level is the harder of two things, never the easier: how long the sentence is, and how hard its words are. A six-character sentence full of HSK 5 vocabulary is not a beginner item, and neither is a twenty-character sentence of HSK 1 words.

Level Sentence Meaning
Beginner
up to 6 characters, HSK 1 words
我爱我的家。
Wǒ ài wǒ de jiā.
I love my family.
Dictionary example
Elementary
up to 9 characters, HSK 2 words
我们八点看电视。
Wǒmen bā diǎn kàn diànshì.
We watch TV at eight o'clock.
Dictionary example
Intermediate
up to 13 characters, HSK 3 words
听到这个消息我很高兴。
Tīngdào zhège xiāoxi wǒ hěn gāoxìng.
I'm happy to hear this news.
Dictionary example
Advanced
up to 18 characters, HSK 4 words
根据老师的要求,我们要写一段话。
Gēnjù lǎoshī de yāoqiú, wǒmen yào xiě yí duàn huà.
As the teacher requires, we have to write a paragraph.
Dictionary example
Expert
the longest sentences and HSK 5–6 words
哎,你怎么才来?
Āi, nǐ zěnme cái lái?
Hey, why are you only getting here now?
Dictionary example

Nothing here was written for a drill

The 400 sentences on this page were not written for it. Every one of them is already published somewhere else on this site, and the reading shown under it is the same reading those pages print — so a sentence you shadow here can be read in context, and this page cannot contradict the rest of the site about what it says.

Sentences that could not be shadowed fairly were left out rather than trimmed to fit: too short to be a sentence, too long to say in one breath, or with a word the dictionary cannot read, which would leave a hole in the middle of the pinyin you are reading from. 10,884 sentences were considered.

What happens to your recording

Your take is held in one block of memory that the next take overwrites. It is never uploaded — not to this site and not to anyone else — and nothing is stored: no file is created, no download link exists, and nothing is written to your browser's storage. Closing the tab is the end of it.

There is no speech-recognition service behind this page and there never will be. Shadowing does not need to know what you said, only what your pitch and your timing did, and that is measured by arithmetic running in your own browser. It is also why the page is free: nothing here is metered by anyone.

The only thing this page fetches is the model clip and the list of sentences. The microphone is released the moment a take ends, so your browser's recording indicator goes out between attempts — which is the only evidence any of this is true that does not require taking our word for it.

Questions people ask

What is shadowing, and why does it work for Chinese?

Shadowing means saying a sentence along with a native recording rather than reading it off a page. It works because the problems that make a learner hard to understand in Chinese are rarely individual sounds — they are tones that flatten out across a sentence and English rhythm laid over the top, neither of which you can hear yourself doing while you are doing it.

Should I use listen-and-repeat or live shadowing?

Listen and repeat first. Live shadowing is the harder and more useful version, but it needs headphones: without them your microphone picks up the model as well as your voice, and what gets measured is the two mixed together. On a phone speaker, use listen and repeat.

Why did it say my take was unreadable?

Because it would rather say nothing than draw something confident and wrong. A pitch tracker fed breath, a whisper or room noise still returns numbers, and a curve drawn from them looks exactly as authoritative as a real one. A take that is silent, under a third of a second, or has no clear pitch in it is reported as unreadable and nothing is drawn.

Which level should I start at?

The one whose sentences you can already read without stopping. Shadowing works when the sentence is in your ear; a sentence you are decoding character by character cannot be said at the speed you just heard, and the pace score will simply tell you so five times in a row — dropping to the slower playback moves that speed, not the standard, so it will still tell you. Move up when a round averages 75%.

Do I still need an app or a course as well?

Yes, and that is not a concession: no single app covers the whole job, which is the one point the 2026 app round-ups agree on. A course sequences your learning and gets you to turn up; this page does the part those courses are most often criticised for leaving thin. To choose the other half, the comparison of fifteen ways to learn Chinese says what each one is best at and what each one is missing.

The rest of the speaking toolkit