For teachers

How to teach pronunciation online

Video lessons strip out the mouth, the room and half the frequency range. Here is a method for correcting pronunciation over video that does not rely on the student being able to see your tongue.

A student says a word. It is wrong. You know it is wrong. You say the word correctly, they repeat it, it is wrong in exactly the same way, and you both silently agree to move on.

That loop is the default state of pronunciation teaching, and video makes it worse. In a room, a student can see your jaw drop, watch where your tongue goes, and hear the full frequency range of your voice. Over video they get a compressed 30 kbps stream that mangles sibilants, a camera framed on your forehead, and a half-second of latency that makes simultaneous repetition impossible.

So the techniques that work in person — model, repeat, model again — degrade badly. What survives is a method built on three things the medium cannot destroy: explicit description, contrast, and recorded evidence.

Why “listen and repeat” fails

Repetition assumes the student can hear the difference. Very often they cannot.

Adult learners perceive foreign sounds through the phoneme categories of their first language. A Japanese speaker hearing English right and light is not hearing two sounds and failing to reproduce one. They are hearing one sound, twice. Asking them to repeat more carefully is asking them to be more accurate about a distinction that, perceptually, is not there.

This is well established — it is the reason categorical perception experiments exist — and it has one enormous practical consequence:

If a student cannot hear a contrast, production practice is premature. Fix perception first.

You can test this in ninety seconds. Say ten words at random from a minimal pair set (ship / sheep, pero / perro, tuli / tuuli) and have the student write down which one they heard. If they score near chance, you have a perception problem, and no amount of “try again” will solve it.

The method

1. Diagnose against the student’s first language

Almost every pronunciation error is predictable from the learner’s L1. You do not need to discover their problems by trial and error; you can look them up.

A Spanish speaker learning English will start words like school with an epenthetic /e/. A French speaker will flatten English stress and drop /h/. A Mandarin speaker will simplify final consonant clusters and devoice final stops. A German speaker will devoice final consonants too — but for a different reason, and it will be more consistent.

Knowing this before the first lesson means you spend the diagnostic phase confirming a hypothesis rather than fishing. Our language pronunciation guides list the specific hard sounds for 29 languages precisely so you can do this in five minutes.

2. Describe the articulation, do not just demonstrate it

This is the single biggest adjustment for teaching over video. Instead of showing, tell — in physical, checkable instructions.

Bad: “It’s more like /ʒ/, listen — /ʒ/.”

Good: “Put your tongue where it is for /s/, then slide it back about a centimetre, and round your lips slightly. Now turn your voice on.”

The second version gives the student something they can verify without hearing themselves accurately. They can feel whether their tongue moved. They can feel whether their lips rounded. Proprioception works over video; audio fidelity does not.

Useful physical checks you can ask for down a video link:

FeatureHow the student checks it themselves
VoicingFingers on the throat — buzz or no buzz
AspirationSheet of paper in front of the mouth — does it move?
NasalityPinch the nose — does the sound stop or change?
Vowel lengthTap the table once per beat while saying it
Lip roundingLook in the self-view window, not at you

That last one matters. Tell students to watch their own video tile. Most of them have never looked at their mouth while speaking a foreign language, and the self-view is the one high-quality video feed in the call.

3. Contrast, always

Never present a sound alone. A sound in isolation has nothing to be measured against, and the student’s brain will assimilate it to the nearest L1 category within about a second.

Present it as a pair:

  • The target sound vs. the sound they are actually producing
  • The target word vs. their L1 near-neighbour word
  • Two target-language words that differ by only that sound

The third is a minimal pair, and it is the most efficient drill in existence because it isolates exactly one variable. If the student gets pero and perro right at 90%, the trill is learned. If they get 50%, it is not, and you know precisely what to work on.

4. Record, then compare

The student’s live perception of their own voice is unreliable — bone conduction changes what they hear, and they are busy producing rather than listening. A recording removes both problems.

The workflow that actually gets used:

  1. Student says the word. You capture it.
  2. Play it back immediately, followed by a model.
  3. Ask: “What is different?” — not “Was that right?”
  4. One adjustment. Record again. Compare the two attempts.

The question matters. “Was that right?” invites a yes/no from someone who cannot judge. “What is different?” forces them to attend to a specific dimension, which is the actual skill you are training.

Two attempts side by side is the moment students usually say “oh” — the first time they hear their own error from outside their own head.

5. Write it down phonetically

Whatever you correct in the lesson evaporates by Thursday unless it exists somewhere the student will look again.

A note that says “work on your R” is useless. A note that says:

rouge /ʁuʒ/ — uvular ʁ, not alveolar r. Gargle position. Compare with rue /ʁy/.

is a practice instruction. The student can act on it alone, six days later, without you.

This is where the IPA earns its keep. You do not need to teach the whole chart. You need the student to know the four or five symbols that represent their personal problems, so a transcription in their notes carries real information instead of being decoration.

A five-minute pronunciation slot

Pronunciation is a motor skill. Motor skills respond to short, frequent practice and barely respond to occasional long practice. Five minutes every lesson beats fifty minutes once a month, comfortably.

A slot that works:

0:00  Perception check — 6 minimal pairs, student writes what they hear
1:30  Review the one sound from last lesson — 3 words
2:30  New target sound — description + physical check
3:30  Production — 5 words, record the best and worst
4:30  Write the target into their notes with a transcription

Do that every lesson for a term and you will have covered every problem sound in the student’s inventory, twice, with recorded evidence of progress. That is more pronunciation work than most learners get in a decade.

Correcting without discouraging

Pronunciation correction is uniquely face-threatening. Grammar errors feel like a knowledge gap; accent errors feel like a personal deficiency, because they are tied up with identity.

Three rules keep it survivable:

Correct the sound, not the person. “That /θ/ came out as /s/” is a technical observation. “Your th is bad” is a verdict.

Correct one thing at a time. A learner producing an English sentence may have six deviations from the model. Naming all six teaches nothing and demoralises thoroughly. Pick the one with the highest intelligibility cost — usually stress placement or a vowel merger, rarely the exotic consonant everyone worries about.

Aim at intelligibility, not nativeness. Most adult learners will keep an accent forever, and that is fine. The goal is that a stranger understands them the first time, without effort. Say this out loud to the student early, because otherwise they will privately measure themselves against a native speaker and privately conclude they are failing.

What to prioritise when you can only fix three things

Ranked by how much they damage intelligibility, roughly across languages:

  1. Word stress and rhythm. A word with the wrong stress is often simply not recognised. This outranks every consonant.
  2. Vowel contrasts that carry meaning in the target language — English tense/lax, Finnish and Italian length, Arabic long/short.
  3. Consonant contrasts with a high functional load — the pairs that separate a lot of real words, not the ones that are merely exotic.

The English th is famously difficult and famously unimportant: substituting /f/ or /t/ costs almost nothing in comprehension. Meanwhile a learner who says CON-tri-bute instead of con-TRI-bute will be misheard every time. Teach the second one first.

The part that should be automatic

Everything above is method, and method is the teacher’s job. What is not the teacher’s job is transcribing a sentence by hand, looking up its phonetic form, typing a translation, and pasting the lot into a document at eleven at night.

That is the half of pronunciation teaching we built Teachee for: you record the sentence during the lesson, and the transcript, the translation and the phonetic transcription appear on the student’s screen while you are still talking about it, ready to become a flashcard. The judgement stays with you. The typing does not.

Frequently asked questions

Can you actually teach pronunciation properly over video?

Yes, but not with the same techniques you would use in a room. Video compression damages exactly the high-frequency information that distinguishes sibilants, and webcam framing hides the jaw and lips. The method that survives the medium is description plus contrast plus recorded evidence — telling the student precisely what to do with their tongue, giving them a minimal pair to test it against, and letting them hear their own recording next to a model.

How much of a lesson should be pronunciation?

Five to ten minutes, every lesson, is far more effective than a dedicated pronunciation lesson once a month. Pronunciation is a motor skill; it responds to frequent short practice and barely responds to occasional long practice.

Should I use the IPA with beginners?

Use it as a labelling tool, not a curriculum. You do not need a student to read the whole chart. You need them to recognise the four or five symbols that represent their personal problem sounds, so that a note in their vocabulary list means something specific.

What if I cannot hear the difference myself?

Then do not correct it. Teachers do more damage guessing than ignoring. Record the student, put the recording next to a native model, and let the comparison do the work — or use a phonetic transcription to check what the target actually is before you claim it is wrong.