For learners
Why you cannot hear the difference
Adults do not fail to reproduce foreign sounds — they fail to receive them. Here is what happens to your ears at ten months old, and what can still be done about it decades later.
A Spanish speaker learning English is asked whether beach and bitch sound different. They listen carefully, twice, and say no.
They are not being inattentive. They are reporting accurately on what reached them.
What happens at ten months
Newborns discriminate essentially every phonetic contrast used by any human language. A Japanese infant distinguishes English /r/ from /l/. An English infant distinguishes Hindi’s dental and retroflex stops. This is measurable with head-turn preference procedures long before any child speaks.
Then it goes away.
Between roughly six and twelve months, infants stop responding to contrasts their surrounding language does not use, and get sharper at the ones it does. This is perceptual narrowing, and it is not damage — it is optimisation. A brain that must extract meaning from noisy speech at conversational speed benefits enormously from discarding distinctions that never carry information.
The cost arrives thirty years later, when you want those distinctions back.
Categorical perception
The adult system does something specific and worth understanding, because it explains why “listen more carefully” is useless advice.
Take a synthesised continuum of sounds between /b/ and /p/, varying voice onset time in even steps. Physically it is a smooth gradient. Perceptually it is not. Listeners hear /b/, /b/, /b/, /b/, then abruptly /p/, /p/, /p/ — with a sharp boundary and almost no ability to discriminate two stimuli that fall on the same side of it, even when the physical difference between them is identical to one that straddles the boundary.
Your perception is not a measuring instrument. It is a classifier, and it discards within-category detail as a design feature.
So when a contrast in another language cuts across one of your categories, the difference is removed before you become aware of it. There is no “listening harder” to be done, because the information is gone by the time it reaches the part of you that could listen harder.
Physical signal: ●───●───●───●───●───●───●───●
(smooth, even steps)
What you hear: [────── B ──────][────── P ──────]
(two boxes, sharp edge, no detail inside)
A foreign contrast that lives inside one box:
[──── B ────]
↑ ↑
these two are the same sound to you
Which contrasts will be hard — predictably
The difficulty is not about the sound in the abstract. It is about the relationship between the new sound and your existing categories. This is the core insight of Best’s Perceptual Assimilation Model and Flege’s Speech Learning Model, and it makes usefully specific predictions.
Two new sounds land in one of your categories → very hard. Both get assimilated to the same box. Japanese /r/ and /l/ for English speakers; English /iː/ and /ɪ/ for Spanish speakers. This is the worst case and the one that needs training.
Two new sounds land in two different categories → easy. You already have the boxes. English speakers have little trouble with Spanish /p/ and /b/.
One new sound lands in one category but badly → moderate. French /y/ assimilates to English /uː/, but poorly enough that learners notice something is off, which helps.
A new sound lands nowhere → often easy. Sounds with no native analogue at all — clicks, some pharyngeals — are sometimes learned faster than near-misses, because there is no category to fight. Counterintuitive and well attested: similar is harder than strange.
That last point has a practical consequence. Learners worry about the exotic sounds and get ambushed by the ones that are almost, but not quite, familiar.
What actually works
The good news is that adult perceptual systems are not frozen. They are plastic and they respond to a specific kind of practice.
Forced choice with immediate feedback. Not passive listening. You must commit to an answer and be told immediately whether it was right. Perceptual learning is error-driven; without the error signal, you spend twenty minutes rehearsing your existing categorisation.
High variability. Many talkers, many words, many phonetic contexts. Train on one speaker and you learn that speaker. Train on six and you form a category that generalises to voices you have never heard — which is the entire point. This finding, high-variability phonetic training, is the single most useful result in the field and the one most often ignored by software that ships one recording per word.
Short and frequent. Ten minutes daily beats an hour weekly, comfortably. This is a perceptual skill and it behaves like one.
Perception before production. Producing a distinction you cannot hear means you cannot evaluate your own output, so errors get reinforced rather than corrected. Fix the ears first. Full protocol in minimal pairs training.
What to expect
The classic /r/–/l/ training studies with Japanese adults are the reference point. Eight to twelve sessions of high-variability training produce reliable improvement, retained months later, with partial transfer to production — participants got measurably better at saying the contrast without any production practice at all.
They also show a ceiling. Adult performance improves substantially and typically stays short of native.
Both halves of that matter, and both should be said to learners out loud:
- Real improvement is available, in weeks, not years.
- Perfect native perception probably is not, and this is not a personal failure.
Learners who are told only the first half quit when they hit the ceiling. Learners told only the second half never start.
Why this changes how you should be taught
If a student mispronounces a sound and the teacher’s response is to model it again, the lesson assumes a production problem. Very often it is a perception problem, and modelling louder does nothing.
The diagnostic takes ninety seconds: play ten words at random from a minimal pair and have the student write down which they heard. Near chance means perception. Near ceiling means production. Two entirely different interventions follow, and picking the wrong one wastes a term.
This diagnostic is also why a shared, precise vocabulary is worth the setup cost — see how to read the IPA. “You are producing /s/ where the target is /θ/” tells a student what to test. “That sounded a bit off” tells them nothing they can act on alone.
Frequently asked questions
Why can't I hear the difference between two foreign sounds?
Your brain sorts incoming speech into the sound categories of your first language, automatically and below awareness. A contrast your language does not use gets collapsed into a single category before it reaches conscious perception, so both sounds genuinely arrive as the same sound.
Can adults learn to hear new speech sounds?
Yes. Training studies consistently show real, durable gains, particularly with high-variability training using many voices and immediate feedback. Adult performance usually stops short of native-level, but the improvement is substantial and it transfers to production.
How long does it take to learn to hear a new contrast?
Two to four weeks of ten minutes daily produces measurable change for most contrasts. Automaticity in fast connected speech takes considerably longer.
Does listening to lots of the language fix this on its own?
Only slowly and unreliably. Passive exposure without feedback lets you keep mis-categorising for years. What drives perceptual learning is forced choice plus immediate correction.