Hearing yourself: the oldest pronunciation trick, now with a score
Somewhere around 1997 I worked out that the sound card in our family PC could change my voice while I was speaking. Not record, then process. Live. I could talk into a cheap microphone and hear myself come back through the speakers half a second later, pitched down, sounding like a stranger who happened to know Polish.
I was supposed to be learning English and German. What I actually did, for hours, was read passages aloud and listen to this stranger read them back.
It felt like messing about. Turns out I had stumbled onto something that phonetics teachers had been doing deliberately for most of a century, and that the research community has since measured fairly carefully.
Nearly thirty years later, in 2026, we finally built it into Taalhammer: you speak a sentence in Learn, and you get your pronunciation back scored word by word and sound by sound. I will show you exactly what that looks like at the end, screenshots and all. First, why this exercise works at all, because the answer is more interesting than "practice makes perfect".
Why your own voice sounds wrong to you
Here is the first thing nobody tells you. When you speak, the sound reaches your inner ear by two routes at once. Some of it travels out of your mouth, through the air, and back into your ear canal, which is the path everyone else hears. The rest travels straight through the bones of your skull.
Bone conducts low frequencies better than air does. So the voice you have listened to your entire life is bassier and rounder than the one that exists outside your head. Press record, play it back, and the bottom drops out. That thin, slightly higher stranger is what the rest of the world has been hearing all along.
This is why the first recording is always unpleasant, and it is also why the exercise works. You can't evaluate a voice you're still producing. Speaking and judging compete for the same attention. Recording splits the job in two: say it now, judge it later, when your mouth isn't busy.
My accidental pitch-shifting made that split wider. If the voice is unfamiliar enough, you stop hearing "me, but wrong" and start hearing "a person saying a Polish vowel in the middle of an English word". The self-recognition gets out of the way, and the mistakes stand out.
Then the university put me in a booth
Fast forward to applied linguistics studies. The department had a phonetics lab: a room of booths, each with headphones and a microphone, and a console at the front where the teacher sat.
You would work through exercises. The teacher could drop into your headphones without warning, say the word properly, listen to you try again, and move on to the next booth. On the screen in front of you there were shapes. Sounds, drawn. You could see that your vowel was in the wrong place before anyone told you.
That room was descended from a long tradition. The first recognisable language lab was set up at the University of Grenoble in 1908, and by the mid-1960s the United States alone had something like 10,000 of them in secondary schools and 4,000 more in universities. Then the funding dried up, the theory behind them fell out of fashion, and most were dismantled.
What made that room valuable was not the technology, which by modern standards was primitive. It was the combination of three things: you produced sound, you saw it represented, and somebody told you specifically what was wrong. Not "your accent needs work". Which sound, in which word.
What the research says, including the parts that do not flatter the idea
Three things are well established, and one of them is a warning.
Pronunciation can be taught, and feedback is the part that does the work. Researchers have pooled dozens of studies on this. Teaching pronunciation clearly beats not teaching it, and when they looked at what separated the effective studies from the weak ones, the same factor kept coming up: whether learners were told how they had done. Longer courses helped. Feedback helped more.
What this means for you: saying a sentence twenty times on your own is worth less than saying it three times with somebody, or something, telling you which bit was off. Repetition without feedback is just repetition.
Training your ear works, and it works better with many voices than with one. There is a well-supported technique with an ugly name, High Variability Phonetic Training, and the whole idea is that you hear a target sound from lots of different speakers instead of one careful teacher. Learners trained this way get clearly better at telling similar sounds apart, they stay better months later, and the skill carries over to words they never trained on.
What this means for you: one narrator is not enough. Hearing the same Dutch g from twenty different mouths is what teaches your ear that they are all the same sound.
And the warning: hearing a sound is not the same as being able to say it. This is the finding people skip, and it is the honest one. When researchers train perception alone and then measure speech, the improvement in speech is real but modest, and sometimes barely there at all. The gap closes when learners are also told how the sound is physically made and then practise producing it.
What this means for you: listening alone will not fix your accent. You have to open your mouth, and then find out what came out.
So the picture isn't "listen and you will speak". It's closer to: listen, know what you are aiming at, say it, find out how close you got, adjust.
That loop is exactly what the booth was for. (If you want the underlying papers: Lee, Jang and Plonsky on teaching pronunciation, Uchihara, Karas and Thomson on training the ear, and this review on how much of it reaches your actual speech.)
Where a machine fits, and where it does not
Speech recognition has been studied as a pronunciation tool for years, and the research is encouraging without being starry-eyed. Three findings are worth knowing before you trust any app, ours included.
It's strong on individual sounds and weak on melody. Machines are good at consonants and vowels, and much weaker at stress, rhythm and intonation. A machine can tell you your Dutch g came out as a k. It is far less useful for telling you that your whole sentence has the wrong shape.
Explicit beats implicit. Feedback that names the problem works considerably better than feedback that merely fails to understand you. Being misheard tells you something is wrong. It does not tell you what.
It needs weeks, not evenings. The studies that showed real gains ran for five weeks and upwards. The ones that ran for an afternoon showed nothing. This is a habit, not a hack.
None of that is a reason to skip the tool. It's a reason to use it for what it's good at, and to keep listening to real speakers for the music of the language.
What this looks like in Taalhammer
This is the feature I wanted at fourteen and got at university, minus the room and the booking sheet.
You are in Learn, working through a card. The answer comes up, the target sentence in the language you are learning. There is a microphone in the bottom corner. Hold it, say the sentence out loud, let go.

A moment later you get:
- A score out of 100 for the whole sentence, with a plain verdict rather than just a number.
- The sentence marked up word by word, so a weak word is obvious at a glance instead of buried in an average.
- A sound by sound breakdown, if you open it. Each phoneme in each word, showing what the target was and what you actually produced. This is the IPA doing the job it was invented for in 1888: one symbol, one distinct sound, no arguing about spelling. If you have never read those symbols, our guide to the International Phonetic Alphabet is the place to start.
- What we heard, written out, so you can see the gap between the sentence you meant and the sentence that arrived.
- Your own recording, playable. The oldest part of the whole exercise, and still the part that teaches most.
- A retry, because the second attempt right after seeing the breakdown is where the learning happens.

In the screenshot above the sentence scored 85. Seven words are underlined green and one, était-il, is red at 54. Open the detail and you can see precisely why: the e came out as an s, the final l drifted towards a y, and an i went missing altogether. That is a different instruction from "your French needs work". It is one word, three sounds, and you can hear your own attempt on the same screen.
The rating buttons stay where they are. Your pronunciation score does not touch your review schedule, deliberately. Scheduling is about whether you remember the sentence. This is about whether you can say it. Mixing them would corrupt both.
It works for 31 languages, which is most of what people actually study with us: Dutch, German, English, Spanish, French, Italian, Portuguese, Polish, Swedish, Danish, Norwegian, Finnish, Czech, Slovak, Croatian, Hungarian, Romanian, Greek, Russian, Ukrainian, Turkish, Arabic, Hebrew, Hindi, Thai, Vietnamese, Indonesian, Filipino, Japanese, Korean and Chinese. Where we cannot score a language reliably, the microphone does not appear at all, which we think is better than a number that means nothing.
How I would actually use it
If you want the version that matches the evidence rather than the version that feels productive:
- Say it before you look at the score. The point is to commit to a pronunciation, not to hedge.
- Open the sound by sound view when a word is red. The word score tells you where to look. The phonemes tell you what to change.
- Play your own recording back at least once. The score is the machine's opinion. The recording is the evidence.
- Retry immediately. One informed second attempt beats five uninformed ones.
- Do not chase 100. Above roughly 80 you are intelligible, and intelligibility is the goal. Native-like is a different and much longer project.
- Keep listening to real people. The machine is good at sounds and poor at melody. Melody comes from input.
The thing I keep coming back to
The sound card, the booth, and the microphone in a flashcard app are all solving one problem: you can't hear yourself properly while you're talking, and you can't fix what you can't hear.
Everything since 1908 has been an attempt to hand you your own voice back, quickly enough that you still remember what you were trying to do with it. What has changed is that you no longer need a room, a timetable, or a teacher with a console to get an answer. You need about four seconds.
I would have loved this at fourteen. The stranger with the deep voice would have been out of a job.