Prática de Shadowing: Advance English Learning | The Mechanics of Fast Speech: Assimilation, Elision, and Intrusion - Aprenda a falar inglês com vídeo

Criando lição...
1
I'll spite.
2
So you study English for five years, right?
3
You memorize the vocabulary.
4
You master all those complicated verb tenses.
5
Oh, yeah.
6
Endless grammar drills.
7
Exactly.
8
And you practice in this quiet room, reading textbooks, where every single word is perfectly separated by, you know, a clean white space.
9
Right.
10
The visual boundaries are so clear on a page.
11
But then you step off a plane in London, or New York or Sydney, and you just try to order a cup of coffee.
12
And that's when the panic sets in.
13
It really does.
14
The barista looks at you and fires off a sentence that sounds like this single chaotic blur of noise.
15
Which is so frustrating because you know you should understand them.
16
Right.
17
I mean, if they wrote that exact same sentence down on a napkin, you would read it perfectly.
18
Instantly.
19
You know all those words.
20
But in the air, the words just seem to crash into each other.
21
Like a simple phrase like, I will see you next week.
22
Somehow it sounds like, I'll see you next week.
23
It's just a wall of sound.
24
It is.
25
And it makes you feel like you are starting from zero.
26
But I want you, the listener, to know that this is a completely universal shock for almost every language learner.
27
It really is.
28
I mean, you spend years training your eyes to understand a language, and you just expect your ears to naturally follow along.
29
Which makes sense, logically.
30
Sure.
31
But written English and spoken English are almost two entirely different systems.
32
We kind of assume that speaking is just reading out loud, but it isn't.
33
Not at all.
34
When native speakers talk at a normal, natural speed, those comfortable spaces between the words, they completely disappear.
35
The physical sounds warp, they shrink, and sometimes they vanish completely.
36
I always think of it like looking at a train standing at a station platform.
37
Oh, I like that analogy.
38
Yeah, so when the train is stopped, you can walk right alongside it.
39
You can see the clear gap between the first train car and the second train car.
40
You can count every single one.
41
Exactly.
42
Reading written English is like looking at that stopped train.
43
The individual train cars are the individual words on the page.
44
The structure is obvious, and you have all the time in the world to examine every piece.
45
But natural spoken English is that exact same train rushing past you at, you know, 100 miles per hour.
46
Right.
47
And at that speed, your eyes cannot process the gaps between the cars anymore.
48
No, the distinct shapes just blur together into one continuous streak of color.
49
The individual cars are still there, obviously, making up the train.
50
But the speed fundamentally changes how you perceive them.
51
And this continuous flowing stream of sound is exactly what linguists call connected speech.
52
So take a couple of seconds right now and think about a time recently when you felt that specific frustration.
53
Yeah, maybe you were watching a movie without subtitles where you were in a meeting with native speakers.
54
And you caught a few words, but the rest just sounded like a stream of water rushing over your head.
55
Try to remember how incredibly fast that felt.
56
Because the initial reaction for most learners in that moment is to blame their own skills.
57
They think, um, my vocabulary isn't good enough or my grammar is too weak.
58
They blame the native speaker.
59
They think the speaker is just mumbling or using some really obscure slang.
60
Exactly.
61
But usually neither of those things is true.
62
The native speaker is using simple words and they are actually pronouncing them correctly.
63
Correctly according to the hidden rules of connected speech, that is.
64
Right.
65
And before we learn these specific rules, we have to understand the fundamental reason why they exist.
66
Because it's not a conspiracy.
67
People don't speak this way to keep language learners out of the club.
68
No, it's not a secret code designed to make listening comprehension difficult.
69
It literally comes down to biology.
70
Biology.
71
So we have to look at the human mouth not just as this magical source of language, but as a physical mechanical engine.
72
That is a very helpful way to frame it.
73
The physical engine of speech.
74
we produce sound using a biological machine called the vocal tract.
75
And like any machine, it's made of parts.
76
Muscle, flesh, cartilage, bone.
77
Right.
78
It includes your lungs pushing the air up, your vocal cords vibrating in your throat, your jaw moving up and down.
79
The lips opening and closing.
80
Exactly.
81
And the tongue, which is actually a remarkably complex and heavy muscle, moving around inside the mouth.
82
And operating any machine made of heavy muscle requires physical energy.
83
Like every time we open our mouths to speak, we're literally burning calories.
84
We are.
85
And every single sound in the English language requires a highly specific, coordinated physical movement inside that vocal tract.
86
Which linguists call articulatory gestures, right?
87
Yes, articulatory gestures.
88
And to really understand how connected speech works, you have to feel these gestures in your own mouth.
89
Well, let's try it.
90
Let's think about the B sound, like in the word boy, and the P sound, like in the word pet.
91
So if you say boy or pet, notice what your lips are doing.
92
Yeah, my top lip and bottom lip have to press firmly together.
93
I hold the air behind my lips for just this tiny fraction of a second, and then I release it.
94
Right.
95
You literally cannot make those specific sounds if your lips remain open.
96
Now compare that lip gesture to a completely different sound.
97
Let's take the K sound in the word cat, or the hard G sound in the word go.
98
Okay, so for cat and go, my lips aren't doing anything at all.
99
They're just open.
100
Exactly.
101
The physical work is happening all the way at the back of your mouth.
102
Yeah, the back of my tongue has to pull upward to block the air near my throat.
103
So notice the huge physical distance between those two gestures.
104
Right, the lips are at the very front of the machine, and the back of the tongue is at the absolute rear.
105
And in a slow, deliberate, robotic sentence, moving between those two positions is easy.
106
But in a continuous, fast stream of conversation...
107
The brain starts sending these movement commands to the muscles very rapidly.
108
Yes.
109
The mouth is already preparing the muscular shape for the next
110
sound before it has even finished the physical movement for the current sound.
111
The muscles are overlapping their jobs.
112
It kind of reminds me of a relay race, you know, where one runner starts sprinting before the other runner has even handed them the baton.
113
That's exactly it.
114
The mouth is anticipating the future.
115
And this overlapping is incredibly clear when we look at laboratory research.
116
Oh, the computer models, right.
117
This is fascinating.
118
It really is.
119
Scientists who study linguistics and biomechanics, they use computers to build digital physical models of the human vocal tract.
120
So they basically program virtual lungs, virtual vocal cords, a virtual tongue, and virtual lips.
121
Yes.
122
And then they input the phonetic commands for a specific sentence.
123
So they literally type in the exact tongue and lip positions for every single word.
124
And when they tell the computer simulation to speak those words very slowly, the audio output sounds perfectly crisp.
125
Perfectly crisp.
126
Every consonant and vowel is clear, exactly like the dictionary pronunciation.
127
But the experiment gets really interesting when they speed up the timing of those commands.
128
Right.
129
When they force the virtual tongue and lips to move at the speed of a normal, fluent human conversation, the physics of the machine just take over.
130
Because the virtual muscles literally cannot move fast enough to hit every perfect position.
131
Exactly.
132
The sounds naturally begin to weaken.
133
Certain consonants change their shape because the virtual tongue just doesn't have the time to travel across the mountain.
134
And some sounds just disappear completely.
135
They do.
136
And the incredible thing is, the researchers did not program connected speech rules into the computer.
137
The computer just naturally generated connected speech simply because it was forced to move mechanical parts efficiently at high speeds.
138
Precisely.
139
When a native speaker says, I'll see you next week, they aren't being lazy.
140
Right.
141
They aren't just ignoring the grammar they learned in school.
142
Not at all.
143
They are acting out of biological efficiency.
144
The human body is governed by this principle called the economy of movement.
145
Economy of movement.
146
So whether you're walking, swimming, or speaking, your brain is constantly trying to use the absolute least amount of physical energy required to achieve the goal.
147
Yes.
148
And in speech, the goal is simply to be understood.
149
So the brain calculates exactly how much phonetic detail it can delete or blur together without losing the core meaning.
150
Now, economy of movement makes sense for any physical action.
151
But it seems to affect English much more aggressively than other languages.
152
It does seem that way, doesn't it?
153
Yeah.
154
I mean, a Spanish speaker talking quickly still seems to pronounce their vowels pretty clearly.
155
But an English speaker talking quickly turns half their vowels into absolute mud.
156
And that difference comes down to the underlying rhythm of the language itself.
157
Languages generally fall into two broad rhythmic categories.
158
Okay, what are they?
159
Well, languages like Spanish, French, and Japanese are generally what we call syllable-timed.
160
Syllable-timed, meaning every syllable gets a relatively equal piece of the pie.
161
Right.
162
If you listen to Spanish, it often sounds very even, almost like a machine gun.
163
Every syllable, whether it's an important noun or a tiny grammar word, receives a roughly equal amount of time and physical energy.
164
Exactly.
165
The rhythm is steady and flat, but English operates on a completely different system.
166
English is a stress-timed language.
167
Okay, so how does a stress-timed rhythm physically work in the mouth?
168
Well, in English, the time it takes to say a sentence does not depend on how many syllables are in the sentence.
169
It depends almost entirely on how many stressed, important words are in the sentence.
170
So the entire rhythm of spoken English is built around landing heavily and loudly on the core, meaning carrying words.
171
Yes, and rushing as quickly as possible through all the little grammatical filler words in between them.
172
I picture it like crossing a really fast moving river by jumping on large stepping stones.
173
That's a great visual.
174
Yeah, so the stepping stones are the stressed words you want to land on them firmly, and the water rushing between the stones represents the unimportant grammar words.
175
Right, you want to spend as little time as possible in the water.
176
You just leap over those gaps to get to the next heavy stone.
177
And to keep that heavy, bouncing rhythm intact, the words in the water, the unimportant words, have to physically shrink.
178
They must take up less time, which means the mouth must spend less energy pronouncing them.
179
Let's actually test this physical sensation.
180
I have a sentence for you, the listener.
181
The sentence is, they are here and they want to eat.
182
Okay, so first, try to pronounce this sentence with equal weight on every single word.
183
Right.
184
Do not let any word shrink.
185
Give every single syllable full energy.
186
Take a second and try it aloud.
187
It sounds incredibly rigid, doesn't it?
188
Totally.
189
They are here and they want to eat.
190
It actually takes a surprising amount of breath and jaw movement to force equal energy into tiny words like are and and.
191
It's exhausting.
192
So now let's apply the English rhythm.
193
The important stepping stones in that sentence are hear, want, and eat.
194
So let all the other words shrink.
195
Try speaking it naturally.
196
They're here and they want to eat.
197
They're here and they want to eat you can actually feel your jaw relax when you do that.
198
Yeah, you can feel your tongue barely moving for the smaller words.
199
And that physical relaxation is the economy of movement in action, because the rhythm demands that we reach the word eat on a specific beat.
200
The words want to literally have to compress into wanna.
201
And that compression is the foundation of our first major phonetic mechanism, which is called weak forms.
202
Weak forms are exactly where the dictionary pronunciation of a word actively betrays the language learner.
203
It really does.
204
So to grasp weak forms, we essentially have to divide the entire English vocabulary into two functional categories, right?
205
Yes.
206
The first group contains what we call the content words.
207
These are your heavy stepping stones.
208
They carry the actual information, the mental picture of the sentence.
209
Right.
210
This group includes nouns, main verbs, adjectives, and adverbs.
211
So if someone walks into a room and just shouts four content words at you, like dog, eat, big sandwich, you instantly have a vivid mental movie of what has happened.
212
Exactly.
213
The words contain heavy content, but the second group contains the function words.
214
And these words carry very little meaning on their own.
215
Their only job is to provide the grammatical architecture that holds the content words together.
216
Yes.
217
This group includes pronouns, prepositions, articles, conjunctions, and auxiliary verbs.
218
Words like an, at, to, can, should, of, for.
219
So if someone walks into a room and just shouts, and, at, to, can, of, at you, you have absolutely no idea what they want.
220
None.
221
The words are structurally necessary, but completely empty of imagery.
222
It's like building a brick wall.
223
The content words, the nouns and verbs, are the heavy, solid bricks.
224
and the function words are the wet cement that you just smear between the bricks.
225
And the cement doesn't need to be thick.
226
It doesn't need to look pretty or perfect.
227
It just needs to be there to hold the bricks in place.
228
Right.
229
So in spoken English, native speakers use the absolute minimum amount of physical effort to lay down that phonetic cement.
230
And the result of that absolute minimum effort is the most common sound in the English language.
231
The schwa.
232
The schwa.
233
The phonetic symbol looks like an upside-down E, and it represents a sound of pure, complete physical rest.
234
To make a schwa, you honestly do almost nothing.
235
You just drop your jaw a fraction of an inch.
236
You let your tongue lie completely flat and relaxed in the center of your mouth.
237
You don't round your lips.
238
You just push a tiny puff of air over your vocal cords to make this short, lazy grunt.
239
Ugh.
240
It is the sound of absolute physical neutrality.
241
And in connected speech, almost all of those grammatical function words lose their clear dictionary bowels,
242
and they just collapse into this lazy schwa sound.
243
Let's actually trace a few of these function words as they collapse.
244
Take the conjunction and.
245
Okay.
246
If a learner reads a text, they learn to pronounce it with a wide, clear mouth shape. And?
247
But in the middle of a fast sentence, a native speaker will not spend the energy to open their mouth that wide for wet cement.
248
No way!
249
A phrase like fish and chips never features the full and.
250
Right.
251
The a sound collapses into a schwa, and the deed sound frequently disappears completely.
252
It just becomes fish and chips.
253
Fish and chips, you barely hear a vowel at all, just a quick nasal hum connecting the two heavy bricks.
254
Exactly.
255
And the preposition of undergoes a very similar collapse.
256
Yeah.
257
The dictionary tells you it has a round vowel. Of.
258
But native speakers reduce it to a schwa, and they often drop the consonant entirely.
259
Like, you don't order a cup of coffee, you order a cup of coffee.
260
Cup of coffee.
261
The of is nothing more than a tiny relaxed breath.
262
What about the word to?
263
By itself, to has a strong rounded vowel sound. To.
264
But as a preposition showing direction in a sentence, the lips just stop rounding.
265
So, I'm going to the store, because I'm going to the store.
266
Going to the store.
267
The lips stay totally relaxed.
268
And honestly, this explains why learners find modal verbs so incredibly difficult to hear in natural speech.
269
Oh, absolutely.
270
Modal verbs like can, could, should, and would.
271
Right, because modals are auxiliary verbs.
272
They are function words.
273
Let's look really closely at the word can.
274
A learner is taught that can sounds like a metal container, like a tin can.
275
But listen to this sentence, she can speak Spanish.
276
If I'm listening as a learner, I fully expect to hear she can speak Spanish.
277
But the native speaker says she ends speak Spanish.
278
The vowel just disappears into the schwa.
279
She can speak.
280
It happens in a fraction of a second.
281
And this weak form is actually incredibly important for comprehension because English speakers use stress to differentiate between positive and negative abilities.
282
Yes.
283
Walk us through that positive versus negative difference because this is a massive roadblock for learners.
284
Okay.
285
So when a native speaker says they cannot do something.
286
They place heavy stress on the negative word, can't.
287
And because it is stressed, the vowel stays strong, full, and clear.
288
I can't swim.
289
Right.
290
But when they can do something, the word can is unstressed wet cement.
291
It shrinks into a weak schwa.
292
I shan't swim.
293
So the difference isn't really about hearing the T at the end of can't.
294
It is almost entirely about the vowel quality.
295
Exactly.
296
Strong vowel means negative.
297
I can't do it.
298
We, schwa vowel means positive.
299
I can't do it.
300
So if a learner is listening desperately for that sound to understand if something is possible or impossible, they will often fail because in fast speech, that tone might be very quiet or even dropped.
301
Right.
302
You have to listen for the rhythm.
303
You have to listen for the schwa.
304
But there's a situation where function words suddenly refuse to shrink.
305
Yeah.
306
Where they refuse to be wet cement and suddenly act like heavy bricks.
307
Yes, because the rules of connected speech are ultimately governed by human intention.
308
We use stress to highlight the specific meaning we want to convey.
309
And usually, function words carry no special meaning.
310
But what happens if you need to clarify a misunderstanding?
311
Right.
312
Like, imagine we're looking at a map, and you ask me if we're flying to London or Paris.
313
And you want to make sure I know we are visiting both cities.
314
Yeah, so I'd say we're flying to London and Paris.
315
Exactly.
316
In that exact moment, the word and is no longer just grammatical structure.
317
It is literally the most vital piece of information in the sentence.
318
Your intention assigns heavy stress to that word.
319
And because it gets heavy stress, it cannot physically shrink.
320
The mouth spends the energy required to make the full dictionary vowel.
321
So you don't say London and Paris, you say London and Paris.
322
We see this with prepositions too.
323
If someone asks, are you looking at the box or in the box?
324
You would reply, I am looking in the box.
325
The word in takes full stress and full clarity.
326
So the dictionary pronunciation isn't wrong.
327
It is just reserved for specific moments of high contrast and emphasis.
328
Let's let the listener test their control over these weak forms.
329
The phrase is a piece of cake and a cup of tea.
330
So think about the heavy bricks in that phrase.
331
Piece, cake, cup, tea.
332
And think about the wet cement.
333
Uh, of, and uh, of.
334
I want you to consciously weaken those cement words.
335
Drop your jaw, relax your tongue, and use the schwa.
336
Take a second and try saying it naturally.
337
It should flow very smoothly with all the energy bouncing on the nouns.
338
A piece of cake and a cup of tea.
339
A piece of cake and a cup of tea.
340
So shrinking sounds into the schwa obviously saves an enormous amount of physical energy.
341
It does.
342
But the vocal tract is ruthlessly efficient.
343
Sometimes just weakening a sound is not enough.
344
Right.
345
Sometimes a specific physical transition between two sounds is simply too difficult, too clunky to perform at high speed.
346
And when a transition is too much work, the brain just cancels the command entirely.
347
The sound is written in the spelling, but it gets thrown in the trash before it ever brutches the air.
348
And this complete disappearance of a sound is called elysian.
349
It is the second major tool in our connected speech framework.
350
And it's important to note, elysian is not just accidental dropping of sounds.
351
it is highly systematic and it's predictable based on the physical shapes of consonants.
352
So which sounds get thrown in the trash most frequently?
353
Elysian overwhelmingly targets two specific consonants, the T sound and the D sound.
354
And to understand why, you have to feel how these sounds are produced, because both T and D are created in the exact same location in the mouth.
355
Right.
356
Take a second to say T and D.
357
T and D.
358
Notice where the tip of your tongue touches.
359
It hits the hard, bumpy ridge right behind your top front teeth.
360
That bumpy area is called the alveolar ridge.
361
The T and D are what we call alveolar stops.
362
So to make them, the tongue must raise, press against the ridge to completely stop the airflow, and then pull away to release the burst of air.
363
It is a very precise, distinct tapping motion.
364
And that tapping motion is perfectly easy to execute when it is surrounded by open vowels.
365
Sure, but English vocabulary is famous for grouping lots of consonants together without any vowels in between, which we call consonant clusters.
366
Words like strength.
367
That is a massive cluster of consonants.
368
It's a workout.
369
It really is.
370
So when a T or a D gets trapped in the middle of a consonant cluster, especially across the boundary between two words, the mouth faces a physical crisis.
371
The tongue simply does not have the milliseconds required to rise up, press the alveolar ridge, and release before it immediately has to move to a completely different position for the next consonant.
372
Let's analyze a slow motion collision to see this.
373
Let's take the phrase, last night.
374
Okay, last night.
375
Let's look at the physical requirements there.
376
The word last ends in an s sound, followed by a t sound.
377
And the word night begins with an n sound.
378
So if a learner tries to pronounce every letter, they are forcing their mouth to perform an s, a T and an N in rapid succession.
379
Last night, I could feel my tongue.
380
It makes the S, then it has to stop everything.
381
Press the ridge for the T, release it, and then press the ridge again for the N.
382
It feels like a physical roadblock.
383
It stops the rhythm dead.
384
It really does.
385
And the brain recognizes that roadblock.
386
It recognizes that making the T in that cluster requires way too much energy, and it disrupts the heavy, stress-timed beat of the sentence.
387
So the brain applies elision.
388
It simply deletes the T command.
389
The tongue stays exactly where it is.
390
The speaker just transitions directly from the S in last to the N in night.
391
The phrase becomes last night.
392
Yeah.
393
Try out last night.
394
Last night.
395
The T is a ghost.
396
It exists on paper, but in the acoustic reality of the air, it is completely absent.
397
What about the D sound?
398
The mechanics are identical.
399
Let's look at the phrase old man.
400
Okay.
401
Old ends in a cluster.
402
L then D.
403
And the next word starts with M.
404
The D is trapped between the L and the M.
405
Right.
406
Old man.
407
To say the D, you have to stop the air completely before starting the humming M.
408
It's just too much work.
409
So the D undergoes elision.
410
The tongue flows straight from the L into the lips, closing for the M, becomes old man.
411
He is an old man.
412
Old man.
413
And this happens constantly with regular verbs in the past tense, right?
414
Oh, all the time.
415
Because regular verbs form the past tense by adding an ed to the spelling, they create massive consonant clusters at the ends of words.
416
Think about the word finished.
417
The base word ends in an she sound.
418
We add the grammar for past tense, which sounds like a t.
419
Finish t.
420
Now put that word next to a noun, he finished the project.
421
Right.
422
Finished ends in sh and t.
423
The starts with a d sound.
424
Shhh, TTA.
425
That is an agonizingly complex sequence for the tongue.
426
So the T gets deleted, it sounds like he finished the project.
427
He finished the project.
428
Now, as a learner, I have to say this triggers massive anxiety.
429
I'm sure it does.
430
If I am speaking and I delete the past tense sound at the end of the word, how will anyone listening to me know that I'm talking about the past?
431
I spent years learning to put that ED there, and now my grammar sounds completely wrong.
432
I completely understand that fear.
433
But this fear comes from treating spoken language like a written exam.
434
Right.
435
In a written sentence, the letter D is visually required to convey the tense.
436
But human communication does not occur in a vacuum.
437
It relies heavily on context and predictive processing in the brain.
438
Predictive processing.
439
So our brains are literally guessing what is coming next.
440
Constantly.
441
If you say, yesterday I walked to the store, yes, you dropped the past tense sound on walk, but the listener's brain already heard the word yesterday.
442
The temporal context of the sentence is already firmly established.
443
Exactly.
444
The brain knows it is the past, so it doesn't need to hear the phonetic T on the verb to understand the grammar.
445
And even without a specific time marker like yesterday, the overall context of the conversation usually provides the clues.
446
Right.
447
The brain of a native listener acts like an automatic grammar correction software.
448
It expects the past tense, so it subjectively hears the past tense, even when the physical sound waves do not actually contain the T or D.
449
That's wild.
450
Context overrules perfect phonetic clarity.
451
But lesion doesn't just happen to T and D, does it?
452
No, it frequently attacks the H sound as well, especially in pronouns.
453
Because the H sound is simply a voiceless breath of air.
454
Right.
455
When pronouns like he, him, his, or her fall into weak, unstressed positions in a sentence.
456
The mouth just doesn't bother pushing that extra breath of air.
457
So if I see the sentence, I told him to give it to her.
458
A learner might carefully pronounce, I told him to give it to her, but a native speaker lets those pronouns shrink and the H sound drops completely.
459
I told him to give it to her.
460
I told him to give it to her.
461
The consonants of the previous words just crash right into the vowels of the pronouns, told him.
462
And this specific elision is actually a really common hurdle for learners whose native language does not naturally contain the H sound.
463
Like French speakers, right.
464
They provide a classic example.
465
Yes.
466
Because the H is a foreign concept in French, they have to exert conscious mental effort to produce it when they learn English.
467
So they train themselves to push that breath of air for house and happy.
468
But because they apply so much conscious focus to it, they often overpronounce it.
469
They will painstakingly pronounce the H in unstressed pronouns in the middle of fast sentences,
470
places where native English speakers would instinctively apply elision and drop it entirely.
471
Ah.
472
So by trying to be too grammatically perfect, they actually end up sounding slightly less fluent.
473
Because it interrupts the natural glide of connected speech.
474
Exactly.
475
We also see illusions swallow entire syllables, by the way.
476
This happens most often with longer words that contain multiple unstressed syllables.
477
The weak, schwa-filled syllables just disappear to save time.
478
Like the word comfortable.
479
It's the perfect example.
480
It has four syllables on paper.
481
Com-fri-ta-ble.
482
But nobody says that.
483
The unstressed middle syllables just collapse.
484
We drop the or completely.
485
The word shrinks to three syllables.
486
Comfortable.
487
Or even too comfortable.
488
What about temperature?
489
Temperature.
490
That middle-er vanishes.
491
Temperature.
492
Three syllables.
493
The same with interesting.
494
It becomes interest- Interesting.
495
Let's have you, the listener, feel the physical relief of elision.
496
The practice sentence is, he kept sending texts to his friends.
497
Notice the consonant clusters there.
498
The word kept ends in P and T.
499
The word texts is a nightmare.
500
K-S-T-S.
501
And notice the weak pronoun his.
502
Your goal is to delete the T in kept.
503
Delete the P in texts and delete the H in his.
504
Allow your mouth to take the lazy, efficient path.
505
Take a moment and try saying it naturally, dropping the roadblocks.
506
You should feel your tongue gliding right over those gaps.
507
He kept sending texts to his friends.
508
He kept sending texts to his friends.
509
So we use weak forms to shrink sounds, and we use elision to delete sounds.
510
But the physical engine of speech occasionally encounters scenarios where a sound cannot be shrunk and it cannot be deleted.
511
Right.
512
Sometimes, two heavy, distinct sounds are forced to sit right next to each other across a word boundary.
513
They crash into each other, and neither one is willing to disappear into the trash.
514
In these situations, the mouth has to negotiate a compromise.
515
The sounds alter their physical shapes, borrowing characteristics from each other so they can connect smoothly.
516
And this shape-shifting process is our third major phonetic tool, assimilation.
517
Assimilation.
518
I like to think of it like mixing wet paint on a canvas.
519
Oh, that's good.
520
How so?
521
Well, if you paint a thick stripe of pure blue, and right next to it you paint the thick stripe of pure yellow, the colors are distinct.
522
But at the exact boundary line where the wet blue touches the wet yellow, the paints bleed into each other.
523
The edge turns green.
524
That is a perfect visualization of what happens in the mouth.
525
The sounds are the wet paint.
526
When they touch at high speeds, they bleed together to create a brand new sound at the boundary.
527
And there are two primary types of this phonetic paint mixing rate.
528
The first is anticipatory assimilation.
529
Anticipatory, meaning the mouth is anticipating or preparing for the future.
530
Precisely.
531
Anticiptory assimilation occurs when the mouth physically prepares for the second sound a fraction of a second too early.
532
Right, so the first sound gets warped because the muscles have already moved into the position required for the next word.
533
Let's track the exact muscular movements for a common example.
534
The word handbag, a bag you hold in your hand.
535
Okay, let's break down the boundary.
536
Hand ends in a D, but we already know through elision that the D in that cluster gets dropped.
537
Right, so acoustically the first word ends in the N sound.
538
And the second word, bag, begins with the B sound.
539
So the transition we need to make is from N to B.
540
Feel where those sounds live in your mouth.
541
Make an N sound.
542
The tip of your tongue touches the alveolar ridge behind your teeth.
543
The air flows through your nose.
544
Now make a B sound.
545
Your tongue does nothing, but your two lips press firmly together.
546
Moving from the tongue on the roof of the mouth all the way forward to the lips pressing together
547
takes a significant fraction of a second.
548
And at a normal conversational pace, the brain essentially says that travel distance is too far.
549
We are about to need closed lips for the B in bag.
550
It just closed the lips early while we are still trying to make the N.
551
But what happens if you try to make an N sound while your lips are clamped shut?
552
Let's try it.
553
You physically cannot do it.
554
With closed lips, the nasal N sound naturally morphs into an M sound.
555
The blue and yellow paint turn green.
556
So handbag becomes ham bag.
557
Ham bag.
558
I am fully pronouncing an M, even though there is no letter M anywhere in the word.
559
The physical reality overrides the spelling, and we see this exact same anticipatory shift with the phrase 10 bikes.
560
10 ends with the alveolar N.
561
Bikes starts with the lip closing B.
562
So the lips close early.
563
10 bikes becomes 10 bikes.
564
He owns 10 bikes.
565
What if the second sound isn't at the lips, but all the way at the back of the throat, like the K or G sound?
566
The anticipation works in reverse.
567
Let's look at the phrase 10 coins.
568
The N requires the tip of the tongue at the front teeth.
569
But the K in coins requires the very back of the tongue to raise up against the soft palate near the throat.
570
That is a massive leap from the front teeth to the back of the throat.
571
So the brain applies anticipatory simulation.
572
It tells the tongue to move to the back of the throat early.
573
It says, don't bother tapping the front teeth for the N.
574
Just stay at the back.
575
Right.
576
And if I make a nasal sound with the back of my tongue raised, I produce the N sound, like at the end of the word sing.
577
Exactly.
578
The alveolar N morphs into the velar N.
579
10 coins bleeds into 10 coins.
580
10 coins.
581
10 coins.
582
The shape-shifting completely erases the original sound.
583
And this doesn't just happen with location.
584
It can happen with the voicing of sounds too.
585
Anticipatory assimilation definitely happens a lot.
586
But the second major type is perhaps even more dramatic.
587
It is called coalescent assimilation.
588
Coalescent.
589
To coalesce means to crash together and fuse into a single, unified entity.
590
Right.
591
In coalescent assimilation, the two neighboring sounds don't just influence each other, they completely destroy each other and merge to create a brand new consonant that wasn't there before.
592
This fusion happens primarily when words ending in T or D crash into words starting with the J sound.
593
The J sound in phonetics being the Y sound, like in the words U or yes.
594
Exactly.
595
When the alveolar stop of a T meets the palatal glide of a Y at high speeds,
596
The immense physical pressure causes the two sounds to burst into a completely new shape.
597
The chi sound.
598
The chi sound.
599
Like in cheese or a church.
600
Give us the most common example of this crash.
601
Oh, it happens millions of times a day in the phrase, nice to meet you.
602
Uh, meet ends in the hard T.
603
U begins with the soft Y.
604
So the tongue hits the alveolar ridge for the T, but immediately tries to slide backward for the Y.
605
The resulting friction creates the chi.
606
Meet plus you coalesces into meet you.
607
Nice to meet you.
608
Nice to meet you.
609
It is basically impossible to say it naturally without creating that chi sound.
610
So what happens when the d sound crashes into the y sound?
611
Well, if t plus y makes chi, then the voice d plus y makes a voice j sound, like in the word jump or judge.
612
So if I look at the phrase would you, would ends a d, you starts with y.
613
They fuse together.
614
Would you becomes would you.
615
Would you like some coffee?
616
Would you, or the phrase did you, did plus you becomes did you, did you go to the store.
617
And a learner reading the subtitle, did you go to the store, will listen desperately for the distinct D and the distinct Y.
618
But the acoustic reality reaching their ear is a loud, clear J sound that exists nowhere on the page.
619
This is exactly why dictation exercises are so vital.
620
You have to train your brain to map the sound of did you back to the spelling of did you.
621
Let's have the listener practice this shape-shifting.
622
The sentence is, don't you want to buy that handbag?
623
Okay, so we have a coalescent assimilation in don't you.
624
The T and the Y need to crash into a T.
625
Don't you.
626
And we have an anticipatory assimilation in handbag.
627
The N needs to turn into an M before the B.
628
Handbag.
629
Try saying it naturally, allowing the sounds to bleed into each other.
630
Give it a try.
631
Don't you want to buy that handbag?
632
Don't you want to buy that handbag?
633
Notice how we also applied weak forms to the word to, turning it into a schwa inside wanna.
634
Right.
635
Because connected speech isn't just one rule at a time.
636
It is all of these tools operating simultaneously, overlapping and interacting every single second you speak.
637
We shrink sounds with weak forms.
638
We delete sounds with delision.
639
We morph sounds with assimilation.
640
All of these tools serve to remove physical barriers.
641
They basically knock down the walls between words to create a smooth flow.
642
But you know, knocking down walls isn't always enough to create a perfect river of sound.
643
Sometimes a gap appears that threatens to break the rhythm entirely.
644
And when faced with this specific type of gap, the mouth doesn't remove effort.
645
It actually spends extra physical energy to build a bridge.
646
Which brings us to the fourth and final phonetic mechanism, linking and intrusion.
647
So what kind of gap forces the mouth to build a phonetic bridge?
648
Well, the vocal tract absolutely despises coming to a sudden, complete halt.
649
It hates starting and stopping the vibration of the vocal cords.
650
And the most disruptive stop occurs when a word ends in a full vowel sound.
651
And the very next word also begins with a full vowel sound.
652
Linguists call this collision of vowels a hiatus.
653
To pronounce two full vowels back-to-back while maintaining clear separation, you have to perform what is called a glottal stop.
654
Right.
655
You have to literally close your throat, stop the airflow, and then forcefully push air out again to start the second vowel.
656
Let's demonstrate a hiatus.
657
The phrase go away, the word go ends with the rounded O vowel.
658
The word away begins with the open O vowel.
659
If you keep them completely separate, you have to choke off the air in the middle.
660
Go.
661
Away.
662
Go.
663
Away.
664
It sounds completely robotic.
665
It sounds like a computer-generated voice from the 1990s.
666
It does.
667
The glottal stop completely shatters the heavy-flowing rhythm of connected speech.
668
So to prevent that harsh robotic stop, native speakers intuitively insert a tiny, invisible consonant between the two vowels.
669
This consonant acts as a ramp, allowing the vocal cords to keep vibrating continuously as they slide from the first word to the second.
670
And these invisible bridges are almost always formed by glides, sound that are halfway between vowels and consonants.
671
Specifically, the W sound and the Y sound.
672
But how does the mouth know which bridge to build?
673
Do we just guess?
674
It isn't a guess at all.
675
The bridge is entirely determined by the physical shape of the lips at the end of the first vowel.
676
Let's look at the W bridge first.
677
The W bridge naturally emerges when the first word ends in a vowel that requires the lips to be rounded.
678
Vowels like the oo in to or the o in go.
679
When I say go, my lips push forward into a tight circle.
680
Exactly.
681
When your lips are already in that tight circle and you try to move to the next vowel without stopping your voice, your lips have to unround.
682
And the physical act of unrounding the lips while making sound automatically produces a w noise.
683
It's a biomechanical byproduct.
684
So in our phrase go away as my lips open from the O to the AH, the W bridge appears naturally.
685
Go away.
686
Please go away.
687
You hear it perfectly.
688
It connects the vowels seamlessly.
689
Consider the phrase who is.
690
The word who ends in the deeply rounded OO vowel.
691
The next word starts with I.
692
The lips open.
693
The bridge forms who is.
694
Who's coming to dinner?
695
The continuous vibration never stops.
696
Now let's contrast that with the Y bridge.
697
This bridge appears when the The first word ends in a vowel that requires the lips to stretch wide like a smile.
698
Vowels like the E in C or the A in they or the I in my.
699
When I say C, the corners of my mouth pull back tightly.
700
Right.
701
When you hold that strict position and immediately try to transition to another vowel while keeping your vocal cords vibrating,
702
the movement of your tongue gliding downward naturally generates a Y sound.
703
Let's test it with the phrase, they are.
704
They ends with the wide smiling I vowel.
705
R starts with the open.
706
Slide from the smile to the open mouth.
707
They are.
708
They're here.
709
It works perfectly.
710
What about the phrase I am.
711
I ends wide.
712
It generates the bridge.
713
I am ready.
714
The W and I bridges are totally invisible in the spelling, but they are absolutely essential to the acoustic flow of fluent English.
715
But there's a third bridge, one that causes intense confusion for learners and honestly fierce debate among native speakers, the R bridge.
716
The linking and intrusive R is one of the most fascinating phenomena in English phonetics.
717
To understand it, we first have to take a brave detour into the historical evolution of English accents.
718
We need to distinguish between rhodic and non-rhodic accents.
719
Rhodicity refers to whether or not a speaker actually pronounces the letter R when they see it in a word.
720
Historically, all English speakers pronounced the R everywhere it appeared.
721
But around the 18th century in southern England, speakers began to drop the R sound if it occurred after a vowel, particularly at the end of a word or syllable.
722
And this dropping of the R became fashionable and spread across much of the British Empire.
723
This created the non-erotic accents we hear today in most of England, Australia, New Zealand, and South Africa.
724
A non-erotic speaker looks at the word car, spelled C-A-R, and pronounces it ca.
725
The R is completely silent.
726
However, North America was largely settled before this trend took over, so most American and Canadian accents remain rhodic.
727
A rhodic speaker sees the word car and pronounces the hard R at the end, car.
728
So how does rhodicity affect the bridges we build between words?
729
It creates a unique situation for non-rhodic speakers.
730
Let's say a speaker from London is saying a sentence where the word car is the last word.
731
That is my car.
732
The R is silent.
733
But what happens if the next word in the sentence begins with a vowel, like the phrase, the car is blue?
734
The non-erotic speaker faces a hiatus.
735
The ah in car crashes into the i in is.
736
To avoid the robotic glottal stop, their brain searches for a bridge.
737
And because the letter r is historically present in the spelling of car, the brain basically resurrects it.
738
The dead r comes back to life purely to act as a bridge.
739
So they say the car is blue.
740
Exactly.
741
Exactly.
742
Even though they normally drop the R, they insert it to link the vowels.
743
This is known as the linking R.
744
It makes structural sense because the letter is sitting right there on the page waiting to be used.
745
But that brings us to the more controversial bridge, the intrusive R.
746
The intrusive R happens when a non-rotic speaker encounters a hiatus between two vowels and their brain inserts an R-bridge,
747
even though there's absolutely no letter R anywhere in the spelling of either word.
748
Wait, they just invent a letter out of thin air?
749
They do.
750
It is an unconscious physical reflex.
751
If a word ends in a low open vowel sound like the swa or ah or ah, and the next word starts with a vowel,
752
the tongue naturally glides up to form an R to cross the gap smoothly.
753
Give us a real world example of this phantom R.
754
Let's look at a famous television genre, law and order.
755
The word law is spelled L-A-W, no R.
756
The word and is spelled A-N-D, no R.
757
But if I listen to a British newsreader, I often hear a very distinct R between those words.
758
Law and order.
759
Law and order.
760
The tongue adds the R because moving from the A in law directly to the A in and feels too awkward.
761
Another common example is the phrase, I saw it.
762
S-A-W.
763
No R.
764
I saw art.
765
I saw art happen.
766
It also happens inside single words when suffixes are added.
767
Take the word drawing.
768
D-R-A-W-I-N-G.
769
The aw vowel meets the I vowel.
770
Many non-erotic speakers will say drawing.
771
I have to point out that if you go onto any language forum online, you will find angry native speakers complaining endlessly about this.
772
Oh, absolutely.
773
They claim that saying law and order or drawing is uneducated,
774
sloppy, and fundamentally bad English because it violates the spelling.
775
So should learners who want a British accent actually try to copy this intrusive R or will they be judged for it?
776
This touches on the endless war between prescriptive grammar and descriptive linguistics.
777
Prescriptivists believe that written spelling is the absolute law and speech must obey the ink on the page.
778
In the mid-20th century, listeners actually wrote furious letters to the BBC demanding that broadcasters be fired for using the intrusive R.
779
They thought it was a corruption of the language.
780
But descriptive linguistics looks at how the human vocal tract actually operates in reality.
781
Phonetics does not care about spelling.
782
The truth is, the intrusive R is a completely natural, systemic, and deeply embedded phonetic strategy in non-rotic accents.
783
Right.
784
It is an unconscious reflex used by highly educated people, prime ministers, and actors to maintain the rhythmic flow of the sentence.
785
So it isn't an error.
786
It is a biological feature of the accent.
787
If a learner is aiming to master a non-rhodic accent like British or Australian English,
788
incorporating the intrusive R will actually make them sound significantly more authentic and fluent.
789
It demonstrates that they have internalized the physical rhythm of the language rather than just reading words off a page.
790
Let's let the listener test their ear for these invisible bridges.
791
I'm going to give you a short phrase.
792
I want you to listen closely to the transition between the two words.
793
See if you can identify which bridge your mouth naturally builds.
794
The W bridge, the Y bridge, or the R bridge.
795
Say the phrase aloud a few times at a normal fast speed.
796
The phrase is two apples, two apples.
797
Take a second.
798
What happens is your lips open from the ooh into two W apples.
799
The W bridge naturally appears.
800
What about this phrase?
801
She arrives, she arrives.
802
Your lips are stretched wide for she.
803
As they relax, the Y sound emerges.
804
She arrives, she arrives.
805
These bridges are the final polish.
806
They're the mortar that seals the gaps between the bricks, ensuring that the heavy bouncing rhythm of English never has to stop for an awkward robotic pause.
807
So let's summarize the immense amount of mechanical engineering we have just covered.
808
The mystery of why spoken English sounds like a chaotic blur is solved by understanding the physical engine of the vocal tract.
809
Right.
810
The mouth seeks efficiency.
811
And to maintain the stress-timed rhythm, it relies on four primary tools.
812
First, we have weak forms.
813
The structural cement words shrink down, collapsing their vowels into the lazy, effortless schwa sound.
814
Second, we have Elysian.
815
When consonant clusters create physical roadblocks, especially with T and D, the brain simply deletes the sound entirely, throwing it in the phonetic trash.
816
Third, we have Assimilation.
817
When sounds crash into each other and cannot be deleted, they shape-shift.
818
They borrow lip or tongue positions from their neighbors, or they fuse together entirely, like T and Y making the Qi sound in Michu.
819
And fourth, we have linking and incursion.
820
When two vowels threaten to cause a robotic glottal stop, the mouth builds invisible consonant bridges, the W, the Y, and the R,
821
to keep the vocal cords vibrating smoothly.
822
Now, if you are a learner at the B2 or C1 level, you already possess a massive vocabulary and a strong grasp of grammar.
823
Hearing all these hidden physical rules might feel overwhelming.
824
It's a lot to take in.
825
You might be wondering how you can possibly calculate all of these assimilations
826
and illusions in real time while trying to hold a conversation.
827
The thought of actively trying to plan a coalescent assimilation while speaking is paralyzing.
828
It really is.
829
So the most practical advice you can take away from this is
830
do not try to manually force these rules into your speaking right now.
831
Right.
832
If you try to consciously engineer your pronunciation word by word, your fluency will just freeze.
833
The primary goal of understanding connected speech is not to change your mouth immediately.
834
It is to change your ears.
835
Focus on listening comprehension first.
836
Exactly.
837
The human brain is a pattern recognition machine.
838
For years, your brain has been listening to native speakers
839
and failing to match the chaotic sounds to the perfect dictionary spelling you memorized.
840
But now that you consciously know these rules exist, your brain has a new blueprint.
841
You will suddenly start noticing the dropped T's in movies.
842
You will hear the W bridges in conversations.
843
Once your ear maps the patterns, your own mouth will slowly, unconsciously begin to adopt them.
844
The speed and smoothness will come naturally over time without forced effort.
845
It changes your relationship with the language completely.
846
You stop feeling like you are failing to hear the words
847
and you start realizing you are successfully hearing the physics of the human body.
848
That is a profound perspective.
849
We spend so much time treating language purely as an abstract intellectual code, a set of invisible grammar rules living in a textbook or in the mind.
850
But the reality is that the literal physical shape of our human bodies, the exact length of the vocal tract, the dense weight of the tongue muscle,
851
the flexibility of our lips directly dictates the rules of pronunciation.
852
The biology governs the language.
853
Language is a physical act.
854
When you listen to connected speech, you aren't just hearing words.
855
You are hearing the biomechanics of a living engine trying to be as efficient as possible.
856
I want to issue a daily challenge for you, the listener.
857
If you truly want to stop being frustrated by fast English, you need to recalibrate your ears to this new blueprint.
858
Your mission is to practice active, forensic listening for just 10 minutes every day.
859
Take a very short audio clip of a native speaker talking naturally.
860
A 10-second clip from a YouTube vlog, an interview, or a movie, but do not look at the subtitles.
861
Listen to that 10-second clip over and over again.
862
Try to write down exactly what you hear, not what you think the grammar should be, but the actual sounds hitting your ear.
863
Where did they drop a D?
864
Where did U turn into 2?
865
Where did the vowels collapse into a schwa?
866
Once you have your phonetic map, compare it to the actual transcript.
867
Find the gaps.
868
10 minutes a day of this intense focused matching will completely rewire how your brain processes spoken English.
869
You will stop expecting perfect separated train cars and you'll learn to read the blur of the fast train rushing past.
870
Keep listening, keep analyzing the physical engine of speech and trust that your ears will adapt.
871
Practice every day.
872
You've got this.

Objetivos de fala deste vídeo

Este vídeo visa ajudar você a superar um marco crucial na fluência: entender a fala rápida de falantes nativos. Muitos aprendizes dominam a gramática e o vocabulário, mas se perdem quando as palavras "se chocam" no discurso natural. Aqui, você descobre por que a língua falada parece uma "mancha caótica de ruído" e como decifrá-la, preparando-se para situações reais, como pedir café em Londres ou Nova York.

Banco de frases

  • "The panic sets in" (O pânico começa)
  • "A single chaotic blur of noise" (Uma única mancha caótica de ruído)
  • "Warp, shrink, and sometimes vanish" (Se deformam, encolhem e às vezes desaparecem)
  • "Rushing past you at 100 miles per hour" (Passando rápido a 160 km/h)
  • "Fundamentally changes how you perceive them" (Muda fundamentalmente como você os percebe)

Corrija seus pontos fracos

O maior desafio aqui é a percepção da fala contínua. A escrita separa as palavras claramente, mas a fala natural elimina espaços, causando assimilação, elisão e intrusão. Para melhorar, pratique o shadowing (imitar a fala em tempo real) com um shadowing site: ouça frases rápidas, pause e repita, tentando reproduzir a fluidez e as mudanças de som. Isso ajuda a treinar o ouvido e a boca a adaptar-se à ritmada da língua. Além disso, use exercícios de shadow speech para notar como "I will see you next week" se transforma em "I'll see you next week" – pequenas mudanças que fazem toda a diferença. Com prática de conversação em inglês e foco na pronúncia, você passará a ver a "trem" da fala rápida não como uma mancha, mas como uma sequência de "vagões" claros. Lembre: a chave é treinar o ouvinte, não apenas o leitor!

O que é a Técnica de Shadowing?

Shadowing é uma técnica de aprendizado de idiomas com base científica, originalmente desenvolvida para o treinamento de intérpretes profissionais. O método é simples, mas poderoso: você ouve áudio em inglês nativo e repete imediatamente em voz alta — como uma sombra seguindo o falante com 1-2 segundos de atraso. Pesquisas mostram melhora significativa na precisão da pronúncia, entonação, ritmo, sons conectados, compreensão auditiva e fluência na fala.