Shadowing Practice: 99 - Learn English Speaking with Video

Les maken...
1
Every statistics course draws the same little square.
2
The truth is one of two things, and your decision is one of two things, so there are four boxes, two of which are mistakes,
3
and the two mistakes get names, type 1 and type 2, printed side by side as though they were two of a kind.
4
They are not, and the difference between them is the whole subject.
5
The box where nothing was going on and you called it real, the false alarm, carries a number you picked yourself before collecting a single observation,
6
and that number is 0.05.
7
The box below it, where something real was there, and you walked straight past it, has no number attached to it anywhere,
8
not on your screen, and not in the paper, and usually not in your head either.
9
So let's go and find the second number, because once you can see it, the sentence we have all written at some point,
10
no significant difference, stops meaning what you thought it meant.
11
Let's make this concrete with about the smallest real decision I can think of.
12
You change the color of a checkout button, you send some traffic to each version,
13
and you want to know whether the new one genuinely converts better or whether you are looking at noise wearing a costume.
14
The null hypothesis is the boring one, that the change did nothing, and it gets to be the starting point,
15
not out of fairness, but because it's the only hypothesis specific enough to make predictions.
16
If the button does nothing, the difference you measure is pure sampling wobble, and dividing that difference by its own standard error,
17
which is simply how much a number like this bounces around from one repeat of the experiment to the next, gives you a statistic whose distribution you already know.
18
It's a bell centered on zero, whose spread is exactly 1.
19
So from here on, one step along that axis means one standard error.
20
Now draw a line.
21
Pass the line you call the result real, short of it you don't.
22
And since you only care about the new button being better, one line on the right is all you need.
23
Put it where 5% of the bell is left stranded beyond it, which lands at 1.645 and shade that sliver in,
24
because area under a curve like this is probability, so the sliver is telling you how often you would cross the line in a world where the button did nothing.
25
That area is alpha, the type 1 error rate.
26
Alpha isn't measured, it's declared, and notice what that buys you, because you just computed it with no data, no knowledge of the true effect,
27
and without running the test at all.
28
So, try the same move on the other error.
29
You want the probability of missing a real effect,
30
which means the probability that your statistic lands short of the line in a world where the button does work,
31
and for that you need to know where the statistic sits in that world.
32
But the alternative hypothesis, as it is almost always written, only says the effect isn't And not zero is not a location.
33
No curve, so no area, so nothing to compute.
34
The only way forward is to stop being vague and name a specific effect out loud.
35
Say the lift is big enough that your statistic now averages three standard errors out instead of zero.
36
Then, there is a second bell, identical to the first, in every respect except where it's centered.
37
And now, the question has an answer.
38
Because beta is the piece of that second bell lying on the wrong side of your line,
39
the region where a real effect turns up looking ordinary enough to wave through.
40
With the line still at 1.645 and the alternative 3 standard errors out,
41
that shaded piece works out to about 9% of the second bell.
42
But 9% came with strings attached, because it was 9% for that effect.
43
Shrink the true lift and the second bell slides back toward the first.
44
The two overlap more and beta climbs.
45
So there is no such thing as the type 2 error rate of a test, only beta at an effect size somebody was willing to commit to,
46
which is exactly why alpha is printed on every table you have ever read and beta on almost none of them.
47
Alpha needs a null, and everybody has one of those lying around, while beta needs a guess about the truth, and nobody wants to write theirs down.
48
Now that both errors have pictures, watch what happens when you get nervous about false alarms, and pull the line further to the right.
49
The tail beyond it shrinks so alpha drops exactly as you wanted, but that line swept through the second bell as well,
50
and everything it passed over went from caught to missed, and the strip it swept is thin under the null and fat under the alternative,
51
since the alternative keeps most of its mass out in exactly that region, so beta climbs by a great deal more than alpha falls.
52
The numbers make the lopsidedness plain, with the alternative still 3 standard errors out.
53
At an alpha of 5%, you miss a real effect about 9% of the time.
54
Tighten alpha to 1%, which most people would call ordinary caution rather than paranoia, and you saved 4 points of false alarm,
55
and paid 16 points of miss, because the miss rate jumps to 25%.
56
Tighten to 1 in 1000, the kind of threshold people reach for when they want to sound rigorous,
57
and you now miss a real effect 54% of the time,
58
which makes a test that's strict worse than a coin flip at noticing the very thing it was built to notice.
59
And nothing about the world changed while we did that,
60
because nobody collected another observation and the button is exactly as good as it always was.
61
So the line is the only thing here you actually control.
62
And the cleanest way to see what you get for putting it somewhere is to flip beta over.
63
1 minus beta is the probability of catching a real effect when there is one, and that is the power of the test.
64
Because beta depended on which alternative you named, power inherits the same dependence.
65
So power isn't a number that belongs to your test the way alpha does.
66
It's a whole curve, with one value for every size of effect the world might be hiding from you.
67
So sweep the true effect from nothing up to enormous and plot what the test does at each point.
68
Choosing which point on that climb you need to be standing at is most of what designing a study actually is.
69
Far out on the right, the second bell has slid clean past the line, and power sits at essentially 100%,
70
because an effect that large is impossible to miss.
71
Walk back towards zero, and the curve sags, and the left-hand end is where it gets interesting.
72
Because when the true effect is exactly zero, the power does not drop to zero, it settles at 5%.
73
It settles at alpha.
74
That isn't a coincidence to memorize, is the definition unfolding, because when the alternative bell sits exactly on top of the null bell,
75
the chance of landing past the line is simply the tail area, and the tail area is what alpha was in the first place.
76
Which brings us to the sentence this whole video is really about.
77
Two teams run the same test on the same button at the same 5%, and the button really does help,
78
by a small amount, the sort of lift that is worth having, and that you would never spot by eye.
79
The first team routes a modest slice of traffic for a few days, and in their sample that true effect is worth one standard error.
80
The second team runs it across the whole site for a month, which is 16 times as much traffic,
81
and since the standard error falls like one over the root of the sample size, 16 times the data cuts that error to a quarter of what it was.
82
So the very same real effect is now worth 4 standard errors.
83
Both teams get a statistic that falls short of the line, and both write down the same sentence,
84
no significant difference, and since we stipulated that the button really works, both of them are wrong.
85
Only one of them was entitled to be.
86
For the first team, the two belts overlapped so heavily that their chance of catching that lift was 26%,
87
well under the 80% usually treated as the minimum worth running.
88
So being wrong was the likely outcome, expected about 3 times in 4,
89
and their sentence carries almost no information about the button at all.
90
For the second team, the chance of catching it was 99%.
91
So being wrong meant drawing a card that comes up about once in a hundred,
92
and their identical sentence is worth something precisely because it was so unlikely to get written.
93
Same words, same alpha, same underlying truth, and one of those sentences is nearly empty while the other is close to decisive.
94
A negative result never says there is no effect.
95
It says that if there is one, it was too small to be seen with the sample you had.
96
And how small that is, lives entirely in the number neither report printed.
97
Everything so far has been a trade, so it's fair to ask whether you can ever have both.
98
You can, but start with what doesn't work, because collecting more data does not lower alpha.
99
Alpha is wherever you put a line, so leave the line alone and quadruple your sample,
100
And your false alarm rate is still exactly 5% since the null bell is standard by construction
101
and doesn't care how much data went into building it.
102
What the extra data moves is the other bell.
103
We just watched the root and rule work against the first team and pointed the other way it works for you.
104
Because four times, the data doubles the distance between the two curves.
105
And distance is exactly what you were short of.
106
Take that alternative sitting 3 standard errors out, where the line at 5% left you missing 9% of real effects,
107
then quadruple the sample so the separation becomes 6, and put the line at 3, which is exactly halfway between the two bells.
108
Halfway makes the two tails mirror images of each other, so alpha and beta come out identical.
109
both of them about 1 in 750.
110
Neither error was traded away for the other.
111
Both of them came down by more than a factor of 30, and the invoice for that came to 4 times the data.
112
The only other way to widen the gap is to make the effect itself bigger.
113
And that is usually a property of the world, rather than a knob on your desk.
114
So the line is a trade and more data is an escape and both of those quietly assume you ran one test.
115
Run several and alpha starts leaking in a way the single picture hides.
116
Because alpha is a rate per test and rates accumulate.
117
Run one test at 5% and you have a 1 in 20 chance of a false alarm, which most people accept without much thought.
118
Run 20 independent tests over the same data set, 20 button variants, or 20 outcomes you happen to record, and the chance
119
that at least one of them lights up in a world where nothing at all is happening is not 5%, it's 64%.
120
A room containing nothing but noise hands, you a significant result more often than not,
121
and every one of those results carries a p-value under 0.05, exactly as alpha promised.
122
The standard repair is to divide the threshold by the number of tests you ran.
123
So 20 tests each get an alpha of a quarter of a percent, and the chance of any false alarm anywhere comes back down to roughly 5%.
124
But look at what you just did to every one of those tests, because you moved each line further out.
125
And by now, we know exactly what that costs.
126
Against that same alternative, three standard errors out, power drops from 91% to 58%,
127
so a study that was properly powered for one comparison becomes underpowered for 20%, without anybody touching the data.
128
Every repair in this video has been somebody deciding which of the two errors to spend,
129
which is why 0.05 deserves rather more suspicion than it normally gets.
130
It isn't derived from anything, it's a convention that hardened into a habit, and adopting it without a thought silently announces
131
that a false alarm is the mistake you would rather avoid in every situation for the rest of your career.
132
Sometimes that is exactly right.
133
If you are deciding whether a new drug reaches the market, a false alarm puts something useless on the shelf,
134
and every study built on top of it inherits the mistake, so holding alpha very tight is defensible,
135
provided you say out loud what you are buying it with, which is a higher chance that a compound that genuinely works dies in a trial too small to see it.
136
And sometimes the posture is backwards.
137
If you are watching for a rare structural failure, a false alarm costs you an inspection,
138
and a wasted afternoon, while a miss costs you the failure itself.
139
So the sensible design drives beta as low as it will go
140
and simply pays for the drizzle of false alarms that comes with it.
141
And sometimes the answer is to refuse the choice, which is what a two-stage design is for,
142
where a cheap screen runs at a deliberately loose alpha so its beta is tiny, and almost nothing real slips past,
143
and everything it flags goes on to an expensive test whose alpha is orders of magnitude tighter.
144
Neither stage escapes the trade.
145
What the pair of them does is spend the loose alpha, where testing is cheap, and the tight alpha,
146
where it is expensive, so that the errors surviving to the end are rare on both sides.
147
So those two boxes were never two of a kind.
148
One is a promise you make before you look, the other is a consequence you inherit and can only pin down by naming what you were hoping to see,
149
and the only question a test really puts to you is which of the two you would rather be wrong about.
150
If you found this helpful, hit that like button, subscribe for more, and drop a comment if you have any questions.
151
See you in the next one.
152
Bye-bye!

Over deze les

Wat is de Shadowing-techniek?

Shadowing is een wetenschappelijk onderbouwde taalleermethode die oorspronkelijk is ontwikkeld voor professionele tolkentraining en gepopulariseerd door polyglot Dr. Alexander Arguelles. De methode is eenvoudig maar krachtig: je luistert naar native Engelse audio en herhaalt het onmiddellijk hardop — als een schaduw die de spreker volgt met slechts 1–2 seconden vertraging. In tegenstelling tot passief luisteren of grammaticadrills, dwingt shadowing je hersenen en mondspieren om echte spraakpatronen tegelijkertijd te verwerken en te reproduceren. Onderzoek toont aan dat het de uitspraaknauwkeurigheid, intonatie, ritme, verbonden spraak, luisterbegrip en spreekvaardigheid aanzienlijk verbetert — waardoor het een van de meest effectieve methoden is voor IELTS Speaking-voorbereiding en echte Engelse communicatie.

Shadowing-techniek: lees de volledige stap-voor-stap-gids →