Shadowing Practice: Introduction to statistical prediction models - Learn English Speaking with Video

Les maken...
1
In this video I will give you a very short introduction to what statistical prediction is and what statistical prediction models are.
2
And my goal of this video is to try to you to understand what a large language model,
3
which is the technology underlying these AI chatbots,
4
what they do and why they sometimes hallucinate or give you incorrect answers.
5
So let's start with a more effective example.
6
This is Claude III, the smallest model,
7
and this is something that I encountered when I was preparing material for a doctoral level entrepreneurship course.
8
And I asked what is the contribution of Israel Kirchner in modern entrepreneurship, and then I asked Claude to explain or give me his biography.
9
And it starts well, most of this is correct.
10
So Israel Kirchner was born in 1930s, and then the biography ends that he passed away at 2021 at the age of 90.
11
Well, he is very much alive still.
12
So if you google his name, you can view that the dead guy is giving a talk on YouTube in 2023.
13
So now the question is, why does the large language model,
14
Claude in this case, give us this incorrect answer that the person who is alive was actually died a couple years ago.
15
To understand why we get incorrect answers from these tools, we need to understand a bit about what they do.
16
So let's talk about statistical prediction a bit.
17
So here's our some data, then we're going to fit a statistical prediction model to this data.
18
So we have sales, a fictional company how many units they sell and this is how much they advertise.
19
This might be something expensive equipment and this is your advertisement spending thousands, let's say that it's over a year.
20
And we can see that if the company doesn't advertise anything they sell about 150,
21
130 something that units and then there is this declining returns to advertisement.
22
So a little bit of advertisement gets you from 150 to 200, but then going beyond 200 is much more difficult.
23
So we have diminishing returns.
24
The simplest way to model this data statistically is to use a regression model.
25
And regression model happens to be the simplest statistical prediction model.
26
The idea of a regression model is that we try to explain this data with a line.
27
And we draw the line through the data.
28
How the line is defined is not important for this talk, but it's a line that describes the data.
29
And this line is defined by two parameters.
30
So these parameters define a model.
31
So we are saying that this beta 0 gives us what is the the value of the dependent variable,
32
the predicted variable, the sales when ad spending is zero,
33
and then beta 1 gives us how many units the sales increases for each thousands of dollars spend on advertisement.
34
In machine learning, you might also see that these are called weights and biases.
35
So the beta one is weight because it tells how much ad spending affects sales,
36
and then beta zero is called the bias because it tells what is the overall level of sales, even if we don't do any advertisement.
37
And how the recursion line here is calculated is that we take the bias 158,
38
we add weight times add spending, and then if we want to generate observations from this model, we add some error.
39
So the idea of error is that there's some variation in the actual values around the predicted line.
40
So this equation without error here gives you the recursion line, and if you want to simulate observations from this data then you would add error.
41
So we can use this line to simulate more data.
42
So if we have, if we know ad spending, then we can just plot more predicted sales values here.
43
This works fairly well as long as you predict within samples.
44
So if you predict something that is, you have data, so our predictions work fairly well between zero and fifty thousand.
45
We are predicting a bit more here than what we are observing.
46
Problems occur when we start to extrapolate.
47
So if we are asking this model to predict how much we would sell if we actually increased our sales spending,
48
our advertisement spending, to a million.
49
Here's the data.
50
So we can predict what the sales would look like if we advertised for a million dollars.
51
Now this model is not very trustworthy for this kind of predictions.
52
First of all, it doesn't consider that advertisement has diminishing returns.
53
So the first dollar is more valuable than the tenth dollar and so on.
54
Also there are probably other constraints that come into play.
55
So So just increasing your advertisement 100-fold will not increase your sales 100-fold.
56
There are other constraints.
57
So the real value might be somewhere closer to 200 or 500 units than 1,500 units.
58
So extrapolation generally is not good in statistical models.
59
At least this line of gross extrapolation.
60
So we need to understand now that we calibrate the model using data, in this case a regression model.
61
It gives us some parameters and that parameters, they give us an equation and we can use the equation to calculate predictions.
62
The predictions are closed with original data as long as the values that we predict with our closed-grid original data.
63
All right, now let's take a look at how this same idea works with text.
64
So this is our training data and we are going to specify a small language model
65
and we try to predict more text like this.
66
And the idea of modeling language is that instead of looking at what is the predicted sales are given advertisements.
67
We look at what is the predicted word given the previous word.
68
And this kind of model will be called biogram, but it's not super important to understand how it works.
69
But the general idea is that when we start to train this kind of a prediction model that works with language, we take a look at, for example, the word the.
70
And the is here, and the is followed by cat three times.
71
It's followed by dog once and mat once.
72
So we can say that when we see the word the, then 60% of the case we have cat as the next word, and 20% of the case we have dog,
73
20% case we have mat.
74
So that gives us kind of like these probabilities and we can start simulating new data from this model.
75
Then after cat we have sat, the and ran.
76
So all of these have 33% probability and so on.
77
So we can calculate the probabilities.
78
This is the word here on this on rows and then columns is the next word.
79
We can start using this model to generate data.
80
And this is what large language models do.
81
So we have an initial, let's call it prompt.
82
So this is what we would write to chat GPT, for example.
83
We just start with the.
84
And then what the model looks at, it looks that, okay, the word is the.
85
When we predict the next word, we have cat, 60% dog, and then 20% mat.
86
And it randomly picks one of these.
87
This has 60% probability 20% and 20% and it picks cat by random.
88
Then after cat we have either ran, sat and the with equal probabilities,
89
the model just picks randomly sat.
90
And after sat we always have on, so that is what the model predict,
91
and so on until it predicts away, which is the last word.
92
and this is how language models work.
93
So we have a word and then we predict what is the next word.
94
Then we take the predicted word and we predict the following word.
95
This is a pretty simple thing and this story doesn't really make sense.
96
The cat sat on the cat, sat on the mat, the cat sat on, the cat ran away.
97
But it looks somewhat similar to the original data.
98
We can make this a lot better if we start to look at broader context.
99
So instead of looking at just previous word, we could take a look at the two previous words or three previous words.
100
We can increase the context length.
101
So this small language model has context length of one word, and then we have 81 parameters because we have 81 probabilities.
102
I just didn't include the zeros in this matrix.
103
This is how ChatGPT works.
104
So it looks what I have written this far, what is the likely next word.
105
ChatGPT, Claude, Gemini, they all work with the same principle, they are just a lot bigger.
106
So instead of 81 parameters like we have here, these models are in the hundreds of billions,
107
probably the trillions of parameters and the context length is not one word, but it is something called a token,
108
which is like a word part or a letter combination.
109
And for example, Claude has 100,000 tokens.
110
That's like multiple books that it looks at when it tries to predict the next word.
111
Now we get to look at why did Claude predict incorrectly.
112
And this is because of XAAAs.
113
So the training data for Claude
114
when it was trained using like most of the internet does not contain information about the death of Kirchner
115
and it's natural because he is still alive.
116
And here we have that the person died, was born in 1930,
117
and the training data also contains lots of biographies of people born in 1930
118
and most of those biographies end with explaining when the person passed.
119
So the language model does not have information about the passing of Kirchner,
120
but it has seen a general pattern that when the biography starts by telling that the person was born in 1930,
121
then the likely ending for the biography is just that the person died somewhere in 2020s for example.
122
So this is the basic idea of statistical prediction.
123
You calibrate the data, calibrate a model based on data, and then you try to predict new observations.
124
In large language models you predict the next word or the next word part.
125
In recursive models you predict a value, for example.

Woordenschat en spreektips bij deze les

Deze spreekles op niveau C1 is gebaseerd op de video “Introduction to statistical prediction models”. Deze woorden komen het vaakst terug: model, word, predict, prediction, cat. Deze video bevat 125 zinnen en 1690 woorden om na te spreken. Het gesproken deel duurt 11:51. De spreker praat in een gelijkmatig tempo van ongeveer 143 woorden per minuut, prettig om te shadowen. 86% van de woorden hoort bij de 3.000 meest gebruikte Engelse woorden; de rest kun je beter vooraf bekijken.

Belangrijke woorden in deze video

De 15 moeilijkste woorden uit de video, met uitspraak en betekenis:

WoordUitspraakBetekenis
predict werkwoord/pɹɪˈdɪkt/voorspellen
prediction zelfstandig naamwoord/pɹɪˈdɪkʃən/voorspelling
advertisement zelfstandig naamwoord/ˈædvɚˌtaɪzmənt/reclame, advertentie
biography zelfstandig naamwoord/baɪˈɑɡɹəfi/biografie
probability zelfstandig naamwoord/ˌpɹɑ.bəˈbɪl.ə.ti/waarschijnlijkheid
regression zelfstandig naamwoord/ɹiːˈɡɹɛʃ.ən/teruggang
calibrate werkwoord/ˈkæl.ɪ.bɹeɪ̯t/kalibreer
advertise werkwoord/ˈædvɚˌtaɪ̯z/adverteren
calculate werkwoord/ˈkælkjʊleɪt/berekenen, uitwerken
incorrect bijvoeglijk naamwoord/ˌɪn.kəˈɹɛkt/incorrect
equation zelfstandig naamwoord/ɪˈkweɪ.ʒən/vergelijking
constraint zelfstandig naamwoord/kənˈstɹeɪnt/beperking, inperking
diminish werkwoord/dɪˈmɪnɪʃ/verkleinen, verminderen
extrapolation zelfstandig naamwoord/ɛkˌstɹæp.əˈleɪ.ʃən/extrapolatie
recursion zelfstandig naamwoord/ɹɪˈkɜː(ɹ)ʒən/recursie

Phrasal verbs die je zult horen

WoordBetekenis
run away werkwoordvluchten, weglopen

Grammatica in deze video

De structuren die de spreker het meest gebruikt, met de exacte woorden uit de video:

StructuurIn de video
Lijdende vorm be + voltooid deelwoord — het gaat om wat er gebeurt, niet om wie het doetwas born · is defined · are called
Present perfect have/has + voltooid deelwoord — iets uit het verleden dat nu nog telthave sat · have written · has seen
Betrekkelijke bijzinnen who / which + zin — extra informatie over een persoon of dingperson who is · away, which is

Uitspraak om op te letten

De spreker gebruikt 6 samentrekkingen en verkorte vormen, zoals doesn't, didn't, don't. Spreek ze kort uit, zoals je ze hoort.

  • De “th”-klanken: trustworthy /ˈtɹʌst.wɜː.ði/, tenth /tɛnθ/
  • De klanken “sh” en “zh”: prediction /pɹɪˈdɪkʃən/, regression /ɹiːˈɡɹɛʃ.ən/, equation /ɪˈkweɪ.ʒən/, diminish /dɪˈmɪnɪʃ/, extrapolation /ɛkˌstɹæp.əˈleɪ.ʃən/
  • Lange woorden — let op de klemtoon: advertisement /ˈædvɚˌtaɪzmənt/, parameter /pəˈɹæm.ə.tɚ/, biography /baɪˈɑɡɹəfi/, probability /ˌpɹɑ.bəˈbɪl.ə.ti/, extrapolation /ɛkˌstɹæp.əˈleɪ.ʃən/

Zo oefen je met deze video

  1. Luister de hele video één keer zonder te spreken en noteer de woorden die je niet kent.
  2. Spreek zin voor zin na op normale snelheid en herhaal elke zin tot je ritme gelijk is aan dat van de spreker.
  3. Neem jezelf op en vergelijk met het origineel; let daarbij op woorden als predict, prediction, advertisement.

Wat is de Shadowing-techniek?

Shadowing is een wetenschappelijk onderbouwde taalleermethode die oorspronkelijk is ontwikkeld voor professionele tolkentraining en gepopulariseerd door polyglot Dr. Alexander Arguelles. De methode is eenvoudig maar krachtig: je luistert naar native Engelse audio en herhaalt het onmiddellijk hardop — als een schaduw die de spreker volgt met slechts 1–2 seconden vertraging. In tegenstelling tot passief luisteren of grammaticadrills, dwingt shadowing je hersenen en mondspieren om echte spraakpatronen tegelijkertijd te verwerken en te reproduceren. Onderzoek toont aan dat het de uitspraaknauwkeurigheid, intonatie, ritme, verbonden spraak, luisterbegrip en spreekvaardigheid aanzienlijk verbetert — waardoor het een van de meest effectieve methoden is voor IELTS Speaking-voorbereiding en echte Engelse communicatie.

Shadowing-techniek: lees de volledige stap-voor-stap-gids →