Pratique du Shadowing: Introduction to statistical prediction models - Apprendre l'anglais à l'oral avec la vidéo

Création de la leçon...
1
In this video I will give you a very short introduction to what statistical prediction is and what statistical prediction models are.
2
And my goal of this video is to try to you to understand what a large language model,
3
which is the technology underlying these AI chatbots,
4
what they do and why they sometimes hallucinate or give you incorrect answers.
5
So let's start with a more effective example.
6
This is Claude III, the smallest model,
7
and this is something that I encountered when I was preparing material for a doctoral level entrepreneurship course.
8
And I asked what is the contribution of Israel Kirchner in modern entrepreneurship, and then I asked Claude to explain or give me his biography.
9
And it starts well, most of this is correct.
10
So Israel Kirchner was born in 1930s, and then the biography ends that he passed away at 2021 at the age of 90.
11
Well, he is very much alive still.
12
So if you google his name, you can view that the dead guy is giving a talk on YouTube in 2023.
13
So now the question is, why does the large language model,
14
Claude in this case, give us this incorrect answer that the person who is alive was actually died a couple years ago.
15
To understand why we get incorrect answers from these tools, we need to understand a bit about what they do.
16
So let's talk about statistical prediction a bit.
17
So here's our some data, then we're going to fit a statistical prediction model to this data.
18
So we have sales, a fictional company how many units they sell and this is how much they advertise.
19
This might be something expensive equipment and this is your advertisement spending thousands, let's say that it's over a year.
20
And we can see that if the company doesn't advertise anything they sell about 150,
21
130 something that units and then there is this declining returns to advertisement.
22
So a little bit of advertisement gets you from 150 to 200, but then going beyond 200 is much more difficult.
23
So we have diminishing returns.
24
The simplest way to model this data statistically is to use a regression model.
25
And regression model happens to be the simplest statistical prediction model.
26
The idea of a regression model is that we try to explain this data with a line.
27
And we draw the line through the data.
28
How the line is defined is not important for this talk, but it's a line that describes the data.
29
And this line is defined by two parameters.
30
So these parameters define a model.
31
So we are saying that this beta 0 gives us what is the the value of the dependent variable,
32
the predicted variable, the sales when ad spending is zero,
33
and then beta 1 gives us how many units the sales increases for each thousands of dollars spend on advertisement.
34
In machine learning, you might also see that these are called weights and biases.
35
So the beta one is weight because it tells how much ad spending affects sales,
36
and then beta zero is called the bias because it tells what is the overall level of sales, even if we don't do any advertisement.
37
And how the recursion line here is calculated is that we take the bias 158,
38
we add weight times add spending, and then if we want to generate observations from this model, we add some error.
39
So the idea of error is that there's some variation in the actual values around the predicted line.
40
So this equation without error here gives you the recursion line, and if you want to simulate observations from this data then you would add error.
41
So we can use this line to simulate more data.
42
So if we have, if we know ad spending, then we can just plot more predicted sales values here.
43
This works fairly well as long as you predict within samples.
44
So if you predict something that is, you have data, so our predictions work fairly well between zero and fifty thousand.
45
We are predicting a bit more here than what we are observing.
46
Problems occur when we start to extrapolate.
47
So if we are asking this model to predict how much we would sell if we actually increased our sales spending,
48
our advertisement spending, to a million.
49
Here's the data.
50
So we can predict what the sales would look like if we advertised for a million dollars.
51
Now this model is not very trustworthy for this kind of predictions.
52
First of all, it doesn't consider that advertisement has diminishing returns.
53
So the first dollar is more valuable than the tenth dollar and so on.
54
Also there are probably other constraints that come into play.
55
So So just increasing your advertisement 100-fold will not increase your sales 100-fold.
56
There are other constraints.
57
So the real value might be somewhere closer to 200 or 500 units than 1,500 units.
58
So extrapolation generally is not good in statistical models.
59
At least this line of gross extrapolation.
60
So we need to understand now that we calibrate the model using data, in this case a regression model.
61
It gives us some parameters and that parameters, they give us an equation and we can use the equation to calculate predictions.
62
The predictions are closed with original data as long as the values that we predict with our closed-grid original data.
63
All right, now let's take a look at how this same idea works with text.
64
So this is our training data and we are going to specify a small language model
65
and we try to predict more text like this.
66
And the idea of modeling language is that instead of looking at what is the predicted sales are given advertisements.
67
We look at what is the predicted word given the previous word.
68
And this kind of model will be called biogram, but it's not super important to understand how it works.
69
But the general idea is that when we start to train this kind of a prediction model that works with language, we take a look at, for example, the word the.
70
And the is here, and the is followed by cat three times.
71
It's followed by dog once and mat once.
72
So we can say that when we see the word the, then 60% of the case we have cat as the next word, and 20% of the case we have dog,
73
20% case we have mat.
74
So that gives us kind of like these probabilities and we can start simulating new data from this model.
75
Then after cat we have sat, the and ran.
76
So all of these have 33% probability and so on.
77
So we can calculate the probabilities.
78
This is the word here on this on rows and then columns is the next word.
79
We can start using this model to generate data.
80
And this is what large language models do.
81
So we have an initial, let's call it prompt.
82
So this is what we would write to chat GPT, for example.
83
We just start with the.
84
And then what the model looks at, it looks that, okay, the word is the.
85
When we predict the next word, we have cat, 60% dog, and then 20% mat.
86
And it randomly picks one of these.
87
This has 60% probability 20% and 20% and it picks cat by random.
88
Then after cat we have either ran, sat and the with equal probabilities,
89
the model just picks randomly sat.
90
And after sat we always have on, so that is what the model predict,
91
and so on until it predicts away, which is the last word.
92
and this is how language models work.
93
So we have a word and then we predict what is the next word.
94
Then we take the predicted word and we predict the following word.
95
This is a pretty simple thing and this story doesn't really make sense.
96
The cat sat on the cat, sat on the mat, the cat sat on, the cat ran away.
97
But it looks somewhat similar to the original data.
98
We can make this a lot better if we start to look at broader context.
99
So instead of looking at just previous word, we could take a look at the two previous words or three previous words.
100
We can increase the context length.
101
So this small language model has context length of one word, and then we have 81 parameters because we have 81 probabilities.
102
I just didn't include the zeros in this matrix.
103
This is how ChatGPT works.
104
So it looks what I have written this far, what is the likely next word.
105
ChatGPT, Claude, Gemini, they all work with the same principle, they are just a lot bigger.
106
So instead of 81 parameters like we have here, these models are in the hundreds of billions,
107
probably the trillions of parameters and the context length is not one word, but it is something called a token,
108
which is like a word part or a letter combination.
109
And for example, Claude has 100,000 tokens.
110
That's like multiple books that it looks at when it tries to predict the next word.
111
Now we get to look at why did Claude predict incorrectly.
112
And this is because of XAAAs.
113
So the training data for Claude
114
when it was trained using like most of the internet does not contain information about the death of Kirchner
115
and it's natural because he is still alive.
116
And here we have that the person died, was born in 1930,
117
and the training data also contains lots of biographies of people born in 1930
118
and most of those biographies end with explaining when the person passed.
119
So the language model does not have information about the passing of Kirchner,
120
but it has seen a general pattern that when the biography starts by telling that the person was born in 1930,
121
then the likely ending for the biography is just that the person died somewhere in 2020s for example.
122
So this is the basic idea of statistical prediction.
123
You calibrate the data, calibrate a model based on data, and then you try to predict new observations.
124
In large language models you predict the next word or the next word part.
125
In recursive models you predict a value, for example.

Vocabulaire et conseils d’expression pour cette leçon

Cette leçon d’expression orale de niveau C1 s’appuie sur la vidéo « Introduction to statistical prediction models ». Les mots qui reviennent le plus souvent : model, word, predict, prediction, cat. Cette vidéo contient 125 phrases et 1690 mots à répéter en shadowing. La partie parlée dure 11:51. Le locuteur parle à un rythme régulier d’environ 143 mots par minute, confortable pour le shadowing. 86 % des mots font partie des 3 000 mots les plus courants en anglais ; le reste mérite d’être vu avant de commencer.

Vocabulaire clé de cette vidéo

Les 15 mots les plus avancés de la vidéo, avec leur prononciation et leur sens :

MotPrononciationSens
predict verbe/pɹɪˈdɪkt/prédire
prediction nom/pɹɪˈdɪkʃən/prédiction
advertisement nom/ˈædvɚˌtaɪzmənt/publicité, pub
parameter nom/pəˈɹæm.ə.tɚ/paramètre
biography nom/baɪˈɑɡɹəfi/biographie
probability nom/ˌpɹɑ.bəˈbɪl.ə.ti/probabilité
regression nom/ɹiːˈɡɹɛʃ.ən/régression
calibrate verbe/ˈkæl.ɪ.bɹeɪ̯t/étalonner, calibrer
simulate verbe/ˈsɪm.jə.lət/simuler
advertise verbe/ˈædvɚˌtaɪ̯z/annoncer
calculate verbe/ˈkælkjʊleɪt/calculer
equation nom/ɪˈkweɪ.ʒən/équation
constraint nom/kənˈstɹeɪnt/contrainte
diminish verbe/dɪˈmɪnɪʃ/réduire
extrapolation nom/ɛkˌstɹæp.əˈleɪ.ʃən/extrapolation

Les verbes à particule que vous entendrez

MotSens
run away verbes'enfuir

La grammaire de cette vidéo

Les structures que le locuteur utilise le plus, avec les mots exacts de la vidéo :

StructureDans la vidéo
Voix passive be + participe passé — l’accent est mis sur ce qui arrive, pas sur qui le faitwas born · is defined · are called
Present perfect have/has + participe passé — une action passée qui compte encore maintenanthave sat · have written · has seen
Propositions relatives who / which + proposition — une précision sur une personne ou une choseperson who is · away, which is

Prononciation à surveiller

Le locuteur utilise 6 contractions et formes réduites, comme doesn't, didn't, don't. Prononcez-les sous leur forme courte, telles que vous les entendez.

  • Les sons « th »: trustworthy /ˈtɹʌst.wɜː.ði/, tenth /tɛnθ/
  • Les sons « sh » et « zh »: prediction /pɹɪˈdɪkʃən/, regression /ɹiːˈɡɹɛʃ.ən/, equation /ɪˈkweɪ.ʒən/, diminish /dɪˈmɪnɪʃ/, extrapolation /ɛkˌstɹæp.əˈleɪ.ʃən/
  • Mots longs — placez bien l’accent: advertisement /ˈædvɚˌtaɪzmənt/, parameter /pəˈɹæm.ə.tɚ/, biography /baɪˈɑɡɹəfi/, probability /ˌpɹɑ.bəˈbɪl.ə.ti/, extrapolation /ɛkˌstɹæp.əˈleɪ.ʃən/

Comment s’entraîner avec cette vidéo

  1. Écoutez la vidéo en entier une fois sans parler et notez les mots que vous ne connaissez pas.
  2. Répétez phrase par phrase à vitesse normale, en reprenant chacune jusqu’à ce que votre rythme corresponde à celui du locuteur.
  3. Enregistrez-vous et comparez avec l’original, en faisant attention à des mots comme predict, prediction, advertisement.

Qu'est-ce que la technique du Shadowing ?

Le Shadowing est une technique d'apprentissage des langues fondée sur la science, développée à l'origine pour la formation des interprètes professionnels. Le principe est simple mais puissant : vous écoutez de l'anglais natif et le répétez immédiatement à voix haute — comme une ombre suivant le locuteur avec un décalage de 1 à 2 secondes. Les recherches montrent une amélioration significative de la précision de la prononciation, de l'intonation, du rythme, des liaisons, de la compréhension orale et de la fluidité.

Technique du shadowing : lire le guide complet étape par étape →