跟读练习: Introduction to statistical prediction models - 通过视频学习英语口语

正在创建课程...
1
In this video I will give you a very short introduction to what statistical prediction is and what statistical prediction models are.
2
And my goal of this video is to try to you to understand what a large language model,
3
which is the technology underlying these AI chatbots,
4
what they do and why they sometimes hallucinate or give you incorrect answers.
5
So let's start with a more effective example.
6
This is Claude III, the smallest model,
7
and this is something that I encountered when I was preparing material for a doctoral level entrepreneurship course.
8
And I asked what is the contribution of Israel Kirchner in modern entrepreneurship, and then I asked Claude to explain or give me his biography.
9
And it starts well, most of this is correct.
10
So Israel Kirchner was born in 1930s, and then the biography ends that he passed away at 2021 at the age of 90.
11
Well, he is very much alive still.
12
So if you google his name, you can view that the dead guy is giving a talk on YouTube in 2023.
13
So now the question is, why does the large language model,
14
Claude in this case, give us this incorrect answer that the person who is alive was actually died a couple years ago.
15
To understand why we get incorrect answers from these tools, we need to understand a bit about what they do.
16
So let's talk about statistical prediction a bit.
17
So here's our some data, then we're going to fit a statistical prediction model to this data.
18
So we have sales, a fictional company how many units they sell and this is how much they advertise.
19
This might be something expensive equipment and this is your advertisement spending thousands, let's say that it's over a year.
20
And we can see that if the company doesn't advertise anything they sell about 150,
21
130 something that units and then there is this declining returns to advertisement.
22
So a little bit of advertisement gets you from 150 to 200, but then going beyond 200 is much more difficult.
23
So we have diminishing returns.
24
The simplest way to model this data statistically is to use a regression model.
25
And regression model happens to be the simplest statistical prediction model.
26
The idea of a regression model is that we try to explain this data with a line.
27
And we draw the line through the data.
28
How the line is defined is not important for this talk, but it's a line that describes the data.
29
And this line is defined by two parameters.
30
So these parameters define a model.
31
So we are saying that this beta 0 gives us what is the the value of the dependent variable,
32
the predicted variable, the sales when ad spending is zero,
33
and then beta 1 gives us how many units the sales increases for each thousands of dollars spend on advertisement.
34
In machine learning, you might also see that these are called weights and biases.
35
So the beta one is weight because it tells how much ad spending affects sales,
36
and then beta zero is called the bias because it tells what is the overall level of sales, even if we don't do any advertisement.
37
And how the recursion line here is calculated is that we take the bias 158,
38
we add weight times add spending, and then if we want to generate observations from this model, we add some error.
39
So the idea of error is that there's some variation in the actual values around the predicted line.
40
So this equation without error here gives you the recursion line, and if you want to simulate observations from this data then you would add error.
41
So we can use this line to simulate more data.
42
So if we have, if we know ad spending, then we can just plot more predicted sales values here.
43
This works fairly well as long as you predict within samples.
44
So if you predict something that is, you have data, so our predictions work fairly well between zero and fifty thousand.
45
We are predicting a bit more here than what we are observing.
46
Problems occur when we start to extrapolate.
47
So if we are asking this model to predict how much we would sell if we actually increased our sales spending,
48
our advertisement spending, to a million.
49
Here's the data.
50
So we can predict what the sales would look like if we advertised for a million dollars.
51
Now this model is not very trustworthy for this kind of predictions.
52
First of all, it doesn't consider that advertisement has diminishing returns.
53
So the first dollar is more valuable than the tenth dollar and so on.
54
Also there are probably other constraints that come into play.
55
So So just increasing your advertisement 100-fold will not increase your sales 100-fold.
56
There are other constraints.
57
So the real value might be somewhere closer to 200 or 500 units than 1,500 units.
58
So extrapolation generally is not good in statistical models.
59
At least this line of gross extrapolation.
60
So we need to understand now that we calibrate the model using data, in this case a regression model.
61
It gives us some parameters and that parameters, they give us an equation and we can use the equation to calculate predictions.
62
The predictions are closed with original data as long as the values that we predict with our closed-grid original data.
63
All right, now let's take a look at how this same idea works with text.
64
So this is our training data and we are going to specify a small language model
65
and we try to predict more text like this.
66
And the idea of modeling language is that instead of looking at what is the predicted sales are given advertisements.
67
We look at what is the predicted word given the previous word.
68
And this kind of model will be called biogram, but it's not super important to understand how it works.
69
But the general idea is that when we start to train this kind of a prediction model that works with language, we take a look at, for example, the word the.
70
And the is here, and the is followed by cat three times.
71
It's followed by dog once and mat once.
72
So we can say that when we see the word the, then 60% of the case we have cat as the next word, and 20% of the case we have dog,
73
20% case we have mat.
74
So that gives us kind of like these probabilities and we can start simulating new data from this model.
75
Then after cat we have sat, the and ran.
76
So all of these have 33% probability and so on.
77
So we can calculate the probabilities.
78
This is the word here on this on rows and then columns is the next word.
79
We can start using this model to generate data.
80
And this is what large language models do.
81
So we have an initial, let's call it prompt.
82
So this is what we would write to chat GPT, for example.
83
We just start with the.
84
And then what the model looks at, it looks that, okay, the word is the.
85
When we predict the next word, we have cat, 60% dog, and then 20% mat.
86
And it randomly picks one of these.
87
This has 60% probability 20% and 20% and it picks cat by random.
88
Then after cat we have either ran, sat and the with equal probabilities,
89
the model just picks randomly sat.
90
And after sat we always have on, so that is what the model predict,
91
and so on until it predicts away, which is the last word.
92
and this is how language models work.
93
So we have a word and then we predict what is the next word.
94
Then we take the predicted word and we predict the following word.
95
This is a pretty simple thing and this story doesn't really make sense.
96
The cat sat on the cat, sat on the mat, the cat sat on, the cat ran away.
97
But it looks somewhat similar to the original data.
98
We can make this a lot better if we start to look at broader context.
99
So instead of looking at just previous word, we could take a look at the two previous words or three previous words.
100
We can increase the context length.
101
So this small language model has context length of one word, and then we have 81 parameters because we have 81 probabilities.
102
I just didn't include the zeros in this matrix.
103
This is how ChatGPT works.
104
So it looks what I have written this far, what is the likely next word.
105
ChatGPT, Claude, Gemini, they all work with the same principle, they are just a lot bigger.
106
So instead of 81 parameters like we have here, these models are in the hundreds of billions,
107
probably the trillions of parameters and the context length is not one word, but it is something called a token,
108
which is like a word part or a letter combination.
109
And for example, Claude has 100,000 tokens.
110
That's like multiple books that it looks at when it tries to predict the next word.
111
Now we get to look at why did Claude predict incorrectly.
112
And this is because of XAAAs.
113
So the training data for Claude
114
when it was trained using like most of the internet does not contain information about the death of Kirchner
115
and it's natural because he is still alive.
116
And here we have that the person died, was born in 1930,
117
and the training data also contains lots of biographies of people born in 1930
118
and most of those biographies end with explaining when the person passed.
119
So the language model does not have information about the passing of Kirchner,
120
but it has seen a general pattern that when the biography starts by telling that the person was born in 1930,
121
then the likely ending for the biography is just that the person died somewhere in 2020s for example.
122
So this is the basic idea of statistical prediction.
123
You calibrate the data, calibrate a model based on data, and then you try to predict new observations.
124
In large language models you predict the next word or the next word part.
125
In recursive models you predict a value, for example.

本课的词汇与口语要点

这节 C1 级别的口语课以视频“Introduction to statistical prediction models”为素材。 视频中反复出现的词有:model, word, predict, prediction, cat。 这段视频共有 125 个句子、1690 个单词可供跟读。 讲话部分时长为 11:51。 说话人语速平稳,每分钟约 143 个词,很适合跟读。 86% 的单词属于英语最常用的 3,000 词,其余的词建议在练习前先学一下。

视频中的重点词汇

视频中最难的 15 个单词,附发音和释义:

单词发音释义
predict 动词/pɹɪˈdɪkt/預言 /预言
prediction 名词/pɹɪˈdɪkʃən/預言 /预言, 預估 /预估
advertisement 名词/ˈædvɚˌtaɪzmənt/廣告 /广告
parameter 名词/pəˈɹæm.ə.tɚ/參數 /参数, 參量 /参量
biography 名词/baɪˈɑɡɹəfi/傳記 /传记
probability 名词/ˌpɹɑ.bəˈbɪl.ə.ti/可能性
regression 名词/ɹiːˈɡɹɛʃ.ən/回歸 /回归, 迴歸 /回归
calibrate 动词/ˈkæl.ɪ.bɹeɪ̯t/校準 /校准, 標定 /标定
simulate 动词/ˈsɪm.jə.lət/模拟
advertise 动词/ˈædvɚˌtaɪ̯z/賣廣告 /卖广告
calculate 动词/ˈkælkjʊleɪt/計算 /计算, 算
incorrect 形容词/ˌɪn.kəˈɹɛkt/不對, 不对
equation 名词/ɪˈkweɪ.ʒən/方程式, 方程
constraint 名词/kənˈstɹeɪnt/約束 /约束
extrapolation 名词/ɛkˌstɹæp.əˈleɪ.ʃən/延拓

视频中出现的短语动词

单词释义
run away 动词逃跑

视频中的语法

说话人最常用的结构,并附上视频中的原话:

结构视频中的用法
被动语态 be + 过去分词 — 强调发生了什么,而不是谁做的was born · is defined · are called
现在完成时 have/has + 过去分词 — 过去发生但与现在仍有关联的事have sat · have written · has seen
定语从句 who / which + 从句 — 补充说明人或事物person who is · away, which is

需要注意的发音

说话人用了 6 次缩略和弱读形式,例如 doesn't, didn't, don't。请按听到的简短形式来说。

  • “th” 音: trustworthy /ˈtɹʌst.wɜː.ði/, tenth /tɛnθ/
  • “sh” 和 “zh” 音: prediction /pɹɪˈdɪkʃən/, regression /ɹiːˈɡɹɛʃ.ən/, equation /ɪˈkweɪ.ʒən/, diminish /dɪˈmɪnɪʃ/, extrapolation /ɛkˌstɹæp.əˈleɪ.ʃən/
  • 长单词——注意重音位置: advertisement /ˈædvɚˌtaɪzmənt/, parameter /pəˈɹæm.ə.tɚ/, biography /baɪˈɑɡɹəfi/, probability /ˌpɹɑ.bəˈbɪl.ə.ti/, extrapolation /ɛkˌstɹæp.əˈleɪ.ʃən/

如何用这段视频练习

  1. 先完整听一遍视频,不要开口,记下不认识的单词。
  2. 用正常速度逐句跟读,每句重复到你的节奏与说话人一致为止。
  3. 录下自己的声音并与原声对比,特别注意 predict, prediction, advertisement 这类单词。

什么是跟读法?

跟读法 (Shadowing) 是一种有科学依据的语言学习技巧,最初开发用于专业口译员的培训,并由多语言者Alexander Arguelles博士普及。这个方法简单而强大:您在听英语母语原声的同时立即大声重复——就像是一个延迟1-2秒紧跟说话者的影子。与被动听力或语法练习不同,跟读法强迫您的大脑和口腔肌肉同时处理并模仿真实的讲话模式。研究表明它能显着提高发音准确性,语调,节奏,连读,听力理解和口语流利度——使其成为雅思口语备考和真实英语交流最有效的方法之一。

影子跟读法: 阅读完整分步指南 →