ฝึกพูดภาษาอังกฤษด้วยเทคนิค Shadowing จากวิดีโอ: Introduction to statistical prediction models

กำลังสร้างบทเรียน...
1
In this video I will give you a very short introduction to what statistical prediction is and what statistical prediction models are.
2
And my goal of this video is to try to you to understand what a large language model,
3
which is the technology underlying these AI chatbots,
4
what they do and why they sometimes hallucinate or give you incorrect answers.
5
So let's start with a more effective example.
6
This is Claude III, the smallest model,
7
and this is something that I encountered when I was preparing material for a doctoral level entrepreneurship course.
8
And I asked what is the contribution of Israel Kirchner in modern entrepreneurship, and then I asked Claude to explain or give me his biography.
9
And it starts well, most of this is correct.
10
So Israel Kirchner was born in 1930s, and then the biography ends that he passed away at 2021 at the age of 90.
11
Well, he is very much alive still.
12
So if you google his name, you can view that the dead guy is giving a talk on YouTube in 2023.
13
So now the question is, why does the large language model,
14
Claude in this case, give us this incorrect answer that the person who is alive was actually died a couple years ago.
15
To understand why we get incorrect answers from these tools, we need to understand a bit about what they do.
16
So let's talk about statistical prediction a bit.
17
So here's our some data, then we're going to fit a statistical prediction model to this data.
18
So we have sales, a fictional company how many units they sell and this is how much they advertise.
19
This might be something expensive equipment and this is your advertisement spending thousands, let's say that it's over a year.
20
And we can see that if the company doesn't advertise anything they sell about 150,
21
130 something that units and then there is this declining returns to advertisement.
22
So a little bit of advertisement gets you from 150 to 200, but then going beyond 200 is much more difficult.
23
So we have diminishing returns.
24
The simplest way to model this data statistically is to use a regression model.
25
And regression model happens to be the simplest statistical prediction model.
26
The idea of a regression model is that we try to explain this data with a line.
27
And we draw the line through the data.
28
How the line is defined is not important for this talk, but it's a line that describes the data.
29
And this line is defined by two parameters.
30
So these parameters define a model.
31
So we are saying that this beta 0 gives us what is the the value of the dependent variable,
32
the predicted variable, the sales when ad spending is zero,
33
and then beta 1 gives us how many units the sales increases for each thousands of dollars spend on advertisement.
34
In machine learning, you might also see that these are called weights and biases.
35
So the beta one is weight because it tells how much ad spending affects sales,
36
and then beta zero is called the bias because it tells what is the overall level of sales, even if we don't do any advertisement.
37
And how the recursion line here is calculated is that we take the bias 158,
38
we add weight times add spending, and then if we want to generate observations from this model, we add some error.
39
So the idea of error is that there's some variation in the actual values around the predicted line.
40
So this equation without error here gives you the recursion line, and if you want to simulate observations from this data then you would add error.
41
So we can use this line to simulate more data.
42
So if we have, if we know ad spending, then we can just plot more predicted sales values here.
43
This works fairly well as long as you predict within samples.
44
So if you predict something that is, you have data, so our predictions work fairly well between zero and fifty thousand.
45
We are predicting a bit more here than what we are observing.
46
Problems occur when we start to extrapolate.
47
So if we are asking this model to predict how much we would sell if we actually increased our sales spending,
48
our advertisement spending, to a million.
49
Here's the data.
50
So we can predict what the sales would look like if we advertised for a million dollars.
51
Now this model is not very trustworthy for this kind of predictions.
52
First of all, it doesn't consider that advertisement has diminishing returns.
53
So the first dollar is more valuable than the tenth dollar and so on.
54
Also there are probably other constraints that come into play.
55
So So just increasing your advertisement 100-fold will not increase your sales 100-fold.
56
There are other constraints.
57
So the real value might be somewhere closer to 200 or 500 units than 1,500 units.
58
So extrapolation generally is not good in statistical models.
59
At least this line of gross extrapolation.
60
So we need to understand now that we calibrate the model using data, in this case a regression model.
61
It gives us some parameters and that parameters, they give us an equation and we can use the equation to calculate predictions.
62
The predictions are closed with original data as long as the values that we predict with our closed-grid original data.
63
All right, now let's take a look at how this same idea works with text.
64
So this is our training data and we are going to specify a small language model
65
and we try to predict more text like this.
66
And the idea of modeling language is that instead of looking at what is the predicted sales are given advertisements.
67
We look at what is the predicted word given the previous word.
68
And this kind of model will be called biogram, but it's not super important to understand how it works.
69
But the general idea is that when we start to train this kind of a prediction model that works with language, we take a look at, for example, the word the.
70
And the is here, and the is followed by cat three times.
71
It's followed by dog once and mat once.
72
So we can say that when we see the word the, then 60% of the case we have cat as the next word, and 20% of the case we have dog,
73
20% case we have mat.
74
So that gives us kind of like these probabilities and we can start simulating new data from this model.
75
Then after cat we have sat, the and ran.
76
So all of these have 33% probability and so on.
77
So we can calculate the probabilities.
78
This is the word here on this on rows and then columns is the next word.
79
We can start using this model to generate data.
80
And this is what large language models do.
81
So we have an initial, let's call it prompt.
82
So this is what we would write to chat GPT, for example.
83
We just start with the.
84
And then what the model looks at, it looks that, okay, the word is the.
85
When we predict the next word, we have cat, 60% dog, and then 20% mat.
86
And it randomly picks one of these.
87
This has 60% probability 20% and 20% and it picks cat by random.
88
Then after cat we have either ran, sat and the with equal probabilities,
89
the model just picks randomly sat.
90
And after sat we always have on, so that is what the model predict,
91
and so on until it predicts away, which is the last word.
92
and this is how language models work.
93
So we have a word and then we predict what is the next word.
94
Then we take the predicted word and we predict the following word.
95
This is a pretty simple thing and this story doesn't really make sense.
96
The cat sat on the cat, sat on the mat, the cat sat on, the cat ran away.
97
But it looks somewhat similar to the original data.
98
We can make this a lot better if we start to look at broader context.
99
So instead of looking at just previous word, we could take a look at the two previous words or three previous words.
100
We can increase the context length.
101
So this small language model has context length of one word, and then we have 81 parameters because we have 81 probabilities.
102
I just didn't include the zeros in this matrix.
103
This is how ChatGPT works.
104
So it looks what I have written this far, what is the likely next word.
105
ChatGPT, Claude, Gemini, they all work with the same principle, they are just a lot bigger.
106
So instead of 81 parameters like we have here, these models are in the hundreds of billions,
107
probably the trillions of parameters and the context length is not one word, but it is something called a token,
108
which is like a word part or a letter combination.
109
And for example, Claude has 100,000 tokens.
110
That's like multiple books that it looks at when it tries to predict the next word.
111
Now we get to look at why did Claude predict incorrectly.
112
And this is because of XAAAs.
113
So the training data for Claude
114
when it was trained using like most of the internet does not contain information about the death of Kirchner
115
and it's natural because he is still alive.
116
And here we have that the person died, was born in 1930,
117
and the training data also contains lots of biographies of people born in 1930
118
and most of those biographies end with explaining when the person passed.
119
So the language model does not have information about the passing of Kirchner,
120
but it has seen a general pattern that when the biography starts by telling that the person was born in 1930,
121
then the likely ending for the biography is just that the person died somewhere in 2020s for example.
122
So this is the basic idea of statistical prediction.
123
You calibrate the data, calibrate a model based on data, and then you try to predict new observations.
124
In large language models you predict the next word or the next word part.
125
In recursive models you predict a value, for example.

เกี่ยวกับบทเรียนนี้

คุณกำลังฝึกภาษาอังกฤษกับ "Introduction to statistical prediction models" ด้วยเทคนิค Shadowing — วิธีที่พัฒนาขึ้นสำหรับการฝึกนักแปลมืออาชีพ

ฟังทีละประโยค สังเกตการเน้นเสียงและการเชื่อมเสียง แล้วพูดตามดังๆ อย่างมั่นใจ ฝึกวันละ 15–30 นาทีจะเห็นผลลัพธ์ที่ชัดเจน

เทคนิค Shadowing คืออะไร?

Shadowing เป็นเทคนิคการเรียนรู้ภาษาที่ได้รับการรับรองทางวิทยาศาสตร์ พัฒนาขึ้นสำหรับการฝึกนักแปลมืออาชีพ วิธีการนี้เรียบง่ายแต่ทรงพลัง: คุณฟังเสียงภาษาอังกฤษจากเจ้าของภาษาและพูดตามทันที — เหมือนเงาที่ตามผู้พูดด้วยช่วงเวลาห่าง 1-2 วินาที การวิจัยแสดงว่าเทคนิคนี้ปรับปรุงความแม่นยำในการออกเสียง ทำนองเสียง จังหวะ การเชื่อมเสียง การฟังเข้าใจ และความคล่องแคล่วในการพูดได้อย่างมีนัยสำคัญ

เทคนิค shadowing: อ่านคู่มือฉบับเต็มทีละขั้นตอน →