跟读练习: Transformers, the tech behind LLMs | Deep Learning Chapter 5 - 通过视频学习英语口语
正在创建课程...
1
The initials GPT stand for Generative Pre-trained Transformer.
2
So that first word is straightforward enough, these are bots that generate new text.
3
Pre-trained refers to how the model went through a process of learning from a massive amount of data,
4
and the prefix insinuates that there's more room to fine-tune it on specific tasks with additional training.
5
But the last word, that's the real key piece.
6
A transformer is a specific kind of neural network, a machine learning model, and it's That's the core invention underlying the current boom in AI.
7
What I want to do with this video
8
and the following chapters is go through a visually driven explanation for what actually happens inside a transformer.
9
We're going to follow the data that flows through it and go step by step.
10
There are many different kinds of models that you can build using transformers.
11
Some models take in audio and produce a transcript.
12
This sentence comes from a model going the other way around, producing synthetic speech just from text.
13
All those tools that took the world by storm in 2022, like Dolly and Midjourney that take in a text description and produce an image, are based on Transformers.
14
Even if I can't quite get it to understand what a pie creature is supposed to be, I'm still blown away that this kind of thing is even remotely possible.
15
And the original Transformer, introduced in 2017 by Google, was invented for the specific use case of translating text from one language into another.
16
But the variant that you and I will focus on, which is the type that underlies tools like ChatGPT, will be a model that's trained to take in a piece of text,
17
maybe even with some surrounding images or sound accompanying it, and produce a prediction for what comes next in the passage.
18
That prediction takes the form of a probability distribution over many different chunks of text that might follow.
19
At first glance, you might think that predicting the next word feels like a very different goal from generating new text.
20
But once you have a prediction model like this, a simple thing you could try to make it generate a
21
longer piece of text is to give it an initial snippet to work with, have it take a random sample from the distribution it just generated, append that sample to the text,
22
and then run the whole process again to make a new
23
prediction based on all the new text including what it just added.
24
I don't know about you, but it really doesn't feel like this should actually work.
25
In this animation, for example, I'm running GPT-2 on my laptop and having it repeatedly predict
26
and sample the next chunk of text to generate a story based on the seed text.
27
And the story just doesn't actually really make that much sense.
28
But if I swap it out for API calls to GPT-3 instead, which is the same basic model, just much bigger, suddenly, almost magically,
29
we do get a sensible story, one that even seems to infer that a pie creature would live in a land of math and computation.
30
process here of repeated prediction and sampling is essentially what's happening
31
when you interact with ChatGPT or any of these other large language models
32
and you see them producing one word at a time.
33
In fact, one feature
34
that I would very much enjoy is the ability to see the underlying distribution for each new word that it chooses.
35
Let's kick things off with a very high-level preview of how data flows through a transformer.
36
We will spend much more time motivating and interpreting and expanding on the details of each step, but in broad strokes.
37
When one of these chat bots generates a given word, here's what's going on under the hood.
38
First, the input is broken up into a bunch of little pieces.
39
These pieces are called tokens, and in the case of text, these tend to be words or little pieces of words or other common character combinations.
40
If images or sound are involved, then tokens could be little patches of that image or little chunks of that sound.
41
Each one of these tokens is then associated with a vector, meaning some list of numbers, which is meant to somehow encode the meaning of that piece.
42
If you think of these vectors as giving coordinates in some very high dimensional space, words with similar meanings tend to land on vectors that are close to each other in that space.
43
This sequence of vectors then passes through an operation that's known as an attention block,
44
and this allows the vectors to talk to each other and pass information back and forth to update their values. For example, the meaning of the word model in the phrase a machine
45
learning model is different from its meaning in the phrase a fashion model.
46
The attention block is what's responsible for figuring out
47
which words in the context are relevant to updating the meanings of which other words, and how exactly those meanings should be updated.
48
And again, whenever I use the word meaning, this is somehow entirely encoded in the entries of those vectors.
49
After that, these vectors pass through a different kind of operation, and depending on the source that you're reading, this will be referred to as a multi-layer perceptron, or maybe a feedforward layer.
50
And here, the vectors don't talk to each other, they all go through the same operation in parallel.
51
And while this block is a little bit harder to interpret, later on we'll talk about how this step is a little bit like asking a long list of questions about each vector,
52
and then updating them based on the answers to those questions.
53
All of the operations in both of these blocks look like a giant pile of matrix multiplications,
54
and our primary job is going to be to understand how to read the underlying matrices.
55
I'm glossing over some details about some normalization steps that happen in between, but this is after all a high-level preview.
56
After that, the process essentially repeats.
57
You go back and forth between attention blocks and multilayer perceptron blocks, until at At the very end, the hope is
58
that all of the essential meaning of the passage has somehow been baked into the very last vector in the sequence.
59
We then perform a certain operation on that last vector
60
that produces a probability distribution over all possible tokens that might come next.
61
And like I said, once you have a tool that predicts what comes next given a snippet of text,
62
you can feed it a little bit of seed text and have it repeatedly play this game predicting what comes next, sampling from the distribution, appending it, and then repeating over and over.
63
Some of you in the know may remember how long before chat GPT came into the scene.
64
This is what early demos of GPT-3 looked like.
65
You would have it autocomplete stories and essays based on an initial snippet.
66
To make a tool like this into a chatbot, the easiest starting point is to have a little bit of text
67
that establishes the setting of a user interacting with a helpful AI assistant, what you would call the system prompt,
68
and then you would use the user's initial question or prompt as the first bit of dialogue,
69
and then you have it start predicting what such a helpful AI assistant would say in response.
70
There is more to say about an added step of training that's required to make this work well, but at a high level this is the general idea. In this chapter you
71
and I are going to expand on the details of what
72
happens at the very beginning of the network at the very end of the network, and I also want to spend a lot of time reviewing some important bits of background knowledge,
73
things that would have been second nature to any machine learning engineer by the time Transformers came around.
74
If you're comfortable with that background knowledge and a little impatient, you could probably feel free to skip to the next chapter, which is going to focus on the attention blocks,
75
generally considered the heart of the Transformer.
76
After that, I want to talk more about these multi-layer perception blocks, how training works, and a number of other details that will have been skipped up to that point.
77
For broader context, these videos are additions to a mini-series about deep learning.
78
And it's okay if you haven't watched the previous ones, I think you can do it out of order.
79
But before diving into Transformers specifically, I do think it's worth making sure that we're on the same page about the basic premise and structure of deep learning.
80
At the risk of stating the obvious, this is one approach to machine learning, which describes any model where you're using data to somehow determine how a model behaves.
81
What I mean by that is, let's say you want a function that takes in an image, and it produces a label describing it, or our example of predicting the next word given a passage of text,
82
or any other task that seems to require some element of intuition and pattern recognition.
83
We almost take this for granted these days, but the idea with machine learning is
84
that rather than trying to explicitly define a procedure for how to do that task in code, which is what people would have done in the earliest days of AI.
85
Instead, you set up a very flexible structure with tunable parameters, like a bunch of knobs and dials,
86
and then, somehow, you use many examples of what the output should look like for a given input to tweak
87
and tune the values of those parameters to mimic this behavior.
88
For example, maybe the simplest form of machine learning is linear regression, where your inputs and your outputs are each single numbers,
89
something like the square footage of a house and its price.
90
And what you want is to find a line of best fit through this data, you know, to predict future house prices.
91
That line is described by two continuous parameters, say the slope and the y-intercept,
92
and the goal of linear regression is to determine those parameters to closely match the data.
93
Needless to say, deep learning models get much more complicated.
94
GPT-3, for example, has not two, but 175 billion parameters.
95
But here's the thing, it's not a given
96
that you can create some giant model with a huge number of parameters without it either grossly overfitting the training data, or being completely intractable to train.
97
Deep learning describes a class of models that in the last couple decades have proven to scale remarkably well.
98
What unifies them is that they all use the same training algorithm, It's called backpropagation, we talked about it in previous chapters.
99
And the context that I want you to have as we go in is
100
that in order for this training algorithm to work well at scale, these models have to follow a certain specific format.
101
And if you know this format going in, it helps to explain many of the choices for how a transformer processes language, which otherwise run the risk of feeling kind of arbitrary.
102
First, whatever kind of model you're making, the input has to be formatted as an array of real numbers.
103
This could simply mean a list of numbers, it could be a two-dimensional array, or very often you deal with higher dimensional arrays, where the general term used is tensor.
104
You often think of that input data as being progressively transformed into many distinct layers,
105
where again each layer is always structured as some kind of array of real numbers, until you get to a final layer which you consider the output.
106
For example, the final layer in our text processing model is a list of numbers representing the probability distribution
107
for all possible next tokens.
108
In deep learning, these model parameters are almost always referred to as weights, and this is because a key feature of these models is
109
that the only way these parameters interact with the data being processed is through weighted sums.
110
You also sprinkle some nonlinear functions throughout, but they won't depend on parameters.
111
Typically though, instead of seeing the weighted sums all naked and written out explicitly like this,
112
you'll instead find them packaged together as various components in a matrix vector product.
113
It amounts to saying the same thing.
114
If you think back to how matrix vector multiplication works, each component in the output looks like a weighted sum.
115
It's just often conceptually cleaner for you and me to think about matrices
116
that are filled with tunable parameters that transform vectors that are drawn from the data being processed.
117
For example, those 175 billion weights in GPT-3 are organized into just under 28,000 distinct matrices.
118
Those matrices in turn fall into eight different categories, and what you
119
and I are going to do is step through each one of those categories to understand what that type does.
120
As we go through, I think it's kind of fun to reference the specific numbers
121
from GPT-3 to count up exactly where those 175 billion come from.
122
Even if nowadays there are bigger and better models, this one has a certain charm as the first large language model to really capture the world's attention outside of ML communities.
123
Also, practically speaking, companies tend to keep much tighter lips around the specific numbers for more modern networks.
124
I just want to set the scene going in
125
that as you peek under the hood to see what happens inside a tool like ChatGPT, almost all of the actual computation looks like matrix vector multiplication.
126
There's a little bit of a risk getting lost in the sea of billions of numbers, but you should draw a very sharp distinction in your mind between the weights of the model
127
which I'll always color in blue or red, and the data being processed, which I'll always color in gray.
128
The weights are the actual brains.
129
They are the things learned during training, and they determine how it behaves.
130
The data being processed simply encodes whatever specific input is fed into the model for a given run, like an example snippet of text.
131
With all of that as foundation, let's dig into the first step of this text processing example, which is to break up the input into little chunks and turn those chunks into vectors.
132
I mentioned how those chunks are called tokens, which might be pieces of words or punctuation, but every now and then in this chapter, and especially in the next one,
133
I'd like to just pretend that it's broken more cleanly into words.
134
Because we humans think in words, this'll just make it much easier to reference little examples and clarify each step.
135
The model has a predefined vocabulary, some list of all possible words, say 50,000 of them, and the first matrix that we'll encounter, known as the embedding matrix,
136
has a single column for each one of these words.
137
These columns are what determines what vector each word turns into in that first step.
138
We label it we, and like all the matrices we see, its values begin random, but they're going to be learned based on data.
139
Turning words into vectors was common practice in machine learning long before transformers, But it's a little weird if you've never seen it before,
140
and it sets the foundation for everything that follows, so let's take a moment to get familiar with it.
141
We often call this embedding a word, which invites you to think of these vectors very geometrically, as points in some high-dimensional space.
142
Visualizing a list of three numbers as coordinates for points in 3D space would be no problem, but word embeddings tend to be much, much higher dimensional.
143
In GPT-3, they have 12,288 dimensions, and as you'll see, it matters to work in a space that has a lot of distinct directions.
144
In the same way that you could take a two-dimensional slice through a 3D space
145
and project all the points onto that slice, for the sake of animating word embeddings that a simple model is giving me,
146
I'm going to do an analogous thing by choosing a three-dimensional slice through this very high-dimensional space
147
and projecting the word vectors down onto that and displaying the results.
148
The big idea here is that as a model tweaks and tunes its weights
149
to determine how exactly words get embedded as vectors during training, it tends to settle on a set of embeddings where directions in the space have a kind of semantic meaning.
150
For the simple word-to-vector model I'm running here, if I run a search for all the words whose embeddings are closest to that of tower,
151
you'll notice how they all seem to give very similar tower-ish vibes.
152
And if you want to pull up some Python and play along at home, this is the specific model that I'm using to make the animations.
153
It's not a transformer, but it's enough to illustrate the idea that directions in the space can carry semantic meaning.
154
A very classic example of this is how if you take the difference between the vectors for woman and man,
155
something you would visualize as a little vector in the space connecting the tip of one to the tip of the other,
156
it's very similar to the difference between king and queen.
157
So let's say you didn't know the word for a female monarch.
158
You could find it by taking King,
159
adding this woman-man direction, and searching for the embeddings closest to that point. At least, kind of.
160
Despite this being a classic example, for the model I'm playing with, the true embedding of Queen is actually a little farther off than this would suggest,
161
presumably because the way Queen is used in training data is not merely a feminine version of King.
162
When I played around, family relations seemed to illustrate the idea much better.
163
The point is, it looks like during training the model found it advantageous to choose embeddings such
164
that one direction in this space encodes gender information.
165
Another example is that if you take the embedding of Italy, and you subtract the embedding of Germany, and then you add that to the embedding of Hitler,
166
you get something very close to the embedding of Mussolini.
167
It's as if the model learned to associate some directions with Italian-ness, and others with World War II Axis leaders.
168
Maybe my favorite example in this vein is how in some models, if you take the difference between Germany and Japan and you add it to sushi,
169
you end up very close to Broadwurst.
170
Also in playing this game of finding nearest neighbors, I was very pleased to see how close Cat was to both Beast and Monster.
171
One bit of mathematical intuition that's helpful to have in mind, especially for the next chapter,
172
is how the dot product of two vectors can be thought of as a way to measure how well they align.
173
Computationally, dot products involve multiplying all the corresponding components and then adding the results,
174
which is good since so much of our computation has to look like weighted sums.
175
Geometrically, the dot product is positive when vectors point in similar directions, it's zero if they're perpendicular,
176
and it's negative whenever they point in opposite directions.
177
For example, let's say you were playing with this model
178
and you hypothesize that the embedding of cats minus cat might represent a sort of plurality direction in this space.
179
To test this, I'm going to take this vector
180
and compute its dot product against the embeddings of certain singular nouns
181
and compare it to the dot products with the corresponding plural nouns.
182
If you play around with this, you'll notice that the plural ones do indeed seem to consistently give higher values than the singular ones, indicating that they align more with this direction.
183
It's also fun how if you take this dot product with the embeddings of the words 1, 2, 3, and so on,
184
they give increasing values, so it's as if we can quantitatively measure how plural the model finds a given word.
185
Again, the specifics for how words get embedded is learned using data.
186
This embedding matrix, whose columns tell us what happens to each word, is the first pile of weights in our model.
187
And using the GPT-3 numbers, the vocabulary size specifically is 50,257, and again, technically this consists not of words per se, but of tokens.
188
And the embedding dimension is 12,288.
189
Multiplying those tells us this consists of about 617 million weights.
190
Let's go ahead and add this to a running tally, remembering that by the end we should count up to 175 billion.
191
In the case of transformers, you really want to think of the vectors in this embedding space as not merely representing individual words.
192
For one thing, they also encode information about the position of that word, which we'll talk about later,
193
but more importantly, you should think of them as having the capacity to soak in context.
194
A vector that started its life as the embedding of the word king, for example, might progressively get tugged and pulled by various blocks in this network,
195
so that by the end it points in a much more specific and nuanced direction
196
that somehow encodes that it was a king who lived in Scotland, and who had achieved his post after murdering the previous king, and who's being described in Shakespearean language.
197
Think about your own understanding of a given word.
198
The meaning of that word is clearly informed by the surroundings, and sometimes this includes context from a long distance away.
199
So in putting together a model that has the ability to predict what word comes next, the goal is to somehow empower it to incorporate context efficiently.
200
To be clear, in that very first step, when you create the array of vectors based on the input text, each one of those is simply plucked out of the embedding matrix,
201
so initially each one can only encode the meaning of a single word without any input from its surroundings.
202
But you should think of the primary goal of this network that it flows through, as being to enable each one of those vectors to soak up a meaning that's much more rich
203
and specific than what mere individual words could represent.
204
The network can only process a fixed number of vectors at a time, known as its context size.
205
For GPT-3, it was trained with a context size of 2048, so the data flowing through the network always looks like this array of 2048 columns,
206
each of which has 12,000 dimensions.
207
This context size limits how much text the transformer can incorporate when it's making a prediction of the next word.
208
This is why long conversations with certain chatbots, like the early versions of ChatGPT, often gave the feeling of the bot kind of losing the thread of conversation as you continued too long.
209
We'll go into the details of attention in due time, but skipping ahead, I want to talk for a minute about what happens at the very end.
210
Remember, the desired output is a probability distribution over all tokens that might come next.
211
For example, if the very last word is professor, and the context includes words like Harry Potter, and immediately preceding we see least favorite teacher,
212
and also if you give me some leeway by letting me pretend that tokens simply look like full words,
213
then a well-trained network that had built up knowledge of Harry Potter would presumably assign a high number to the word Snape.
214
This involves two different steps.
215
The first one is to use another matrix
216
that maps the very last vector in that context to a list of 50,000 values, one for each token in the vocabulary.
217
Then there's a function that normalizes this into a probability distribution.
218
It's called Softmax, and we'll talk more about it in just a second.
219
But before that, it might seem a little bit weird to only use this last embedding to make a prediction, when after all, in that last step,
220
there are thousands of other vectors in the layer just sitting there with their own context-rich meanings.
221
This has to do with the fact that in the training process, it turns out to be much more efficient
222
if you use each one of those vectors in the final
223
layer to simultaneously make a prediction for what would come immediately after it.
224
There's a lot more to be said about training later on, but I just want to call that out right now.
225
This matrix is called the unembedding matrix, and we give it the label W U.
226
Again, like all the weight matrices we see, its entries begin at random, but they are learned during the training process.
227
Keeping score on our total parameter count, this unembedding matrix has one row for each word in the vocabulary, and each row has the same number of elements as the embedding dimension.
228
It's very similar to the embedding matrix, just with the order swapped, so it adds another 617 million parameters to the network,
229
meaning our count so far is a little over a billion.
230
A small, but not wholly insignificant fraction of the 175 billion that we'll end up with in total.
231
As the very last mini-lesson for this chapter, I want to talk more about the softmax function, since it makes another appearance for us once we dive into the attention blocks.
232
The idea is that if you want a sequence of numbers to act as a probability distribution, say, a distribution over all possible next words, then each value has to be between 0 and 1,
233
and you also need all of them to add up to 1.
234
However, if you're playing the deep learning game, where everything you do looks like matrix vector multiplication,
235
the outputs that you get by default don't abide by this at all.
236
The values are often negative, or much bigger than 1, and they almost certainly don't add up to 1.
237
Softmax is the standard way to turn an arbitrary list of numbers into a valid distribution in such a way
238
that the largest values end up closest to 1, and the smaller values end up very close to 0.
239
That's all you really need to know.
240
But if you're curious, the way that it works is to first raise e to the power of each of the numbers, which means you now have a list of positive values,
241
and then you can take the sum of all those positive values and divide each term by that sum, which normalizes it into a list that adds up to 1.
242
You'll notice that if one of the numbers in the input is meaningfully bigger than the rest, then in the output the corresponding term dominates the distribution,
243
so if you were sampling from it you'd almost certainly just be picking the maximizing input.
244
But it's softer than just picking the max in the sense that when other values are similarly large, they also get meaningful weight in the distribution,
245
and everything changes continuously as you continuously vary the inputs.
246
In some situations, like when ChatGPT is using this distribution to create a next word,
247
There's room for a little bit of extra fun by adding a little extra spice into this function, with a constant t thrown into the denominator of those exponents.
248
We call it the temperature, since it vaguely resembles the role of temperature in certain thermodynamics equations, and the effect is that when t is larger,
249
you give more weight to the lower values, meaning the distribution is a little bit more uniform, and if t is smaller, then the bigger values will dominate more aggressively.
250
Where in the extreme, setting t equal to zero means all of the weight goes to that maximum value.
251
For example, I'll have GPT-3 generate a story with the seed text, once upon a time there was A, but I'm going to use different temperatures in each case.
252
Temperature zero means that it always goes with the most predictable word, and what you get ends up being kind of a trite derivative of Goldilocks.
253
A higher temperature gives it a chance to choose less likely words, but it comes with a risk.
254
In this case, the story starts out a bit more originally, about a young web artist from South Korea, but it quickly degenerates into nonsense.
255
Technically speaking, the API doesn't actually let you pick a temperature bigger than two.
256
There is no mathematical reason for this, it's just an arbitrary constraint imposed, I suppose, to keep their tool from being seen generating things that are too nonsensical.
257
So if you're curious the way this animation is actually working, is I'm taking the 20 most probable NEXT tokens that GPT-3 generates, which seems to be the maximum they'll give me,
258
and then I tweak the probabilities based on an exponent of 1 fifth.
259
As another bit of jargon, in the same way that you might call the components of the output of this function probabilities, people often refer to the inputs as logits.
260
Or some people say logits, some people say logits, I'm going to say logits.
261
So for instance, when you feed in some text, you have all these word embeddings flow through the network, and you do this final multiplication with the unembedding matrix,
262
machine learning people would refer to the components in that raw unnormalized output as the logits for the next word prediction.
263
A lot of the goal with this chapter was to lay the foundations for understanding the attention mechanism, Karate Kid wax on wax off style.
264
You see, if you have a strong intuition for word embeddings, for softmax, for how dot products measure similarity,
265
and also the underlying premise that most of the calculations have to look like matrix multiplication with matrices full of tunable parameters.
266
Then, understanding the attention mechanism, this cornerstone piece and the whole modern boom in AI, should be relatively smooth.
267
For that, come join me in the next chapter.
268
As I'm publishing this, a draft of that next chapter is available for review by Patreon supporters.
269
A final version should be up in public in a week or two, it usually depends on how much I end up changing based on that review.
270
In the meantime, if you want to dive into attention and if you want to help the channel out a little bit, it's there waiting.
📺 同频道
✨ 推荐视频
关于本课
您正在使用跟读技巧通过视频"Transformers, the tech behind LLMs | Deep Learning Chapter 5"练习英语口语和发音。
每天练习15到30分钟,将显著提高您的英语流利度和发音准确度。
什么是跟读法?
跟读法 (Shadowing) 是一种有科学依据的语言学习技巧,最初开发用于专业口译员的培训,并由多语言者Alexander Arguelles博士普及。这个方法简单而强大:您在听英语母语原声的同时立即大声重复——就像是一个延迟1-2秒紧跟说话者的影子。与被动听力或语法练习不同,跟读法强迫您的大脑和口腔肌肉同时处理并模仿真实的讲话模式。研究表明它能显着提高发音准确性,语调,节奏,连读,听力理解和口语流利度——使其成为雅思口语备考和真实英语交流最有效的方法之一。






















![You’ll stop using ChatGPT after listening to this | Jonathan Pageau [ARC 2026]](https://img.youtube.com/vi/yZUuKzDQSsI/mqdefault.jpg)


