Shadowing Practice: How might LLMs store facts | Deep Learning Chapter 7 - Learn English Speaking with Video

Creating lesson...
1
If you feed a large language model the phrase, Michael Jordan plays the sport of blank, and you have it predict what comes next, and it correctly predicts basketball,
2
this would suggest that somewhere, inside its hundreds of billions of parameters, it's baked in knowledge about a specific person and his specific sport.
3
And I think in general, anyone who's played around with one of these models has the clear sense that it's memorized tons and tons of facts.
4
So a reasonable question you could ask is, how exactly does that work? work, and where do those facts live?
5
Last December, a few researchers from Google DeepMind posted about work on this question, and they were using this specific example of matching athletes to their sports.
6
And although a full mechanistic understanding of how facts are stored remains unsolved, they had some interesting partial results,
7
including the very general high -level conclusion that the facts seem to live inside a specific part of these networks,
8
known fancifully as the Multilayer Perceptrons, or MLPs for short.
9
In the last couple of chapters, you and I have been digging into the details behind Transformers, the architecture underlying large language models, and also underlying a lot of other modern AI.
10
In the most recent chapter, we were focusing on a piece called Attention.
11
And the next step, for you and me, is to dig into the details of what happens inside these multi -layer perceptrons, which make up the other big portion of the network.
12
The computation here is actually relatively simple, especially when you compare it to a tension.
13
It boils down essentially to a pair of matrix multiplications with a simple something in between.
14
However, interpreting what these computations are doing is exceedingly challenging.
15
Our main goal here is to step through the computations and make them memorable, but I'd like to do it in the context of showing a specific example of how one of these blocks could,
16
at least in principle, store a concrete fact.
17
Specifically, it'll be storing the fact that Michael Jordan plays basketball.
18
I should mention the layout here is inspired by a conversation I had with one of those DeepMind researchers, Neil Nanda.
19
For the most part, I will assume that you've either watched the last two chapters, or otherwise you have a basic sense for what a transformer is.
20
But refreshers never hurt, so here's the quick reminder of the overall flow.
21
You and I have been studying a model that's trained to take in a piece of text and predict what comes next.
22
That input text is first broken into a bunch of tokens, which means little chunks that are typically words or little pieces of words,
23
and each token is associated with a high -dimensional vector, which is to say a long list of numbers.
24
This sequence of vectors then repeatedly passes through two kinds of operation.
25
Attention, which allows the vectors to pass information between one another, and then the multi -layer perceptrons, the thing that we're to dig into today.
26
And also there's a certain normalization step in between.
27
After the sequence of vectors has flowed through many many different iterations of both of these blocks,
28
by the end, the hope is that each vector has soaked up enough information, both from the context ,
29
and also from the general knowledge that was baked into the model weights through training,
30
that it can be used to make a prediction of what token comes next.
31
One of the key ideas that I want you to have in your mind is
32
that all of these vectors live in a very very high dimensional space, and when you think about that space, different directions can encode different kinds of meaning.
33
So a very classic example that I like to refer back to is how
34
if you look at the embedding of woman and subtract the embedding of man, and you take that little step and you add it to another masculine noun,
35
something like uncle, you land somewhere very very close to the corresponding feminine noun.
36
In this sense, this particular direction encodes gender information.
37
The idea is
38
that many other distinct directions in this super high -dimensional space could correspond to other features
39
that the model might want to represent.
40
In a transformer, these vectors don't merely encode the meaning of a single word, though.
41
As they flow through the network, they imbibe a much richer meaning based on all the context around them, and also based on the model's knowledge.
42
Ultimately, each one needs to encode something far far beyond the meaning of a single word, since it needs to be sufficient to predict what will come next.
43
We've already seen how attention blocks let you incorporate context, but a majority of the model parameters actually live inside the MLP blocks,
44
and one thought for what they might be doing is that they offer extra capacity to store facts.
45
Like I said, the lesson here is going to center on the concrete toy example of how exactly it could store the fact
46
that Michael Jordan plays basketball.
47
Now this toy example is going to require that you and I make a couple of assumptions about that high dimensional space.
48
First, we'll suppose that one of the directions represents the idea of a first name Michael.
49
And then another, nearly perpendicular direction, represents the idea of the last name Jordan.
50
And then yet a third direction will represent the idea of basketball.
51
So specifically what I mean by this is
52
if you look in the network and you pluck out one of the vectors being processed, if its dot product with this first name Michael direction is 1,
53
that's what it would mean for the vector to be encoding the idea of a person with that first name.
54
Otherwise, that dot product would be zero or negative, meaning the vector doesn't really align with that direction.
55
And for simplicity, let's completely ignore the very reasonable question of what it might mean if that dot product was bigger than one.
56
Similarly, its dot product with these other directions would tell you whether it represents the last name Jordan, or basketball.
57
So let's say a vector is meant to represent the full name, Michael Jordan, then its dot product with both of these directions would have to be one.
58
Since the text, Michael Jordan, spans two different tokens, this would also mean we have to assume
59
that an earlier attention block has successfully passed information to the second of these two vectors
60
so as to ensure that it can encode both names.
61
With all of those as the assumptions, let's now dive into the meat of the lesson.
62
What happens inside a multilayer perceptron?
63
You might think of this sequence of vectors flowing into the block.
64
And remember, each vector was originally associated with one of the tokens from the input text.
65
What's going to happen is that each individual vector from that sequence goes through a short series of operations, we'll unpack them in just a moment,
66
and at the end we'll get another vector with the same dimension.
67
That other vector is going to get added to the original one that flowed in, and that sum is the result flowing out.
68
This sequence of operations is something you apply to every vector
69
in the sequence associated with every token in the input and it all happens in parallel.
70
In particular, the vectors don't talk to each other in this step.
71
They're all kind of doing their own thing.
72
And for you and me, that actually makes it a lot simpler because it means
73
if we understand what happens to just one of the vectors through this block, we effectively understand what happens to all of them.
74
When I say this block is going to encode the fact that Michael Jordan plays basketball, what I mean is that if a vector flows in that encodes first name Michael and last name Jordan,
75
then this sequence of computations will produce something that includes that direction basketball, which is what will add on to the vector in that position.
76
The first step of this process looks like multiplying that vector by a very big matrix, no surprises there, this is deep learning.
77
And this matrix, like all of the other ones we've seen, is filled with model parameters that are learned from data, which you might think of as a bunch of knobs
78
and dials that get tweaked and tuned to determine what the model behavior is.
79
One nice way to think about matrix multiplication is to imagine each row of that matrix as being its own vector,
80
and taking a bunch of dot products between those rows and the vector being processed, which I'll label as E for embedding.
81
For example, suppose that very first row happened to equal this first name Michael direction that we're presuming exists.
82
That would mean that the first component in this output, this dot product right here, would be 1 if that vector encodes the first name Michael,
83
and 0 or negative otherwise.
84
Even more fun, take a moment to think about what it would mean
85
if that first row was this first name Michael plus last name Jordan direction.
86
And for simplicity, let me go ahead and write that down as m plus j.
87
Then taking a dot product with this embedding e, things distribute really nicely, so it looks like m dot e plus j dot e,
88
and notice how that means the ultimate value would be 2 if the vector encodes the full name, Michael Jordan, and otherwise it would be 1 or something smaller than 1.
89
And that's just one row in this matrix.
90
You might think of all of the other rows as, in parallel, asking some other kinds of questions, probing at some other sorts of features of the vector being processed.
91
Very often this step also involves adding another vector to the output, which is full of model parameters learned from data.
92
This other vector is known as the bias.
93
For our example, I want you to imagine that the value of this bias in that very first component is negative 1,
94
meaning our final output looks like that relevant dot product, but minus 1.
95
You might very reasonably ask why I would want you to assume that the model has learned this, and in a moment you'll see why it's very clean
96
and nice if we have a value here which is positive
97
if and only if our vector encodes the full name Michael Jordan, and otherwise it's 0 or negative.
98
The total number of rows in this matrix, which is something like the number of questions being asked, in the case of GPT -3, whose numbers we've been following,
99
is just under 50 ,000.
100
In fact, it's exactly four times the number of dimensions in this embedding space.
101
That's a design choice, you could make it more, you could make it less, but having a clean multiple tends to be friendly for hardware.
102
Since this matrix full of weights maps us into a higher dimensional space, I'm going to give it the shorthand.. w up.
103
I'll continue labeling the vector we're processing as e, and let's label this bias vector as b up and put that all back down in the diagram.
104
At this point, a problem is that this operation is purely linear, but language is a very non -linear process.
105
If the entry that we're measuring is high for Michael plus Jordan, it would also necessarily be somewhat triggered by Michael plus Phelps, and also Alexis plus Jordan,
106
despite those being unrelated conceptually.
107
What you really want is a simple yes or no for the full name.
108
So the next step is to pass this large intermediate vector through a very simple non -linear function.
109
A common choice is one that takes all of the negative values and maps them to zero, and leaves all of the positive values unchanged.
110
And continuing with the deep learning tradition of overly fancy names, this very simple function is often called the rectified linear unit, or or ReLU for short,
111
here's what the graph looks like.
112
So taking our imagined example where this first entry of the intermediate vector is 1
113
if and only if the full name is Michael Jordan and 0 or negative otherwise, after you pass it through the ReLU,
114
you end up with a very clean value where all of the 0 and negative values just get clipped to 0.
115
So this output would be 1 for the full name Michael Jordan and 0 otherwise.
116
In other words, it very directly mimics the behavior of an AND gate.
117
Often models will use a slightly modified function that's called the JLU, which has the same basic shape, it's just a bit smoother.
118
But for our purposes, it's a little bit cleaner if we only think about the ReLU.
119
Also, when you hear people refer to the neurons of a transformer, they're talking about these values right here.
120
Whenever you see that common neural network picture with a layer of dots
121
and a bunch of lines connecting to the previous layer, which we had earlier in this series, that's typically meant to convey this combination of a linear step,
122
a matrix multiplication, followed by some simple term -wise non -linear function like a ReLU.
123
You would say that this neuron is active whenever this value is positive, and that it's inactive if that value is zero.
124
The next step looks very similar to the first one.
125
You multiply by a very large matrix, and you add on a certain bias term.
126
In this case, the number of dimensions in the output is back down to the size of that embedding space, so I'm going to go ahead and call this the down projection matrix.
127
And this time, instead of thinking of things row by row, it's actually nicer to think of it column by column.
128
You see, another way
129
that you can hold matrix multiplication in your head is to imagine taking each column of the matrix
130
and multiplying it by the corresponding term in the vector that it's processing and adding together all of those rescaled columns.
131
The reason it's nicer to think about this way is because here the columns have the same dimension as the embedding space, so we can think of them as directions in that space.
132
For instance, we will imagine that the model has learned to make
133
that first column into this basketball direction that we suppose exists.
134
What that would mean is that when the relevant neuron in that first position is active, we'll be adding this column to the final result.
135
But if that neuron was inactive, if that number was zero, then this would have no effect.
136
And it doesn't just have to be basketball.
137
The model could also bake into this column many other features
138
that it wants to associate with something that has the full name Michael Jordan.
139
And at the same time, all of the other columns in this matrix are telling you what will be added to the final result
140
if the corresponding neuron is active.
141
And if you have a bias in this case, it's something that you're just adding every single time, regardless of the neuron values.
142
You might wonder what's that doing.
143
As with all parameter -filled objects here, it's kind of hard to say exactly.
144
Maybe there's some bookkeeping that the network needs to do, but you can feel free to ignore it for now.
145
Making our notation a little more compact again, I'll call this big matrix W down, and similarly call that bias vector B down, and put that back into our diagram.
146
Like I previewed earlier, what you do with this final result is add it to the vector
147
that flowed into the block at that position, and that gets you this final result.
148
So for example, if the vector flowing in encoded both first name Michael and last name Jordan,
149
then because this sequence of operations will trigger that AND gate, it will add ON the basketball direction so what pops out will encode all of those together.
150
And remember, this is a process happening to every one of those vectors in parallel.
151
In particular, taking the GPT -3 numbers, it means that this block doesn't just have 50 ,000 neurons in it,
152
it has 50 ,000 times the number of tokens in the input.
153
So that is the entire operation, two matrix products each with a bias added, and a simple clipping function in between.
154
Any of you who watched the earlier videos of the series
155
will recognize this structure as the most basic kind of neural network that we studied there.
156
In that example, it was trained to recognize handwritten digits.
157
Over here, in the context of a transformer for a large language model, this is one piece in a larger architecture,
158
and any attempt to interpret what exactly it's doing is heavily
159
intertwined with the idea of encoding information into vectors of a high -dimensional embedding space.
160
That is the core lesson, but I do want to step back and reflect on two different things.
161
The first of which is a kind of bookkeeping, and the second of which involves a very thought provoking fact about higher dimensions
162
that I actually didn't know until I dug into Transformers.
163
In the last two chapters, you and I started counting up the total number of parameters in GPT -3 and seeing exactly where they live, so let's quickly finish up the game here.
164
I already mentioned how this up projection matrix has just under 50 ,000 rows, and that each row matches the size of the embedding space,
165
which for GPT -3 is 12 ,288.
166
Multiplying those together, it gives us 604 million parameters just for that matrix.
167
And the down projection has the same number of parameters, just with a transposed shape.
168
So together, they give about 1 .2 billion parameters.
169
The bias vector also accounts for a couple more parameters, but it's a trivial proportion of the total, so I'm not even going to show it.
170
In GPT -3, this sequence of embedding vectors flows through not one,
171
but 96 distinct MLPs, so the total number of parameters devoted to all of these blocks adds up to about 116 billion.
172
This is around two -thirds of the total parameters in the network, and when you add it to everything that we had before for the attention blocks, the embedding, and the unembedding,
173
you do indeed get that grand total of 175 billion as advertised.
174
It's probably worth mentioning there's another set of parameters associated with those normalization steps that this explanation has skipped over, but like the bias vector,
175
they account for a very trivial proportion of the total.
176
As to that second point of reflection, you might be wondering if this central toy example we've been spending
177
so much time on reflects how facts are actually stored in real large language models.
178
It is true that the rows of that first matrix can be thought of as directions in this embedding space,
179
and that means the activation of each neuron tells you how much a given vector aligns with some specific direction.
180
It's also true that the columns of
181
that second matrix tell you what will be added to the result if that neuron is active.
182
Both of those are just mathematical facts.
183
However, the evidence does suggest that individual neurons very rarely represent a single clean feature like Michael Jordan.
184
And there may actually be a very good reason this is the case, related to an idea floating around interpretability researchers these days known as superposition.
185
This is a hypothesis
186
that might help to explain both why the models are especially hard to interpret and also why they scale surprisingly well.
187
The basic idea is that if you have an n -dimensional space
188
and you want to represent a bunch of different features using directions that are all perpendicular to one another in that space, you know, that way if you add a component in one direction,
189
it doesn't influence any of the other directions, then the maximum number of vectors you can fit is only n, the number of dimensions.
190
To a mathematician, actually, this is the definition of dimension.
191
But where it gets interesting is if you relax that constraint a little bit and you tolerate some noise.
192
Say you allow those features to be represented by vectors that aren't exactly perpendicular, they're just nearly perpendicular, maybe between 89 and 91 degrees apart.
193
If we were in two or three dimensions, this makes no difference.
194
That gives you hardly any extra wiggle room to fit more vectors in, which makes it all the more counterintuitive that for higher dimensions, the answer changes dramatically.
195
I can give you a really quick
196
and dirty illustration of this using some scrappy python that's going to create a list of 100 dimensional vectors,
197
each one initialized randomly, and this list is going to contain 10 ,000 distinct vectors, so 100 times as many vectors as there are dimensions.
198
This plot right here shows the distribution of angles between pairs of these vectors.
199
So because they started at random, those angles could be anything from 0 to 180 degrees, but you'll notice that already, even just for random vectors,
200
there's this heavy bias for things to be closer to 90 degrees.
201
Then what I'm going to do is run a certain optimization process
202
that iteratively nudges all of these vectors so that they try to become more perpendicular to one another.
203
After repeating this many different times, here's what the distribution of angles looks like.
204
We have to actually zoom in on it here
205
because all of the possible angles between pairs of vectors sit inside this narrow range between 89 and 91 degrees.
206
In general, a consequence of something known as the Johnson -Lindenstrauss Lemma is
207
that the number of vectors you can cram into a space
208
that are nearly perpendicular like this grows exponentially with the number of dimensions.
209
This is very significant for large language models, which might benefit from associating independent ideas with nearly perpendicular directions.
210
It means that it's possible for it to store many, many more ideas than there are dimensions in the space that it's allotted.
211
This might partially explain why model performance seems to scale so well with size.
212
A space that has 10 times as many dimensions can store way, way more than 10 times as many independent ideas.
213
And this is relevant not just to that embedding space where the vectors flowing through the model live,
214
but also to that vector full of neurons in the middle of that multilayer perceptron that we just studied.
215
That is to say, at the sizes of GPT -3, it might not just be probing at 50 ,000 features,
216
but if it instead leveraged this enormous added capacity by using nearly perpendicular directions of the space, it could be probing at many,
217
many more features of the vector being processed.
218
But if it was doing that, what it means is that individual features aren't going to be visible as a single neuron lighting up.
219
It would have to look like some specific combination of neurons instead, a superposition.
220
For any of you curious to learn more, a key relevant search term here is sparse autoencoder, which is a tool
221
that some of the interpretability people use to try to extract what the true features are even
222
if they're very superimposed on all these neurons.
223
I'll link to a couple really great anthropic posts all about this.
224
At this point, we haven't touched every detail of a transformer, but you and I have hit the most important points.
225
The main thing that I want to cover in a next chapter is the training process.
226
On the one hand, the short answer for how training works is that it's all backpropagation, and we covered backpropagation in a separate context with earlier chapters in the series.
227
But there is more to discuss, like the specific cost function used for language models, the idea of fine -tuning using reinforcement learning with human feedback,
228
and the notion of scaling laws.
229
Quick note for the active followers among you, there are a number of non -machine learning related videos
230
that I'm excited to sink my teeth into before I make
231
that next chapter so it might be a while but I do promise it'll come in due time.

Learn English with Videos: Dive into LLMs and Boost Your Skills

Watching videos about complex topics like large language models (LLMs) isn’t just for tech fans—it’s a great way to improve English pronunciation and practice listening to natural, conversational speech. The video we’re exploring breaks down how LLMs store facts, using simple examples like Michael Jordan and basketball. It’s perfect for IELTS speaking practice because it mixes everyday vocabulary with technical terms, helping you get comfortable with diverse language.

Top 5 Phrases for Daily Communication (and Beyond!)

  • "Baked in knowledge" – Meaning: Information that’s deeply embedded. Example: "My phone’s baked-in knowledge of my habits makes it easy to use."
  • Imbibe a richer meaning – Meaning: To absorb deeper understanding. Example: "Reading books helps you imbibe a richer meaning of the world."
  • High-dimensional space – (Technical but useful!) Meaning: A complex system with many variables. Example: "Social media exists in a high-dimensional space of user interactions."
  • Pluck out – Meaning: To pick or select. Example: "I plucked out the best idea from the meeting notes."
  • Tweaked and tuned – Meaning: Adjusted carefully. Example: "The recipe was tweaked and tuned until it tasted perfect."

Step-by-Step Shadowing Guide for This Video

The shadowing technique is a fantastic way to improve English speaking practice. Here’s how to use it with this video:

  1. Listen once, then pause: Watch a 10-second clip, then repeat it out loud immediately. Focus on matching the speaker’s rhythm and stress (e.g., "multi-layer perceptrons" has stress on "multi" and "perceptrons").
  2. Break down tricky terms: Words like "transformers" or "embedding" might feel hard. Say them slowly, emphasizing each syllable, then speed up. This helps with improve English pronunciation.
  3. Copy the tone: The speaker uses a conversational tone, even with complex ideas. Mimic their casual yet clear delivery—this is key for IELTS speaking practice.
  4. Practice with context: After shadowing, summarize the clip in your own words. This connects the language to meaning, making it easier to remember.

By using this video for learn English with videos practice, you’ll not only understand LLMs but also build confidence in speaking English fluently. Keep shadowing, and you’ll notice progress in no time!

What is the Shadowing Technique?

Shadowing is a science-backed language learning technique originally developed for professional interpreter training and popularized by polyglot Dr. Alexander Arguelles. The method is simple but powerful: you listen to native English audio and immediately repeat it out loud — like a shadow following the speaker with just a 1–2 second delay. Unlike passive listening or grammar drills, shadowing forces your brain and mouth muscles to simultaneously process and reproduce real speech patterns. Research shows it significantly improves pronunciation accuracy, intonation, rhythm, connected speech, listening comprehension, and speaking fluency — making it one of the most effective methods for IELTS Speaking preparation and real-world English communication.

Shadowing technique: read the full step-by-step guide →