Shadowing Practice: MIT Introduction to Deep Learning | 6.S191 - Learn English Speaking with Video

Creando lección...
1
ALEXANDER AMINI - Good morning, everyone.
2
Or actually, good afternoon, everyone.
3
My name is Alexander Amini, and thank you all for joining us today in MIT 6S191 Introduction to Deep Learning.
4
We're super excited to welcome you to this class.
5
My name is Alexander.
6
I'll be your instructor along with Ava, who you'll hear from later today as well with the second lecture.
7
This is a one -week boot camp on everything deep learning.
8
So you'll go from the very beginnings of neural networks all
9
the way to some of the most recent LLM advances that we've had, even up to last week, even.
10
We'll be covering the field.
11
The pace of this field is really, truly remarkable, right?
12
And that's one thing
13
that we'll keep touching back on throughout the course in a lot of detail
14
that we want to see actually not just what the current state of the art is, but also remember how we got there.
15
And I think it's really easy for us to become desensitized
16
actually with a lot of the state of art with AI
17
because there are so many amazing things happening and we really are on this exponential trend.
18
It's really hard for us to actually remember where we were just a few years ago.
19
So I think what better way to do that than to see it.
20
So exactly one decade ago, we started teaching the class.
21
We created this class nine years ago.
22
We started teaching it nine years ago.
23
One decade ago, in 2015, this was the state of VR in terms of facial image generation.
24
So generating facial images.
25
Fast forward a couple years to 2018, you can already see a massive leap in quality
26
and then fast forward a couple more years and you see
27
that extend into a new dimension of time into generating videos as well.
28
In fact, this was a video that we generated for the class in 2020 to actually introduce the class.
29
If you guys haven't seen this video, you should watch it online.
30
It's on YouTube.
31
It went a bit viral a few years ago when we started the class.
32
Now, in the past few years, we've seen this revolution expand also into natural language as well.
33
So in 2022, we had ChatGPT3 released and launched to the public, followed shortly by GPT4 in 2023.
34
Now, the jump between 3 .5
35
and 4 was oftentimes regarded as this massive leap in AI capabilities when we went from GPT3 .5 to GPT4,
36
because really everybody felt that the capabilities were vastly improved between these two models
37
and everybody could feel it even non -technical folks could really
38
feel the capabilities come through now what you couldn't really see
39
or what you couldn't feel was what was happening behind all of those giant models right
40
so gpt4 was this incredible model
41
but this is what you don't see powering gpt4 are these massive and five of course as well,
42
are these massive data centers, GPU data centers.
43
They cost hundreds of millions of dollars to build and tons of infrastructure to maintain and to manage these different data centers.
44
Now, these are for large language models.
45
We've also had this revolution of AI computing capability on small form factors as well.
46
So just like we have large language models powered in the cloud,
47
we've also started to to see small language models go directly down onto edge devices.
48
Now traditionally this gap between large language models and small language models has been very large in a capability sense.
49
We couldn't really accomplish huge jumps in capability in small language world that we've seen in the large language world.
50
But let me show you actually what's happening right now.
51
So I'm gonna switch over my screen to this screen
52
which is actually running here
53
so this is actually live on this phone running right here you can see everything that's happening on the phone
54
so what I'll do is I'll show you actually a language model running entirely on this phone right now
55
so we can ask you questions this should look very familiar to what you have seen with GPT's
56
so how are you doing we can ask you very simple questions to start now the thing
57
that you'll notice here
58
if I swipe down from the top this model is running
59
entirely on airplane mode there's no Wi -Fi connected to this
60
phone this is entirely offline everything that's being computed is running
61
on the device itself now of course we can ask it more complex questions as well
62
so here's a question about finding the zeros in this equation it looks simple at first glance
63
but actually if you read it it's actually asking about it in Z7 space.
64
Does anybody know what Z7 space is?
65
Anybody?
66
Yes, exactly, exactly.
67
So it's in module seven, the finite space of a module seven.
68
We can ask it this question.
69
It can immediately know the answer.
70
You can check the answer.
71
It is indeed correct.
72
We can also ask it to maybe like give me more explanation.
73
Right.
74
Okay.
75
So it's able to do this now.
76
Now, let me go back out and let me actually show you another demonstration.
77
Now, these models are also permeating beyond text language as well.
78
We can ask questions directly in audio language.
79
Now, these are models that are not converting from text to audio, audio to text.
80
These are pure audio models that take audio in and directly output audio tokens out.
81
So I can ask you the question like, can you give me a step -by -step guide on how to give a really compelling talk?
82
Sure.
83
Start by clarifying your main message and who you're speaking to, then sketch a simple outline with an intro, a few key points, and a strong conclusion.
84
Next, craft a few opening and closing lines that grab attention and leave a lasting impression.
85
And practice your delivery so you can adjust your pace in body language.
86
Finally, rehearse your talk, get feedback, and tweaks any weak spots before the event.
87
Right.
88
So what have we seen?
89
Right.
90
So this is an end -to -end audio model.
91
It means that the latency is extremely low is what you'll probably notice there.
92
Because number one, you don't send anything to the cloud.
93
So you don't have to wait for internet times.
94
Everything runs on device.
95
And you're directly processing audio.
96
So it's extremely responsive.
97
These are models that can actually understand your emotion.
98
Because they don't convert audio to text.
99
When you convert to text, you lose emotion.
100
So you directly capture the emotion in the speech, and you can directly convey the emotion back to the user.
101
Now, this was a model actually developed by Liquid AI.
102
It released about one week ago.
103
It's multimodal.
104
It can process all of these different modalities, text, audio, and vision.
105
But there have been a ton of these model releases in the past few years, actually, since GPT first launched.
106
And this problem that traditionally was having between large language models, the gap between large language models and small language models has always been around, right?
107
This was actually a very special release for the community because of this context on the left -hand side.
108
So I'll read it out just for everybody to see.
109
So it's pretty crazy when you think about it.
110
One of the biggest wow moments of GPT was the update from 3 .5 to 4.
111
And now this open source model that anybody can download
112
and put on their phone with only two billion parameters is
113
better than the original GPT -4 model in terms of every benchmark basically that comes out, right?
114
And this is a model, again, that you can run entirely on your phone.
115
It's open source.
116
You can fine -tune it on your own prem, right?
117
So that's the pace of the field.
118
That's why this class is happening at such an exciting time.
119
So let's dive into it, right?
120
This is what we'll learn in the class.
121
Let's dive in by, first of all, laying some foundation of what exactly is deep learning and to do
122
that we need to first understand what is intelligence
123
because deep learning at the heart of deep learning is ai right
124
so what is ai first of all ai is the practice
125
of building artificial algorithms to process information in order to inform
126
some future decisions right at its core that's what we call intelligence in fact even
127
And artificial intelligence is just the ability for a machine to do that same ability, process information to inform a future decision.
128
Now, machine learning is nothing more than a subset of AI
129
that specifically does this without explicitly programming a computer on the steps of how to process that information.
130
It's able to learn this from data, right?
131
That's the differentiating point between AI and machine learning.
132
And now deep learning is nothing more than a subset of machine learning, which focuses on the use of neural networks,
133
specifically deep neural networks, to do this task of learning from data to inform future decisions and other tasks.
134
So we'll circle around this theme of teaching computers to learn a task directly from raw data.
135
That really is what this whole class is all about, and we'll learn about different models to do that, different ways, and different algorithms to accomplish that main objective.
136
This class is split between both technical lectures as well as hands -on project software labs.
137
We'll have several new updates this year, especially towards the second half of the week, where we start to cover a lot more of the new advances in AI
138
and deep learning and we'll conclude also with some guest lectures
139
and some project presentation competitions that will give everybody the opportunity to win some awesome prizes.
140
Do you want to come up with the TED Talk?
141
Yes.
142
Also, John has graciously offered to host the top winners of the project presentation for TEDxMIT talks.
143
So who has not given a TEDxMIT talk?
144
Probably most of this audience.
145
So this is your opportunity to win a spot on the big stage
146
and give a TEDxMIT talk as part of the project pitch competitions.
147
the software labs are an excellent way to get hands -on
148
and actually learn how to deploy
149
and build what we learn about in the technical lectures directly with your own code
150
and programming we offer software labs in both tentaflow and pytorch we'll talk more about this later today
151
okay so how do the project labs work right
152
so every day we have a dedicated software lab that that helps mirror exactly what you learn in the technical lectures.
153
Also, you learn this in the project labs as well.
154
And every day covers on both sides.
155
So starting with lab one today, where you'll be learning how to build a music generation model that can learn how to generate a certain style of music.
156
Lab two, which is tomorrow, is going to be focused on computer vision and facial detection systems.
157
And then lab three is a lab dedicated on large language modeling.
158
You'll fine -tune your own large language model.
159
You'll evaluate it.
160
You'll see the whole process end -to -end from training all the way to evaluation and deployment.
161
And finally, on the last day of this class, we'll cover a final project pitch competition.
162
Three to five minutes, very Shark Tank -like style.
163
You'll get on stage.
164
You can work individually or in groups.
165
We'll talk more about the details for this a bit later throughout the class.
166
This class has many amazing resources to help throughout the week.
167
Please post to Piazza.
168
I think everybody here should have received a link for the Piazza, but please reach out to us if you haven't.
169
We also have some incredible TAs who you'll meet throughout the week and will help support if there's any questions.
170
All of these slides, by the way, are online on the website, so feel free to grab any of these links or details from online as well.
171
And, of course, we have an incredible team myself and Ava are your main instructors, but then we also have an incredible team of TAs
172
and we also have some incredible guest lectures lined up for Thursday
173
and Friday as well that you'll definitely not want to miss.
174
And finally, just a huge thanks and call out to our sponsors who without their support, this class is not possible at all.
175
Okay, so now let's start with some of the fun stuff and let's start by asking ourselves,
176
first of all, a question that probably is pretty self -explanatory at this point, right? which is, why do we actually care about deep learning so much, right?
177
Why are you all here?
178
Why are you so excited about this topic?
179
And the answer really boils down to what I believe is one thing, right?
180
Traditional machine learning algorithms really require a level of hand engineering,
181
human -guided hand engineering to teach computers how to perform a particular task.
182
Deep learning is so exciting because it enables us to learn those rules
183
that traditionally are driven by human engineering now by computers and by data, right?
184
So the key idea here is to learn those features directly from data and to do that in a hierarchical fashion, right?
185
So if we want to learn how to detect faces, for example, right, you may do that as a human, how would you detect faces, right?
186
And think from a first principles way, you would probably start by looking at an image
187
and first detecting the lines in the image
188
and composing those lines together to see
189
if there are certain corners in certain parts of the image
190
that resemble a face and then putting those corners
191
and different configurations together to see if you can configure an eye
192
or a nose or a mouth and
193
if you detect any of those things then it's even more inkling
194
that you probably are looking at a face right deep learning
195
algorithms are extremely hierarchical in the same way they build up features from very low level features like lines
196
and edges to higher -level shapes like corners and curves, all the way up to mid -level objects and higher -level objects as well.
197
Now, applying the fundamental building blocks of how to learn those hierarchical representations is actually something
198
that is not new and not very recent at all.
199
A lot of these core algorithms have existed since the 1950s, 1970s.
200
They really took off a lot.
201
So why are we really now seeing this resurgence?
202
And there are really three answers to this.
203
One is big data.
204
The data in the world has never been more prevalent than today.
205
And these algorithms live off of data.
206
This is what they are feeding off of, right?
207
So they really need and require massive amounts of data that the world has now been able to provide.
208
The second is hardware.
209
Companies like AMD and NVIDIA, of course, as well, are providing new GPUs that are accelerating parallel computing.
210
and accelerating the development and the deployment of these algorithms
211
and finally software that you'll get hands -on experience throughout this class
212
that actually democratizes the ability to create train and deploy any of these models even at relatively large scales
213
okay
214
so let's start with the fundamental the most fundamental building block of what makes up pretty much every neural network
215
and that is the single neuron right also known as the perceptron Right,
216
so this is really the most basic building block that's very important to understand when starting deep learning.
217
Now, what is the idea of a single perceptron?
218
At its core, it's actually extremely simple.
219
Let's start by first talking about the forward propagation of information through this neuron.
220
So if you have information and you're passing it through a neuron, we can define a set of the inputs to that neuron as X1 to Xn,
221
Or here we call it XM, excuse me.
222
Okay, so we have M inputs, one to M, and this one neuron takes all of these inputs and it has to create some output.
223
This is the goal of this neuron.
224
Now, how do we do this?
225
How does this neuron actually take those M inputs and compute its output?
226
It does this by taking each of those inputs, multiplying every input with a corresponding weight, what's called a weight.
227
Every input has a corresponding weight, So x1 corresponds to weight 1, x2 corresponds to weight 2, you multiply every weight with its corresponding input,
228
and then you add up all of the answers, and then you pass this through to the output.
229
With one more exception, which is that you also pass it through what's called an activation function.
230
Here it's denoted as g.
231
So you multiply your inputs with your weights, add up everything into one number, and then pass it through some transformation,
232
which we'll talk about in a second and that is your final answer.
233
That's one number that comes out of this neuron from M inputs that go in.
234
Okay.
235
There's one more detail I left out of this, which is that we can also have a, what's called a bias term to this neuron as well.
236
So in addition to our M inputs on the left, we also have one more weight, W0, which is called our bias.
237
Now our bias is what can allow us to shift shift this function up or down, right?
238
And I'll show a graphical example of this in a second.
239
But the equation is still the same.
240
We take every input, multiply it by corresponding weights, add a bias, pass through our non -linearity.
241
So now that x and w are actually just lists of numbers, each of which are m -dimensional,
242
we can actually write this in vector form, in linear algebra form.
243
So the output now, y, is just obtained by taking the dot product between x and w,
244
adding this bias, and passing through a non -linearity.
245
This simplifies the equation, at least notation -wise, quite a lot from the previous slide.
246
Now, you might be wondering about this activation function that I mentioned previously.
247
I said this is a non -linear function.
248
What does that mean?
249
It just means that you're transforming something from the x -axis that can be any real number, anything from negative infinity to infinity,
250
into a new real number that sometimes is bounded, sometimes not, but that transformation is not linear.
251
So here's one example of a nonlinear activation function.
252
This is called the sigmoid activation function.
253
It transforms any number on the x -axis into a new number between 0 and 1, right?
254
In fact, there are many types of nonlinear activation functions that you'll get exposure with throughout this class.
255
The sigmoid activation, which you see on the left -hand side, is just one example of very common activation functions in neural networks.
256
The central theme that links all of these together, though, is that they're all nonlinear.
257
The activation function with sigmoid on the left is very commonly used for things like probabilities, because it always outputs between 0 and 1, which is the space that probabilities live in.
258
On the right -hand side, you have activation functions that output between 0 and positive infinity,
259
which is great for introducing non -negativity constraints into your models as well.
260
And it's very simple because it's actually just two linear functions with a non -linearity between them.
261
Okay, so now why do we actually even need activation functions?
262
Why not just directly use the dot product that comes from our weights and our inputs?
263
The point of an activation function is precisely to introduce nonlinearities into our model.
264
And why do we care about this?
265
It's because in real life...
266
Real life is highly nonlinear.
267
It's highly complex and dynamic.
268
And let me show you an example.
269
So here's an example of a data set, just a two -dimensional data set, not even extremely complex.
270
But if I asked you to draw a line separating the green points from the red points in this data set, it would actually be really hard.
271
In fact, there's no line that perfectly separates the green points from the red points.
272
And that's because this data is nonlinear, even though it's just in two -dimensional space.
273
It's not in high -dimensional space, but it's still nonlinear
274
and very complex now imagine you're dealing with real world data that's extremely high dimensional
275
and also non -linear definitely you need non -linearities in your model to handle this
276
if i tell you you can make a curved line this problem becomes very easy right
277
and that's what it means to have non -linear activation functions
278
we're trying to learn a mapping a decision to classify between red points
279
and green points but we allow our model to draw curves in the edges in this decision space between the two points.
280
Let's understand this with maybe a simple example, and we can walk through it together.
281
So let's assume now that we have a neural network that was trained with two weights.
282
So W1 and W2 is 3 and negative 2, just as an example.
283
We trained it.
284
This is the result that we got.
285
And we want to pass in a new input to this model.
286
So we have two inputs to this model, X1 and X2 we obtain the output of this neuron by,
287
as we saw before, computing a dot product between our X and our W.
288
Our W is already trained, so we can plug in those two weights directly into this vector here, adding our bias, and then passing through non -linearity.
289
Now, if you look at this, what's inside of the activation function, this is nothing more than a two -dimensional line as a function of x1 and x2 are two inputs.
290
We can plot this line, right?
291
This is a trained neural network, or it's a trained neuron at least.
292
We can plot this line and observe what this decision boundary looks like in the space between x1 and x2.
293
So for any new point, x1 and x2, if I want to pass in an input to this neuron, I can actually plot this point.
294
So let's say I pass in a new input, negative 1 positive 2 is the xy coordinate of this point.
295
I can plot it in this space and I can see
296
that it actually lies on the left side of this decision boundary.
297
What does it mean to lie on the left side of the decision boundary?
298
Well, let's actually compute it mathematically first and then we'll see.
299
So if we plug in those numbers to the equation, we'll get 1 plus 3 times negative one
300
which is negative three minus two times two
301
which is negative four right add up all those together you get negative six right that's a negative number
302
that means that we're on the left side of the decision boundary
303
when we pass that through our non -linearity
304
which is our sigmoid function being a very negative number
305
that means that we're going to be further on the left on the
306
under 50 under 0 .5 on the sigmoid function so where it's a very small number close to zero 0 .00.
307
In fact, we can actually draw a hyperplane between these two parts of the space
308
and say anything that lies on the left side of the decision boundary will always be negative, right?
309
It will always be negative before the activation function.
310
It will always be less than 0 .5 after the activation function, if it's sigmoid activation.
311
And anything on the right is going to be positive before the activation function.
312
Now, this is really nice and convenient
313
that we can visualize this trained neuron in this way here
314
that for any input we can directly visualize exactly where it lives in this decision space
315
but of course this is only a two -dimensional space it
316
has only two neurons in our neuron excuse me two parameters
317
in our neuron in our neuron right it's extremely small right modern neural networks have billions of parameters
318
so imagine this plot now with a billion dimensional space this is what we would be looking at
319
so of course that's not possible right
320
so this is why this is a helpful exercise to explain
321
but of course we need to actually scale without having this exact visual
322
but building up more mathematical intuition so let's do exactly
323
that let's scale this idea now into larger networks
324
and let's start by taking that one neuron
325
and actually building a neural network out of it to seeing how this all comes together, right?
326
So this is our original neuron picture, okay?
327
Now, if there's a few things that you remember from this class, this is probably the most important thing, is that you should remember how a single neuron works,
328
and that's by doing a dot product, adding a bias, and applying a non -linearity.
329
It's really three steps, dot product, add a bias, and apply a non -linearity.
330
That equation is the core to so many different parts of deep learning.
331
Now, let's simplify the diagram a little bit now that we got that foundation settled, right?
332
I'll remove all of the weight labels and we'll just have every line, you can remember
333
that every line will have both an input coming in as well as a weight associated to that line, okay?
334
Now, z here, what's shown as z, is the result of that dot product plus a bias, right?
335
This is what's happening before the non -linearity, right?
336
So the final output is simply Y of G, which is our non -linearity, of Z.
337
OK, now let's define a multi -output neural network now.
338
To do this, we just have two neurons instead of one.
339
They will both see the same inputs, but now they will have their own independent weights.
340
So each neuron has its own three weights.
341
It sees the same three inputs, but it now can output two different answers because it has two different sets of weights.
342
Okay, so now with this mathematical understanding, let's see if we can now build our first layer, neural network layer, entirely from scratch.
343
So let's do this by first initializing a matrix called W.
344
This is going to be a matrix that contains all of our weights for all of our neurons in that one layer.
345
Okay, so we can create now this matrix of weights
346
which is going to be dimensionality of the input space by
347
the dimensionality of our output space right now the bias is also important to initialize here and then
348
when we create our call function our forward pass function this
349
is going to show actually how we can multiply do the
350
dot product of our inputs with our weights okay right here add the bias
351
apply the non -linearity and that's our answer right it's the same equation again
352
okay we can do this in tensorflow we can also do it in pytorch
353
and you'll notice actually between these two it's very very similar
354
type of language a lot of parallels basically just different names for a lot of things
355
but the overall frameworks are extremely similar
356
same foundations we'll initialize the weights
357
and the biases in the forward pass we'll compute the output
358
of this layer by computing a matrix multiplication of our inputs with our weights adding a bias
359
and then applying our non -linearity okay now luckily both tensorflow
360
and pytorch have actually defined that exact layer for us already
361
so you don't actually need to to write
362
that code yourself you would just call the layer in tensorflow
363
it's called the dense layer in pytorch it's called a linear
364
layer they do the exact same thing they do the matrix multiply add a bias apply a non -linearity
365
so we can just call it instead right we call it with the number of units
366
or the number of neurons in each layer okay let's keep
367
building up this abstraction step by step that's a single layer neural network.
368
This is one where we can have a single hidden layer now.
369
So this is actually two layers.
370
Now we can see how we can do this transformation from input space to the hidden layer space with one layer, and then hidden layer space to output space with another layer.
371
So now we actually have two weight matrices, right?
372
The first weight matrix on the left -hand side, W1, converts the inputs to the hidden layer.
373
Second weight matrix is W2, which goes to the output, Each of these are separate weight matrices, and they actually have different sizes as well,
374
because your output shape is different than your hidden layer shape.
375
Now, if we look at a single unit, let's zoom back in now to a single unit within the hidden layer.
376
Let's take, for example, this one, Z2.
377
This is just the second unit in our hidden layer of this model.
378
This is the same perceptron that we saw before, nothing special here.
379
We compute its output, Z2, by taking a dot product of all inputs with the weights corresponding to that neuron,
380
adding its bias, and applying a non -linearity.
381
If we took a different node, like Z3, for example, it would look also the exact same, except its weights would be different than Z2's weights.
382
The inputs are the same, but they see different weights.
383
Okay, so this picture looks a bit messy, so let's abstract even further so we can keep building up this intuition.
384
So now I'm going to remove all the lines
385
and just put these symbols here in cases wherever we have these fully connected layers.
386
Everything that connects from input to output with the matrix multiply in between.
387
And again, we can actually stack these fully connected
388
or linear layers on top of each other with nonlinearities between them to make sure that we have the nonlinearities.
389
across our network as well
390
and here's an example of how to do this you basically stack the same layers
391
that we saw before now in these sequential blocks sequential just
392
basically allows you to stack them okay now finally the the point
393
that we've been waiting for is how to go from these
394
neural networks to deep neural networks there's nothing more to this
395
than just stacking a bunch of linear layers on top of
396
each other a deep neural network is nothing more than a neural network
397
that has more than usually three layers is what people can conceptually say, right?
398
So this is a network where the final output is computed by going deeper
399
and deeper and deeper across each of these layers to compute its final output at the end of the network.
400
And again, to no surprise, this is done in the same abstraction code as before, just having a sequential block and a bunch of linear layers within that block.
401
Okay, this is awesome.
402
So now we have this good intuition of how to build
403
up from a single neuron all the way to a layer to a neural network
404
and to compose these in the forward pass.
405
Let's take a quick look of how we can apply this to a particular problem
406
that it's it's a good starting problem for everybody in this class
407
which is should will I pass this class right
408
so this is a question it has a yes no answer
409
and there's a lot of data for this question actually
410
because we've taught this class for a long time
411
so we have a lot of data from past students who have taken this class here's an example
412
so we have an x and a y -axis
413
x is the number of lectures that you attend
414
and the y -axis is the number of hours
415
that you spend on the final project and we also have individual data points
416
which represent past students who have taken this class you can see where they live in this space
417
and then we have you you fall right here at four or five, right?
418
You've attended four lectures and you spent five hours on your final project.
419
And the question is, will you pass this class based off of all of the other data that you've seen, right?
420
Now, how can we do this?
421
We have a neural network that can do this.
422
Let's try and actually feed it into our neural network.
423
We have two inputs, which are the exact two axes that you saw before, number of classes, number of lectures, the number of hours one input is four the other input
424
is five we can feed these two inputs into our neural network
425
and we can see that it predicts an answer of 0 .1
426
or 10 probability
427
that you'll pass this class very bad i know okay who
428
knows why the the network is wrong in this case yes exactly okay Okay, so this is a randomly initialized network.
429
It is effectively the same as a brand new baby, right?
430
It has never seen the world before.
431
It has never seen any of the data before.
432
So the step that's missing here is that we've just done a forward pass through the model.
433
Now we have to actually teach the model, right?
434
And the key part of this is that every time that we have a prediction, the model predicts point one, there's also a ground truth answer for every prediction that we have in our training set.
435
And the ground truth answer here is true, right?
436
You did actually pass this class, right?
437
So there's a deviation here.
438
There's a difference between the 0 .1 that the model predicted and the 1 that it should be answering.
439
Now, to train this network, we need to show it this difference.
440
We need to tell it when it gets the answer wrong
441
so it can learn to move closer towards the correct answer the next time it sees a 4 .5, right?
442
If it sees 4 .5 again, it should now improve on its answer and should predict something closer to a 1.
443
We do this by quantifying what's called a loss.
444
A loss is nothing more than just the error, this deviation term between the prediction and the ground truth.
445
Smaller losses means that the model is outputting a higher quality answer.
446
Larger losses means that the answer is more erroneous.
447
So let's assume, of course, that the data is not just from one student, but we have a data set of many, many students right?
448
This is coming from now a loss over the average of each of those losses
449
that we have in our data set, right?
450
So when training a neural network, oftentimes we're not going to actually minimize the loss on a single data point, but minimize it across the entire data set, right?
451
And that's to have higher quality outputs.
452
Now, if we look at this problem in the case of classification
453
a yes no type of problem then this is where the
454
answers will typically predict a probability a probability of being a yes versus probability of being a no
455
and that's why we predict between zero
456
and one this ties us back to the sigmoid function from before right
457
because those outputs are zero and one outputs
458
so we can train it in this way right
459
and we can use losses like the cross entropy loss that actually support matching those two distributions of probability against each other.
460
Now, you might be asking yourselves, okay, what if you didn't want to predict a probability distribution, like a yes -no question, but rather a continuous value, like what's the temperature going to be?
461
Or what score will I actually get on the class rather than will I just pass or fail it?
462
In those cases, you care about predicting continuous values instead, not probability distributions.
463
So you actually want to output and have losses that encourage deviations that are also continuous.
464
In those cases, you may consider things like mean squared errors or L1 errors,
465
things like this, that simply compare the predicted value minus the ground truth value and just minimizing that absolute difference.
466
Okay, now let's put all of this loss information together into the problem of actually finding the neural network weights
467
that are optimally suited for any particular data set or any particular task.
468
So we know that we want to find a neural network that looks roughly like this form, right?
469
We want it to minimize the loss on our entire training dataset.
470
And this means that we want to find the Ws, the weights, that minimize the loss here,
471
the J, the J of W, right?
472
J of W is our empirical average loss across our entire dataset of N data points.
473
okay so remember again
474
that w is nothing more than just a collection of many
475
many weights across all of the different layers in your neural
476
network everything from layer zero the first layer all the way to layer n
477
which is your last layer now our loss function is also nothing more than just a function
478
that outputs a single scaler at the end right we take
479
the average loss the average error across every data point compute its average
480
and that is our loss it's just one number at the end of the day
481
and that one number is a function of our weights
482
so on on here you can see a visualization of let's
483
say again a two parameter neural network just two parameters w0
484
and w1 for any configuration of those two parameters we have a particular value of the loss which is the height, okay?
485
There are some weights that encourage a really good loss, very low.
486
There are other weights that are really bad losses, right?
487
Very high on the curve here, right?
488
And the goal is that we want to find the neural network that has the two weights, W0 and W1, that have the lowest loss.
489
Now, how do we do that?
490
So, this is called training the neural network.
491
We start by training the neural network.
492
We just randomly initialize it.
493
We pick a W0 and a W1 to start with.
494
Let's say we start here at this black point.
495
We compute its loss.
496
And then we compute the gradient.
497
The gradient tells us the direction of the slope at that location, right?
498
It doesn't tell us the slope everywhere.
499
It just tells us at this very local location, which way is pointing up.
500
But of course we want to go down.
501
So we actually take the negative of the gradient and take a small step down.
502
And then we repeat this process.
503
We compute the new gradient at the new point, and we take a negative step in the opposite direction of
504
which way the gradient is pointing until we finally get to the bottom.
505
And this will actually be guaranteed to converge to a bottom.
506
May not be the lowest bottom, right?
507
Because it actually depends on where you start.
508
You may not find yourself in the absolute lowest, but you will find a bottom and you'll eventually converge there.
509
So we can summarize this algorithm.
510
This is known as gradient descent
511
and we can write it roughly in pseudocode it looks like this we start by initializing our weights randomly
512
and then we repeat the following process until convergence we compute the gradient at the weights
513
that we're currently in
514
and then we take a small negative step of our gradient with our weights
515
so we change the gradients to step in the opposite way
516
that the gradient tells us
517
and then we keep repeating this until there's convergence until our weights are not really changing anymore
518
and then finally we call those final weights the trained network
519
of the model um in code right this also is is uh is doable in code we can do this
520
while looping in a
521
while loop where we compute the loss we can compute the gradient of our loss with respect to our weights
522
and then compute that same update formula
523
again now a very important line here is to look at this term right here the gradient right now i
524
glossed over this this point
525
but actually the gradient this is a key part of this algorithm
526
because computing this tells us exactly how to update our weights on every part of the iteration
527
but I never actually told you how to compute this so let's talk about actually computing the gradient
528
for a neural network
529
and this is a process known as back propagation we already
530
learned about forward propagation of information through the model now we'll learn about back propagation
531
which is how to actually compute the gradient of the model
532
we'll start with a very simple neural network to illustrate this example
533
and this is actually probably the simplest neural network
534
that you can have which is just a single neuron with a single output as well
535
so x goes to a single neuron goes to a single output goes to a loss.
536
So let's visualize what this would look like
537
if we wanted to compute the gradient of our loss with respect to our weight w2,
538
which tells us how a small change in w2 affects our loss j.
539
Okay, so let's write it out as a derivative.
540
So this is the derivative of j with respect to w2, right?
541
Now to compute this, we can use the chain rule, We can decompose dj, dw2 into two parts,
542
the first part being dj, dy, times dy, dw2.
543
This is just an application of the chain rule in linear algebra.
544
Now, this is only possible because y is only dependent on the previous layer.
545
Let's suppose now we wanted to compute the gradients of the weight before
546
that, which is w1 not w2
547
so let's actually change this to dj of d w1 now
548
here this is dy dw1 we all we again need to apply a chain rule again to compute this right
549
so we expand it out one more time
550
that part turns into dy dz1 followed by dz1 dw1 right
551
and if you had a deeper neural network you would just keep repeating this process as you go deeper
552
and deeper into the model and
553
and it just repeats like this right you can repeat this process by propagating those gradients
554
that you compute from the end from the output all the
555
way back in in the in the topology of the neural network all the way back to the input
556
and this will allow you to determine how every single weight
557
impacts your final loss that's exactly what the gradient shows you it shows you
558
if you move in some direction how will the loss change?
559
Will it go up or down?
560
Now, that's the backprop algorithm.
561
In theory, it's just an application of the chain rule, right?
562
It's nothing more than the chain rule. In practice, this is a lot more complex
563
because computing the gradient of neural networks does not look as clean as what we saw in that previous slide, right?
564
In practice, neural network loss topologies look highly non -convex, right?
565
There are a lot of different minimum, and it's highly dependent on the initializations that you pick in these models
566
and the regularizations that you apply to make sure that you can find actually this minimum.
567
So this was a really cool paper from many years ago, 2017, where they actually tried to visualize what these high dimensional neural network loss landscapes actually look like.
568
Of course, this is a two dimensional representation of something that is, you know, millions of dimensions.
569
So you need to take it with a grain of salt.
570
But even projecting it down into two dimensions, you can see that this is a highly nonlinear landscape.
571
Now, recall the update equation that we defined during gradient descent.
572
And remember this parameter that we also didn't talk so much about, right?
573
I said we take a small step in the opposite direction of the gradient.
574
What I didn't mention is, you know, how big is this step?
575
This step here is referred to as eta, right, this small symbol here.
576
This is called the learning rate this is the rate at
577
which you follow the gradient in practice setting this learning rate it's just a number
578
that you set in practice setting this number is also quite difficult right
579
because if you set something
580
that is too small then you actually never really reach some
581
of the the big minimums right this is a minimum
582
but the model gets stuck in it even though it's not the biggest one right you would actually
583
and if you set something too large you can overshoot
584
and your loss can also explode and you never find the minimum, either minimum.
585
So ideally you want to use learning rates
586
that are basically just large enough that you can skip over these kind of fake local minimums,
587
but get stuck in these global or close to global minimums to converge there.
588
So how do we deal with this?
589
How do we actually set the learning rates?
590
One option is just to try a bunch of different learning rates and see what works best.
591
And actually this is not as crazy of an option as it sounds like
592
but can you do something smarter than this how about we design adaptive learning rates
593
that adapt to the shape of our lost landscape
594
that actually is a much smarter idea because it means
595
that we can now increase or decrease our learning rate depending on how large the gradient is at that location.
596
When you have a very large gradient, maybe you wanna take a bit of a smaller step.
597
If you have a small gradient, maybe you want to adapt or have some momentum that can carry you through these local minimums and pass over them.
598
And that's exactly what a lot of different optimizers have come out with.
599
So we saw the most basic version of stochastic gradient descent previously, but there are a bunch of other variations of gradient descent
600
algorithms called atom at a delta at a grad rms prop etc
601
that you'll get a lot of practice with throughout your software labs
602
and in general these are actually much better types of optimizers for the reasons especially around setting learning rates
603
and making sure that you can actually find these minimum
604
and these are widely studied in practice right in general you will very rarely use the vanilla stochastic gradient descent,
605
SGD algorithm, you'd probably almost always use
606
and change this line to something like Atom or another adaptive learning rate scheduler or optimizer.
607
Awesome.
608
Okay.
609
I'll continue now in just talking about some more practical considerations, especially now towards the end of the technical part of the lecture, this first lecture.
610
I want talk about some of the practical considerations
611
that you should have when actually training neural networks
612
so let's go back to the gradient descent algorithm for a
613
second we just saw this this algorithm here we saw how we can compute the gradient
614
but remember this is done through back propagation this is a really expensive process
615
because you have to go entirely back
616
and compute derivatives for every single parameter in your network going
617
from the output all the way back to the input
618
and you have to do this for every single data point in your data set right not just one data point, of course.
619
Now, in most real -life problems, it's not feasible to do this on every training iteration.
620
You simply would not want to do that over every data point in your data set.
621
So instead, we're going to define what's called stochastic gradient descent, SGD.
622
So instead of computing the gradient over your entire data set, let's compute the gradient over just one point.
623
Pick a random point
624
and compute the gradient with respect to that one point and step in the opposite direction of that one point.
625
Now, that is much, much faster to compute, obviously, on big data sets.
626
Computing a gradient on one point is going to be way more efficient than the entire data set, but it's also going to be way more noisy.
627
So there's a trade -off there, of course.
628
Now, the middle ground, what is the middle ground?
629
Is that you do this over what's called a mini -batch.
630
So you compute some batch size of data points, and you compute the gradient with respect to
631
that batch size batch sizes can range from anything very small
632
to like a few data points all the way up to you know
633
like most commonly people use things like 32 is a very
634
common batch size you can scale it up in large language
635
models we have batch sizes of you know millions right it depends on the type of data sets
636
that you're using
637
but in general you always want to pick something much smaller than the size of your entire data set
638
because you don't want to be going through your entire data set every time.
639
Now, B is normally not that large for most problems.
640
Like I said, this is usually on the order of tens, hundreds, and even this gives you a pretty good estimate of the gradient for most problems.
641
This increases the accuracy that's going to be much more accurate than SGD, stochastic gradient descent, that's only looking at one point,
642
but it will also be much more efficient than computing it on the whole data set.
643
So it's this nice balance, and you can actually control the balance yourself by controlling the size of the batch.
644
Now, the last topic I want to touch through is overfitting, right?
645
This is known, a very well -known problem in the field of machine learning, but it extends also, obviously, into deep learning as well.
646
Now, what is overfitting?
647
Overfitting is nothing more than the process of learning too much into your data, right?
648
Looking too much into it and overfitting on the details of
649
that data set that do not generalize outside of the particular training data that you've seen.
650
So here's a good example, right?
651
If you start on the right -hand side and here's your training data set,
652
the optimal line that fits this data is not going to follow all of the intricate details of the data.
653
But on the same side, on the other, excuse me, on the other side, there's also underfitting as well, where you don't learn a function that's expressive enough
654
that captures all of the intricacies of the data
655
so for example a linear function on the left side is going to be underfitting on this problem
656
because the problem is a non -linear type of relationship
657
but on the right hand side it's also too high the
658
the learned solution on the right hand side is too high dimensional
659
and too overfitting to to capture the generalization of the of
660
the problem the ideal is somewhere in the middle right where you don't have too much complexity
661
and you don't read into too much of the details so that when you see brand new data
662
a test set or in deployment you can actually extend your solution to have accurate predictions in
663
that regime as well so to address this problem how do you go
664
and toggle between these two sides of the coin it's called regularization right regularization is the technique
665
that allows us to constrain how much we want to underfit versus overfit on a particular problem
666
And this is effectively allowing us to discourage the learning of very complex models.
667
This is how we're able to learn billions of parameter models on actually data sets,
668
but not actually memorizing all of the details in the data sets or allowing us to generalize beyond those data sets.
669
And this will allow us to improve the ability of our model to perform even on unseen data.
670
The most popular form of regularization, there are many different ways to regularize depending on the type of model, but I'll talk about one of the most popular ways.
671
This is the idea of dropout.
672
Dropout is a stochastic method.
673
It means that it's probabilistic.
674
Let's revisit this chart, this schematic that we saw earlier in the lecture from a neural network.
675
In dropout, all we do is
676
that we randomly drop out some activations of the hidden layers with some probability that we define.
677
Okay, so let's assume that we set a probability like 0 .5, 50%.
678
All this will mean is
679
that we will randomly pick 50 % of the neurons in our two hidden layers
680
and randomly kill the outputs from coming out of those neurons.
681
So what does that mean?
682
That this neuron here has to be resilient to sometimes receiving inputs from this neuron, but sometimes not receiving it from this neuron, right?
683
And then on other iterations, it may receive it from the other set of neurons.
684
So what this encourages is actually the ability that on every iteration, every neuron is seeing different pathways through the network.
685
This is forcing it to not rely on any one pathway
686
and not memorize any one pathway too much through its forward propagation.
687
And that's actually able to allow it to generalize as well
688
because it has to find these relationships that span across different pathways.
689
Yes?
690
Are the same activations set to zero for the entire mini -batch?
691
Same activations are set to zero for the entire mini -batch, but then on the next mini -batch, you have a new set of activations, a new set of neurons that are set to zero.
692
So you reset on every batch.
693
Awesome.
694
Okay.
695
So, you repeat that in every iteration.
696
Let me show you one more technique for regularization as well that goes beyond dropout.
697
Dropout is an architectural level of regularization.
698
This is going to be beyond architectural level.
699
So, let's assume for any architecture, you can have the ability to stop training once you start to overfit.
700
It sounds easy.
701
Let's see how you might actually do this, right?
702
This is called early stopping.
703
So we already have this definition of overfitting,
704
which is simply that we start to perform better on our training set than we are on our testing set, right?
705
That's effectively what it means to overfit.
706
So we're no longer generalizing our learnings from training into testing, right?
707
Our training is continuing to improve, but our testing data is just getting worse, right?
708
So we can do this by actually having two separate data sets, one that we train on and one that we test on.
709
And constantly throughout training, we're evaluating both data sets.
710
Now, in the beginning, both lines just decreased together, which is excellent.
711
It means our model is learning.
712
It's moving from that underfit regime to a more fit regime.
713
And it continues to have this trend.
714
Eventually, though, the network loss will start to plateau
715
and the testing loss will start to increase right this is actually the regime where we start to see overfitting
716
and this pattern typically will continue for the rest of training
717
now the question here is where would you actually you train
718
the model for all of these training iterations you save the
719
model at these checkpoints across the x -axis you can now
720
take this checkpoint right you've saved all the checkpoints you look at the curve
721
and you actually say okay this is the checkpoint
722
that I want to use
723
because after this checkpoint yes my training loss continues to get better
724
but it's not actually translating to my testing data set
725
so this is the one that i actually should be using
726
and we can see
727
that you know after the early stopping checkpoint things do get worse on your testing data
728
so you wouldn't actually want to apply those checkpoints this is such a very powerful idea
729
because it's very general this approach you can really apply this to any type of model
730
and you can really track these these curves right it doesn't
731
require architectural modifications uh you just monitor the loss of your
732
neural network over time on two different data sets one that you train on one that you test done.
733
Awesome.
734
Okay.
735
I'll conclude the first lecture now, just summarize quickly what we've covered.
736
So we started with the most fundamental building blocks of neural networks, which was just a single neuron.
737
We've scaled that up to single layers and multi -layer perceptrons, multi -layered neural networks.
738
And we've learned how we could build complex hierarchical learning machines that go from input to output across a hierarchy of abstractions.
739
And finally, we addressed a lot of the practical sides of actually training these models, evaluating them, picking the best model and so on.
740
In the next lecture, we're going to be hearing from Ava on covering sequence models, right?
741
This is a really exciting lecture, especially in today's world, because sequence models power GPTs, right?
742
Everything, a lot of things are sequences in today's world, right?
743
You think of text as a sequence of words, audio is a sequence of waveforms, video is a sequence of images, right?
744
A lot of things are sequences, and what we saw today is only covering non -sequential data so far in my lecture.
745
In the next lecture, we'll see how we can extend that into sequences of data.
746
So we'll take a five -minute break, and then we'll just set up, and then we'll continue from there.
747
Thank you.

¿Qué aprenderás?

Este video es una excelente oportunidad para mejorar tu práctica de speaking para el IELTS y desarrollar habilidades clave. Primero, aprenderás a seguir un discurso estructurado, ideal para organizar ideas en presentaciones o respuestas extensas. Segundo, practicarás la comprensión de vocabulario técnico (como "modelos de lenguaje" o "centros de datos") en contextos reales, lo que es fundamental para el examen. Y tercero, mejorarás tu capacidad de escuchar y repetir frases fluidamente, una habilidad esencial para la técnica de shadowing.

Escucha estos sonidos

En el diálogo, observa la conexión de palabras comunes en el inglés hablado. Por ejemplo, "a lot of" se reduce a "alotof" y "you'll" se enlaza con "go" en "you'll go". También escucha la reducción de "it's" a "its" en frases como "it's really easy". Estas características de la lengua hablada son cruciales para entender y ser entendido en situaciones cotidianas y en exámenes como el IELTS.

Háblalo como un nativo

Para imitar el ritmo y el énfasis del hablante, usa la técnica de shadowing (o "shadowspeak"): reproduce el video en pequeños fragmentos, deténlo y repítelo inmediatamente, intentando igualar su tono y velocidad. Presta atención a cómo el instructor enfatiza palabras clave como "remarkable", "exponential" o "massive" para dar énfasis a sus ideas. También, nota su pausa después de frases como "right?" o "actually", que le dan naturalidad al discurso. Con práctica, tu práctica de inglés hablado se volverá más fluida y natural, lo que te ayudará a sobresalir en el IELTS y en cualquier situación comunicativa.

¿Qué es la Técnica de Shadowing?

Shadowing es una técnica de aprendizaje de idiomas respaldada por la ciencia, desarrollada originalmente para la formación de intérpretes profesionales y popularizada por el políglota Dr. Alexander Arguelles. El método es simple pero poderoso: escuchas audio en inglés nativo y lo repites en voz alta de inmediato, como una sombra que sigue al hablante con solo 1-2 segundos de retraso. A diferencia de la escucha pasiva o los ejercicios de gramática, el shadowing obliga a tu cerebro y músculos de la boca a procesar y reproducir simultáneamente patrones de habla reales. Las investigaciones muestran que mejora significativamente la precisión de la pronunciación, la entonación, el ritmo, el habla conectada, la comprensión auditiva y la fluidez al hablar, convirtiéndola en una de las metodologías más efectivas para la preparación del IELTS Speaking y la comunicación en inglés en el mundo real.

Técnica de shadowing: lee la guía completa paso a paso →