Pratica di Shadowing: The GELU, SiLU and SwiGLU activation functions, clearly explained!!! - Impara a parlare inglese con i video

Creazione lezione...
1
Sometimes drop out is all you need.
2
Hooray!
3
StatQuest.
4
Hello, I'm Josh Starmer and welcome to StatQuest.
5
Today we're going to talk about the GelU, SilU, and Swiglu activation functions, and they're going to be clearly explained.
6
This StatQuest is brought to you by the letters A, B, and C.
7
A always, B be, C curious, always be curious.
8
Too long, didn't watch.
9
If you're already familiar with the REL U activation function, then the GAL U activation function,
10
which sounds like it could be some sort of expletive, but is short for Gaussian Error Linear Unit, and the CIL U activation function,
11
which is short for Sigmoid Linear Unit and thus sound sort of like an oxymoron, are almost the same.
12
One relatively subtle difference is that the RELU outputs zero for all negative input values.
13
And the GELU and CELU have this little dip for input values between negative two and zero.
14
Another subtle difference is that the RELU has this angular bend at zero.
15
And the Gell-U and Sill-U have a softer, differentiable curve at zero.
16
Lastly, for all positive input values, the Rell-U just outputs the original input value.
17
While the Gell-U and Sill-U have a slight curve
18
that makes the output values just a little smaller than the positive input values.
19
In contrast, the SWIGLU, which stands for Swish Gated Linear Unit, which sounds like a random jumble of words,
20
can look like a smooth version of the RELU, or it can look like pretty much anything else it wants to look like.
21
That said, despite the glaring differences in the shapes that the SWIGLU can take on, it's closely related to the GELU and CILU.
22
Anyway, as a result of these relatively subtle differences for the Gelu and Siyu,
23
and the relatively extreme differences for the Swiglu,
24
these activation functions do not tend to overfit training data as much as the Relu.
25
In other words, if all these dots were the training data,
26
then a large neural network using the Relu activation function would have a strong tendency to overfit the training data.
27
In contrast, a large neural network using the GelU,
28
SilU, or Swiglu activation functions does not have as strong a tendency to overfit the training data.
29
Thus, this new class of activation functions tend to result in neural networks that do a better job making predictions after training.
30
Baaam.
31
But Josh, where do these new activation functions get their different shapes from?
32
And how did anyone come up with them to begin with?
33
Well, just for you, Squatch, and anyone else who's interested, I made this whole stat quest just to answer those very questions.
34
Hooray!
35
Way back in the days before 2010, when you built a neural network, you usually did it using sigmoid activation functions.
36
At the time, sigmoid activation functions seemed like a great idea because their output is constrained to be between 0 and 1,
37
which represented the outputs of actual neurons either not firing or firing.
38
And the curve reflected how neurons transition from one state to the other.
39
Lastly, the sigmoid activation function is continuous and differentiable for all input values, making it a reasonable choice for backpropagation.
40
And, for simple neural networks like this one, sigmoid activation functions work great.
41
However, once people started scaling up their neural networks to have more hidden layers and more activation functions in each hidden layer,
42
it became harder and harder to train the model.
43
Thus, when we use the sigmoid activation function, we can't build very large neural networks.
44
The reason for this size limitation has to do with the derivative of the sigmoid shape, which is needed to fit the neural network to data.
45
Specifically, as the input for the sigmoid activation function gets further and further from zero,
46
the derivative gets closer and closer to zero.
47
And a super small value for the derivative means gradient descent, or whatever else we're using to fit the neural network to training data,
48
takes super small steps towards the optimal parameter values.
49
And that means finding the optimal parameter values can take a long, long time.
50
As a result, if these were the training data, then, ideally, we'd want to fit a squiggly shape like this to it.
51
However, because we can't easily train a large neural network using the sigmoid activation function, we might get a squiggly shape like this instead.
52
The squiggly shape from the sigmoid underfits the training data.
53
And thus, the squiggly shape created with the sigmoid activation function doesn't fit the data as well as we might like.
54
So, around 2010, the sigmoid activation function was replaced by the rel U, this bent shape.
55
And the rel U activation function proved to be fundamentally important to making deep neural networks trainable.
56
Ultimately, in 2017, the REL-U achieved peak awesomeness by being part of the first transformer.
57
Bam.
58
The REL-U was great because, among other things, it was super easy to compute.
59
If the input value, the x-axis coordinate, was negative, then the output, the corresponding y-axis coordinate, was 0.
60
And if the input value was positive, then the output, the corresponding y-axis value, was the exact same value.
61
Another advantage was that the derivative was defined to either be 0 or 1.
62
Having a derivative of 1 for large input values helped simplify backpropagation
63
and made it easier to optimize the weights and biases in the model.
64
Unlike the sigmoid activation function that had a super small derivative for input values far from zero,
65
the ReLU has a derivative of one regardless of how much greater the input value is compared to zero.
66
And that means, at least for large input values, the ReLU can optimize parameter values relatively quickly.
67
Boom!
68
Nope, not yet.
69
Unfortunately, the RELU has its own problem.
70
First, remember that we had this training data
71
and we showed the ideal squiggly shape that we wanted our neural network to fit to the data.
72
And, unfortunately, when we used the sigmoid activation function, instead of the ideal squiggly shape,
73
we got a squiggly shape that underfit the training data.
74
Well, the problem with using the RELU activation function is
75
that now we can create a huge neural network that makes a squiggly shape that overfits the training data.
76
In other words, the large neural networks
77
that we can create with the RELU activation function often overcorrect for the problem we had with the sigmoid.
78
Can't you just make a smaller neural network with the RELU?
79
Great question, Squatch.
80
In theory, we could do that, but in practice, it's next to impossible to figure out the optimal size for a neural network.
81
I mean, imagine you started with a network that had 100,000 activation functions.
82
Are you going to then test to see if 99,999 activation functions is better?
83
And then test 99,998?
84
Etc?
85
Okay, I get your point.
86
It would take way too much time.
87
So, ideally, it would be nice if we could have something
88
that didn't underfit the data like a small neural network using the sigmoid,
89
and didn't overfit the data like a large neural network that uses the ReLU.
90
One early solution to this problem was something called dropout.
91
During training, when we are fitting a neural network to a dataset,
92
Dropout randomly selects a subset of activation functions and removes them from the network.
93
And then applies backpropagation to this smaller network to update the weights and biases.
94
Once the weights and biases are updated,
95
it goes back to the original neural network and selects another random subset of activation functions to remove from the training step.
96
And then applies backpropagation to this smaller network to update the weights and biases.
97
And this process is repeated for the duration of training.
98
As a result of using Dropout during training, instead of training one big neural network,
99
we effectively train lots of smaller networks.
100
Thus, after training, when it comes time to make predictions with the entire neural network,
101
we're essentially averaging over all of the smaller neural networks that we created during training.
102
And the model averaging from the rail U combined with dropout results in a squiggly shape
103
that does a better job approximating the ideal shape.
104
Bam.
105
Nope, not yet.
106
Unfortunately, combining the rail U with dropout has its own problems.
107
The first problem is that dropout is an extra thing that we need to add to the training stage.
108
And the second, perhaps more important problem,
109
is that dropout is completely independent of the inputs to each activation function in the network. In other words,
110
dropout will remove an activation function from a neural network regardless of whether the input signal is large or small.
111
Ugh, how do we fix that?
112
Well, that's where the Galu, Silu, and Swiglu activation functions come in.
113
BAM!
114
Remember, at the start of this stat quest, we saw what the Galu, Silu, and Swiglu activation functions looked like.
115
And we learned that they did a better job fitting a squiggly shape to the data.
116
Well, now we'll learn where these activation functions come from.
117
And to do this, we'll We'll learn how to build them from scratch.
118
We'll start with a super simple neural network that just has one input, one output, and nothing else.
119
Josh, is that even a neural network?
120
It's enough to show how the Galu and Silu activation functions were created.
121
Okay.
122
Now, even though this neural network is ridiculously simple, let's imagine we apply dropout to it.
123
If we apply dropout to this neural network, then, sometimes, an input value, like negative 2, will make it to the output.
124
And sometimes, the input value will not make it to the output, and the output value will be 0.
125
Now, if we apply dropout 50% of the time, then we can calculate the average output value
126
by multiplying the top output value by 0.5 and multiplying the bottom output value by 0.5 and adding those two products together.
127
In this case, the average output value is negative 1.
128
Bam.
129
Now let's see what happens when we use the input value x to determine the probability
130
that we will avoid dropout p of x.
131
Specifically, we'll use this blue squiggle to convert the input value x into a probability
132
that we will avoid dropout p of x.
133
Note, if you're dying to know more details about this blue squiggle,
134
It's the cumulative distribution function for a standard Gaussian curve.
135
However, the important thing is that we have a squiggle that goes from 0 to 1.
136
In this case, if the input, x, is a negative number, then the probability that we'll avoid dropout, p of x, will be closer to 0.
137
And if the input is a positive number, the probability that we will avoid dropout will be closer to 1.
138
For example, if the input value is negative 2, then, according to this blue squiggle,
139
the probability we will avoid dropout is 0.02.
140
So we plug 0.02 in for P of X, the probability for avoiding dropout.
141
And we plug in 1 minus 0.02, or 0.98, into the probability that we will use dropout.
142
Note, because the probabilities we avoid dropout, 0.02, or use dropout, 0.98,
143
are different, we're now calculating a weighted average of the output values.
144
Anyway, in this case, when the input value is negative 2, the weighted average is negative 0.04.
145
So we can put a dot at
146
on a graph that has input values on the x-axis and the weighted average of the output values on the y-axis.
147
Now, if the input value is negative 1, then the probability we will avoid dropout is 0.16.
148
So we plug 0.16 into the probability for avoiding dropout.
149
And we plug 1 minus 0.16 or 0.84 into the probability that we will use dropout.
150
And the weighted average is negative 0.16.
151
So we can put a dot at negative 1 comma negative 0.16.
152
Likewise, we can put dots at 0 comma 0.
153
1, 0.84, and 2, 1.96.
154
And, in general, when we allow the probability for avoiding dropout to be determined by this blue squiggle, we end up with this function.
155
And now, instead of plugging numbers into this neural network with dropout, we can use this new function as an activation function.
156
In other words, this neural network with dropout and probabilities can be simplified to this neural network with a single activation function.
157
And that means that our new activation function has the effects of input-dependent dropout built right into it.
158
Because it was derived from the weighted averages that input-dependent dropout creates.
159
Lastly, because this blue squiggle is the cumulative distribution function for a Gaussian distribution,
160
this activation function is called GalU for Gaussian Error Linear Unit.
161
Double bam!
162
Now that we know how the GalU activation function can be derived entirely from input-dependent dropout,
163
Let's go back to the larger neural network to derive the general equation for the gelu and associated activation functions.
164
First, let's just plug in x into the input.
165
Now, when we just do the math, we get...
166
x times p of x plus 0 times 1 minus p of x.
167
And because the second term is multiplied by zero, just like an activation function that gets dropped out, it goes away.
168
And we're left with the input value, x, times the probability we avoid dropout, P . In other words,
169
input values from negative infinity to positive infinity times their corresponding
170
y-axis coordinate on a curve like this give us an activation function function.
171
Now that we have a general equation for creating activation functions with input dependent dropout,
172
one cool thing we can do is swap out the blue squiggle
173
that we're using for P of X and get a related, but still different, activation function.
174
For example, remember the sigmoid activation function we started with long ago, but ultimately replaced with ReLU?
175
Well, compared to the cumulative distribution function for a Gaussian curve, the sigmoid shape, the orange squiggle, isn't that different.
176
when we multiply input values by their corresponding y-axis coordinates on the orange squiggle, we get this new activation function.
177
And since this new activation function is based on the sigmoid shape,
178
this activation function is called the SilU for sigmoid linear unit.
179
When we plot the GelU and the SilU activation functions on the same graph, we see that the differences are not super extreme.
180
And just as we might expect by the similarity in the shapes, both can work in similar contexts.
181
That said, as of this quest creation, it seems that Google favors the GelU for its AI models.
182
Noted.
183
Now let's talk about the SWEGLU activation function.
184
To do that, let's go back to our super simple neural network with Dropout.
185
And make it fancy.
186
First, we'll add a weight.
187
Josh, that's not very fancy.
188
Don't worry, Squatch, we'll get there.
189
Anyway, now when we plug a value, X, into the input, we get WX on the output on top.
190
And, because we still have dropout on the bottom, we have zero in the output down there.
191
Now we plug the output, wx, into the sigmoid probability function.
192
And that gives us a weighted average like before.
193
And, if we stopped right now, we'd have something very close to an activation function called swish.
194
The big difference between the weighted average we have now and the swish is the swish just has the input value,
195
x, outside of the probability function instead of wx.
196
So, instead, we'll call our weighted average swishish.
197
Also, while we're at it, most people, including the maintainers of the PyTorch documentation,
198
just use SILU to refer to both SILU and swish.
199
Anyway, regardless of what it's called, we're not going to stop here.
200
Instead, we'll add another copy of our input value, x, and multiply it by a different weight, v.
201
Lastly, we multiply vx by the weighted average, the swish-ish thing, to get the final output value, y.
202
Whoa, that's pretty fancy.
203
Agreed.
204
Because this bypass multiplication term, Vx, which is used by another activation function called a gated linear unit,
205
or glue, is combined with something that is swish-ish, we call the whole thing the swiglu activation function.
206
Boom!
207
These additional weights and multiplication add a lot more flexibility to the shapes that we can get for the activation functions.
208
But Josh, why would this be better?
209
Well, to quote the authors of the SuiGlu activation function, we offer no explanation as to why these architectures seem to work.
210
Small bam.
211
That said, I have two ideas for why they work.
212
First, the SuiGlu requires a lot more computation than the ReLU.
213
As a result, we tend to make smaller neural networks with Swiglu.
214
And a smaller neural network may not overfit the training data as much as a larger one.
215
Second, and perhaps more importantly, throughout the history of neural networks, one recurring theme is that the more we,
216
as people, get out of the way and don't tell the network what to do, the better they perform.
217
And with the Swiglu activation function, with its extra trainable parameters, we're essentially saying, Neural network,
218
you figure out what shape you need for the activation function.
219
Anyway, since it seems to work, the Swiglu variation is commonly used in AI, especially AI models created by Meta.
220
In summary, the idea for these activation functions is that they are all based on the weighted average of dropout.
221
And, if anything, the big lesson here is that neural networks are still an area of very active experimentation.
222
People are trying new things all the time and seeing what happens.
223
If it works, then they publish their new ideas.
224
Triple bam!
225
Now it's time for some shameless self-promotion.
226
If you want to review statistics, machine learning, and AI offline, check out the StatQuest PDF study guides and my best-selling books on machine learning,
227
neural networks and AI, and statistics at StatQuest.org.
228
There's something for everyone.
229
Hooray!
230
We've made it to the end of another exciting StatQuest.
231
If you like this StatQuest and want to see more, please subscribe.
232
And if you want to support StatQuest, consider contributing to my Patreon campaign, becoming a channel member, buying one or two of my original songs or a t-shirt or a hoodie,
233
or just donate.
234
The links are in the description below.
235
Alright, until next time, quest on!

Lo scenario: Imparare l'inglese con i video su funzioni di attivazione

Immagina di guardare un video in inglese su argomenti tecnici, come le funzioni di attivazione GelU, SiLU e SwiGLU. Il narratore parla in modo chiaro ma veloce, con espressioni colloquiali e terminologia specifica. Questo scenario è perfetto per esercitare l'ascolto e la pronuncia, poiché combina lingua tecnica con un tono accessibile. Ecco come sfruttarlo al massimo per la pratica di conversazione in inglese e il shadowing in inglese.

Frasi utili e collocazioni da memorizzare

  • "Too long, didn't watch": Espressione informale per dire che qualcosa è troppo lungo e non l'hai guardato interamente.
  • "closely related to": Legato strettamente a (utile per confrontare concetti, come nel video: "SWIGLU è closely related to GELU e SiLU").
  • "as a result of": A causa di (es. "as a result of these differences, the model performs better").
  • "scaling up": Ampliare (riferito a reti neurali, ma usabile in contesti generali: "scaling up a project").
  • "take small steps towards": Fare piccoli passi verso (es. "gradient descent takes small steps towards the optimal solution").

La tua sfida di shadowing: "Shadowspeak" in azione

Per esercitarti con il shadowspeak, segui questi passaggi: 1. Riproduci un frammento del video (5-10 secondi) e ascolta attentamente l'intonazione e il ritmo. 2. Ripeti immediatamente dopo, cercando di imitare la pronuncia e la velocità. 3. Focus su espressioni come "Hooray!" o "Baaam!" per catturare il tono emotivo. 4. Ripeti il processo finché non senti che la tua replica è fluida. Questo esercizio migliora non solo la pronuncia, ma anche la comprensione auditiva in contesti reali. Provalo ora con il video: è il modo più efficace per imparare l'inglese con i video!

La grammatica di questo video

Le strutture che chi parla usa di più, con le parole esatte del video:

StrutturaNel video
Frasi condizionali if + frase, will/would + verbo — una condizione e il suo risultatoif we stopped right now, we'd have
Forma passiva be + participio passato — conta ciò che accade, non chi lo fais brought · is constrained · is needed

Cos'è la tecnica dello Shadowing?

Shadowing è una tecnica di apprendimento delle lingue supportata da studi scientifici, originariamente sviluppata per la formazione dei traduttori professionisti e resa popolare dal poliglotta Dr. Alexander Arguelles. Il metodo è semplice ma potente: ascolti un audio in inglese di madrelingua e lo ripeti immediatamente ad alta voce — come un'ombra che segue il parlante con un ritardo di solo 1–2 secondi. A differenza dell'ascolto passivo o degli esercizi di grammatica, lo shadowing costringe il tuo cervello e i muscoli della bocca a elaborare e riprodurre simultaneamente i modelli di discorso reale. La ricerca dimostra che migliora significativamente la precisione della pronuncia, l'intonazione, il ritmo, il discorso connesso, la comprensione dell'ascolto e la fluidità del parlato — rendendolo uno dei metodi più efficaci per la preparazione alla prova di speaking dell'IELTS e per la comunicazione reale in inglese.

Tecnica dello shadowing: leggi la guida completa passo dopo passo →