शैडोइंग अभ्यास: Mixture of Experts Explained Visually: How Trillion-Parameter Models Actually Work - वीडियो के साथ अंग्रेजी बोलना सीखें
पाठ बनाया जा रहा है...
1
If you look at the architecture of recent open-weight models like DeepSeq,
2
Kimi, GLM, and Quen, they all share one key design, mixture of experts.
3
DeepSeq V3 has 671 billion parameters.
4
Kimi K3 has 2.8 trillion.
5
Yet for any single token, only a tiny fraction of those parameters actually run.
6
That's the key idea behind mixture of experts.
7
scale total model capacity dramatically without increasing active compute.
8
The trick is replacing one giant monolithic network with multiple specialized subnetworks called experts.
9
For each token, only a small subset of these experts is activated.
10
This turns dense computation into sparse execution, giving you the intelligence of a massive model at the speed and cost of a much smaller one.
11
By the end of this video, you'll understand how mixture of experts works, the main problems it creates during training, how it executes efficiently on GPUs,
12
and how modern MOE models solve them.
13
A transformer block has two main components, self-attention and the feed-forward network.
14
Self-Attention handles token-to-token communication across the sequence.
15
For example, when the model reads it, attention looks back at animal to gather context.
16
Because every token compares against surrounding words, this step remains dense even in MOE models.
17
After that, we pass the output through a residual connection, adding it back to the original hidden representation.
18
Then, layer normalization keeps the combined representation stable before it's passed into the feedforward network.
19
This part of the transformer works completely differently.
20
It processes each token individually with no token interaction and holds over two-thirds of the model's total parameters.
21
So this is where we want to make the computation sparse and why mixture of experts replaces the feedforward network.
22
Let's take a single token, tired, and see how the feedforward network processes it.
23
We'll call this hidden vector x.
24
It has the same dimension as the model's hidden size, dModel.
25
To expand it into the hidden layer, we multiply x by the up-projection matrix and add a bias vector.
26
This projects the token into a higher dimensional space, called dff, the feed forward dimension.
27
In the original transformer paper, dff is four times the model dimension, giving the layer more space to learn diverse features.
28
Each row in the up-projection matrix corresponds to a single hidden neuron.
29
You can think of each neuron as responding to a learned feature.
30
When Vector X enters, the dot product measures how strongly the token aligns with that neuron's pattern.
31
One neuron might detect emotional state, and another detects math concepts, and another detects grammar.
32
When the word tired passes through, neurons detecting emotional state and fatigue fire strongly with large positive scores.
33
Meanwhile, neurons looking for code syntax score negative or zero because those features aren't present in our input token.
34
These raw scores then pass through a non-linear activation function.
35
For simplicity, we'll use RELU, the activation function from the original transformer paper.
36
RELU sets negative scores to zero, creating a sparse token-specific pattern.
37
Finally, we project this activation pattern back to the model dimension using the down projection matrix and bias.
38
This combines the active features into an updated token vector Y.
39
Since tokens are independent, this step runs in parallel across the whole batch.
40
To make the model memorize more facts and learn complex concepts, we can expand this hidden layer.
41
Wider hidden layers give the network more neurons to detect a broader set of features.
42
Here's the problem.
43
In a dense feedforward network, every token multiplies against every parameter.
44
Scale the parameter count by 10x and compute scales by 10x.
45
That quickly gets too expensive.
46
Notice what's happening.
47
Activations are naturally sparse.
48
A coding token lights up completely different neurons than a medical token.
49
For any single token, most neurons sit at zero, so running the full dense matrix for every token wastes huge amounts of compute.
50
So how do we exploit this sparsity?
51
Think of it like a university.
52
Imagine a university with dozens of specialized departments.
53
If you have a question about quantum mechanics, you go to the physics department.
54
If you have a question about medieval history, you go to the history department.
55
You don't ask every department on campus.
56
That would be completely inefficient.
57
Language models work the same way.
58
Different tokens only need a small subset of parameters.
59
That's why we split the massive feedforward network into smaller subnetworks called experts.
60
In real models, expert specialization isn't labeled by humans.
61
It emerges automatically during training.
62
But labeling them here makes the concept clear.
63
Each token is dynamically routed to only the experts it needs.
64
To decide which experts should process each token, we introduce a small neural network called the router.
65
First, it takes the token's hidden representation coming directly from the residual stream
66
and multiplies it by a learnable weight matrix to produce a routing logit for each expert.
67
Softmax then turns these logits into probabilities across all experts.
68
For instance, for this specific token, expert 2 receives an 80% routing score
69
while expert 1 gets just three percent meaning expert two is far more relevant
70
and will carry most of the weight in the final output
71
once routed each expert independently processes the token to generate its
72
own candidate output we then scale each of these outputs by
73
the router's corresponding confidence score finally we sum them together to form the final result but
74
if every expert processes every token we're still doing the full
75
computation to solve this we introduce sparsity we keep only the top scoring experts
76
and set the gating weights of all remaining experts to zero.
77
That way, only chosen experts run, keeping active compute low even as total parameters scale.
78
There are a few ways we can handle sparse routing.
79
One direction is top one routing, pioneered by the switch transformer.
80
This is the most aggressive masking strategy.
81
We retain only the single highest scoring expert and zero out everything else.
82
While incredibly efficient, it introduces a severe structural limitation.
83
Real-world data is inherently messy and complex.
84
A single token might encapsulate multiple concepts simultaneously, such as mathematics and programming.
85
Forcing that token through a single expert severely bottlenecks the model's expressive capacity.
86
To overcome this, we use top-k routing, where the router selects the top-k experts for each token.
87
This has become the standard paradigm for modern large language models.
88
For instance, MixedRoll routes to the top two experts, while DeepSeqv3 routes to the top eight experts.
89
This lets the token combine insights from multiple experts at once, blending their outputs based on the router's confidence.
90
While it uses slightly more compute than top one routing, it is far more efficient than running the full dense network.
91
But training these networks is where things get tricky.
92
The router initializes with random weights, so its early routing decisions are essentially random.
93
Suppose Expert 1 happens to receive slightly more tokens.
94
Under backpropagation, an expert only receives gradient updates when tokens are routed to it.
95
If it sits idle, its parameters never update and it learns nothing.
96
Because Expert 1 processes more tokens, it accumulates more gradient updates,
97
so it learns faster and becomes better at minimizing the loss than its peers.
98
The router quickly learns that routing to Expert 1 yields a lower error, so in subsequent forward passes, it assigns even more tokens to it.
99
This creates a positive feedback loop called Expert Collapse.
100
As training runs, the router keeps picking the same expert.
101
Soon, one expert handles everything while the rest stay idle, wasting most of the experts we added.
102
But we can see this clearly in a deep seek ablation study.
103
The expert shown in green gradually takes over until it's receiving almost all of the tokens.
104
This happens because standard router training only optimizes the language modeling loss,
105
the cross entropy loss that measures how well the model predicts the MEX token.
106
But this loss doesn't care how tokens are distributed across the experts.
107
One simple idea is to add another loss function, the auxiliary load balancing loss.
108
To calculate this balancing penalty, we first need to measure the actual token distribution.
109
For each expert, we simply take the number of tokens routed to it
110
and divide it by the total tokens in the batch.
111
This gives us the token fraction, a measure of how much traffic that specific expert handled.
112
Now we look at the router's confidence.
113
We compute the average routing probability for that same expert across the entire batch.
114
This tells us, on average, how strongly the router favored this expert, regardless of whether the token was actually assigned to it.
115
We then multiply this actual token fraction by the average routing probability for each expert.
116
Summing these products across all experts and scaling by the total number of experts gives us our auxiliary load balancing loss.
117
Intuitively, this penalizes the model when the router is highly confident in an expert that is already receiving too much traffic.
118
In an ideal scenario where routing is perfectly balanced, say every expert receives 25% of the traffic,
119
the auxiliary loss evaluates to its theoretical minimum of 1.
120
Conversely, in a collapsed state where a single expert hoards all the tokens,
121
the loss spikes significantly, sending a strong penalty signal to correct the router's behavior.
122
During training, we combine these two objectives.
123
Our total loss becomes the primary language modeling plus the auxiliary load balancing loss scaled by a hyperparameter alpha.
124
We are now simultaneously teaching the model to predict accurately and to utilize its experts more evenly.
125
This alpha parameter controls the trade-off between prediction quality and expert utilization.
126
We keep it small, typically around 0.01,
127
so the language loss still dominates while the router gets a small push to distribute tokens more evenly across the experts.
128
And if we look at the results here, we can see the difference.
129
Without the auxiliary loss, the routing becomes concentrated on a few experts.
130
With the balancing penalty, the tokens are distributed much more evenly across the experts.
131
The auxiliary loss effectively stops expert collapse, but it introduces a new problem into to our objective function.
132
The network now has to balance two conflicting goals.
133
Because these gradients optimize different objectives, they produce opposing parameter updates during back propagation.
134
The primary cross entropy loss produces gradients that optimize the router for prediction accuracy.
135
At the same time, the auxiliary loss produces gradients that encourage a more balanced distribution of tokens across the experts.
136
Because these gradients optimize conflicting objectives, they produce opposing parameter updates during back propagation,
137
which degrades routing quality and ultimately hurts final model performance.
138
To avoid this gradient interference, DeepSeq V3 removes the auxiliary loss from the training objective.
139
It enforces load balancing by applying a dynamic bias directly to the routing scores during the forward pass.
140
Standard routers calculate a routing score for each expert by taking the dot product between the token's hidden representation
141
and the corresponding router weight vector.
142
DeepSeq V3 uses a different gating mechanism.
143
It drops the softmax entirely and applies a sigmoid to the router's output, producing an independent gating score for each expert.
144
Then, before top case election, it adds a dynamic bias to each expert's gating score.
145
Here's the important part.
146
This isn't a normal learn parameter.
147
Unlike a typical bias, it isn't updated through back propagation.
148
Instead, DeepSeq updates it directly after each training step using a simple heuristic.
149
The bias adjustment is directly proportional to the load error, the difference between the target load and the expert's actual load for that batch.
150
If an expert's actual load is higher than the target load, its bias is decreased by a small step size,
151
making it less likely to receive tokens in the next batch.
152
On the flip side, if an expert's actual load is below the target load, its bias is increased by that step size,
153
making it more likely to receive tokens until the distribution becomes balanced.
154
While DeepSeq keeps bias updates separate from backpropagation, the logic is heuristic.
155
Add or subtract the fixed step based on observed expert load.
156
This works fine for smaller models, but moving toward extreme sparsity like Kimi K3, this step size is way too slow.
157
Moonshot AI's recently released open-weight model has nearly 3 trillion parameters.
158
Kimi K3 scales to 896 experts, yet activates only 16 of them for any given token.
159
This means each expert is expected to receive less than 2% of the tokens.
160
At this level of sparsity, a fixed bias update might not react quickly enough when the token distribution changes.
161
For example, if the update step is too small, an inactive expert could take thousands of iterations before its bias is high enough to make the top 16.
162
Until then, it gets very few or even no tokens, so it barely gets a chance to learn.
163
On the flip side, if the step size is too large, the bias overcorrects.
164
An expert can quickly go from receiving too few tokens to receiving too many, causing the routing decisions to become unstable.
165
Kimi K3 avoids these incremental updates altogether.
166
Instead, it computes the bias directly from the token distribution in each batch.
167
Here's a simple example.
168
Take a batch of 1,000 tokens and sort their scores from lowest to highest.
169
Since each expert handles 2% of the tokens, we place the cutoff at the 98th percentile.
170
Let's say that cutoff comes out to positive 1.2.
171
This is the expert's quantile.
172
We subtract it from the per-token cutoff, alpha, to get the expert's bias.
173
Because top-k routing depends only on relative ranking, we can set alpha to zero without changing which experts are selected.
174
The bias then effectively shifts each expert's scores according to its own token distribution, helping align the experts to the same target workload.
175
Since an overutilized expert tends to produce higher routing scores, its 98th percentile is also higher, in our example, positive 1.2.
176
Subtracting this positive quantile gives the expert a negative bias.
177
This applies a targeted penalty to curb its token allocation.
178
On the flip side, an underutilized expert has lower scores, giving it a negative quantile, like negative 3.
179
Subtracting this negative quantile gives the expert a positive bias, boosting its token allocation.
180
So, unlike DeepSeq's reactive approach, which updates the bias through small,
181
fixed steps, KimiK3 computes the required bias directly from the batch score distribution in a single step.
182
In large models like KimiK3, all the experts can't fit on a single GPU, so they're spread across multiple GPUs.
183
When a token is routed to an expert, that token has to be sent to the GPU where that expert is located.
184
The bias helps balance this traffic across the GPUs.
185
Once the routing decision is made, the bias is discarded.
186
The final mixture weights are computed using the original scores before the bias was added, keeping output representations unaffected.
187
This lets the router balance the workload without changing what each expert learns.
188
But as models scale up, another problem arises.
189
training stability.
190
Since routing decisions rely on softmax, if raw logits grow too large, the calculation quickly becomes unstable.
191
The router multiplies the token input by its weight matrix, producing raw routing logits for each expert.
192
Softmax turns these logits into probabilities by exponentiating them and dividing by their sum.
193
That denominator is the normalization term, z of x.
194
As training progresses, the router becomes more confident, driving logits higher.
195
Say a logit reaches 12.
196
To save memory and compute, LLMs use mixed precision training, like FP16, where numbers are capped at 65,504.
197
Exponentiating a logit of 12 produces over 160,000, exceeding FP16's limit.
198
This causes Z of X to overflow to infinity, making softmax produce NAN and destabilizing training.
199
To prevent this during the forward pass, we subtract the constant C from every logit before exponentiating.
200
Softmax is translation invariant, so this shift doesn't alter the output probabilities at all.
201
In practice, C is the maximum logit.
202
Subtracting it shifts the highest logit to zero.
203
This keeps every exponentiated value between zero and one, preventing overflow without changing routing decisions.
204
But logit shifting only fixes the forward pass.
205
The underlying logits can still grow over time, making optimization increasingly unstable.
206
To solve this, we introduce router Z loss.
207
This directly penalizes the normalization term Z of X to keep logic growth firmly under control.
208
Because the normalization term Z of X grows exponentially, we first take its logarithm to compress it onto a numerically stable scale.
209
We then square it to guarantee a strictly positive loss that aggressively penalizes massive outliers.
210
Substituting Z of X gives the complete router Z loss formula.
211
Z loss acts as a soft penalty, only stepping in if raw logits grow excessively large.
212
Integrating Z loss into the total objective combines task loss, load balancing penalty, and router Z loss.
213
The hyperparameter beta controls the penalty strength.
214
If it's too small, the logits can still grow too large.
215
If it's too large, the penalty begins to interfere with the model's main training objective.
216
The routing problem is mostly solved.
217
The challenge now is
218
that we can't just run an MOE model on a GPU the same way we run a dense model.
219
GPUs are built for massive parallelism, so running experts one by one leaves most of the hardware idle.
220
To fix this, we run all experts together using batched matrix multiplication.
221
Under the hood, each expert is simply a standard feedforward network, just size smaller so multiple experts can run efficiently.
222
Tokens first pass through the up projection layer, followed by the activation function, and finally through the down projection layer to produce the final output.
223
We'll focus on the up projection layer, where the primary compute bottleneck occurs.
224
We can run them all at once by combining their weight matrices into a single block diagonal matrix.
225
We call this expert capacity.
226
Under top 1 routing, that's just total tokens divided by the number of experts, scaled by our capacity factor.
227
Say we have a batch of 75 tokens routed across 3 experts.
228
At a capacity of 1, we allocate space for 25 tokens per expert.
229
But routing is dynamic and skewed.
230
In our example, Expert 1 gets 35 tokens, while Experts 2 and 3 receive 25 and 20 tokens.
231
While Expert 1 overshoots its capacity, Expert 3 only gets 20 tokens, leaving 5 slots completely empty.
232
That's the biggest drawback of fixed allocation.
233
Since Expert 1 only has room for 25 tokens, the remaining 10 tokens overflow.
234
These excess tokens are dropped and bypass expert computation, flowing directly to the residual stream, which can degrade overall model performance.
235
One quick workaround is to increase the capacity factor.
236
Now, Expert 1 can handle all 35 tokens without dropping any.
237
The catch is that batched matrix multiplication requires uniform shapes,
238
so experts 2 and 3 must also allocate 35 rows even though they only have 25 and 20 tokens.
239
These extra rows are just zero padding, wasting GPU memory and compute.
240
So whether it's dropping tokens or wasting padding, both problems stem from the exact same issue,
241
forcing every expert matrix into identical shapes.
242
To fix this bottleneck without dropping tokens or wasting memory, MegaBlocks introduces non-padded block-sparse matrix multiplication.
243
With MegaBlocks, matrix dimensions don't have to be uniform.
244
It represents both the input and weight matrices as variable-sized blocks that match each expert's actual token count.
245
35 tokens for expert 1, 25 for expert 2, and 20 for expert 3.
246
Notice how all that wasted padding is completely gone.
247
We get the exact same output without dropping tokens or wasting GPU memory.
248
To execute these variable-sized token batches efficiently, MegaBlocks uses block tiling.
249
It partitions each expert's workload into small, fixed-sized blocks.
250
These blocks are sized to match GPU hardware tiles, typically 16 or 32 tokens per block.
251
Take Expert 1.
252
Its 35 tokens are split into two full 16-token blocks plus a partial block with three tokens.
253
Expert 2's 25 tokens are split into two blocks.
254
Expert 3's 20 tokens also form two blocks, giving us seven uniform blocks in total.
255
The GPU no longer sees unbalanced experts.
256
It processes these seven uniform blocks in one continuous stream.
257
The only remaining padding occurs inside partial blocks.
258
We only have to pad a few entries inside one tile, making the overhead negligible.
259
But there's another way to handle load balancing altogether.
260
Instead of having tokens choose their experts, ExpertChoice flips this around.
261
The experts choose which tokens they process.
262
As before, multiplying the input tokens by the routing weight matrix yields the routing score matrix.
263
Under traditional token choice routing, each token evaluates its row to identify its top-scoring experts.
264
Expert choice completely reverses this logic.
265
Instead, each expert evaluates its column and picks the tokens with the highest routing scores.
266
Say we take eight tokens and four experts at a capacity factor of 1.0.
267
Each expert gets a fixed capacity of two tokens and picks the top two tokens from its column.
268
For example, expert one evaluates the first column and picks tokens 1 and 3.
269
Similarly, experts 2, 3, and 4 evaluate their columns and select their top two tokens.
270
By assigning an equal number of tokens to every expert, load balancing is built into the architecture, eliminating the need for auxiliary losses.
271
Computation also adapts to each token.
272
Important tokens can be selected by multiple experts, while simpler tokens are selected once or skipped entirely.
273
But expert choice still relies on hard top-case selection.
274
A token is either picked or ignored.
275
Soft MOE removes that boundary completely.
276
Instead of assigning discrete tokens to experts, every expert receives a soft, weighted blend of the entire batch,
277
making routing fully differentiable, with zero drop tokens.
278
For example, expert 1 doesn't get token A by itself.
279
It gets a mix of token A, token B, token C, and every other token in the batch, with each one contributing by a different amount.
280
Expert 1 processes that mixture, and the output is distributed back to each token based on its original contribution.
281
Another direction is heterogeneous MOE, where experts don't all have to be the same size.
282
Some experts are larger and more capable, while others are lightweight.
283
The router can send complex tokens to the larger experts and simpler tokens to the smaller ones, using the model's parameters more efficiently.
284
So far, we've seen how standard routing works well, but modern architectures push this idea further.
285
They separate the knowledge that's useful across many tokens into a shared pathway, while making the routed experts more specialized.
286
DeepSeq keeps a shared expert active for every token, bypassing the router.
287
It handles common patterns and knowledge, while the routed experts can focus on more specific knowledge.
288
It also changes the size of those routed experts.
289
Instead of using a few large experts, it splits them into smaller micro-experts, a technique called fine-grained expert segmentation.
290
This vastly expands the number of possible expert combinations.
291
With 16 large experts selecting 2, there's 120 possible combinations, but splitting
292
that capacity into 64 smaller experts selecting 8 keeps compute the same while combinations jump to over 4 billion,
293
giving the router much more flexibility.
294
If we look at DeepSeq's ablation study,
295
combining shared experts with fine-grained segmentation consistently improves benchmark performance without increasing inference cost.
296
Moonshot AI also uses shared experts in Kimi K3, running two shared experts alongside its standard routed experts.
297
With 896 standard experts and two shared experts, Kimi K3 reaches 2.8 trillion total parameters.
298
Yet, for each token, it only activates 16 routed experts, using a fraction of its compute, while shared experts handle common language patterns.
299
Now let's put this routing into a full transformer layer.
300
First we turn our input text into tokens.
301
Each token gets mapped to an ID, and an embedding table converts those IDs into vectors.
302
Together these form our input matrix X.
303
We pass matrix X into multi-head self-attention, letting each token gather context from all the others, producing our context-aware matrix A.
304
Next, we add the original input X directly back to A through a skip connection to keep training stable.
305
Then, layer normalization stabilizes the values, giving us our normalized matrix.
306
After attention comes the feedforward network.
307
In a standard transformer, this is one massive set of weights.
308
If this block has 48 billion parameters, every token must run through all 48 billion parameters.
309
It's brute force.
310
With a sparse MOE layer, we split that block into 8 6 billion parameter experts.
311
Using top two routing, each token only activates two experts, running through 12 billion parameters instead of all 48.
312
The router evaluates each token and picks its top experts.
313
For example, HART routes to experts A and D, while PUMPs goes to B and E.
314
Every other token follows the same path.
315
Each token picks its top two experts, weights their outputs by the router scores, and combines them into matrix Y.
316
Then, a second skip connection adds Y back to our normalized matrix.
317
Finally, another layer normalization gives us matrix Z, which feeds straight into the next layer.
318
And that's how an MOE transformer layer works.
319
This is how we scale to trillion parameter models while only spending a fraction of the compute per token.
320
Thanks for watching.
321
If you like this video, please like and subscribe, and I'll see you in the next video.
✨ अनुशंसित वीडियो
इस पाठ के बारे में
आप "Mixture of Experts Explained Visually: How Trillion-Parameter Models Actually Work" के साथ Shadowing तकनीक का उपयोग करके अपनी अंग्रेजी का अभ्यास कर रहे हैं।
शैडोइंग तकनीक क्या है?
शैडोइंग (Shadowing) एक विज्ञान-समर्थित भाषा सीखने की तकनीक है जो मूल रूप से पेशेवर दुभाषिया प्रशिक्षण के लिए विकसित की गई थी। विधि सरल लेकिन शक्तिशाली है: आप मूल अंग्रेज़ी ऑडियो सुनते हैं और तुरंत इसे ज़ोर से दोहराते हैं — जैसे वक्ता की छाया 1-2 सेकंड की देरी से। शोध से पता चलता है कि यह उच्चारण सटीकता, स्वर, लय, जुड़ी हुई ध्वनियाँ, सुनने की समझ और बोलने की प्रवाहशीलता में काफ़ी सुधार करता है।











