Prática de Shadowing: World Models explained in 10min.. - Aprenda a falar inglês com vídeo
Criando lição...
1
If I flip this coin, you know that the odds are going to be 50-50.
2
And you know that as a fact without using fancy notations because you grew up observing the laws of physics over time.
3
But unlike us, large language models don't have this luxury because LLMs don't have a simulated environment to test out its theory.
4
The closest thing LLMs have to that is reasoning models that use chain of thought to reason its thinking.
5
So therein lies the question, are large language models inherently flawed in their ability to grasp the physical world?
6
And how exactly does world models actually overcome this?
7
Welcome to Kilbride's Code, where every second counts.
8
Quick shout out to BiCloud, more on him later.
9
On a recent podcast with Dario Amodei, he talked about the current approach of pre-training in large English models.
10
Unlike humans, LLMs are trained using trillions and trillions of tokens over months in time.
11
And in comparison, humans not only develop a lot slower, but they also experience the world in modalities other than trillions of text tokens.
12
For example, we know that the laws of physics govern the physical world, and our senses observe and act in the physical world,
13
and we map our understanding about the physical world as we embody them in our own brain.
14
But pure Lars English models are trained only using text, which is the highest abstraction that describe the physical world.
15
So if LLMs never really experienced a coin flip, how does it really know whether a coin flip by law
16
of large numbers will eventually converge to 50-50 without really observing them in the physical world.
17
Around 2018, there was a resurgence of what's called world model.
18
And world models approach things very differently than LLMs in how it modeled its understanding.
19
What if instead of feeding an AI model streams and streams of text tokens,
20
we train them to essentially simulate the physical world in its own brain that best represent the physical world?
21
As you can see, being able to pull this kind of thing requires a thorough understanding of the laws of physics
22
and cause and effect that best mimic the physical world in its own model.
23
So how exactly does a world model do this?
24
The original paper by David Ha uses three main components.
25
You start with an environment and introduce a vision model that essentially observes this environment.
26
The vision model's underlying architecture uses variational autoencoder
27
that is trained to essentially compress what it observes visually into a lower dimension in latent space.
28
What this compression does is that it trains the vision model to extract only the important features and throw out the rest.
29
So now that we've set up our vision model that observes the physical world, we now need a model that can process this information.
30
They use MDN RNN as the underlying architecture.
31
And since recurrent neural networks are really expressive in being able to keep track of all of its previous hidden states,
32
What this architecture allows is to store all the previous hidden states that it saw in the past,
33
which allows the model to then predict based on what it saw before and based on what it's currently seeing.
34
And the MDN RNN architecture works very similar to this app Sketch RNN, where once I start drawing a picture of my dog Logan,
35
it'll take over and actually finish my drawing based on its prediction of what the rest of the drawing would look like.
36
The final piece is the controller model where it samples from NDN RNN's output, as well as the original outputs from the vision model.
37
And the sole purpose of the controller model is just to make an action, like passing butter or moving left and right or picking up objects.
38
As you can see, these three components in a world model will learn to interact with the physical world
39
and map out what it thinks the physical world looks like after seeing the cause
40
and effect of the real world
41
and eventually its own representation of the world will be sufficient that you can just sever the actual environment.
42
What this allows is for the agent to be trained solely
43
by the simulation of the world model's inner mapping of the real world to train the agent to do whatever you want.
44
And the assertion being made here is
45
that this kind of modeling fits much closer to how humans think and gets us much closer to AGI than LLMs do.
46
And the initial result was pretty promising where the world model
47
was able to learn how to drive on this randomly generated track by staying on the road,
48
while only using less than 5 million parameters in total size when we sum up all its parts.
49
So the real question here is, can world models actually scale to human-level capabilities or even more?
50
And here's a quick shout-out to ByCloud.
51
If you want to learn more about the theory behind AI and AI research, check out Bycloud's intuitive AI that is full of learning materials.
52
You can start from beginner level to understand all the way from how tokens work to embeddings,
53
encodings, and attention mechanisms that power most large language models today.
54
He really mixes in good illustrations while giving an easy-to-read narrative on how the technology actually works intuitively.
55
You don't need to have a deep math background.
56
It'll just read like a novel where you can sequentially learn from the beginning
57
or just use it as a supplemental tool on areas that you're curious about.
58
He goes through different pre-training and post-training mechanism here, as well as more advanced concepts like LoRa.
59
For you guys, BuyCloud has given away a 40% discount on the yearly plan using the coupon called Caleb.
60
Link in the description below.
61
One of the biggest reasons why large language models that we use today are so popular is because LLMs scaled quite beautifully.
62
While world models are still domain-specific to certain jobs, most LLMs are what's called foundation models,
63
which basically means that we can use a generic LLM like GPT 5.2
64
or OPUS 4.6 to do many downstream tasks beyond simple chat, but run deep research, software development, manage our computers, and more.
65
But ever since the resurgence of role models in 2018, we also had many iterations and different flavors of role models too.
66
One of the biggest proponents of role models is Yen LeCun, who worked on role models in Meta as he contributed to JEPA-based models.
67
He later left Meta to start his own company called AMI, or Advanced Machine Intelligence that is seeking up to $5 billion valuation.
68
And Yen has been patronizing LLMs for quite a while now,
69
saying how LLMs simply don't understand the physical world beyond its autoregressive nature that predicts token after token.
70
But we know that language actually contains more than what Yen probably gives credit for.
71
Languages not only represent the physical world in words, they also contain grammars that dictate meanings,
72
figures of speech that provide abstract understanding of the physical world, and the individual units of words like nouns,
73
adjectives, adverbs, they all reveal facts about the physical world.
74
But meanwhile, the gap between pure LLMs and world models have also been blurred as of around 2023, as models like GPT-4 from OpenAI,
75
Gemini 1 from Google introduced what's called multi-modality,
76
where vision language models that can also perceive images using cross-attention help LLMs also perceive, so to speak.
77
And conversely, we also have VLA or Vision Language Action, which is a type of a world model that uses vision transformers with LLMs to create action tokens.
78
And this kind of setup is what powers Neo, the humanoid that was released back in October 2025 that became viral back then.
79
But many people still criticize multimodal LLMs because, well, at the core, it still is an LLM that lacks spatial awareness about the physical world.
80
And Fei Fei Li wanted to demonstrate spatial intelligence through her startup World Labs back in September 2024, which raised more than $230 million dollars.
81
World Labs released a product called Marble that essentially creates Gaussian splats
82
that produce millions and millions of these particles to interact with that are quite beautiful to work with.
83
So just like how we saw in the original paper, what you're seeing here is a world model's representation of the actual environment
84
but mapped in its own model that we're able to see here.
85
But what's different with Marble in comparison to traditional world models is the absence of controllers
86
that actually grapple with the physics of the generated world.
87
This kind of sentiment is expressed by Pim, who founded General Intuition in October 2025,
88
that's trying to create a closer representation of a world model that can actually interact with games and simulations.
89
Google has also been a huge contributor in this space.
90
When we look at SEMA in March 2024, and SEMA 2 in November 2025, and of course, their most recent Genie 3,
91
that creates a hyper-realistic world that we can move around in.
92
As you can see, the world model's depiction of the physical world is what allows us to generate AI videos on Sora, train cars in a simulation,
93
and align robots in factories.
94
NVIDIA also has a hand in this by providing an open-source platform called Cosmos, which is a world foundation model.
95
Similar to how foundation models and LLMs allow a generic model to be used across many downstream tasks,
96
NVIDIA's Cosmos provide tools for developers more upstream, where we can use three pre-trained models to cover various use cases,
97
mostly around data augmentation and data generation, for downstream training like custom post-training like autonomous vehicles, robots, and video agents.
98
Now that we covered the landscape of world models architecture and different approaches that are taken by other companies, we are still left with this rather philosophical question.
99
Do LLMs really demonstrate thinking and understanding like humans do, or does it even matter that they actually think like humans?
100
Do you think that LLMs and world models are mutually exclusive, or do they just solve different problems?
101
What is the best way to augment intelligence artificially?
102
What do you think?
✨ Vídeo recomendado
Sobre esta lição
Você está praticando inglês com "World Models explained in 10min.." usando a técnica de Shadowing.
O que é a Técnica de Shadowing?
Shadowing é uma técnica de aprendizado de idiomas com base científica, originalmente desenvolvida para o treinamento de intérpretes profissionais. O método é simples, mas poderoso: você ouve áudio em inglês nativo e repete imediatamente em voz alta — como uma sombra seguindo o falante com 1-2 segundos de atraso. Pesquisas mostram melhora significativa na precisão da pronúncia, entonação, ritmo, sons conectados, compreensão auditiva e fluência na fala.











