Luyện nói tiếng Anh bằng Shadowing qua video: World Models explained in 10min..

Đang tạo bài học...
1
If I flip this coin, you know that the odds are going to be 50-50.
2
And you know that as a fact without using fancy notations because you grew up observing the laws of physics over time.
3
But unlike us, large language models don't have this luxury because LLMs don't have a simulated environment to test out its theory.
4
The closest thing LLMs have to that is reasoning models that use chain of thought to reason its thinking.
5
So therein lies the question, are large language models inherently flawed in their ability to grasp the physical world?
6
And how exactly does world models actually overcome this?
7
Welcome to Kilbride's Code, where every second counts.
8
Quick shout out to BiCloud, more on him later.
9
On a recent podcast with Dario Amodei, he talked about the current approach of pre-training in large English models.
10
Unlike humans, LLMs are trained using trillions and trillions of tokens over months in time.
11
And in comparison, humans not only develop a lot slower, but they also experience the world in modalities other than trillions of text tokens.
12
For example, we know that the laws of physics govern the physical world, and our senses observe and act in the physical world,
13
and we map our understanding about the physical world as we embody them in our own brain.
14
But pure Lars English models are trained only using text, which is the highest abstraction that describe the physical world.
15
So if LLMs never really experienced a coin flip, how does it really know whether a coin flip by law
16
of large numbers will eventually converge to 50-50 without really observing them in the physical world.
17
Around 2018, there was a resurgence of what's called world model.
18
And world models approach things very differently than LLMs in how it modeled its understanding.
19
What if instead of feeding an AI model streams and streams of text tokens,
20
we train them to essentially simulate the physical world in its own brain that best represent the physical world?
21
As you can see, being able to pull this kind of thing requires a thorough understanding of the laws of physics
22
and cause and effect that best mimic the physical world in its own model.
23
So how exactly does a world model do this?
24
The original paper by David Ha uses three main components.
25
You start with an environment and introduce a vision model that essentially observes this environment.
26
The vision model's underlying architecture uses variational autoencoder
27
that is trained to essentially compress what it observes visually into a lower dimension in latent space.
28
What this compression does is that it trains the vision model to extract only the important features and throw out the rest.
29
So now that we've set up our vision model that observes the physical world, we now need a model that can process this information.
30
They use MDN RNN as the underlying architecture.
31
And since recurrent neural networks are really expressive in being able to keep track of all of its previous hidden states,
32
What this architecture allows is to store all the previous hidden states that it saw in the past,
33
which allows the model to then predict based on what it saw before and based on what it's currently seeing.
34
And the MDN RNN architecture works very similar to this app Sketch RNN, where once I start drawing a picture of my dog Logan,
35
it'll take over and actually finish my drawing based on its prediction of what the rest of the drawing would look like.
36
The final piece is the controller model where it samples from NDN RNN's output, as well as the original outputs from the vision model.
37
And the sole purpose of the controller model is just to make an action, like passing butter or moving left and right or picking up objects.
38
As you can see, these three components in a world model will learn to interact with the physical world
39
and map out what it thinks the physical world looks like after seeing the cause
40
and effect of the real world
41
and eventually its own representation of the world will be sufficient that you can just sever the actual environment.
42
What this allows is for the agent to be trained solely
43
by the simulation of the world model's inner mapping of the real world to train the agent to do whatever you want.
44
And the assertion being made here is
45
that this kind of modeling fits much closer to how humans think and gets us much closer to AGI than LLMs do.
46
And the initial result was pretty promising where the world model
47
was able to learn how to drive on this randomly generated track by staying on the road,
48
while only using less than 5 million parameters in total size when we sum up all its parts.
49
So the real question here is, can world models actually scale to human-level capabilities or even more?
50
And here's a quick shout-out to ByCloud.
51
If you want to learn more about the theory behind AI and AI research, check out Bycloud's intuitive AI that is full of learning materials.
52
You can start from beginner level to understand all the way from how tokens work to embeddings,
53
encodings, and attention mechanisms that power most large language models today.
54
He really mixes in good illustrations while giving an easy-to-read narrative on how the technology actually works intuitively.
55
You don't need to have a deep math background.
56
It'll just read like a novel where you can sequentially learn from the beginning
57
or just use it as a supplemental tool on areas that you're curious about.
58
He goes through different pre-training and post-training mechanism here, as well as more advanced concepts like LoRa.
59
For you guys, BuyCloud has given away a 40% discount on the yearly plan using the coupon called Caleb.
60
Link in the description below.
61
One of the biggest reasons why large language models that we use today are so popular is because LLMs scaled quite beautifully.
62
While world models are still domain-specific to certain jobs, most LLMs are what's called foundation models,
63
which basically means that we can use a generic LLM like GPT 5.2
64
or OPUS 4.6 to do many downstream tasks beyond simple chat, but run deep research, software development, manage our computers, and more.
65
But ever since the resurgence of role models in 2018, we also had many iterations and different flavors of role models too.
66
One of the biggest proponents of role models is Yen LeCun, who worked on role models in Meta as he contributed to JEPA-based models.
67
He later left Meta to start his own company called AMI, or Advanced Machine Intelligence that is seeking up to $5 billion valuation.
68
And Yen has been patronizing LLMs for quite a while now,
69
saying how LLMs simply don't understand the physical world beyond its autoregressive nature that predicts token after token.
70
But we know that language actually contains more than what Yen probably gives credit for.
71
Languages not only represent the physical world in words, they also contain grammars that dictate meanings,
72
figures of speech that provide abstract understanding of the physical world, and the individual units of words like nouns,
73
adjectives, adverbs, they all reveal facts about the physical world.
74
But meanwhile, the gap between pure LLMs and world models have also been blurred as of around 2023, as models like GPT-4 from OpenAI,
75
Gemini 1 from Google introduced what's called multi-modality,
76
where vision language models that can also perceive images using cross-attention help LLMs also perceive, so to speak.
77
And conversely, we also have VLA or Vision Language Action, which is a type of a world model that uses vision transformers with LLMs to create action tokens.
78
And this kind of setup is what powers Neo, the humanoid that was released back in October 2025 that became viral back then.
79
But many people still criticize multimodal LLMs because, well, at the core, it still is an LLM that lacks spatial awareness about the physical world.
80
And Fei Fei Li wanted to demonstrate spatial intelligence through her startup World Labs back in September 2024, which raised more than $230 million dollars.
81
World Labs released a product called Marble that essentially creates Gaussian splats
82
that produce millions and millions of these particles to interact with that are quite beautiful to work with.
83
So just like how we saw in the original paper, what you're seeing here is a world model's representation of the actual environment
84
but mapped in its own model that we're able to see here.
85
But what's different with Marble in comparison to traditional world models is the absence of controllers
86
that actually grapple with the physics of the generated world.
87
This kind of sentiment is expressed by Pim, who founded General Intuition in October 2025,
88
that's trying to create a closer representation of a world model that can actually interact with games and simulations.
89
Google has also been a huge contributor in this space.
90
When we look at SEMA in March 2024, and SEMA 2 in November 2025, and of course, their most recent Genie 3,
91
that creates a hyper-realistic world that we can move around in.
92
As you can see, the world model's depiction of the physical world is what allows us to generate AI videos on Sora, train cars in a simulation,
93
and align robots in factories.
94
NVIDIA also has a hand in this by providing an open-source platform called Cosmos, which is a world foundation model.
95
Similar to how foundation models and LLMs allow a generic model to be used across many downstream tasks,
96
NVIDIA's Cosmos provide tools for developers more upstream, where we can use three pre-trained models to cover various use cases,
97
mostly around data augmentation and data generation, for downstream training like custom post-training like autonomous vehicles, robots, and video agents.
98
Now that we covered the landscape of world models architecture and different approaches that are taken by other companies, we are still left with this rather philosophical question.
99
Do LLMs really demonstrate thinking and understanding like humans do, or does it even matter that they actually think like humans?
100
Do you think that LLMs and world models are mutually exclusive, or do they just solve different problems?
101
What is the best way to augment intelligence artificially?
102
What do you think?

Từ vựng và ghi chú luyện nói cho bài học này

Bài luyện nói trình độ C1 này dựa trên video “World Models explained in 10min.”. Những từ được nhắc lại nhiều nhất trong bài: model, Llm, physical, vision, language. Video này có 102 câu và 1715 từ để luyện shadowing. Phần lời nói dài 9:52. Người nói nói với tốc độ tự nhiên, khoảng 174 từ mỗi phút, gần với hội thoại hằng ngày. Chỉ 81% số từ nằm trong 3.000 từ tiếng Anh thông dụng nhất, nên từ vựng khá khó.

Từ vựng quan trọng trong video

15 từ khó nhất trong video, kèm phiên âm và nghĩa:

TừPhiên âmNghĩa
observe động từ/əbˈzɜːv/quan sát
interact động từ/ɪn.təˈɹækt/tương tác
spatial tính từ/ˈspeɪ.ʃəl/không gian
shout động từ/ʃaʊt/kêu la, la hét
predict động từ/pɹɪˈdɪkt/dự báo
adjective danh từ/ˈæd͡ʒ.ɪk.tɪv/tính từ, hình dung từ
encoding danh từ/ɪnˈkoʊdɪŋ/biên mã
iteration danh từ/ˌɪt.əˈɹeɪ.ʃən/phép lặp
latent tính từ/ˈleɪ.tənt/ngầm, ngấm ngần
transformer danh từ/tɹænsˈfoɹməɹ/máy biến thế, máy biến áp
beginner danh từ/bəˈɡɪnɚ/người bắt đầu, người mới học
conversely trạng từ/kənˈvɝsli/ngược lại
noun danh từ/naʊn/danh từ
criticize động từ/ˈkɹɪtɪsaɪz/chỉ trích
govern động từ/ˈɡʌv.ən/cai trị, quản trị

Cụm động từ bạn sẽ nghe

TừNghĩa
check out động từtính tiền, trả phòng
pick up động từnhặt, lượm

Ngữ pháp trong video

Những cấu trúc người nói dùng nhiều nhất, kèm đúng cụm từ trong video:

Cấu trúcTrong video
Câu bị động be + quá khứ phân từ — nhấn vào việc xảy ra, không phải người làmare trained · is trained · be trained
Thì hiện tại hoàn thành have/has + quá khứ phân từ — việc đã xảy ra nhưng còn liên quan đến hiện tạiwe've set · has given · have also been blurred
Mệnh đề quan hệ who / which + mệnh đề — thêm thông tin về người hoặc vậttext, which is · Action, which is · Cosmos, which is

Phát âm cần chú ý

Người nói dùng 10 dạng rút gọn, ví dụ don't, it'll, you're. Hãy nói theo dạng ngắn đúng như bạn nghe.

  • Âm “th”: therein /ˈðɛəɹˈɪn/, thorough /ˈθʌɹoʊ/
  • Âm “sh” và “zh”: simulation /ˌsɪm.jəˈleɪ.ʃn̩/, spatial /ˈspeɪ.ʃəl/, shout /ʃaʊt/, abstraction /æbˈstɹæk.ʃn̩/, depiction /dɪˈpɪkʃən/
  • Từ dài — đặt trọng âm cho đúng: iteration /ˌɪt.əˈɹeɪ.ʃən/, sequentially /səˈkwɛntʃəli/, supplemental /ˌsʌplɪˈmɛntəl/, parameter /pəˈɹæm.ə.tɚ/, intuitive /ɪnˈtjuːɪtɪv/

Cách luyện với video này

  1. Nghe hết video một lần, chưa cần nói, và ghi lại những từ bạn chưa biết.
  2. Bắt đầu ở tốc độ 0,75×, nói đuổi từng câu, rồi quay lại tốc độ bình thường khi đã quen.
  3. Ghi âm giọng mình rồi so với bản gốc, chú ý các từ như observe, interact, spatial.

Phương Pháp Shadowing Là Gì?

Shadowing là kỹ thuật học ngôn ngữ có cơ sở khoa học, ban đầu được phát triển cho chương trình đào tạo phiên dịch viên chuyên nghiệp và được phổ biến rộng rãi bởi nhà đa ngôn ngữ học Dr. Alexander Arguelles. Nguyên lý cốt lõi đơn giản nhưng cực kỳ hiệu quả: bạn nghe tiếng Anh của người bản xứ và lặp lại to ngay lập tức — như một "cái bóng" (shadow) đuổi theo người nói với độ trễ chỉ 1–2 giây. Khác với luyện ngữ pháp hay học từ vựng bị động, Shadowing buộc não bộ và cơ miệng phải đồng thời xử lý và tái tạo ngôn ngữ thực tế. Các nghiên cứu khoa học xác nhận phương pháp này cải thiện đáng kể phát âm, ngữ điệu, nhịp điệu, nối âm, kỹ năng nghe và độ lưu loát khi nói — đặc biệt hiệu quả cho người luyện IELTS Speaking và muốn giao tiếp tiếng Anh tự nhiên như người bản ngữ.

Phương pháp shadowing: đọc hướng dẫn từng bước đầy đủ →