シャドーイング練習: World Models explained in 10min.. - 動画で英語スピーキングを学ぶ

レッスンを作成中...
1
If I flip this coin, you know that the odds are going to be 50-50.
2
And you know that as a fact without using fancy notations because you grew up observing the laws of physics over time.
3
But unlike us, large language models don't have this luxury because LLMs don't have a simulated environment to test out its theory.
4
The closest thing LLMs have to that is reasoning models that use chain of thought to reason its thinking.
5
So therein lies the question, are large language models inherently flawed in their ability to grasp the physical world?
6
And how exactly does world models actually overcome this?
7
Welcome to Kilbride's Code, where every second counts.
8
Quick shout out to BiCloud, more on him later.
9
On a recent podcast with Dario Amodei, he talked about the current approach of pre-training in large English models.
10
Unlike humans, LLMs are trained using trillions and trillions of tokens over months in time.
11
And in comparison, humans not only develop a lot slower, but they also experience the world in modalities other than trillions of text tokens.
12
For example, we know that the laws of physics govern the physical world, and our senses observe and act in the physical world,
13
and we map our understanding about the physical world as we embody them in our own brain.
14
But pure Lars English models are trained only using text, which is the highest abstraction that describe the physical world.
15
So if LLMs never really experienced a coin flip, how does it really know whether a coin flip by law
16
of large numbers will eventually converge to 50-50 without really observing them in the physical world.
17
Around 2018, there was a resurgence of what's called world model.
18
And world models approach things very differently than LLMs in how it modeled its understanding.
19
What if instead of feeding an AI model streams and streams of text tokens,
20
we train them to essentially simulate the physical world in its own brain that best represent the physical world?
21
As you can see, being able to pull this kind of thing requires a thorough understanding of the laws of physics
22
and cause and effect that best mimic the physical world in its own model.
23
So how exactly does a world model do this?
24
The original paper by David Ha uses three main components.
25
You start with an environment and introduce a vision model that essentially observes this environment.
26
The vision model's underlying architecture uses variational autoencoder
27
that is trained to essentially compress what it observes visually into a lower dimension in latent space.
28
What this compression does is that it trains the vision model to extract only the important features and throw out the rest.
29
So now that we've set up our vision model that observes the physical world, we now need a model that can process this information.
30
They use MDN RNN as the underlying architecture.
31
And since recurrent neural networks are really expressive in being able to keep track of all of its previous hidden states,
32
What this architecture allows is to store all the previous hidden states that it saw in the past,
33
which allows the model to then predict based on what it saw before and based on what it's currently seeing.
34
And the MDN RNN architecture works very similar to this app Sketch RNN, where once I start drawing a picture of my dog Logan,
35
it'll take over and actually finish my drawing based on its prediction of what the rest of the drawing would look like.
36
The final piece is the controller model where it samples from NDN RNN's output, as well as the original outputs from the vision model.
37
And the sole purpose of the controller model is just to make an action, like passing butter or moving left and right or picking up objects.
38
As you can see, these three components in a world model will learn to interact with the physical world
39
and map out what it thinks the physical world looks like after seeing the cause
40
and effect of the real world
41
and eventually its own representation of the world will be sufficient that you can just sever the actual environment.
42
What this allows is for the agent to be trained solely
43
by the simulation of the world model's inner mapping of the real world to train the agent to do whatever you want.
44
And the assertion being made here is
45
that this kind of modeling fits much closer to how humans think and gets us much closer to AGI than LLMs do.
46
And the initial result was pretty promising where the world model
47
was able to learn how to drive on this randomly generated track by staying on the road,
48
while only using less than 5 million parameters in total size when we sum up all its parts.
49
So the real question here is, can world models actually scale to human-level capabilities or even more?
50
And here's a quick shout-out to ByCloud.
51
If you want to learn more about the theory behind AI and AI research, check out Bycloud's intuitive AI that is full of learning materials.
52
You can start from beginner level to understand all the way from how tokens work to embeddings,
53
encodings, and attention mechanisms that power most large language models today.
54
He really mixes in good illustrations while giving an easy-to-read narrative on how the technology actually works intuitively.
55
You don't need to have a deep math background.
56
It'll just read like a novel where you can sequentially learn from the beginning
57
or just use it as a supplemental tool on areas that you're curious about.
58
He goes through different pre-training and post-training mechanism here, as well as more advanced concepts like LoRa.
59
For you guys, BuyCloud has given away a 40% discount on the yearly plan using the coupon called Caleb.
60
Link in the description below.
61
One of the biggest reasons why large language models that we use today are so popular is because LLMs scaled quite beautifully.
62
While world models are still domain-specific to certain jobs, most LLMs are what's called foundation models,
63
which basically means that we can use a generic LLM like GPT 5.2
64
or OPUS 4.6 to do many downstream tasks beyond simple chat, but run deep research, software development, manage our computers, and more.
65
But ever since the resurgence of role models in 2018, we also had many iterations and different flavors of role models too.
66
One of the biggest proponents of role models is Yen LeCun, who worked on role models in Meta as he contributed to JEPA-based models.
67
He later left Meta to start his own company called AMI, or Advanced Machine Intelligence that is seeking up to $5 billion valuation.
68
And Yen has been patronizing LLMs for quite a while now,
69
saying how LLMs simply don't understand the physical world beyond its autoregressive nature that predicts token after token.
70
But we know that language actually contains more than what Yen probably gives credit for.
71
Languages not only represent the physical world in words, they also contain grammars that dictate meanings,
72
figures of speech that provide abstract understanding of the physical world, and the individual units of words like nouns,
73
adjectives, adverbs, they all reveal facts about the physical world.
74
But meanwhile, the gap between pure LLMs and world models have also been blurred as of around 2023, as models like GPT-4 from OpenAI,
75
Gemini 1 from Google introduced what's called multi-modality,
76
where vision language models that can also perceive images using cross-attention help LLMs also perceive, so to speak.
77
And conversely, we also have VLA or Vision Language Action, which is a type of a world model that uses vision transformers with LLMs to create action tokens.
78
And this kind of setup is what powers Neo, the humanoid that was released back in October 2025 that became viral back then.
79
But many people still criticize multimodal LLMs because, well, at the core, it still is an LLM that lacks spatial awareness about the physical world.
80
And Fei Fei Li wanted to demonstrate spatial intelligence through her startup World Labs back in September 2024, which raised more than $230 million dollars.
81
World Labs released a product called Marble that essentially creates Gaussian splats
82
that produce millions and millions of these particles to interact with that are quite beautiful to work with.
83
So just like how we saw in the original paper, what you're seeing here is a world model's representation of the actual environment
84
but mapped in its own model that we're able to see here.
85
But what's different with Marble in comparison to traditional world models is the absence of controllers
86
that actually grapple with the physics of the generated world.
87
This kind of sentiment is expressed by Pim, who founded General Intuition in October 2025,
88
that's trying to create a closer representation of a world model that can actually interact with games and simulations.
89
Google has also been a huge contributor in this space.
90
When we look at SEMA in March 2024, and SEMA 2 in November 2025, and of course, their most recent Genie 3,
91
that creates a hyper-realistic world that we can move around in.
92
As you can see, the world model's depiction of the physical world is what allows us to generate AI videos on Sora, train cars in a simulation,
93
and align robots in factories.
94
NVIDIA also has a hand in this by providing an open-source platform called Cosmos, which is a world foundation model.
95
Similar to how foundation models and LLMs allow a generic model to be used across many downstream tasks,
96
NVIDIA's Cosmos provide tools for developers more upstream, where we can use three pre-trained models to cover various use cases,
97
mostly around data augmentation and data generation, for downstream training like custom post-training like autonomous vehicles, robots, and video agents.
98
Now that we covered the landscape of world models architecture and different approaches that are taken by other companies, we are still left with this rather philosophical question.
99
Do LLMs really demonstrate thinking and understanding like humans do, or does it even matter that they actually think like humans?
100
Do you think that LLMs and world models are mutually exclusive, or do they just solve different problems?
101
What is the best way to augment intelligence artificially?
102
What do you think?

このレッスンの語彙とスピーキングのポイント

このC1レベルのスピーキングレッスンは、動画「World Models explained in 10min.」を教材にしています。 繰り返し出てくる語は次のとおりです:model, Llm, physical, vision, language。 この動画には、シャドーイング用の文が102文、単語が1715語あります。 音声の長さは9:52です。 話す速さは1分あたり約174語で、日常会話に近い自然なペースです。 英語の頻出3,000語に含まれる単語は81%だけなので、語彙は難しめです。

この動画の重要語彙

動画の中で特に難しい単語15語を、発音と意味つきで紹介します。

単語発音意味
token 名詞/ˈtoʊkən/トークン
observe 動詞/əbˈzɜːv/守る, 従う
downstream 形容詞/ˌdaʊnˈstɹiːm/下流
simulation 名詞/ˌsɪm.jəˈleɪ.ʃn̩/真似, 模擬
interact 動詞/ɪn.təˈɹækt/相互作用する, 相呼応する
resurgence 名詞復活
simulate 動詞/ˈsɪm.jə.lət/シミュレートする
perceive 動詞/pɚˈsiv/知覚する, 認める
spatial 形容詞/ˈspeɪ.ʃəl/空間的, 空間の
shout 動詞/ʃaʊt/叫ぶ
predict 動詞/pɹɪˈdɪkt/予言する
adjective 名詞/ˈæd͡ʒ.ɪk.tɪv/形容動詞
augment 動詞/ɔɡˈmɛnt/増やす
compress 動詞/ˈkɑmpɹɛs/圧縮する
encoding 名詞/ɪnˈkoʊdɪŋ/エンコーディング, 文字コード

動画に出てくる句動詞

単語発音意味
check out 動詞試験する, 検査する
grow up 動詞成長する, 大人に成る
pick up 動詞拾う, 拾い上げる
take over 動詞引き受ける
throw out 動詞/ˈθɹəʊ ˌaʊt/捨てる

この動画の文法

話し手がよく使っている文型を、動画の実際の表現とともに紹介します。

文型動画での表現
受動態 be + 過去分詞 — 誰がするかより、何が起きるかに焦点を当てるare trained · is trained · be trained
現在完了形 have/has + 過去分詞 — 過去の出来事が今も関係しているwe've set · has given · have also been blurred
関係詞節 who / which + 節 — 人や物について情報を加えるtext, which is · Action, which is · Cosmos, which is

注意したい発音

話し手はdon't, it'll, you'reなど、短縮形や弱形を10回使っています。聞こえたとおりの短い形で発音しましょう。

  • 「th」の音: therein /ˈðɛəɹˈɪn/, thorough /ˈθʌɹoʊ/
  • 「sh」と「zh」の音: simulation /ˌsɪm.jəˈleɪ.ʃn̩/, spatial /ˈspeɪ.ʃəl/, shout /ʃaʊt/, abstraction /æbˈstɹæk.ʃn̩/, depiction /dɪˈpɪkʃən/
  • 長い単語(アクセントの位置に注意): iteration /ˌɪt.əˈɹeɪ.ʃən/, sequentially /səˈkwɛntʃəli/, supplemental /ˌsʌplɪˈmɛntəl/, parameter /pəˈɹæm.ə.tɚ/, intuitive /ɪnˈtjuːɪtɪv/

この動画での練習方法

  1. まず声を出さずに動画を最後まで聞き、知らない単語をメモします。
  2. まず0.75倍速で一文ずつシャドーイングし、慣れてきたら通常の速度に戻します。
  3. 自分の声を録音して元の音声と比べます。token, observe, downstreamなどの単語に特に注意しましょう。

シャドーイングとは?英語上達に効果的な理由

シャドーイング(Shadowing)は、もともとプロの通訳者養成プログラムで開発された言語学習法で、多言語習得者として知られるDr. Alexander Arguelles によって広く普及されました。方法はシンプルですが非常に効果的:ネイティブスピーカーの英語を聞きながら、1〜2秒の遅延で声に出してすぐに繰り返す——まるで「影(shadow)」のように話者を追いかけます。文法ドリルや受動的なリスニングと異なり、シャドーイングは脳と口の筋肉が同時にリアルタイムで英語を処理・再現することを強制します。研究により、発音精度、抑揚、リズム、連音、リスニング力、そして会話の流暢さが大幅に向上することが確認されています。IELTSスピーキング対策や自然な英語コミュニケーションを目指す方に特におすすめです。

シャドーイングのやり方: ステップ別の完全ガイドを読む →