Luyện nói tiếng Anh bằng Shadowing qua video: Most devs don't understand how LLM tokens work

Đang tạo bài học...
1
So many devs are working with LLMs these days, and they don't really know about the fundamentals.
2
I was teaching an AI workshop recently in Poland, and I asked people to put up their hands if they knew what a token was.
3
And I don't know whether it's my AI bubble on X or something, but I was pretty shocked because only about a third of the people in the workshop put up their hand.
4
And so I wanted to put together a token deep dive for those of you who are using LLMs, but kind of need to catch up on what's going on under the hood.
5
And of course, we're going to explain all of these
6
I promise I will never show a line of Python on this channel, that is for sure.
7
Let's start with a diagram here.
8
Tokens are the currency of LLMs.
9
When you send an input of "Hello World" to an LLM, that gets broken down into its constituent tokens.
10
If you send "Hello World" to OpenAI, that is three tokens billed at an infinitesimally small amount per 1k tokens.
11
This is for GPT -5, I think.
12
Then the LLM will do some thinking and produce some output tokens. to the input tokens.
13
So to get the total amount you spent, you take that number, divide it by a thousand and times it by a cent.
14
You do the same with the input tokens, you add them together and you've got your total cost for that API call.
15
I'm going to demonstrate that using the ASDK.
16
For those who are not familiar with the ASDK, I've got a great free tutorial on my site AIHero, which I'll link up there.
17
Now we're going to generate some text using Claude 3 .5 Haiku, a nice cheap model from Anthropic.
18
And we're going to send it "Hello World".
19
We're then going to log things.
20
We're going to log out the text that it produced and also the usage.
21
This usage, which we get from the AASDK, is the amount of tokens that we used.
22
We can see that it replied with "Hello, how are you doing today?
23
Is there anything I can help you with?" which used 20 output tokens, but bizarrely 11 input tokens.
24
That's pretty weird considering we only prompted it with like two words: "Hello world".
25
Let's do the same thing again, but this time we're going to call Google.
26
We'll use Gemini 2 .0 flashlight and we'll log out 11 and 20,
27
but in the Google version we only used 4 input tokens and 11 output tokens.
28
So this feels a bit like a mystery if you don't know how tokens work.
29
What is the relationship between tokens and text?
30
And why does the same prompt sent to two different models end up with a different number of tokens?
31
Okay, let's reveal the mystery and talk about what tokens actually are.
32
Every LLM has a different token vocabulary.
33
And these tokens are all subwords gets assigned a number and the number is the token.
34
When we call an LLM it actually encodes the text that we send it into tokens.
35
So hello world gets split up into its largest individual tokens in the vocabulary.
36
In this case that might mean hello space world and exclamation mark
37
and then it identifies the numbers in the vocabulary that match that word or subword.
38
We can actually take a look at this in code in TypeScript because OpenAI's token.
39
We're using O200KBASE, which is the tokenizer for the model GPT -40.
40
I've also got here some text at input .md.
41
"The wise owl of moonlight forest, where ancient trees stretch their branches toward the starry sky." We're reading that big chunk of text into memory here,
42
and we're using the GPT -40 tokenizer to encode a bunch of tokens, which are just an array of numbers.
43
Then we'll compare the input length in characters and the number is nearly 2300 characters,
44
but the number of tokens is less than 500.
45
And we end up with just a bunch of numbers here.
46
If we go ahead and test this with something simpler, like Hello World, then here the content length is 12 and the number of tokens is only 3.
47
These tokens then are what gets fed into the LLM while it does its processing.
48
Let's look at a diagram of the full LLM process end -to -end.
49
We pass in Hello World, it gets split into the largest chunks that it knows about in the
50
by just looking up the number assigned to them in the token vocabulary, then the LLM does some thinking here and it outputs some more tokens.
51
In other words, all the computation here is being done on numbers, not on text.
52
And these output tokens then get decoded back into text, joined together and you get the final output.
53
So that's the process.
54
And while you think that the LLM is dealing with text, it's actually dealing with these numeric representations of chunks of text.
55
we can decode these tokens here using tokenizer .decode.
56
Decode is just a function that takes in an array of numbers, tokens, and returns a single concatenated string.
57
And when we run this, that is exactly what we get.
58
And by the way, if you're digging this, then I have something cooking for you on aihero .dev.
59
And I've got a newsletter where I post breakdowns like this, which will be in the description below.
60
So we know what tokens are now, but we don't really know why they differ between model providers.
61
To understand this, we need to know how these vocabularies even get built and why different model providers do them differently.
62
Here's the process for how tokenizers get trained.
63
Let's imagine you've got a corpus of text here.
64
We're going to use an extremely small corpus, just the single line, the cat sat on the mat.
65
But in reality, this would be gigabytes, terabytes of data.
66
The same data, in fact, usually that the model itself is trained on.
67
Now, we can build an extremely simple tokenizer just from this one sentence.
68
We can take this piece of text and extract out the unique characters in it.
69
I've created a character level tokenizer that takes in a dataset and returns a tokenizer.
70
And yes, it's in a class.
71
TypeScript classes are good, folks.
72
I've been saying this for a while.
73
We'll get to this implementation in a minute, but let's actually look at the usage.
74
We have our dataset of the cat sat on the mat.
75
Again, in a real application this might be gigabytes of text.
76
Then we grab our tokenizer by instantiating a character level tokenizer, and I'm passing in an input of cat sat mat.
77
We're then encoding the input, turning it into tokens, and logging out some information about the tokens.
78
When we run this we can see that the input length, the number of characters in catSatMAT is 11.
79
Then the number of tokens that we produce at the end is also 11.
80
This is because we only have characters as tokens.
81
If we remove all these logs and just log something called tokenizer .mapping we can see our entire vocabulary of tokens.
82
And we can see there are only 10 tokens in the entire vocabulary.
83
Space is 3, n is 8, o is 7.
84
So that means always going to be the same as the number of characters.
85
And this is not good because the more tokens we have in the input, the more work it's going to be for our LLM to process them.
86
The vocabulary size is really, really important.
87
Let's say we only have about a thousand tokens in our vocabulary.
88
If we take in the word understanding, it might produce five tokens: un, de, st, and ing.
89
If we bump that up to 50 ,000,
90
then we might get larger 200k we might split it into just two tokens.
91
And two tokens is a lot more efficient for the LLM to process than five tokens.
92
So these are the trade -offs that different models and different model providers make.
93
By the way, you can't just scale this to infinity too, because the larger your vocabulary, the bigger the model needs to be in order to house that vocabulary.
94
And of course, the more memory it takes to execute.
95
So then we know that our implementation of character level of tokens.
96
So how do we get from character level to adding some more subwords?
97
Well, we identify different groups that commonly occur together.
98
For instance, th occurs in the and the.
99
He occurs in the and the as well.
100
And then at occurs in cat, sat and mat.
101
I've got an example of this called the subword level tokenizer.
102
Now, I've basically vibe coded this, so don't look too closely at the implementation.
103
But let's take a peek at the usage down below.
104
We have our same data set and our same
105
And let's take a look at how many tokens we get back from the input of cat, sat, and mat.
106
We see here that the input cat, sat, mat is 11 and we get 8 tokens back.
107
We have achieved this by calculating some subwords in the dataset.
108
Let's look at this by logging out tokenizer .mapping.
109
And we can see here that we now have 15 tokens.
110
At, the, and her are all in as our e space and t space have each been assigned their tokens.
111
This means our encode function is a little more complicated possible subword so that it can tokenize the input text efficiently.
112
The implementation is below if you want to take a closer look.
113
In our implementation, we only go one level deep.
114
And so we end up with groupings that are only two characters long.
115
But real tokenizers identify larger and larger groups by identifying groups of groups.
116
So for instance, in this example, th and he often appear together.
117
And so a group of three tokens, the, makes sense.
118
Finally, then we end up with a set of numbers
119
that are each assigned to the individual token token as we saw in our implementation.
120
Finally, let's look at how a tokenizer behaves when it encounters an unusual word.
121
This is O Frabjous Day from the Lewis Carroll poem.
122
Frabjous is just a word that Lewis Carroll made up, so it's not going to be very frequent in the dataset.
123
When we run this through O200K, Frabjous is actually split up into four separate tokens here.
124
And so unusual words are split up into more tokens than words that occur more frequently in the dataset.
125
This also means that if you're querying the LLM in a language that's not present in the dataset much, it's likely to break it up into more tokens because it's not used to those combinations of letters.
126
This is true for spoken languages, but also for coding languages.
127
And this is yet another advantage that more commonly used coding languages have in the AI era.
128
It's fewer tokens to send 20 lines of JavaScript than it is to send 20 lines of Haskell or something.
129
So let's sum up here.
130
Tokens are the currency of LLMs, and you are charged by the token.
131
But different model providers will treat these inputs differently and produce different numbers of tokens.
132
text into tokens.
133
You take a piece of text, you split it into the largest tokens you can find, and then you turn them into the numbers that you have them stored in the vocabulary.
134
Decoding is much simpler even, you just take the numbers, you find the relevant chunks in your vocabulary, and you join them together.
135
And the whole LLM process involves an encoding step, then the LLM thinks, produces some output tokens, which then get decoded into text.
136
So tokens are really what the LLM does its thinking in.
137
If you dug this video, then I have plenty more content like this coming.
138
I've been full -time thinking about AI plus TypeScript for a year now.
139
And if you're interested in learning more about AI or building with AI, especially, then this channel is for you.
140
I personally think that TypeScript is the language for building AI -powered applications, whereas Python is probably the thing you want to use if you're building models.
141
hero .dev if you want to follow along.
142
I have something dropping there very, very soon.
143
If you like this explanation, then what should I cover next?
144
Please comment below.
145
Thanks for following along, and I will see you very soon.

Từ vựng và ghi chú luyện nói cho bài học này

Video này có 145 câu và 2089 từ để luyện shadowing. Phần lời nói dài 10:57. Người nói nói nhanh, khoảng 191 từ mỗi phút, nên sẽ có nhiều chỗ nối âm và nuốt âm. Chỉ 80% số từ nằm trong 3.000 từ tiếng Anh thông dụng nhất, nên từ vựng khá khó.

Từ vựng quan trọng trong video

15 từ ít gặp trong video, kèm phiên âm và nghĩa:

  • vocabulary /vəˈkæb.jə.lə.ɹi/ (danh từ) — từ vựng. A usually alphabetized and explained collection of words e.g. of a particular field, or prepared for a specific purpose, often for learning.
  • output /ˈaʊtpʊt/ (danh từ) — đầu ra. Production; quantity produced, created, or completed.
  • occur /əˈkɜː/ (động từ) — xảy ra. To happen or take place.
  • chunk /t͡ʃʌŋk/ (danh từ) — mảnh, mẩu. A part of something that has been separated; a generally squat, thick, irregular piece of something, e.g. wood or stone.
  • log /lɔɡ/ (danh từ) — gỗ tròn. The trunk of a dead tree, cleared of branches.
  • mystery /ˈmɪs.t(ə.)ɹi/ (danh từ) — bí ẩn, huyền bí. Something secret or unexplainable; an unknown.
  • currency /ˈkʌɹ.ən.si/ (danh từ) — tiền tệ, ngoại tệ. Money or other items used to facilitate transactions.
  • workshop /ˈwɝk.ʃɑp/ (danh từ) — xưởng. A room, especially one which is not particularly large, used for manufacturing or other light industrial work.
  • diagram /ˈdaɪ.ə.ɡɹæm/ (danh từ) — giản đồ, biểu đồ. A plan, drawing, sketch or outline to show the function or operation of something, or to show the relationships between the parts of a whole.
  • corpus /ˈkɔɹpəs/ (danh từ) — ngữ liệu. A collection of written or spoken texts.
  • complicated /ˈkɑm.plɪˌkeɪ.tɪd/ (tính từ) — phức tạp. Difficult or convoluted.
  • shock /ʃɔk/ (danh từ) — choáng, ngạc nhiên. A sudden, heavy impact.
  • compare /kəmˈpɛɚ/ (động từ) — so sánh, so. To assess the similarities and differences between two or more things [“to compare X with Y”]. Having made the comparison of X with Y, one might have found it…
  • reveal /ɹɪˈviːl/ (động từ) — để lộ, tiết lộ. To uncover; to show and display that which was hidden.
  • string /stɹɪŋ/ (danh từ) — chuỗi. A long, thin and flexible structure made from threads twisted together.

Cụm động từ bạn sẽ nghe

  • log out (động từ) — đăng xuất. To exit a user account in a computer system, so that one is not recognized until signing in again.
  • go on /ˈɡoʊˌɒn/ (động từ) — tiếp tục. To continue in extent.
  • look up /ˌlʊk ˈʌp/ (động từ) — ngước. To have better prospects.

Phát âm cần chú ý

Người nói dùng 38 dạng rút gọn, ví dụ we're, I've, don't. Hãy nói theo dạng ngắn đúng như bạn nghe.

  • Âm “sh” và “zh”: implementation /ˌɪmplɪmənˈteɪʃən/, unusual /ʌnˈjuːʒ(u)əl/, workshop /ˈwɝk.ʃɑp/, shock /ʃɔk/, explanation /ˌɛkspləˈneɪʃən/
  • Từ dài — đặt trọng âm cho đúng: vocabulary /vəˈkæb.jə.lə.ɹi/, implementation /ˌɪmplɪmənˈteɪʃən/, unusual /ʌnˈjuːʒ(u)əl/, differently /ˈdɪf.ə.ɹənt.li/, complicated /ˈkɑm.plɪˌkeɪ.tɪd/

Cách luyện với video này

  1. Nghe hết video một lần, chưa cần nói, và ghi lại những từ bạn chưa biết.
  2. Bắt đầu ở tốc độ 0,75×, nói đuổi từng câu, rồi quay lại tốc độ bình thường khi đã quen.
  3. Ghi âm giọng mình rồi so với bản gốc, chú ý các từ như vocabulary, output, occur.

Ngữ pháp trong video

Những cấu trúc người nói dùng nhiều nhất, kèm đúng cụm từ trong video:

Cấu trúcTrong video
Câu bị động be + quá khứ phân từ — nhấn vào việc xảy ra, không phải người làmbeing done · is trained · been assigned
Thì hiện tại hoàn thành have/has + quá khứ phân từ — việc đã xảy ra nhưng còn liên quan đến hiện tạiI've created · have achieved · I've been
Mệnh đề quan hệ who / which + mệnh đề — thêm thông tin về người hoặc vậtthose who are · tokens, which are · this, which will

Phương Pháp Shadowing Là Gì?

Shadowing là kỹ thuật học ngôn ngữ có cơ sở khoa học, ban đầu được phát triển cho chương trình đào tạo phiên dịch viên chuyên nghiệp và được phổ biến rộng rãi bởi nhà đa ngôn ngữ học Dr. Alexander Arguelles. Nguyên lý cốt lõi đơn giản nhưng cực kỳ hiệu quả: bạn nghe tiếng Anh của người bản xứ và lặp lại to ngay lập tức — như một "cái bóng" (shadow) đuổi theo người nói với độ trễ chỉ 1–2 giây. Khác với luyện ngữ pháp hay học từ vựng bị động, Shadowing buộc não bộ và cơ miệng phải đồng thời xử lý và tái tạo ngôn ngữ thực tế. Các nghiên cứu khoa học xác nhận phương pháp này cải thiện đáng kể phát âm, ngữ điệu, nhịp điệu, nối âm, kỹ năng nghe và độ lưu loát khi nói — đặc biệt hiệu quả cho người luyện IELTS Speaking và muốn giao tiếp tiếng Anh tự nhiên như người bản ngữ.

Phương pháp shadowing: đọc hướng dẫn từng bước đầy đủ →