Prática de Shadowing: Most devs don't understand how LLM tokens work - Aprenda a falar inglês com vídeo
Criando lição...
1
So many devs are working with LLMs these days, and they don't really know about the fundamentals.
2
I was teaching an AI workshop recently in Poland, and I asked people to put up their hands if they knew what a token was.
3
And I don't know whether it's my AI bubble on X or something, but I was pretty shocked because only about a third of the people in the workshop put up their hand.
4
And so I wanted to put together a token deep dive for those of you who are using LLMs, but kind of need to catch up on what's going on under the hood.
5
And of course, we're going to explain all of these
6
I promise I will never show a line of Python on this channel, that is for sure.
7
Let's start with a diagram here.
8
Tokens are the currency of LLMs.
9
When you send an input of "Hello World" to an LLM, that gets broken down into its constituent tokens.
10
If you send "Hello World" to OpenAI, that is three tokens billed at an infinitesimally small amount per 1k tokens.
11
This is for GPT -5, I think.
12
Then the LLM will do some thinking and produce some output tokens. to the input tokens.
13
So to get the total amount you spent, you take that number, divide it by a thousand and times it by a cent.
14
You do the same with the input tokens, you add them together and you've got your total cost for that API call.
15
I'm going to demonstrate that using the ASDK.
16
For those who are not familiar with the ASDK, I've got a great free tutorial on my site AIHero, which I'll link up there.
17
Now we're going to generate some text using Claude 3 .5 Haiku, a nice cheap model from Anthropic.
18
And we're going to send it "Hello World".
19
We're then going to log things.
20
We're going to log out the text that it produced and also the usage.
21
This usage, which we get from the AASDK, is the amount of tokens that we used.
22
We can see that it replied with "Hello, how are you doing today?
23
Is there anything I can help you with?" which used 20 output tokens, but bizarrely 11 input tokens.
24
That's pretty weird considering we only prompted it with like two words: "Hello world".
25
Let's do the same thing again, but this time we're going to call Google.
26
We'll use Gemini 2 .0 flashlight and we'll log out 11 and 20,
27
but in the Google version we only used 4 input tokens and 11 output tokens.
28
So this feels a bit like a mystery if you don't know how tokens work.
29
What is the relationship between tokens and text?
30
And why does the same prompt sent to two different models end up with a different number of tokens?
31
Okay, let's reveal the mystery and talk about what tokens actually are.
32
Every LLM has a different token vocabulary.
33
And these tokens are all subwords gets assigned a number and the number is the token.
34
When we call an LLM it actually encodes the text that we send it into tokens.
35
So hello world gets split up into its largest individual tokens in the vocabulary.
36
In this case that might mean hello space world and exclamation mark
37
and then it identifies the numbers in the vocabulary that match that word or subword.
38
We can actually take a look at this in code in TypeScript because OpenAI's token.
39
We're using O200KBASE, which is the tokenizer for the model GPT -40.
40
I've also got here some text at input .md.
41
"The wise owl of moonlight forest, where ancient trees stretch their branches toward the starry sky." We're reading that big chunk of text into memory here,
42
and we're using the GPT -40 tokenizer to encode a bunch of tokens, which are just an array of numbers.
43
Then we'll compare the input length in characters and the number is nearly 2300 characters,
44
but the number of tokens is less than 500.
45
And we end up with just a bunch of numbers here.
46
If we go ahead and test this with something simpler, like Hello World, then here the content length is 12 and the number of tokens is only 3.
47
These tokens then are what gets fed into the LLM while it does its processing.
48
Let's look at a diagram of the full LLM process end -to -end.
49
We pass in Hello World, it gets split into the largest chunks that it knows about in the
50
by just looking up the number assigned to them in the token vocabulary, then the LLM does some thinking here and it outputs some more tokens.
51
In other words, all the computation here is being done on numbers, not on text.
52
And these output tokens then get decoded back into text, joined together and you get the final output.
53
So that's the process.
54
And while you think that the LLM is dealing with text, it's actually dealing with these numeric representations of chunks of text.
55
we can decode these tokens here using tokenizer .decode.
56
Decode is just a function that takes in an array of numbers, tokens, and returns a single concatenated string.
57
And when we run this, that is exactly what we get.
58
And by the way, if you're digging this, then I have something cooking for you on aihero .dev.
59
And I've got a newsletter where I post breakdowns like this, which will be in the description below.
60
So we know what tokens are now, but we don't really know why they differ between model providers.
61
To understand this, we need to know how these vocabularies even get built and why different model providers do them differently.
62
Here's the process for how tokenizers get trained.
63
Let's imagine you've got a corpus of text here.
64
We're going to use an extremely small corpus, just the single line, the cat sat on the mat.
65
But in reality, this would be gigabytes, terabytes of data.
66
The same data, in fact, usually that the model itself is trained on.
67
Now, we can build an extremely simple tokenizer just from this one sentence.
68
We can take this piece of text and extract out the unique characters in it.
69
I've created a character level tokenizer that takes in a dataset and returns a tokenizer.
70
And yes, it's in a class.
71
TypeScript classes are good, folks.
72
I've been saying this for a while.
73
We'll get to this implementation in a minute, but let's actually look at the usage.
74
We have our dataset of the cat sat on the mat.
75
Again, in a real application this might be gigabytes of text.
76
Then we grab our tokenizer by instantiating a character level tokenizer, and I'm passing in an input of cat sat mat.
77
We're then encoding the input, turning it into tokens, and logging out some information about the tokens.
78
When we run this we can see that the input length, the number of characters in catSatMAT is 11.
79
Then the number of tokens that we produce at the end is also 11.
80
This is because we only have characters as tokens.
81
If we remove all these logs and just log something called tokenizer .mapping we can see our entire vocabulary of tokens.
82
And we can see there are only 10 tokens in the entire vocabulary.
83
Space is 3, n is 8, o is 7.
84
So that means always going to be the same as the number of characters.
85
And this is not good because the more tokens we have in the input, the more work it's going to be for our LLM to process them.
86
The vocabulary size is really, really important.
87
Let's say we only have about a thousand tokens in our vocabulary.
88
If we take in the word understanding, it might produce five tokens: un, de, st, and ing.
89
If we bump that up to 50 ,000,
90
then we might get larger 200k we might split it into just two tokens.
91
And two tokens is a lot more efficient for the LLM to process than five tokens.
92
So these are the trade -offs that different models and different model providers make.
93
By the way, you can't just scale this to infinity too, because the larger your vocabulary, the bigger the model needs to be in order to house that vocabulary.
94
And of course, the more memory it takes to execute.
95
So then we know that our implementation of character level of tokens.
96
So how do we get from character level to adding some more subwords?
97
Well, we identify different groups that commonly occur together.
98
For instance, th occurs in the and the.
99
He occurs in the and the as well.
100
And then at occurs in cat, sat and mat.
101
I've got an example of this called the subword level tokenizer.
102
Now, I've basically vibe coded this, so don't look too closely at the implementation.
103
But let's take a peek at the usage down below.
104
We have our same data set and our same
105
And let's take a look at how many tokens we get back from the input of cat, sat, and mat.
106
We see here that the input cat, sat, mat is 11 and we get 8 tokens back.
107
We have achieved this by calculating some subwords in the dataset.
108
Let's look at this by logging out tokenizer .mapping.
109
And we can see here that we now have 15 tokens.
110
At, the, and her are all in as our e space and t space have each been assigned their tokens.
111
This means our encode function is a little more complicated possible subword so that it can tokenize the input text efficiently.
112
The implementation is below if you want to take a closer look.
113
In our implementation, we only go one level deep.
114
And so we end up with groupings that are only two characters long.
115
But real tokenizers identify larger and larger groups by identifying groups of groups.
116
So for instance, in this example, th and he often appear together.
117
And so a group of three tokens, the, makes sense.
118
Finally, then we end up with a set of numbers
119
that are each assigned to the individual token token as we saw in our implementation.
120
Finally, let's look at how a tokenizer behaves when it encounters an unusual word.
121
This is O Frabjous Day from the Lewis Carroll poem.
122
Frabjous is just a word that Lewis Carroll made up, so it's not going to be very frequent in the dataset.
123
When we run this through O200K, Frabjous is actually split up into four separate tokens here.
124
And so unusual words are split up into more tokens than words that occur more frequently in the dataset.
125
This also means that if you're querying the LLM in a language that's not present in the dataset much, it's likely to break it up into more tokens because it's not used to those combinations of letters.
126
This is true for spoken languages, but also for coding languages.
127
And this is yet another advantage that more commonly used coding languages have in the AI era.
128
It's fewer tokens to send 20 lines of JavaScript than it is to send 20 lines of Haskell or something.
129
So let's sum up here.
130
Tokens are the currency of LLMs, and you are charged by the token.
131
But different model providers will treat these inputs differently and produce different numbers of tokens.
132
text into tokens.
133
You take a piece of text, you split it into the largest tokens you can find, and then you turn them into the numbers that you have them stored in the vocabulary.
134
Decoding is much simpler even, you just take the numbers, you find the relevant chunks in your vocabulary, and you join them together.
135
And the whole LLM process involves an encoding step, then the LLM thinks, produces some output tokens, which then get decoded into text.
136
So tokens are really what the LLM does its thinking in.
137
If you dug this video, then I have plenty more content like this coming.
138
I've been full -time thinking about AI plus TypeScript for a year now.
139
And if you're interested in learning more about AI or building with AI, especially, then this channel is for you.
140
I personally think that TypeScript is the language for building AI -powered applications, whereas Python is probably the thing you want to use if you're building models.
141
hero .dev if you want to follow along.
142
I have something dropping there very, very soon.
143
If you like this explanation, then what should I cover next?
144
Please comment below.
145
Thanks for following along, and I will see you very soon.
📺 Mesmo canal
✨ Vídeo recomendado
Sobre esta lição
Você está praticando inglês com "Most devs don't understand how LLM tokens work" usando a técnica de Shadowing.
O que é a Técnica de Shadowing?
Shadowing é uma técnica de aprendizado de idiomas com base científica, originalmente desenvolvida para o treinamento de intérpretes profissionais. O método é simples, mas poderoso: você ouve áudio em inglês nativo e repete imediatamente em voz alta — como uma sombra seguindo o falante com 1-2 segundos de atraso. Pesquisas mostram melhora significativa na precisão da pronúncia, entonação, ritmo, sons conectados, compreensão auditiva e fluência na fala.






















