跟读练习: Natural Language Processing - Tokenization (NLP Zero to Hero - Part 1) - 通过视频学习英语口语
正在创建课程...
1
Hi, and welcome to this series on Zero2Hero for Natural Language Processing using TensorFlow.
2
If you're not an expert on AI or ML, don't worry.
3
We're taking the concepts of NLP and teaching them from first principles.
4
In this first lesson, we'll talk about how to represent words in a way that a computer can process them, with a view to later training a neural network that can understand their meaning.
5
This process is called tokenization.
6
So let's take a look.
7
Consider the word listen, as you can see here.
8
It's made up of a sequence of letters.
9
These letters can be represented by numbers using an encoding scheme.
10
A popular one called ASCII has these letters represented by these numbers.
11
This bunch of numbers can then represent the word listen.
12
But the word silent has the same letters, and thus the same numbers, just in a different order.
13
So it makes it hard for us to understand sentiment of a word just by the letters in it.
14
So it might be easier, instead of encoding letters, to encode words.
15
Consider the sentence, I love my dog.
16
So what would happen if we start encoding the words in this sentence instead of the letters in each words?
17
So for example, the word I could be one.
18
And then the sentence, I love my dog, could be one, two, three, four.
19
Now, if I take another sentence, for example, I love my cat, how would we encode it?
20
Now we see I love my has already been given one, two, three.
21
So all I need to do is encode a cat.
22
I'll give that the number five.
23
And now if we look at the two sentences, they are one, two, three, four and one, two, three,
24
five, which already show some form of similarity between them.
25
And it's a similarity you'd expect because they're both about loving a pet.
26
Given this method of encoding sentences into numbers, now let's take a look at some code to achieve this for us.
27
This process, as I mentioned before, is called tokenization, and there's an API for that.
28
We'll look at how to use it with Python.
29
So here's your first look at some code to tokenize these sentences.
30
Let's go through it line by line.
31
First of all, we'll need the tokenizer APIs, and we can get these from TensorFlow Keras like this.
32
We can represent our sentences as a Python array of strings like this.
33
It's simply the I love my dog and I love my cat that we saw earlier.
34
Now the fun begins.
35
I can create an instance of a tokenizer object.
36
The numWords parameter is the maximum number of words to keep.
37
So instead of, for example, just these two sentences, imagine if we had hundreds of books to tokenize.
38
But we just want the most frequent 100 words in all of that.
39
This would automatically do that for us when we do the next step.
40
And that's to tell the tokenizer to go through all the text and then fit itself to them like this.
41
The full list of words is available as the tokenizer's word index property.
42
So we can take a look at it like this and then simply print it out.
43
The result will be this dictionary showing the key being the word and the value being the token for that word.
44
So, for example, my has a value of three.
45
The tokenizer is also smart enough to catch some exceptions.
46
So, for example, if we updated our sentences to this by adding a third sentence, noting that dog here is followed by an exclamation mark.
47
The nice thing is that the tokenizer is smart enough to spot this and not create a new token.
48
It's just dog.
49
And you can see the results here.
50
There's no token for dog exclamation, but there is one for dog.
51
And there's also a new token for the word you.
52
If you want to try this out for yourself, I've put the code in a colab here.
53
Take it for a spin and experiment.
54
You've now seen how words can be tokenized and the tools in TensorFlow that handle that tokenization for you.
55
Now that your words are represented by numbers like this, you'll next need to represent your sentences by sequences of numbers in the correct order.
56
You'll then have data ready for processing by a neural network to understand or maybe even generate new text.
57
You'll see the tools that you can use to manage this sequencing in the next episode.
58
So don't forget to hit that subscribe button.
✨ 推荐视频
为什么用这个视频练习口语?
这个讲解自然语言处理的视频不仅能帮你了解AI知识,更是英语口语练习的好素材。视频中 speaker 用清晰的逻辑和简单的词汇讲解“标记化”,语速适中,适合影子跟读(shadow speak)。通过模仿他的发音、语调和节奏,你能同时提升听力理解和口语流畅度。尤其是他对专业术语的解释方式,能让你学会用简单英语表达复杂概念,这对日常交流和职场表达都很有帮助。
语境中的语法与表达
视频里有不少实用表达值得关注:
- “If you're not an expert on... don't worry.” 这是典型的安慰句式,在口语中用来缓解对方压力,可模仿用于日常对话。
- “Consider the word... as you can see here.” 讲解时常用的引导语,能让表达更有条理,适合在分享观点或解释事物时使用。
- “The tokenizer is smart enough to... ” 用“be smart enough to”强调某物的功能,比直接说“can”更生动,可用于描述工具或技术的优势。
常见发音陷阱
练习时要注意这些易出错的地方:
- “tokenization” 这个专业术语的发音,重音在第二个音节,不要读成“to-KEN-i-za-tion”。可通过影子跟读反复模仿,纠正发音。
- “sentiment” 中的“ti”发 /ʃ/ 音,而非 /tɪ/,容易被误读。视频中 speaker 发音清晰,可仔细听并重复。
- “exclamation” 一词较长,注意每个音节的连贯性,避免吞音。通过提高英语发音的练习,可增强对长单词的掌控力。
视频中的语法
说话人最常用的结构,并附上视频中的原话:
| 结构 | 视频中的用法 |
|---|---|
| 被动语态 be + 过去分词 — 强调发生了什么,而不是谁做的 | is called · can be represented · been given |
| 现在完成时 have/has + 过去分词 — 过去发生但与现在仍有关联的事 | has already been given · I've put · You've now seen |
什么是跟读法?
跟读法 (Shadowing) 是一种有科学依据的语言学习技巧,最初开发用于专业口译员的培训,并由多语言者Alexander Arguelles博士普及。这个方法简单而强大:您在听英语母语原声的同时立即大声重复——就像是一个延迟1-2秒紧跟说话者的影子。与被动听力或语法练习不同,跟读法强迫您的大脑和口腔肌肉同时处理并模仿真实的讲话模式。研究表明它能显着提高发音准确性,语调,节奏,连读,听力理解和口语流利度——使其成为雅思口语备考和真实英语交流最有效的方法之一。











