Luyện nói tiếng Anh bằng Shadowing qua video: Natural Language Processing - Tokenization (NLP Zero to Hero - Part 1)

Đang tạo bài học...
1
Hi, and welcome to this series on Zero2Hero for Natural Language Processing using TensorFlow.
2
If you're not an expert on AI or ML, don't worry.
3
We're taking the concepts of NLP and teaching them from first principles.
4
In this first lesson, we'll talk about how to represent words in a way that a computer can process them, with a view to later training a neural network that can understand their meaning.
5
This process is called tokenization.
6
So let's take a look.
7
Consider the word listen, as you can see here.
8
It's made up of a sequence of letters.
9
These letters can be represented by numbers using an encoding scheme.
10
A popular one called ASCII has these letters represented by these numbers.
11
This bunch of numbers can then represent the word listen.
12
But the word silent has the same letters, and thus the same numbers, just in a different order.
13
So it makes it hard for us to understand sentiment of a word just by the letters in it.
14
So it might be easier, instead of encoding letters, to encode words.
15
Consider the sentence, I love my dog.
16
So what would happen if we start encoding the words in this sentence instead of the letters in each words?
17
So for example, the word I could be one.
18
And then the sentence, I love my dog, could be one, two, three, four.
19
Now, if I take another sentence, for example, I love my cat, how would we encode it?
20
Now we see I love my has already been given one, two, three.
21
So all I need to do is encode a cat.
22
I'll give that the number five.
23
And now if we look at the two sentences, they are one, two, three, four and one, two, three,
24
five, which already show some form of similarity between them.
25
And it's a similarity you'd expect because they're both about loving a pet.
26
Given this method of encoding sentences into numbers, now let's take a look at some code to achieve this for us.
27
This process, as I mentioned before, is called tokenization, and there's an API for that.
28
We'll look at how to use it with Python.
29
So here's your first look at some code to tokenize these sentences.
30
Let's go through it line by line.
31
First of all, we'll need the tokenizer APIs, and we can get these from TensorFlow Keras like this.
32
We can represent our sentences as a Python array of strings like this.
33
It's simply the I love my dog and I love my cat that we saw earlier.
34
Now the fun begins.
35
I can create an instance of a tokenizer object.
36
The numWords parameter is the maximum number of words to keep.
37
So instead of, for example, just these two sentences, imagine if we had hundreds of books to tokenize.
38
But we just want the most frequent 100 words in all of that.
39
This would automatically do that for us when we do the next step.
40
And that's to tell the tokenizer to go through all the text and then fit itself to them like this.
41
The full list of words is available as the tokenizer's word index property.
42
So we can take a look at it like this and then simply print it out.
43
The result will be this dictionary showing the key being the word and the value being the token for that word.
44
So, for example, my has a value of three.
45
The tokenizer is also smart enough to catch some exceptions.
46
So, for example, if we updated our sentences to this by adding a third sentence, noting that dog here is followed by an exclamation mark.
47
The nice thing is that the tokenizer is smart enough to spot this and not create a new token.
48
It's just dog.
49
And you can see the results here.
50
There's no token for dog exclamation, but there is one for dog.
51
And there's also a new token for the word you.
52
If you want to try this out for yourself, I've put the code in a colab here.
53
Take it for a spin and experiment.
54
You've now seen how words can be tokenized and the tools in TensorFlow that handle that tokenization for you.
55
Now that your words are represented by numbers like this, you'll next need to represent your sentences by sequences of numbers in the correct order.
56
You'll then have data ready for processing by a neural network to understand or maybe even generate new text.
57
You'll see the tools that you can use to manage this sequencing in the next episode.
58
So don't forget to hit that subscribe button.

Video Này Dành Cho Ai?

Video này phù hợp cho những người học tiếng Anh ở mọi trình độ, đặc biệt là những bạn muốn nâng cao khả năng nghe và nói thông qua việc luyện tập với các đoạn hội thoại tự nhiên. Nếu bạn là người mới bắt đầu, đừng lo lắng, vì nội dung video sẽ giúp bạn hiểu về quy trình xử lý ngôn ngữ tự nhiên một cách dễ dàng, từ đó có thể áp dụng vào việc giao tiếp hàng ngày. Những người học có mục tiêu phát âm chuẩn xác và hiểu sâu về ngữ nghĩa cũng sẽ tìm thấy nhiều thông tin hữu ích trong video này.

Các Từ & Thành Ngữ Đáng Chú Ý

  • Listen: Nghe
  • Silent: Im lặng (có cùng chữ cái với từ "listen", nhưng không mang ý nghĩa giống nhau)
  • I love my dog: Tôi yêu chó của mình (một câu đơn giản thể hiện cảm xúc yêu thương)
  • Tokenization: Quá trình phân tách (các từ hoặc câu thành các đơn vị nhỏ hơn để máy tính dễ xử lý)

Những từ và cụm từ này không chỉ giúp bạn mở rộng từ vựng mà còn nâng cao khả năng hiểu biết về cách mà ngôn ngữ được sử dụng trong cuộc sống hàng ngày. Hãy thử lặp lại chúng nhiều lần để ghi nhớ sâu hơn.

Cách Để Phát Âm Đúng

Để phát âm chính xác như trong video, hãy chú ý đến một số điểm nhấn trong cách nói của người diễn thuyết. Họ thường sử dụng ngữ điệu linh hoạt, nhấn mạnh vào những từ quan trọng, điều này giúp người nghe dễ dàng nhận biết được thông điệp chính. Hãy luyện nghe nói qua video nhiều lần và áp dụng phương pháp shadow speak để cải thiện khả năng phát âm của bạn. Tức là, bạn sẽ nghe và lặp lại ngay sau khi người nói phát âm, theo kịp nhịp điệu và âm sắc của họ.

Việc luyện tập như vậy không chỉ giúp bạn nhớ từ vựng mà còn giúp bạn tự tin hơn khi giao tiếp. Hãy nhớ rằng, việc luyện nghe nói qua video là một phương pháp hiệu quả để cải thiện kỹ năng ngôn ngữ, giúp bạn cảm nhận rõ hơn về cách diễn đạt tự nhiên trong tiếng Anh.

Ngữ pháp trong video

Những cấu trúc người nói dùng nhiều nhất, kèm đúng cụm từ trong video:

Cấu trúcTrong video
Câu bị động be + quá khứ phân từ — nhấn vào việc xảy ra, không phải người làmis called · can be represented · been given
Thì hiện tại hoàn thành have/has + quá khứ phân từ — việc đã xảy ra nhưng còn liên quan đến hiện tạihas already been given · I've put · You've now seen

Phương Pháp Shadowing Là Gì?

Shadowing là kỹ thuật học ngôn ngữ có cơ sở khoa học, ban đầu được phát triển cho chương trình đào tạo phiên dịch viên chuyên nghiệp và được phổ biến rộng rãi bởi nhà đa ngôn ngữ học Dr. Alexander Arguelles. Nguyên lý cốt lõi đơn giản nhưng cực kỳ hiệu quả: bạn nghe tiếng Anh của người bản xứ và lặp lại to ngay lập tức — như một "cái bóng" (shadow) đuổi theo người nói với độ trễ chỉ 1–2 giây. Khác với luyện ngữ pháp hay học từ vựng bị động, Shadowing buộc não bộ và cơ miệng phải đồng thời xử lý và tái tạo ngôn ngữ thực tế. Các nghiên cứu khoa học xác nhận phương pháp này cải thiện đáng kể phát âm, ngữ điệu, nhịp điệu, nối âm, kỹ năng nghe và độ lưu loát khi nói — đặc biệt hiệu quả cho người luyện IELTS Speaking và muốn giao tiếp tiếng Anh tự nhiên như người bản ngữ.

Phương pháp shadowing: đọc hướng dẫn từng bước đầy đủ →