ฝึกพูดภาษาอังกฤษด้วยเทคนิค Shadowing จากวิดีโอ: Natural Language Processing - Tokenization (NLP Zero to Hero - Part 1)

กำลังสร้างบทเรียน...
1
Hi, and welcome to this series on Zero2Hero for Natural Language Processing using TensorFlow.
2
If you're not an expert on AI or ML, don't worry.
3
We're taking the concepts of NLP and teaching them from first principles.
4
In this first lesson, we'll talk about how to represent words in a way that a computer can process them, with a view to later training a neural network that can understand their meaning.
5
This process is called tokenization.
6
So let's take a look.
7
Consider the word listen, as you can see here.
8
It's made up of a sequence of letters.
9
These letters can be represented by numbers using an encoding scheme.
10
A popular one called ASCII has these letters represented by these numbers.
11
This bunch of numbers can then represent the word listen.
12
But the word silent has the same letters, and thus the same numbers, just in a different order.
13
So it makes it hard for us to understand sentiment of a word just by the letters in it.
14
So it might be easier, instead of encoding letters, to encode words.
15
Consider the sentence, I love my dog.
16
So what would happen if we start encoding the words in this sentence instead of the letters in each words?
17
So for example, the word I could be one.
18
And then the sentence, I love my dog, could be one, two, three, four.
19
Now, if I take another sentence, for example, I love my cat, how would we encode it?
20
Now we see I love my has already been given one, two, three.
21
So all I need to do is encode a cat.
22
I'll give that the number five.
23
And now if we look at the two sentences, they are one, two, three, four and one, two, three,
24
five, which already show some form of similarity between them.
25
And it's a similarity you'd expect because they're both about loving a pet.
26
Given this method of encoding sentences into numbers, now let's take a look at some code to achieve this for us.
27
This process, as I mentioned before, is called tokenization, and there's an API for that.
28
We'll look at how to use it with Python.
29
So here's your first look at some code to tokenize these sentences.
30
Let's go through it line by line.
31
First of all, we'll need the tokenizer APIs, and we can get these from TensorFlow Keras like this.
32
We can represent our sentences as a Python array of strings like this.
33
It's simply the I love my dog and I love my cat that we saw earlier.
34
Now the fun begins.
35
I can create an instance of a tokenizer object.
36
The numWords parameter is the maximum number of words to keep.
37
So instead of, for example, just these two sentences, imagine if we had hundreds of books to tokenize.
38
But we just want the most frequent 100 words in all of that.
39
This would automatically do that for us when we do the next step.
40
And that's to tell the tokenizer to go through all the text and then fit itself to them like this.
41
The full list of words is available as the tokenizer's word index property.
42
So we can take a look at it like this and then simply print it out.
43
The result will be this dictionary showing the key being the word and the value being the token for that word.
44
So, for example, my has a value of three.
45
The tokenizer is also smart enough to catch some exceptions.
46
So, for example, if we updated our sentences to this by adding a third sentence, noting that dog here is followed by an exclamation mark.
47
The nice thing is that the tokenizer is smart enough to spot this and not create a new token.
48
It's just dog.
49
And you can see the results here.
50
There's no token for dog exclamation, but there is one for dog.
51
And there's also a new token for the word you.
52
If you want to try this out for yourself, I've put the code in a colab here.
53
Take it for a spin and experiment.
54
You've now seen how words can be tokenized and the tools in TensorFlow that handle that tokenization for you.
55
Now that your words are represented by numbers like this, you'll next need to represent your sentences by sequences of numbers in the correct order.
56
You'll then have data ready for processing by a neural network to understand or maybe even generate new text.
57
You'll see the tools that you can use to manage this sequencing in the next episode.
58
So don't forget to hit that subscribe button.

คุณจะได้เรียนรู้อะไรบ้าง

ในการสนทนานี้ คุณจะได้เรียนรู้ทักษะการพูดที่สำคัญหลายอย่าง เช่น การเข้าใจแนวคิดของ tokenization ซึ่งเป็นกระบวนการที่ช่วยให้เราสามารถแปลงคำและประโยคเป็นตัวเลข เพื่อให้คอมพิวเตอร์สามารถประมวลผลได้ และยังมีการใช้โค้ดใน Python ที่ช่วยในการ ปรับปรุงการออกเสียงภาษาอังกฤษ ผ่านการทำให้เข้าใจความหมายของคำแต่ละคำได้ดียิ่งขึ้น นอกจากนี้ คุณจะยังได้เรียนรู้การเปรียบเทียบประโยคที่คล้ายกัน และเห็นถึงความสำคัญของการเลือกใช้คำที่เหมาะสมในการสื่อสารอีกด้วย

ฟังเสียงเหล่านี้ให้ดี

ในการสนทนานี้ มีการใช้เสียงที่เชื่อมโยงกันและการลดเสียงในคำพูดที่สามารถฟังได้ชัดเจน เมื่อคุณฟังการสนทนา คุณอาจสังเกตเห็นว่าเสียงบางเสียงเชื่อมกันอย่างไร เช่น การเชื่อมเสียงระหว่างคำ "I love" ที่ฟังดูราบเรียบและมีจังหวะที่เป็นเอกลักษณ์ นอกจากนี้ยังมีการลดเสียงในคำบางคำ เช่น "my" ที่อาจจะไม่ออกเสียงอย่างชัดเจนในบริบทของการสนทนา การจับเสียงเหล่านี้ให้ได้จะช่วยให้คุณเข้าใจการพูดของเจ้าของภาษาได้ดีขึ้น และนำไปสู่การ ปรับปรุงการออกเสียงภาษาอังกฤษ ของคุณเอง

พูดให้เหมือนเจ้าของภาษา

การเลียนแบบจังหวะและการเน้นเสียงของผู้พูดในคลิปนี้จะช่วยให้คุณมีความมั่นใจในการพูดมากขึ้น ลองฟังจังหวะของเสียงและพยายามพูดตามอย่างช้าๆ โดยเริ่มจากการสังเกตว่าคำไหนมีการเน้นเสียงมากกว่าคำอื่นๆ การสะท้อนเสียงของผู้พูด หรือที่เรียกว่า shadowspeak จะช่วยให้คุณฝึกการออกเสียงได้อย่างถูกต้อง และทำให้การพูดของคุณมีชีวิตชีวามากขึ้น นอกจากนี้ การเข้าใจวิธีการพูดที่เป็นธรรมชาติจะช่วยให้การสื่อสารของคุณมีประสิทธิภาพมากยิ่งขึ้น คุณอาจลองบันทึกเสียงของตัวเองเมื่อคุณฝึกพูดตาม เพื่อให้คุณสามารถเปรียบเทียบกับต้นฉบับและพัฒนาทักษะได้อย่างรวดเร็ว

ไวยากรณ์ในวิดีโอนี้

โครงสร้างที่ผู้พูดใช้บ่อยที่สุด พร้อมคำพูดจริงจากวิดีโอ:

โครงสร้างในวิดีโอ
Passive voice be + กริยาช่อง 3 — เน้นสิ่งที่เกิดขึ้น ไม่ใช่ผู้กระทำis called · can be represented · been given
Present perfect have/has + กริยาช่อง 3 — เหตุการณ์ในอดีตที่ยังเกี่ยวข้องกับปัจจุบันhas already been given · I've put · You've now seen

เทคนิค Shadowing คืออะไร?

Shadowing เป็นเทคนิคการเรียนรู้ภาษาที่ได้รับการรับรองทางวิทยาศาสตร์ พัฒนาขึ้นสำหรับการฝึกนักแปลมืออาชีพ วิธีการนี้เรียบง่ายแต่ทรงพลัง: คุณฟังเสียงภาษาอังกฤษจากเจ้าของภาษาและพูดตามทันที — เหมือนเงาที่ตามผู้พูดด้วยช่วงเวลาห่าง 1-2 วินาที การวิจัยแสดงว่าเทคนิคนี้ปรับปรุงความแม่นยำในการออกเสียง ทำนองเสียง จังหวะ การเชื่อมเสียง การฟังเข้าใจ และความคล่องแคล่วในการพูดได้อย่างมีนัยสำคัญ

เทคนิค shadowing: อ่านคู่มือฉบับเต็มทีละขั้นตอน →