Luyện nói tiếng Anh bằng Shadowing qua video: Gaussian Naive Bayes, Clearly Explained!!!

Đang tạo bài học...
1
Beep boop boop beep boop boop beep boop.
2
StatQuest.
3
Hello, I'm Josh Starmer and welcome to StatQuest.
4
Today we're going to talk about Gaussian Naive Bays, and it's going to be clearly explained.
5
Note.
6
This stat quest assumes that you are already familiar with the main ideas behind multinomial Naive Bays.
7
If not, check out the quest!
8
The link is in the description below.
9
This StatQuest also assumes that you are familiar with the log function, the normal or Gaussian distribution,
10
and the difference between probability and likelihood.
11
If not, check out the quests.
12
The links are in the description below.
13
Imagine we wanted to predict if someone would love the 1990 movie Troll 2 or not.
14
So we collected data from people that love Troll 2, and from people that do not love Troll 2.
15
We measured the amount of popcorn they ate each day.
16
How much soda pop they drank.
17
and how much candy they ate.
18
The mean for popcorn for the people who love Troll 2 is 24.
19
and the standard deviation is 4.
20
And a Gaussian, or normal distribution, with mean equals 24 and standard deviation equals 4 looks like this.
21
Likewise, the average amount of popcorn for people who do not love Troll 2 is 4.
22
and the standard deviation is 2.
23
And that corresponds to this Gaussian, or normal distribution.
24
Now we calculate the mean and standard deviation for Soda Pop for people that love Troll 2.
25
and draw the corresponding Gaussian distribution.
26
Then we do the same thing for the people that do not love Troll 2.
27
Lastly, we draw the Gaussian distributions for candy.
28
Gaussian Naive Bayes is named after the Gaussian distributions that represent the data in the training dataset.
29
Now someone new shows up.
30
and says they eat 20 grams of popcorn.
31
and drink 500 milliliters of soda pop, and eat 25 grams of candy every day.
32
Let's use Gaussian Naive Bayes to decide if they love Troll 2 or not.
33
The first thing we do is make an initial guess that they love Troll 2.
34
This guess can be any probability that we want, but a common guess is estimated from the training data.
35
For example, since 8 of the 16 people in the training data loved Troll 2, the initial guess will be 0 .5.
36
So we'll put that up here so we don't forget.
37
Likewise, the initial guess for Does Not Love Troll 2 is 0 .5.
38
So let's put that here so we don't forget.
39
Oh no, it's the dreaded terminology alert.
40
The initial guesses are called prior probabilities.
41
Now, the score for Love's Troll 2 is...
42
The initial guess that the person loves Troll 2, times the likelihood that they eat 10 grams of popcorn given that they love Troll 2.
43
Note: the likelihood is the y -axis coordinate on the curve that corresponds to the x -axis coordinate.
44
and we multiply that by the likelihood that they drink 500 milliliters of soda pop given that they love Troll 2.
45
times the likelihood that they eat 25 grams of candy given that they love Troll 2.
46
The initial guess that someone loves Troll 2 is 0 .5.
47
The likelihood for popcorn is 0 .06.
48
The likelihood for soda pop is 0 .004, And the likelihood for candy is...
49
A really, really small number.
50
When we get really, really small numbers, it's a good idea to take the log of everything to prevent something called underflow.
51
The general idea of Underflow is:
52
Every computer has a limit to how close a number can
53
get to zero before it can no longer accurately keep track of that number.
54
When a number gets smaller than that limit, we run into underflow problems and errors occur.
55
So we use the log function to avoid underflow.
56
Any log will do, but the natural log, or log base E, is the most commonly used log in statistics and machine learning.
57
So we take the log of everything, and the log turns the multiplication into the sum of the individual logs.
58
The log base E of 0 .5 is...
59
-0 .69 The log of 0 .06 is negative 2 .8.
60
The log of 0 .004 is -5 .5.
61
And the log of this really, really small number is negative 115.
62
Now we just add this up.
63
And we get negative 124.
64
So the log of the Love's Troll 2 score is negative 124.
65
BAM!
66
Now let's calculate the score for Not Loving Troll 2.
67
We start with the initial guess that someone does not love Troll 2.
68
times the likelihood that they eat 20 grams of popcorn given that they do not love Troll 2.
69
times the likelihood that they drink 500 milliliters of soda pop, times the likelihood that they eat 25 grams of candy.
70
So let's plug in the numbers.
71
Beep, boop, beep.
72
Boop.
73
and take the log of everything.
74
and that turns the multiplication into the sum of logs.
75
Now we just do the math.
76
Beep, boop, boop, boop.
77
and we get -48.
78
And since the score for Does Not Love Troll 2 is greater than the score for Loves Troll 2,
79
We will classify this person as someone who does not love Troll 2.
80
Double bam!
81
Note: When we look at the raw data, it almost looks like we should have classified this person as someone who loves Troll 2.
82
After all, they ate a lot more popcorn than the average person who doesn't love Troll 2.
83
and they drank as much soda as the average person who loves Troll 2.
84
However, the big thing is that they ate a lot more candy than the people who loved Troll 2.
85
and the log of the likelihoods for candy are way different.
86
And this difference is what made us classify the new person as someone who does not love Troll 2.
87
In other words, Candy can have a much larger say in whether
88
or not someone loves Troll 2 than popcorn and soda pop.
89
And this means we might only need candy to make classifications.
90
We can use cross -validation to help us decide which things, popcorn, soda pop, and /or candy, help us make the best classifications.
91
Shameless self -promotion.
92
If you don't already know about cross -validation, check out the quest.
93
The link is in the description below.
94
Triple bam.
95
Oh no, it's another shameless self -promotion.
96
One awesome way to support StatQuest is to purchase the Gaussia Naive Bayes StatQuest Study Guide.
97
It has everything you need to study for an exam or job interview.
98
It's seven pages of total awesomeness.
99
And while you're there, check out the other StatQuest study guides.
100
There's something for everyone!
101
Hooray!
102
We've made it to the end of another exciting stat quest!
103
If you like this StatQuest and want to see more, please subscribe.
104
And if you want to support StatQuest, consider contributing to my Patreon campaign,
105
becoming a channel member, buying one or two of my original songs or a t -shirt or a hoodie, or just donate.
106
The links are in the description below.
107
Alright, until next time, quest on!

Từ vựng và ghi chú luyện nói cho bài học này

Bài luyện nói trình độ C1 này dựa trên video “Gaussian Naive Bayes, Clearly Explained!!!”. Những từ được nhắc lại nhiều nhất trong bài: Troll, log, likelihood, candy, boop. Video này có 107 câu và 1118 từ để luyện shadowing. Phần lời nói dài 9:25. Người nói giữ tốc độ đều, khoảng 119 từ mỗi phút, vừa sức để nói đuổi theo. Chỉ 77% số từ nằm trong 3.000 từ tiếng Anh thông dụng nhất, nên từ vựng khá khó.

Từ vựng quan trọng trong video

14 từ khó nhất trong video, kèm phiên âm và nghĩa:

TừPhiên âmNghĩa
popcorn danh từ/ˈpɑp.kɔɹn/bắp rang, bỏng ngô
beep danh từ/biːp/bíp
gram danh từ/ˈɡɹæm/gam
classify động từ/ˈklæs.əˌfaɪ/phân loại
probability danh từ/ˌpɹɑ.bəˈbɪl.ə.ti/xác suất, khả năng
multiplication danh từ/ˌmʌltɪplɪˈkeɪʃən/phép nhân
shameless tính từ/ˈʃeɪ̯mlɪs/trơ trẽn
coordinate động từ/koʊˈɔɹ.dəˌneɪt/điều phối, hợp tác
calculate động từ/ˈkælkjʊleɪt/tính toán, tính
axis danh từ/ˈæksɪs/trục
multiply động từ/ˈmʌltɪplaɪ/nhân
subscribe động từ/səbˈskɹaɪb/đăng kí
donate động từ/ˈdoʊˌneɪt/quyên góp
predict động từ/pɹɪˈdɪkt/dự báo

Cụm động từ bạn sẽ nghe

TừNghĩa
check out động từtính tiền, trả phòng

Ngữ pháp trong video

Những cấu trúc người nói dùng nhiều nhất, kèm đúng cụm từ trong video:

Cấu trúcTrong video
Mệnh đề quan hệ who / which + mệnh đề — thêm thông tin về người hoặc vậtpeople who love · people who do · someone who does
Câu bị động be + quá khứ phân từ — nhấn vào việc xảy ra, không phải người làmis named · is estimated · are called

Phát âm cần chú ý

Người nói dùng 9 dạng rút gọn, ví dụ don't, doesn't, I'm. Hãy nói theo dạng ngắn đúng như bạn nghe.

  • Âm “sh” và “zh”: deviation /ˌdiː.viˈeɪʃən/, multiplication /ˌmʌltɪplɪˈkeɪʃən/, shameless /ˈʃeɪ̯mlɪs/, validation /ˌvæl.əˈdeɪ.ʃən/
  • Từ dài — đặt trọng âm cho đúng: deviation /ˌdiː.viˈeɪʃən/, probability /ˌpɹɑ.bəˈbɪl.ə.ti/, multiplication /ˌmʌltɪplɪˈkeɪʃən/, validation /ˌvæl.əˈdeɪ.ʃən/, coordinate /koʊˈɔɹ.dəˌneɪt/

Cách luyện với video này

  1. Nghe hết video một lần, chưa cần nói, và ghi lại những từ bạn chưa biết.
  2. Nói đuổi từng câu ở tốc độ bình thường, lặp lại mỗi câu đến khi nhịp của bạn khớp với người nói.
  3. Ghi âm giọng mình rồi so với bản gốc, chú ý các từ như popcorn, beep, gram.

Phương Pháp Shadowing Là Gì?

Shadowing là kỹ thuật học ngôn ngữ có cơ sở khoa học, ban đầu được phát triển cho chương trình đào tạo phiên dịch viên chuyên nghiệp và được phổ biến rộng rãi bởi nhà đa ngôn ngữ học Dr. Alexander Arguelles. Nguyên lý cốt lõi đơn giản nhưng cực kỳ hiệu quả: bạn nghe tiếng Anh của người bản xứ và lặp lại to ngay lập tức — như một "cái bóng" (shadow) đuổi theo người nói với độ trễ chỉ 1–2 giây. Khác với luyện ngữ pháp hay học từ vựng bị động, Shadowing buộc não bộ và cơ miệng phải đồng thời xử lý và tái tạo ngôn ngữ thực tế. Các nghiên cứu khoa học xác nhận phương pháp này cải thiện đáng kể phát âm, ngữ điệu, nhịp điệu, nối âm, kỹ năng nghe và độ lưu loát khi nói — đặc biệt hiệu quả cho người luyện IELTS Speaking và muốn giao tiếp tiếng Anh tự nhiên như người bản ngữ.

Phương pháp shadowing: đọc hướng dẫn từng bước đầy đủ →