쉐도잉 연습: Every Machine Learning Model Explained in 15 minutes - 영상으로 영어 말하기 배우기

레슨 만드는 중...
1
In the next 15 minutes, we are going to build a clear and simple overview of the most important machine learning algorithms,
2
so you can understand what they do and how to decide which one fits your problem.
3
I am not going to overwhelm you with equations, but simply give you intuitions behind each algorithm.
4
By the end, you will have a clear picture of what machine learning is.
5
Machine learning is a part of artificial intelligence, where computers learn patterns from data.
6
Instead of telling the computer exactly what to do in every situation, we train it using examples, and it learns how to make decisions on its own.
7
Broadly, machine learning is divided into four main categories,
8
supervised learning, unsupervised learning, reinforcement learning, and semi -supervised learning.
9
In supervised learning, we are provided with a data set which has input variables also called independent variables or features
10
and a known output variable also called a label
11
or target we train the model using examples where we already
12
know the correct answers the model learns from labeled examples
13
and then predicts labels for new data for example predicting house prices based on size and location is supervised learning.
14
Classifying emails as spam or not spam is also supervised learning.
15
Then identifying whether the images of a cat or a dog is also supervised learning.
16
In unsupervised learning, we do not have labeled outputs in the data set.
17
We only have inputs.
18
The algorithm must find structure or patterns in the data on its own. For example, grouping customers into segments based on purchasing behavior without knowing the
19
groups in advance is unsupervised learning it is like giving a child a pile of images
20
and asking them to group similar ones together without telling them
21
what the categories are now within supervised learning there are two major types of models
22
that we develop namely regression and classification models in regression the goal is to predict a continuous numeric value,
23
for example, predicting the price of a house.
24
In classification, the goal is to predict a category or class, such as spam versus not spam.
25
The most basic regression algorithm is linear regression.
26
It tries to fit a straight line that best describes the relationship between input variables and the output.
27
It does this by minimizing the square differences between predicted values and actual values.
28
A simple example of a linear relationship could be the connection between a person's height and their shoe size.
29
If we collect data from many people, a linear regression model might discover that for every one unit increase in shoe size,
30
the person is on average about two inches taller.
31
In this case, we are fitting a straight line that best explains how height changes with shoe size.
32
Of course, real life is rarely that simple.
33
We can improve the model by including more features such as gender, age, or ethnicity.
34
Instead of using just one input variable, we now use multiple variables to better predict height.
35
In fact, many advanced machine learning algorithms,
36
including neural networks, are extensions of this basic idea of learning relationships between inputs and outputs.
37
Now for classification, the most basic algorithm is logistic regression.
38
Instead of fitting a straight line for numeric prediction, it uses a sigmoid curve to estimate probabilities of belonging to a class.
39
For example, suppose we want to predict whether a person belongs to category A or category B using height and weight.
40
A logistic regression model does not simply give a yes or no answer.
41
Instead, it calculates a probability.
42
For instance, given a height of 180 centimeters and a weight of 75 kilograms,
43
the model might predict a probability of 0 .8 that the person belongs to category A.
44
That means there is an 80 % chance according to the model.
45
If the probability is greater than 0 .5, we usually classify the person as category A.
46
If it is less than 0 .5, we classify them as category B, depending on what A and B are.
47
Another simple but powerful algorithm is K, nearest neighbors, or KNN.
48
The most interesting thing about KNN is that it does not try to learn any equation like linear regression,
49
and it does not try to draw a boundary like logistic regression.
50
It simply stores the training data.
51
Imagine we already have a data set of many people with known gender labels.
52
Now a new person comes with a height of 175 centimeters and weighs 70 kilograms.
53
If we choose K equal to 5, the algorithm will look at the five closest people in the data set, in terms of height and weight.
54
Suppose among those five nearest neighbors, three are male
55
and two are female then by majority vote the model predicts
56
male it literally says you are most similar to these five people
57
and most of them are male
58
so I classify you as male now comes the important part
59
which is choosing K
60
if K is too small say K equals one then the model will simply copy the nearest data point.
61
This makes it extremely sensitive to noise.
62
If that one neighbor is unusual or an outlier, the prediction will also be unusual.
63
This is called overfitting, which means the model memorizes the training data too closely and does not generalize well.
64
On the other hand, if K is too large, then the model averages over too many points.
65
It ignores local structure and becomes too smooth.
66
This is called underfitting where the model becomes too simple and loses important patterns,
67
so the real art in KNN is selecting the right value of K.
68
Usually we try different values and test which one performs best on validation data.
69
Next up we have support vector machines or SVM.
70
Imagine you are trying to classify animals based on weight and nose length into two groups, dogs and elephants.
71
If you plot the data points on a graph, you may see that dogs lie on one side and elephants on the other.
72
Many lines could separate them, but SVM does not choose just any line.
73
It chooses the line that leaves the maximum possible distance between the two classes.
74
This distance is called the margin.
75
But why maximize the margin?
76
Because a larger margin means the boundary is more robust.
77
If a new data point comes in slightly noisy or slightly shifted, a wide margin makes it less likely to be misclassified.
78
The points that lie closest to this boundary are called support vectors.
79
Interestingly, once the boundary is found, only these support vectors matter for defining it.
80
The rest of the data could disappear and the boundary would remain the same.
81
That makes SVM memory efficient and elegant.
82
Now what if the data is not linearly separable?
83
Suppose the classes are arranged in a circular pattern where one class lies inside a circle and the other outside.
84
A straight line cannot separate them.
85
This is where kernel functions come in.
86
Kernels allow SVM to implicitly transform the data into a higher
87
dimensional space where a non -linear separation turns into a linear separation.
88
You can imagine lifting the data into three dimensions, where what looked like a circle in two dimensions becomes separable by a flat plane.
89
This trick is called the kernel trick, and it makes SVM powerful for complex non -linear problems.
90
Next up, we have the Naive Bayes algorithm.
91
It is a classification algorithm based on probability and Bayes' theorem.
92
I will not explain it here as I have already made a detailed video on the same.
93
Check it out later.
94
Decision Tree A decision tree is one of the most intuitive machine learning algorithms.
95
It works by splitting the data step by step, using a sequence of yes or no questions.
96
At each step, the algorithm chooses the question that best separates the data.
97
For example, when predicting whether a patient is high -risk or low -risk, the first question might be, is age greater than 50?
98
Depending on the answer, the data is split into two groups, and the process continues.
99
This creates a tree -like structure with branches and final decision points called leaves.
100
The goal of a decision tree is to make the leaves as pure as possible.
101
Purity means that most data points in a leaf belong to the same class.
102
For classification tasks, this means minimizing misclassified points.
103
For regression tasks, it means minimizing prediction error within each leaf.
104
Although a single decision tree is easy to understand and interpret, it can sometimes overfit the data and become too sensitive to small changes.
105
To make decision trees more powerful and stable, we use ensemble methods.
106
An ensemble method combines many simple models to create a stronger overall model.
107
One popular ensemble technique is called bagging, and a famous example of bagging is the random forest algorithm.
108
Instead of training just one decision tree, we train many trees on different random subsets of the data.
109
Each tree sees a slightly different version of the dataset, which makes them diverse.
110
In a random forest, each tree makes its own prediction,
111
and the final output is determined by majority vote in classification or averaging in regression.
112
Additionally, each tree only considers a random subset of features when making splits.
113
This randomness reduces correlation between trees
114
and prevents them from all making the same mistakes boosting is another powerful ensemble technique
115
but it works differently from random forests instead of training trees
116
independently in parallel boosting trains them sequentially each new tree focuses on correcting the mistakes made by the previous trees
117
over time many weak learners combine to form a strong learner famous boosting algorithms include gradient boosting and XG boost,
118
which often achieve very high accuracy but require careful tuning to avoid overfitting.
119
Then we have neural networks.
120
In simple regression, we directly map input features to an output using a formula.
121
Neural networks add one or more hidden layers between the input and the output.
122
These hidden layers contain many interconnected nodes, often called neurons.
123
Instead of manually deciding which features are important,
124
the network uses calculus and linear algebra to learn useful internal features called weights and biases automatically from the data.
125
Each layer transforms the data slightly and passes it to the next layer.
126
This makes neural networks flexible and capable of modeling complex relationships.
127
When we stack multiple hidden layers, we get deep learning.
128
With multiple layers, the network can learn increasingly abstract representations of the data.
129
For example, in image recognition, the image goes in and the system first finds simple shapes and small features like lines,
130
curves, and patterns, for example, stripes.
131
Then it combines these features to recognize bigger parts, like a zebra's striped body, or a horse's smooth shape.
132
Finally, it puts everything together and decides whether the image is a zebra, a horse, or something else.
133
This is why deep learning has been so successful in tasks like image recognition, speech recognition, and natural language processing.
134
Now let us jump to unsupervised learning.
135
In unsupervised learning, one of the most common tasks is clustering.
136
In clustering, we do not have labels, and we are simply trying to discover natural groupings or clusters in the data.
137
A very popular algorithm for this is K means clustering.
138
I will not explain this as I have already made a detailed video on the same.
139
Check it out later.
140
The next type of algorithm is dimensionality reduction,
141
which focuses on simplifying data by reducing the number of features while keeping as much useful information as possible.
142
Large data sets often have many correlated or redundant features which can slow down models and introduce noise.
143
For example, do we really need a high resolution picture to identify the cat in the picture?
144
Nope.
145
Therefore, algorithms like Principle Component Analysis, or PCA, solves this by finding new directions in the data that capture the maximum variance.
146
So far, we discussed supervised learning and unsupervised learning.
147
But there are two more important categories you should know, semi -supervised learning and reinforcement learning.
148
Semi -supervised learning is a mix of supervised and unsupervised learning.
149
For example, imagine you have 10 ,000 medical images, but only 500 are labeled by doctors.
150
A semi -supervised algorithm uses the small labeled portion to guide learning, while also extracting structure from the large unlabeled portion.
151
This approach is especially useful when labeling data is difficult or costly.
152
Then, reinforcement learning is completely different from the other three.
153
Here, the algorithm does not learn from labeled examples.
154
Instead, it learns by interacting with an environment and receiving rewards or penalties.
155
Think of training a dog.
156
You do not give it labeled data sets.
157
You reward good behavior and discourage bad behavior.
158
Over time, the dog learns what actions lead to rewards.
159
In reinforcement learning, an agent takes actions, observes outcomes, receives rewards, and updates its strategy to maximize total future reward.
160
This is the core idea behind game -playing AI, robotics, self -driving systems, and decision -making systems.
161
If you enjoyed this video, please don't forget to like, share, and subscribe to our channel. So good!

이 레슨의 어휘와 말하기 포인트

이 C1 수준 말하기 레슨은 영상 “Every Machine Learning Model Explained in 15 minutes”을(를) 바탕으로 합니다. 가장 자주 반복되는 단어는 다음과 같습니다: learning, algorithm, model, tree, example. 이 영상에는 섀도잉할 문장 161개와 단어 2176개가 있습니다. 말하는 구간의 길이는 15:55입니다. 화자는 분당 약 137단어의 일정한 속도로 말해서 섀도잉하기에 편한 속도입니다. 영어에서 가장 많이 쓰이는 3,000단어에 속하는 단어가 76%뿐이라 어휘가 어려운 편입니다.

이 영상의 핵심 어휘

영상에서 가장 어려운 단어 15개를 발음, 뜻과 함께 정리했습니다.

단어발음뜻
algorithm 명사/ˈælɡəɹɪðm̩/알고리즘, 알고리듬
regression 명사/ɹiːˈɡɹɛʃ.ən/회귀
predict 동사/pɹɪˈdɪkt/예언하다, 예측하다
ensemble 명사/ˌɑnˈsɑm.bəl/앙상블
prediction 명사/pɹɪˈdɪkʃən/예언, 예측
probability 명사/ˌpɹɑ.bəˈbɪl.ə.ti/확률
spam 명사/spæm/스팸, 스팸 메일
minimize 동사/ˈmɪn.ɪˌmaɪz/최소화(最小化)하다
vector 명사/ˈvɛktɚ/벡터
kilogram 명사/ˈkɪləɡɹæm/킬로그램, 키로
kernel 명사/ˈkɝ.nəl/핵심
maximize 동사/ˈmæksəmaɪz/최대화(最大化)하다
equation 명사/ɪˈkweɪ.ʒən/방정식
dimension 명사/daɪˈmɛn.ʃən/치수
elephant 명사/ˈɛlɪfənt/코끼리

이 영상의 문법

화자가 가장 많이 쓰는 문형을 영상 속 실제 표현과 함께 정리했습니다.

문형영상 속 표현
수동태 be + 과거분사 — 누가 하는지보다 무슨 일이 일어나는지에 초점is divided · are provided · is supervised
관계절 who / which + 절 — 사람이나 사물에 대한 추가 정보set which has · overfitting, which means · dataset, which makes
현재완료 have/has + 과거분사 — 과거의 일이 지금도 관련이 있을 때have already made · has been

주의할 발음

  • “th” 소리: algorithm /ˈælɡəɹɪðm̩/, theorem /ˈθiərəm/, ethnicity /ɛθˈnɪsɪti/
  • “sh”와 “zh” 소리: regression /ɹiːˈɡɹɛʃ.ən/, prediction /pɹɪˈdɪkʃən/, equation /ɪˈkweɪ.ʒən/, dimension /daɪˈmɛn.ʃən/, validation /ˌvæl.əˈdeɪ.ʃən/
  • 긴 단어 — 강세 위치에 주의: reinforcement /ˌɹiːɪnˈfɔːsmənt/, probability /ˌpɹɑ.bəˈbɪl.ə.ti/, intuitive /ɪnˈtjuːɪtɪv/, validation /ˌvæl.əˈdeɪ.ʃən/, manually /ˈmænj(u)əliː/

한국어 화자가 어려워하는 소리:

  • /f/ — ㅍ(/p/)으로 바꾸지 말고 윗니를 아랫입술에 대기: classify /ˈklæs.əˌfaɪ/, reinforcement /ˌɹiːɪnˈfɔːsmənt/, elephant /ˈɛlɪfənt/, transform /tɹænsˈfɔɹm/, graph /ɡɹæf/
  • /v/ — ㅂ(/b/)과 구별하기: supervise /ˈsuː.pə.vaɪz/, vector /ˈvɛktɚ/, intuitive /ɪnˈtjuːɪtɪv/, validation /ˌvæl.əˈdeɪ.ʃən/
  • /z/ — ㅈ이 아니라 성대를 울리는 /s/: supervise /ˈsuː.pə.vaɪz/, minimize /ˈmɪn.ɪˌmaɪz/, maximize /ˈmæksəmaɪz/, noisy /ˈnɔɪzi/

이 영상으로 연습하는 방법

  1. 먼저 말하지 않고 영상을 끝까지 듣고 모르는 단어를 적어 둡니다.
  2. 보통 속도로 한 문장씩 섀도잉하고, 화자의 리듬과 맞을 때까지 반복합니다.
  3. 자신의 목소리를 녹음해 원본과 비교하고, algorithm, regression, predict 같은 단어에 특히 주의합니다.

쉐도잉이란? 영어 실력을 빠르게 키우는 과학적 방법

쉐도잉(Shadowing)은 원래 전문 통역사 훈련을 위해 개발된 언어 학습 기법으로, 다언어 학자인 Dr. Alexander Arguelles에 의해 대중화된 방법입니다. 핵심 원리는 간단하지만 매우 강력합니다: 원어민의 영어를 들으면서 1~2초의 짧은 지연으로 즉시 소리 내어 따라 말하는 것——마치 '그림자(shadow)'처럼 화자를 따라가는 것입니다. 문법 공부나 수동적인 청취와 달리, 쉐도잉은 뇌와 입 근육이 동시에 실시간으로 영어를 처리하고 재현하도록 훈련합니다. 연구에 따르면 이 방법은 발음 정확도, 억양, 리듬, 연음, 청취력, 말하기 유창성을 크게 향상시킵니다. IELTS 스피킹 준비와 자연스러운 영어 소통을 원하는 분들에게 특히 효과적입니다.

섀도잉 방법: 단계별 전체 가이드 읽기 →