シャドーイング練習: 11. Introduction to Machine Learning - 動画で英語スピーキングを学ぶ
レッスンを作成中...
1
The following content is provided under a Creative Commons license.
2
Your support will help MIT OpenCourseWare continue to offer high quality educational resources for free.
3
To make a donation or to view additional materials from hundreds of MIT courses, visit MIT OpenCourseWare at ocw .mit .edu.
4
No. Okay.
5
Welcome back.
6
You know, it's that time of term when we're all kind of doing this.
7
So let me see if I can get a few smiles by simply noting to you that two weeks from today
8
is the last class.
9
Should be worth at least a little bit of a smile, right?
10
Professor Gutag is smiling.
11
He likes that idea.
12
You're almost there.
13
What are we doing for the last couple of lectures?
14
We're talking about linear regression, and I just want to remind you.
15
This was the idea of, I have some experimental data, case of a spring where I put different weights on, I have measured displacements.
16
And regression was giving us a way of deducing a model to fit that data.
17
And in some cases, it was easy.
18
We knew, for example, it was going to be a linear model.
19
We found the best line that would fit that data.
20
In some cases, we said we could use validation
21
to actually let us explore to find the best model that would fit it, whether it's a linear, a cubic, Sorry, a linear, a quadratic, a cubic, some higher order thing.
22
So we'd be using that to deduce something about a model.
23
That's a nice segue into the topic for the next three lectures, the last big topic of the class, which is machine learning.
24
And I'm going to argue, you can debate whether that's actually an example of learning, but it has many of the elements that we want to talk about when we talk about machine learning.
25
So as always, there's a reading assignment.
26
Chapter 22 of the book gives you a good start on this, and it will follow up with other pieces.
27
And I want to start by basically outlining what we're going to do.
28
And I'm going to begin by saying, as I'm sure you're aware, this is a huge topic.
29
I've listed just five subjects in Course 6 that all focus on machine learning.
30
And that doesn't include other subjects where learning is a central part.
31
So natural language processing, computational biology, computer vision, robotics, all rely today heavily on machine learning,
32
and you'll see those in those subjects as well.
33
So we're not going to compress five subjects into three lectures, but what we are going to do is give you the introduction.
34
We're going to start by talking about the basic concepts of machine learning, the idea of having examples and how do you talk about features representing those examples,
35
how do you measure distances between them, and use the notion of distance to try and group similar things together as a way of doing machine learning.
36
And we're going to look, as a consequence, of two different standard ways of doing learning.
37
One we call classification methods.
38
Example we're going to see there's something called k -nearest -neighbor.
39
And the second class called clustering methods.
40
Classification works well when I have what we would call labeled data.
41
I know labels on my examples, and I'm going to use that to try and define classes that I can learn.
42
And clustering working well when I don't have labeled data, and we'll see what that means in a couple of minutes.
43
But we're going to give you an early view of this.
44
Unless Professor Guttag changes his mind, we're probably not going to show you the current, really sophisticated machine learning methods, like convolutional neural nets or deep learning, things you'll read about in the news.
45
But you're going to get a sense of what's behind those by looking at what we do
46
when we talk about learning algorithms.
47
Before I do it, I want to point out to you just how prevalent this is.
48
And I'm going to admit with my gray hair, I started working in AI in 1975 when machine learning was a pretty simple thing to do.
49
And it's been fascinating to watch over 40 years, the change.
50
And if you think about it, just think about where you see it.
51
AlphaGo, machine learning based system from Google that beat a world class level Go player.
52
Chess has already been conquered by computers for a while.
53
Go now belongs to computers.
54
Best Go players in the world are computers.
55
I'm sure many of you use Netflix, any recommendation system.
56
Netflix, Amazon, pick your favorite.
57
It uses a machine learning algorithm to suggest things for you.
58
In fact, you've probably seen it on Google, right?
59
The ads that pop up on Google are coming from a machine learning algorithm that's looking at your preferences.
60
Scary thought about drug discovery, character recognition.
61
The post office does character recognition of handwritten characters using a machine learning algorithm and a computer vision system behind it.
62
You probably don't know this company.
63
It's actually an MIT spinoff called Two Sigma.
64
It's a hedge fund in New York.
65
They heavily use AI and machine learning techniques.
66
And two years ago, their fund returned a 56.
67
percent return.
68
I wish I'd invested in the fund.
69
I don't have the kinds of millions you need, but that's an impressive return, 56 percent return on your money in one year.
70
Last year, they didn't do quite as well, but they do extremely well using machine learning techniques.
71
Siri, another great MIT company called Mobileye
72
that does computer vision systems with a heavy machine learning component
73
that is used in assistive driving and will be used in completely autonomous driving.
74
It will do things like kick in your brakes if you're closing too fast on the car in front of you,
75
which is going to be really bad for me because I drive like a Bostonian and it would be kicking in constantly.
76
Face recognition.
77
Facebook uses this.
78
Many other systems do to recognize, both detect and recognize faces.
79
IBM Watson, cancer diagnosis.
80
These are all just examples of machine learning being used everywhere, and it really is.
81
I've only picked nine. So what is it?
82
I'm going to make an obnoxious statement.
83
You're now used to that.
84
I'm going to claim that, you know, you could argue that almost every computer program learns something.
85
But the level of learning really varies a lot.
86
So if you think back to the first lecture in 6 .0001, we showed you Newton's method for computing square roots.
87
And you could argue, you'd have to stretch it, but you could argue that that method learned something about how to compute square roots.
88
In fact, you could generalize it to roots of any order power.
89
But I really didn't learn.
90
I really had to program it.
91
Think about last week, when we talked about linear regression.
92
That starts to feel a little bit more like a learning algorithm, because what did we do?
93
We gave you a set of data points, mass displacement data points, and then we showed you how the computer could essentially fit a curve to that data point.
94
And it was in some sense learning a model for that data that it could then use to predict
95
behavior in other situations.
96
And that's getting closer to what we would like when we think about a machine learning algorithm.
97
We'd like to have a program that can learn from experience, something that it can then use to deduce new facts.
98
Now, it's been a problem in AI for a very long time.
99
And I love this quote is from a gentleman named Art Samuel.
100
1959 is the quote, in which he says, his definition of machine learning is the field of study that gives computers the ability to learn without being explicitly programmed.
101
And I think many people would argue he wrote the first such program.
102
It learned from experience.
103
In his case, it played checkers.
104
Kind of shows you how the field has progressed.
105
We started with checkers, we got to chess, we now do go.
106
But it played checkers, it beat national level players.
107
Most importantly, it learned to improve its methods by watching how it did in games
108
and then inferring something to change what it thought about as it did that.
109
Samuel did a bunch of other things.
110
I've just highlighted one you may see in a follow -on course.
111
He invented what's called alpha -beta pruning, which is a really useful technique for doing search.
112
The idea is, how can we have the computer learn without being explicitly programmed?
113
And one way to think about this is to think about the difference between how we would normally program
114
and what we would like from a machine learning algorithm.
115
Normal programming, I know you're not convinced there's such a thing as normal programming, but if you think of traditional programming, what's the process?
116
I write a program that I input to the computer so that it can then take data and produce some appropriate output.
117
And the square root finder really sits there, right?
118
I wrote code for using Newton message to find a square root.
119
And then it gave me the process of, given any number, I'll give you the square root.
120
But if you think about what we did last time, it was a little different.
121
And in fact, in a machine learning approach, the idea is that I'm going to give the computer output.
122
I'm going to give it examples of what I want the program to do.
123
Labels on data, characterizations of different classes of things.
124
And what I want the computer to do is, given that characterization of output and data, I wanted that machine learning algorithm to actually produce for me a program.
125
A program that I can then use to infer new information about things.
126
And that creates, if you like, a really nice loop.
127
Where I can have the machine learning algorithm learn the program which I can then use to solve some other problem.
128
That would be really great if we could do it.
129
And as I suggested, that curve fitting algorithm is a simple version of that.
130
It learned a model for the data, which I could then use to label any other instances of the data
131
or predict what I would see in terms of spring displacement as I changed the masses.
132
So that's the kind of idea we're going to explore.
133
If we want to learn things, we could also ask, so how do you learn?
134
And how should a computer learn?
135
Well, for you as a human, there are a couple of possibilities.
136
This is the boring one.
137
This is the old style way of doing it, right?
138
Memorize facts.
139
Memorize as many facts as you can and hope that we ask you on the final exam instances of those facts, as opposed to some other facts you haven't memorized.
140
This is, if you think way back to the first lecture, an example of declarative knowledge.
141
Statements of truth.
142
Memorize as many as you can.
143
Have Wikipedia in your back pocket.
144
A better way to learn is to be able to infer, to deduce new information from old.
145
And if you think about this, this gets closer to what we called imperative knowledge, ways to deduce new things.
146
Now, in the first cases, we built that in when we wrote that program to do square roots.
147
But what we'd like in a learning algorithm is to have much more like that generalization idea.
148
We're interested in extending our capabilities to write programs that can infer useful information from implicit patterns in the data.
149
So not something explicitly built in, like that comparison of weights and displacements, but actually implicit patterns in the data
150
and have the algorithm figure out what those patterns are
151
and use those to generate a program you can use to infer new data about objects, about spring displacements, whatever it is you're trying to do.
152
OK, so the idea then, the basic paradigm that we're going to see is we're going to give the system some training data, some observations.
153
We did that last time with just the spring displacements.
154
We're going to then try and have a way to figure out how do we write code, how do we write a program, a system that will infer something about the process that generated the data.
155
And then from that, we want to be able to use that to make predictions about things we haven't seen before.
156
So again, I want to drive home this point.
157
If you think about it, the spring example fit that model.
158
I gave you a set of data, spatial deviations relative to mass of spaces.
159
For different masses, how far did the spring move?
160
I then inferred something about the underlying process.
161
In the first case, I said, I know it's linear, but let me figure out what the actual linear equation is.
162
What's the spring constant associated with it?
163
And based on that result, I got a piece of code I could use to predict new displacements.
164
So it's got all of those elements, training data, inference engine, and then the ability to use that to make new predictions.
165
But that's a very simple kind of learning setting.
166
So the more common one is what I'm going to use as an example, which is when I give you a set of examples, those examples have some data associated with them, some features,
167
and some labels
168
that for each example I might say this is a particular
169
kind of thing this other one's another kind of thing
170
and what I want to do is figure out how to do inference on labeling new things
171
so it's not just what's the displacement of the mass it's actually a label
172
and I'm going to use one of my favorite examples I'm a big New England Patriots fan
173
if you're not my apologies
174
but I'm going to use football players
175
so I'm going to show you in a second I'm going to give give you a set of examples of football players.
176
The label is the position they play.
177
And the data, well, it could be lots of things.
178
We're going to use height and weight.
179
But what we want to do is then see how would
180
we come up with a way of characterizing the implicit pattern of how does weight
181
and height predict the kind of position this player could play.
182
And then come up with an algorithm that will predict the position of new players.
183
We do the draft for next year.
184
Where do we want them to play?
185
That's the paradigm.
186
All right, set of observations, potentially labeled, potentially not.
187
Think about how do we do inference to find a model, and then how do we use that model to make predictions.
188
What we're going to see, and we're going to see multiple examples today, is that that learning can be done in one of two very broad ways.
189
The first one's called supervised learning, and in that case, for every new example I give you as part of the training data, I have a label on it.
190
I know the kind of thing it is, and what I'm going to do is look for how do I find a rule
191
that would predict the label associated with unseen input based on those examples.
192
It's supervised because I know what the labeling is.
193
Second kind, if this is supervised, the obvious other one is called unsupervised.
194
In that case, I'm just going to give you a bunch of examples, but I don't know the labels associated with them.
195
I'm going to just try and find what are the natural ways to group those examples together into different models.
196
And in some cases, I may know how many models are there.
197
In some cases, I may want to just say, what's the best grouping I can find?
198
OK, what I'm going to do today is not a lot of code.
199
I was expecting cheers for that, John, but I didn't get them.
200
Not a lot of code.
201
What I'm going to do is show you, basically, the intuitions behind doing this learning.
202
And I'm going to start with my New England Patriots example.
203
So here are some data points about current Patriots players.
204
And I've got two kinds of positions.
205
I've got receivers, and I have linemen.
206
And each one is just labeled by the name, the height in inches, and the weight in pounds, five of each.
207
If I plot those on a two -dimensional plot, this is what I get.
208
Okay, no big deal.
209
What am I trying to do?
210
I'm trying to learn, are there characteristics that distinguish the two classes from one another?
211
And in the unlabeled case, all I have are just a set of examples.
212
So what I want to do is decide what makes two players similar, with the goal of seeing,
213
can I separate this distribution into two or more natural groups.
214
Similar is a distance measure.
215
It says, how do I take two examples with values or features associated, and we decide how far apart are they?
216
And in the unlabeled case, the simple way to do it is to say, if I know that there are at least k groups there.
217
In this case, I'm going to tell you there are two different groups there.
218
How could I decide how best to cluster things together so
219
that all the examples in one group are close to each other, all the examples in the other group are close to each other, and they're reasonably far apart.
220
There are many ways to do it.
221
I'm going to show you one.
222
It's a very standard way, and it works basically as follows.
223
If all I know is that there are two groups there, I'm going to start by just picking two examples as my exemplars.
224
Pick them at random.
225
Actually, at random is not great.
226
I don't want to pick too close to each other.
227
I'm going to try and pick them far apart.
228
But I pick two examples as my exemplars.
229
And for all the other examples in the training data, I say, which one is it closest to?
230
What I'm going to try and do is create clusters with the property
231
that the distances between all of the examples in that cluster are small.
232
The average distance is small.
233
And see if I can find clusters that gets the average distance for both clusters as small as possible.
234
This algorithm works by picking two examples, clustering all the other examples by simply saying put it in the group to which it's closest to that example.
235
Once I've got those clusters, I'm going to find the median element of that group.
236
Not mean, but median.
237
What's the one closest to the center?
238
And treat those as my exemplars, and repeat the process.
239
And I'll just do it either some number of times or until I don't get any change in the process.
240
So it's clustering based on distance.
241
And we'll come back to distance in a second.
242
So here's what would happen with my football players.
243
If I just did this based on weight, there's the natural dividing line.
244
And it kind of makes sense.
245
These three are obviously clustered.
246
And again, it's just on this axis.
247
They're all down here.
248
These seven are at a different place.
249
There's a natural dividing line there.
250
If I were to do it based on height, not as clean.
251
This is what my algorithm came up with as the best dividing line here, meaning meaning that these four, again, just based on this axis, are close together.
252
These six are close together.
253
But it's not nearly as clean.
254
And that's part of the issue we'll look at is, how do I find the best clusters?
255
If I use both height and weight, I get that, which is actually kind of nice, right?
256
Those three cluster together.
257
They're near each other in terms of just distance in the plane.
258
Those seven are near each other that there's a nice natural dividing line through here.
259
And in fact, that gives me a classifier.
260
This line is the equidistant line between the centers of those two clusters,
261
meaning any point along this line is the same distance to the center of that group as it is to that group.
262
And so any new example, if it's above the line, I would say gets that label.
263
If it's below the line, gets that label.
264
In a second, we'll come back to look at how do we measure the distances.
265
But the idea here is pretty simple.
266
I want to find groupings near each other and far apart from the other group.
267
Now, suppose I actually knew the labels on these players.
268
I shouldn't call them objects.
269
They're people, these players.
270
These are the receivers.
271
Those are the linemen.
272
And for those of you who are football fans, you can figure out, right, those are the two tight ends.
273
They're much bigger.
274
I think that's Bennett and that's Gronk, if you're really a big Patriots fan.
275
But those are tight ends is those are wide receivers.
276
And it's going to come back in a second.
277
But there are the labels.
278
Now what I want to do is say, if I could take advantage of knowing the labels, how would I divide these groups up?
279
And that's kind of easy to see.
280
Basic idea in this case is, if I've got labeled groups in that feature space, what I want to do is find a subsurface that naturally divides that space.
281
Now, subsurface is a fancy word.
282
It says, in the two -dimensional case, I want to know what's the best line.
283
If I can find a single line
284
that separates all the examples with one label from all the examples of the second label.
285
We'll see that if the examples are well separated, this is easy to do, and it's great.
286
But in some cases, it's going to be more complicated because some of the examples may be very close to one another.
287
And that's going to raise a problem that you saw last lecture.
288
I want to avoid overfitting.
289
I don't want to create a really complicated surface to separate things.
290
And so we may have to tolerate a few incorrectly labeled things if we can't pull it out.
291
And as you already figured out, in this case, with the labeled data, there's the best fitting line, right there.
292
Anybody over 280 pounds is going to be a great lineman.
293
Anybody under 280 pounds is more likely to be a receiver.
294
OK, so I've got two different ways of trying to think about doing this labeling.
295
I'm going to come back to both of them in a second.
296
Now suppose I add in some new data.
297
I want to label new instances.
298
Now, these are actually players of a different position.
299
These are running backs.
300
But I say all I know about is receivers and linemen.
301
I get these two new data points.
302
I'd like to know, are they more likely to be a receiver or a lineman?
303
And there's the data for these two gentlemen.
304
So if I go back to now plotting them, oh, you notice one of the issues.
305
So there are my linemen.
306
The red ones are my receivers the two black dots are the two running backs.
307
Notice right here, it's going to be really hard to separate those two examples from one another.
308
They are so close to each other, and that's going to be one of the things we have to trade off.
309
But if I think about using what I learned as a classifier, with unlabeled data, there were my two clusters.
310
Now you see, oh, I've got an interesting example.
311
This new example, I would say, is clearly more like a receiver than a lineman.
312
But that one there, unclear.
313
Almost exactly lies along that dividing line between those two clusters.
314
And that would either say, I want to rethink the clustering, or I want to say, you know what?
315
As I know, maybe there aren't two clusters here, maybe there are three.
316
And I want to classify them a little differently.
317
So I'll come back to that.
318
On the other hand, if I'd used the labeled data, there was my dividing line.
319
This is really easy.
320
Both of those new examples are clearly below the dividing line.
321
They are clearly examples that I would categorize as being more like receivers than they are like linemen.
322
And I know it's a football example.
323
If you don't like football, pick another example.
324
But you get the sense of why I can use the data in the labeled case
325
and the unlabeled case to come up with different ways of building the clusters.
326
So what we're going to do over the next two and a half lectures
327
is look at how can we write code to learn that way of separating things out.
328
We're going to learn models based on unlabeled data.
329
That's the case where I don't know what the labels are, by simply trying to find ways to cluster things together nearby and then use the clusters to assign labels to new data.
330
And we're going to learn models by looking at labeled data
331
and seeing how do we best come up with a way of separating, with a line or a plane or a collection of lines, examples from one group from examples of the other group,
332
with the acknowledgment that we want to avoid overfitting.
333
We don't want to create a really complicated system.
334
And as a consequence, we're going to have to make some trade -offs between what we call false positives and false negatives.
335
But the resulting classifier can then label any new data by just deciding where you are with respect of that separating line.
336
So here's what you're going to see over the next two and a half lectures.
337
Every machine learning method has five essential components.
338
We need to decide what's the training data and how are we going to evaluate the success of that system.
339
Already seen some examples of that.
340
We need to decide how are we going to represent each instance that we're giving it.
341
I happen to choose height and weight for football players but I might have been better off to pick average speed, or I don't know, arm length, something else.
342
How do I figure out what are the right features?
343
And associated with that, how do I measure distances between those features?
344
How do I decide what's close and what's not close?
345
Maybe it should be different in terms of weight versus height, for example.
346
I need to make that decision.
347
And those are the two things we're going to show you examples of today, how to go through that.
348
Starting next week, Professor Gutteig is going to show you how you take those
349
and actually start building more detailed versions of measuring clustering,
350
measuring similarities, to find an objective function that you want to minimize to decide what is the best cluster to use,
351
and then what is the best optimization method you want to use to learn that model.
352
So let's start talking about features.
353
I've got a set of examples, labeled or not.
354
I need to decide what is it about those examples that's useful.
355
to use when I want to decide what's close to another thing or not.
356
And one of the problems is if it was really easy, it would be really easy.
357
Features don't always capture what you want.
358
I'm going to belabor that football analogy, but why did I pick height and weight?
359
Because it was easy to find.
360
But what is, if you work for the New England Patriots, what is the thing that you really look for when you're asking what's the right feature?
361
It's probably some other combination thing.
362
So you, as a designer, have to say, what are the features I want to use?
363
A quote, by the way, is from one of the great statisticians of the 20th century, which I think captures it well.
364
So, feature engineering, as you as a programmer, comes down to deciding both what are the features I want to measure in that vector that I'm going to put together,
365
and how do I decide relative ways to weight it?
366
So John and Anna and I could have made our chore this, not chore, wrong term, John, our job this term really easy.
367
If we had sat down at the beginning of the term and said, you know, we've taught this course many times.
368
We've got data from, I don't know, John, thousands of students, probably over this time.
369
Let's just build a little learning algorithm.
370
It takes a set of data and predicts your final grade.
371
You don't have to come to class.
372
You don't have to go through all the problems.
373
We'll just predict your final grade.
374
Wouldn't that be nice?
375
Make our job a little easier, and you may or may not like that idea.
376
But I could think about predicting that grade.
377
Now, why am I telling you this example?
378
I was trying to see if I could get a few smiles.
379
I saw a couple of them there.
380
But think about the features.
381
What would I measure?
382
Actually, I'll put this on John because it's his idea.
383
What would he measure?
384
Well, GPA is probably not a bad predictor performance.
385
You do well in other classes, you're likely to do well in this class.
386
I'm going to use this one very carefully.
387
Prior programming experience is at least a predictor, but it is not a perfect predictor.
388
Those of you who haven't programmed before in this class, you can still do really well in this class.
389
But it's an indication that you've seen other programming languages.
390
On the other hand, I don't believe in astrology.
391
So I don't think the month in which you're born, the astrological sign under which you were born, has probably anything to do with how well you'd program.
392
I doubt that eye color has anything to do with how well you'd program.
393
You get the idea.
394
Some features matter, others don't.
395
Now, I could just throw all the features in and hope
396
that the machine learning algorithm sorts out those it wants to keep from those it doesn't.
397
But I remind you of that idea of overfitting.
398
If I do that, there is the danger that it will find some correlation between birth month, eye color, and GPA.
399
And that's going to lead to a conclusion that we really don't like.
400
By the way, in case you're worried, I can assure you that Stu Schmillan, the dean of admissions department, does not use machine learning to pick you.
401
He actually looks at a whole bunch of things is because it's not easy to replace him with a machine yet.
402
So what this says is we need to think about how do we pick the features.
403
And mostly what we're trying to do is to maximize something called the signal to noise ratio.
404
Maximize those features that carry the most information and remove the ones that don't.
405
So I want to show you an example of how you might think about this.
406
I want to label reptiles.
407
I want to come up with a way of labeling animals as, are they a reptile or not?
408
And I give you a single example.
409
With a single example, you can't really do much.
410
But from this example, I know that a cobra, it lays eggs.
411
It has scales.
412
It's poisonous.
413
It's cold -blooded.
414
It has no legs.
415
And it's a reptile.
416
So I could say my model of a reptile is, well, I'm not certain.
417
I don't have enough data yet.
418
But if I give you a second example, and it also happens to be egg -laying, have scales, poisonous, cold -blooded node legs.
419
There's my model, right?
420
Perfectly reasonable model.
421
Whether I design it or a machine learning algorithm would do it says, if all of these are true, label it as a reptile.
422
OK?
423
And now I give you a boa constrictor.
424
Ah, it's a reptile, but it doesn't fit the model.
425
And in particular, it's not egg -laying, and it's not poisonous.
426
So I've got to refine the model, or the algorithm's got to refine the model.
427
And this, I want to remind you, is looking at the features.
428
So I started out with five features.
429
This doesn't fit.
430
So probably what I should do is reduce it.
431
I'm going to look at scales.
432
I'm going to look at cold -blooded.
433
I'm going to look at legs.
434
That captures all three examples.
435
Again, if you think about this in terms of clustering, all three of them would fit with that.
436
Okay, now I give you another example.
437
Chicken.
438
I don't think it's a reptile.
439
In fact, I'm pretty sure it's not a reptile.
440
And it nicely still fits this model, right?
441
Because while it has scales, which you may not realize, it's not cold -blooded and it has legs.
442
So it is a negative example that reinforces the model.
443
Sounds good?
444
And now I'll give you an alligator.
445
It's a reptile, and no fudge, right?
446
It doesn't satisfy the model.
447
Because while it does have scales, and it is cold -blooded, it has legs.
448
I'm almost done with the example, but you see the point.
449
Again, I've got to think about how do I refine this.
450
And I could by saying, all right, let's make it a little more complicated.
451
Has scales, cold -blooded, zero or four legs, I'm going to say it's a reptile.
452
Now, I give you the dart frog.
453
Not a reptile, it's an amphibian.
454
And that's nice, because it still satisfies it.
455
So it's an example outside of the cluster that says, no scales, not cold -blooded, but happens to have four legs, it's not a reptile, that's good.
456
And then I give you a python, right?
457
I mean, there has to be a python in here.
458
Oh, come on, at least groan at me when I say that.
459
There has to be a python here and I give you
460
that in a salmon and now I am in trouble
461
because look at scales look at cold -blooded look at legs
462
I can't separate them on those features there's no way to come up with a way
463
that will correctly say that the Python is a reptile
464
and the salmon is not and so there's no easy way to add in
465
that rule and probably my best thing is to simply go back to just two features There's scales and cold -blooded.
466
And basically say, if something has scales and it's cold -blooded, I'm going to call it a reptile.
467
If it doesn't have both of those, I'm going to say it's not a reptile.
468
It won't be perfect.
469
It's going to incorrectly label the salmon.
470
But I've made a design choice here that's important.
471
And the design choice is that I will have no false negatives.
472
What that means is there's not going to be any instance of something that's not a reptile
473
that I'm going to call a reptile.
474
I may have some false positives.
475
Sorry, I did that the wrong way.
476
A false negative says everything that's not a reptile, I'm going to categorize that direction.
477
I may have some false positives in that I may have a few things that I will incorrectly label as a reptile.
478
And in particular, whoops, sorry, salmon is going to be an instance of that.
479
This trade -off of false positives and false negatives is something
480
that we worry about as we think about it because there's no perfect way, in many cases, to separate the data.
481
And if you think back to my example of the New England Patriots, that running back and that wide receiver were so close together in height and weight, there's no way I'm going to be able to separate them apart.
482
And I just have to be willing to decide how many false positives or false negatives do I want to tolerate.
483
Once I've figured out what features to use, which is good, then I have to decide about distance.
484
How do I compare two feature vectors?
485
I'm going to say vector because there will be multiple dimensions to it.
486
But how do I decide how to compare them?
487
Because I want to use the distances to figure out either how to group things together
488
or how to find a dividing line that separates things apart.
489
So one of the things I have to decide is which features.
490
I also have to decide the distance.
491
And finally, I may want to decide how to weigh relative importance of different dimensions in the feature vector.
492
Some may be more valuable than others in making that decision.
493
And I want to show you an example of that.
494
So let's go back to my animals.
495
I started off with a feature vector that actually had five dimensions to it.
496
It was egg -laying, cold -blooded, has scales, I forget what the other one was, a number of legs.
497
So one of the ways I could think about this is saying I've got four binary features
498
and one integer feature associated with each animal.
499
And one way to learn to separate out reptiles from non -reptiles is to measure the distance between pairs of examples
500
and use that distance to decide what's near each other and what's not.
501
And as we said before, it'd either be used to cluster things or to find a classifier surface that separates them.
502
So here's a simple way to do it.
503
For each of these examples, I'm going to just let true be one, false be zero.
504
So the first four are either zeros or ones, and the last one's the number of legs.
505
And now I could say, all right, how do I measure distances between animals, or anything else, but these kinds of feature vectors.
506
Here we're going to use something called the Minkowski metric, or the Minkowski difference.
507
Given two vectors and a power p, we basically take the absolute value of the difference between each of the components of the vector,
508
raise it to the pth power, take the sum, and take the pth root of that.
509
So let's do the two obvious examples is if p is equal to 1, I just measure the absolute distance between each component, add them up, and that's my distance.
510
It's called the Manhattan metric.
511
The one you've seen more, the one we saw last time, if p is equal to 2, this is Euclidean distance, right?
512
It's the sum of the squares of the differences of the components.
513
Take the square root.
514
Take the square root because it makes it have certain properties of a distance.
515
That's the Euclidean distance.
516
So now, if I want to measure difference between these two, here's the question.
517
Is this circle closer to the star or closer to the cross?
518
Unfortunately, I put the answer up here.
519
But it differs depending on the metric I use.
520
Euclidean distance, well, that's square root of 2 times 2.
521
So it's about 2 .8.
522
And that's 3.
523
So in terms of just standard distance in the plane, we would say that these two are closer than those two are.
524
Manhattan distance.
525
Why is it called that?
526
Because you can only walk along the avenues and the streets.
527
Manhattan distance would basically say this is one, two, three, four units away.
528
This is one, two, three units away.
529
And under Manhattan distance, this is closer, this pairing is closer than that pairing is.
530
Now you're used to thinking Euclidean, we're going to use that, but this is going to be important when we think about how are we comparing distances between between these different pieces.
531
So typically, we'll use Euclidean.
532
We're going to see Manhattan actually has some values.
533
So if I go back to my three examples, boy, that's a gross slide, isn't it?
534
But there we go.
535
Rattlesnake, boa constrictor, and dart frog.
536
There's the representation.
537
I can ask, what's the distance between them?
538
In the handout for today, we've given you a little piece of code that would do that.
539
And if I actually run through it, I get actually a nice little result.
540
Here are the distances between those vectors using Euclidean metric.
541
Oops, sorry, I'm going to come back to them.
542
But you can see the two snakes nicely are reasonably close to each other, whereas the dart frog is a fair distance away from that.
543
Nice, right?
544
That's a nice separation that says there's a difference between these two.
545
OK, now I throw in the alligator.
546
Sounds like a Dungeons and Dragons game.
547
I throw in the alligator, and I want to do the same comparison.
548
And I don't get nearly as nice a result.
549
Because now it says, as before, the two snakes are close to each other.
550
But it says that the dart frog and the alligator are much closer under this measurement
551
than either of them is to the other.
552
And to remind you, the alligator and the two snakes, I would like to be close to one another and a distance away from the frog,
553
because I'm trying to classify reptiles versus not. So what happened here?
554
Well, this is a place where the feature engineering is going to be important.
555
Because in fact, the alligator differs from the frog in three features.
556
But, and only in two features from, say, the boa constrictor.
557
But one of those features is the number of legs.
558
And there, while on the binary axes, the difference is between a 0 and 1, Here it can be between 0 and 4.
559
So that is weighing the distance a lot more than we would like.
560
The legs dimension is too large, if you like.
561
How would I fix this?
562
This is actually, I would argue, a natural place to use Manhattan distance.
563
Why should I think that the difference in the number of legs
564
or the number of legs difference is more important than whether it has scales or not?
565
Why should I think that measuring that distance, Euclidean -wise, makes sense?
566
They are really completely different measurements.
567
And in fact, I'm not going to do it, but if I ran Manhattan metric on this, it would get the alligator much closer to the snakes, exactly because it differs only in two features, not three.
568
The other way I could fix it would be to say, I'm letting too much weight be associated with the difference in the number of legs.
569
So let's just make it a binary feature here.
570
Either it doesn't have legs, or it does have legs.
571
Run the same classification, and now you see the snakes and the alligator are all close to each other.
572
Whereas the dart frog, not as far away as it was before, but there's a pretty natural separation, especially using that number between them.
573
What's my point?
574
Choice of features matters.
575
Throwing Too many features in may, in fact, give us some overfitting.
576
And in particular, deciding the weights that I want on those features has a real impact.
577
And you as a designer, a programmer, have a lot of influence in how you think about using those.
578
So feature engineering really matters.
579
How you pick the features, what you use, is going to be important.
580
OK.
581
The last piece of this, then, is we're going to look at some examples where we give you data, We've got features associated with them.
582
We're going to, in some cases, have them labeled, in other cases, not.
583
And we know how now to think about how do we measure distances between them.
584
John?
585
You probably didn't intend to say weights of features.
586
You intended to say other scales.
587
Sorry, the scales of not though.
588
Thank you, John.
589
No, I did.
590
I take that back.
591
I did not mean to say weights of features.
592
I meant to say the scale of the dimension is going to be important here.
593
Thank you for the for the amplification and correction.
594
You're absolutely right.
595
Weights in a different way, as we'll see next time.
596
And we're going to see next time why we're going to use weights a different way.
597
So rephrase it, block that out of your mind.
598
We're going to talk about scales, and the scale on the axes as being important here.
599
And we already said, we're going to look at two different kinds of learning, labeled and unlabeled, clustering and classifying.
600
And I want to just finish up by showing you two examples of that, how we would think about them algorithmically, and we'll look at them in more detail next time.
601
As we look at it, I want to remind you the things that are going to be important to you.
602
How do I measure distance between examples?
603
What's the right way to design that?
604
What is the right set of features to use in that vector?
605
And then, what constraints do I want to put on the model?
606
In the case of unlabeled data, how do I decide how many clusters I want to have?
607
Because I can give you a really easy way to do clustering.
608
If I give you 100 examples, I say, They build 100 clusters.
609
Every example is its own cluster.
610
Distance is really good.
611
It's really close to itself.
612
But it does a lousy job of labeling things on it.
613
So I have to think about, how do I decide how many clusters?
614
What's the complexity of that separating service?
615
How do I basically avoid the overfitting problem, which I don't want to have?
616
So just to remind you, we've already seen a little version of this, the clustering method.
617
This is a standard way to do it, simply repeating what we had on an earlier slide.
618
If I want to cluster it into groups, I start by saying, how many clusters am I looking for?
619
Pick an example I take as my early representation.
620
For every other example in the training data, put it to the closest cluster.
621
Once I've got those, find the median, repeat the process.
622
And that led to that separation.
623
Now, once I've got it, I like to validate it.
624
And in fact, I should have said this better, those two clusters came without looking looking at the two black dots.
625
Once I put the black dots in, I'd like to validate, how well does this really work?
626
And that example there is really not very encouraging.
627
It's too close.
628
So that's a natural place to say, OK, what if I did this with three clusters?
629
That's what I get.
630
I like that.
631
That has a really nice cluster up here.
632
The fact that the algorithm didn't know the labeling is irrelevant.
633
There's a nice grouping of five.
634
There's a nice grouping of four.
635
And there's a nice grouping of three in between.
636
And in fact, if I looked at the distance, the average distance between examples in each of these clusters, it is much tighter than in that example.
637
And so that leads to then the question of, should I look for four clusters?
638
Question, please.
639
Is the overlap between the two clusters not an issue?
640
Yeah, the question is, is the overlap between the two clusters a problem?
641
No. I just drew it here so I could let you see where those pieces are, but in fact, if you like, the boundary is that the center is there,
642
those three points are all closer to that center than they are to that center.
643
So the fact that they overlap is a good question, it's just the way I happen to draw them.
644
I should really draw these not as circles, but as some little bit more convoluted surface.
645
Having done three, I could say, should I look for four?
646
Well, those points down there, as I've already said, are an example where it's going to be be hard to separate them out.
647
And I don't want to overfit, because the only way to separate those out is going to be to come up with a really convoluted cluster, which I don't like.
648
Let me finish with showing you one other example from the other direction.
649
Suppose I give you labeled examples.
650
So again, the goal is I've got features associated with each example.
651
They're going to have multiple dimensions on it.
652
But I also know the label associated with them.
653
And I want to learn what is the best way to come up with a rule
654
that will let me take new examples and assign them to the right group.
655
A number of ways to do this.
656
You can simply say, I'm looking for the simplest surface that will separate those examples.
657
In my football case that we're in the plane, what's the best line that separates them, which turns out to be easy?
658
I might look for a more complicated surface.
659
And we're going to see an example in a second where maybe it's a sequence of line segments that separates them out, because there's not just one line that does the separation.
660
As before, I want to be careful.
661
If I make it too complicated, I may get a really good separator, but I overfit to the data.
662
And you're going to see next time, I'm going to highlight it here, there's a third way which will lead to almost the same kind of result called k -nearest neighbors.
663
And the idea here is I've got a set of labeled data, and what I'm going to do is for every new example,
664
say, find the k, say the five closest labeled examples, and take a vote.
665
if three out of five or four out of five or five out of five of those labels are the same, I'm going to say it's part of that group.
666
And if I have less than that, I'm going to leave it as unclassified.
667
And that's a nice way of actually thinking about how to learn them.
668
And let me just finish by showing you an example.
669
And I won't use football players on this one.
670
I'll use a different example.
671
I'm going to give you some voting data.
672
I think this is actually simulated data.
673
But these are a set of voters in the United States with their preference.
674
They tend to vote Republican.
675
They tend to vote Democrat.
676
And the two categories are their age and how far away they live from Boston.
677
Whether those are relevant or not, I don't know.
678
But there are just two things I'm going to use to classify them.
679
And I'd like to say, how would I fit a curve to separate those two classes?
680
I'm going to keep half the data to test.
681
I'm going to use half the data to train.
682
So if this is my training data, I could say, what's the best line that separates these?
683
And I don't know about best, but here are two examples.
684
This solid line has the property that all the Democrats are on one side.
685
Everything on the other side is a Republican, but there are some Republicans on this side of the line.
686
I can't find a line that completely separates these as I did with the football players.
687
But there's a decent line to separate them.
688
Here's another candidate, that dashed line, has the property that on the right side you've got, boy I don't think this is deliberate John Wright,
689
but on the right side you've got almost all Republicans, seems perfectly appropriate, one Democrat, but there's a pretty good separation there,
690
and on the left side you've got a mix of things
691
but most of the Democrats are on the left side of that line.
692
The fact that left and right correlates with distance from Boston is completely irrelevant here
693
but it has a nice punch to it.
694
Irrelevant, but not accidental.
695
But not accidental.
696
Thank you.
697
All right, so now the question is, how would I evaluate these?
698
How do I decide which one's better?
699
And I'm simply going to show you very quickly some examples.
700
First one is to look at what's called the confusion matrix.
701
What does that mean?
702
It says, for one of these classifiers, for example, the solid line, here are the predictions based on the solid line of whether they would be more likely to be Democrat or Republican.
703
And here's the actual label.
704
Same thing for the dashed line.
705
And that diagonal is important, because those are the correctly labeled results, right?
706
It correctly, in the dashed line case, solid line case gets all of the correct labelings of the Democrats.
707
It gets half of the Republicans right, but it has some where it's actually a Republican, but it labels it as a Democrat.
708
That we'd like to be really large.
709
And in fact, it leads to a natural measure called the accuracy, which is, just to go back to that, we say that these are true positives,
710
meaning I labeled it as being an instance and it really is.
711
These are true negatives.
712
I labeled it as not being an instance and it really isn't.
713
And then these are the false positives.
714
I labeled it as being an instance and not, and these are the false negatives.
715
I labeled it as not being an instance and And it is.
716
And an easy way to measure it is to look at the correct labels over all of the labels, the true positives and the true negatives, the ones I got right.
717
And in that case, both models come up with a value of 0 .7.
718
So which one's better?
719
Well, I should validate that, and I'm going to do that in a second by looking at other data.
720
We could also ask, could we find something with less training error?
721
This is only getting 70 % right.
722
Not great.
723
Well, here's a more complicated model, and this is where you start getting worried about overfitting.
724
Now what I've done is I've come up with a sequence of lines that separate them.
725
So everything above this line I'm going to say is a Republican.
726
Everything below this line I'm going to say is a Democrat.
727
So I'm avoiding that one, I'm avoiding that one, I'm still capturing many of the same things.
728
And in this case, I get 12 true positives, 13 true negatives, and only 5 false positives.
729
And that's kind of nice.
730
You can see the five.
731
It's those five red ones down there.
732
Its accuracy is 0 .833.
733
And now if I apply that to the test data, I get an OK result.
734
It has an accuracy of about 0 .6.
735
I could use this idea to try and generalize and say, could I come up with a better model?
736
And you're going to see that next time.
737
There could be other ways in which I measure this, and I want to use this as the last example.
738
Another good measure we use is called PPV,
739
positive predictive value, which is how many true positives do I come up with out of all the things I labeled positively.
740
And in the solid model, in the dashed line, I get values about 0 .57.
741
The complex model on the training data is better, and the testing data is even stronger.
742
And finally, two other examples are called sensitivity and specificity.
743
Sensitivity basically tells you what percentage did I correctly find, and specificity said what percentage did I correctly reject.
744
And I show you this because this is where the trade -off comes in.
745
I could make sense that sensitivity is how many did I correctly label out of those
746
that I both correctly labeled and incorrectly labeled as being negative.
747
How many of them did I correctly labeled as being the kind that I want.
748
I can make sensitivity one.
749
Label everything as the thing I'm looking for.
750
Great.
751
Everything's correct.
752
But the specificity will be 0, because I'll have a bunch of things incorrectly labeled.
753
I could make the specificity 1.
754
Reject everything.
755
Say nothing is an instance.
756
True negatives goes to 0, but the sensitivity.
757
Sorry, true negatives goes to 1, and I'm in a great place there, but my sensitivity goes to 0.
758
I've got a trade -off.
759
As I think about the machine learning algorithm I'm using and my choice of that classifier,
760
I'm going to see a trade -off where I can increase specificity at the cost of sensitivity or vice versa.
761
And you'll see a nice technique called ROC, or Receiver Operator Curve, that gives you a sense of how you want to deal with that.
762
And with that, we'll see you next time.
763
We'll take your question offline, if you don't mind, because I've run over time.
764
But we'll see you next time, where Professor Goodtag will show you examples of this.
📺 同じチャンネル
✨ おすすめ動画
ビデオの背景とコンテキスト
このビデオはMITの講義からのもので、機械学習の入門内容が話されています。講師のERIC GRIMSON教授は、線形回帰から機械学習へと話題を移し、その重要性や応用分野(自然言語処理、コンピュータビジョンなど)を説明しています。また、機械学習の基本概念として「ラベル付きデータを使う分類法」と「ラベルなしのクラスタリング法」に触れ、実例としてAlphaGoも挙げられています。日常的な会話ではないですが、学术的な英語のリスニングやスピーキングの練習に最適です。
日常会話に使えるトップ5フレーズ
- Welcome back.(おかえりなさい。)- 講義やミーティングの再開時に使えます。
- You're almost there.(もう少しですよ。)- 相手を励ますときに便利です。
- Let me see if I can...(...できるか試してみます。)- 提案や試みを表す柔らかい表現。
- As I'm sure you're aware,(ご存知の通り、)- 既知情報を述べる時の前置き。
- We're not going to...(私たちは...しません。)- 明確な意思表示に使えます。
シャドーイングのステップバイステップガイド
このビデオは学术的な内容で、スピードが比較的ゆっくりですが、専門用語が多いので注意が必要です。IELTS スピーキング対策や英語シャドーイングの練習には以下の方法がおすすめです。
- 1回目:リスニング中心 - 全体の流れを理解し、難しい単語(例:linear regression, clustering)をメモします。
- 2回目:シャドーイング開始 - 1文ずつ止めながら、shadow speechのように真似します。アクセントやイントネーションに注目しましょう。
- 3回目:スピードを合わせる - ビデオを再生しながら、ほぼ同時に発音します。shadowspeaksのようなリズムを掴むのがポイントです。
- 4回目:録音して比較 - 自分の発音を録音し、講師との違いを確認します。専門用語の発音(例:convolutional neural nets)も正確になるように練習します。
この方法でshadowing siteでの練習も効果的です。学术的な英語に慣れることで、IELTSのスピーキングやリスニングのスコアアップにも繋がります。
シャドーイングとは?英語上達に効果的な理由
シャドーイング(Shadowing)は、もともとプロの通訳者養成プログラムで開発された言語学習法で、多言語習得者として知られるDr. Alexander Arguelles によって広く普及されました。方法はシンプルですが非常に効果的:ネイティブスピーカーの英語を聞きながら、1〜2秒の遅延で声に出してすぐに繰り返す——まるで「影(shadow)」のように話者を追いかけます。文法ドリルや受動的なリスニングと異なり、シャドーイングは脳と口の筋肉が同時にリアルタイムで英語を処理・再現することを強制します。研究により、発音精度、抑揚、リズム、連音、リスニング力、そして会話の流暢さが大幅に向上することが確認されています。IELTSスピーキング対策や自然な英語コミュニケーションを目指す方に特におすすめです。















