Prática de Shadowing: Why AI art struggles with hands - Aprenda a falar inglês com vídeo

Criando lição...
1
You're called to create.
2
A post -apocalyptic giraffe astronaut?
3
Generated.
4
Genghis Khan playing a guitar solo.
5
Pixel art?
6
Generated.
7
A man holding a delicious apple?
8
Ah, what's with his hands?
9
Why can't AI art make hands?
10
It doesn't matter what AI art model you use, if you have a man holding a delicious apple, his hands will look weird holding it.
11
Why is this so hard?
12
Seems easy enough, right?
13
We've got this weird situation where AI art can instantly make.
14
Abraham Lincoln dressed like glam David Bowie, but struggles with a woman holding a cell phone.
15
This isn't just a weird glitch.
16
The struggle of AI art with hands can actually teach you something bigger about how AI art works.
17
I mean, what is so hard about this?
18
I asked an artist who has taught thousands of people how to draw hands from imagination.
19
Before someone becomes or starts training to be an artist, like officially training, it's pattern recognition.
20
You just grow up seeing a whole bunch of hands and you start knowing what hands look like.
21
You learn how things look by living in the world and recognizing patterns.
22
An AI is similar but has key differences.
23
Imagine an AI is like you but trapped in a museum from birth.
24
All the machine has to learn from are the pictures and the little placards on the side.
25
Apple.
26
A red apple on a brown table.
27
That's like the images it sees from the web and the descriptions that go with them.
28
It's similar to how you learn but locked in that museum.
29
If you want to understand an apple, you can rotate it in your hand.
30
You can watch it whenever you want.
31
If AI wants to understand an apple, it has to find another picture of an apple in the museum.
32
Pattern recognition has allowed AI and people to draw decent apples.
33
But the processes differ.
34
You start training to become an artist and now you're like, okay now I have to learn the rules and that's where it becomes very different from how AI is learning.
35
Artists in order to draw something complicated we tend to simplify things into basic forms and so
36
when you look at a hand you pretty much have the
37
big blocky part of the palm right you have the front you have the back
38
and then you have the thickness and you can pretty much just make
39
that into like a square with some thickness to it.
40
Then an artist can add all the style and texture and detail they want.
41
AI works differently.
42
Look at this hand.
43
The shapes are bizarre, but the AI has done a great job showing the light and texture here.
44
Remember, the AI knows how things look, but not how they work.
45
So these patterns in pixels are easy to understand.
46
It never learned, however, that fingers don't really bend like this.
47
It doesn't simplify to forms.
48
Remember, it's trapped in the museum.
49
So it is just trying to guess where hand -like pixels should be without knowing how hands work, like we do.
50
But listen, I find this kind of dissatisfying.
51
I mean, I'm basically just saying that AI can't draw hands because it's not a person.
52
But AI also doesn't know anything about construction, and it can still make a beautiful skyscraper in New York City.
53
So to understand this better, I spoke to two people who have worked with generative art models.
54
Ilun Du is a grad student whose heart is in robotics, but, you know, AI art is like a big deal now, so he got pulled into it.
55
Because of how popular these models have been in generative art, I've also been like reading a bit on that.
56
And I talked to Roy Shilkrat, who has a super varied resume, but has been teaching about generative art since 2018.
57
Good students that come in that are trying to break those models, take them to the next level.
58
Talking to them helped me figure out three big reasons.
59
Not every reason, but three big reasons that hands are tough for AI art models.
60
The data size and quality, the way hands act, and the low margin for error.
61
For the data size, let's go back to the museum idea.
62
The museum the robot hangs out in, it has a ton of rooms dedicated to faces, but not so many rooms for hands.
63
That means it has less to learn from.
64
Just as an example, available datasets like Flickr HQ has 70 ,000 faces.
65
70 ,000!
66
And this popular one annotates 200 ,000 pics of celebrity faces for lots of details, like eyeglasses or pointy noses.
67
There are There are a ton of great hand datasets that can really understand hands, like this one with 11 ,000 hands,
68
but these may not have been used to train the AI that makes art.
69
That data scarcity combines with the quality and complexity of the data.
70
Hand data in the art museum isn't yet annotated to show how they work, like the celebrity's pointy noses.
71
What they say is there's an image and there is a person in the image and the person is holding an umbrella.
72
You don't give the machine a lot of clues saying, this is a person holding the umbrella.
73
The thumb is going from one side of the handle and the fingers are curled.
74
And then the thumb is covering the index finger, but not the other ones.
75
All that is made worse because hands do lots of things compared to, say, faces.
76
So there's a pretty common portrait photo face.
77
There are a lot of these photos online.
78
And I think that everything's very well -centered, right?
79
Eyes are always around here.
80
There's always this order.
81
That's not true of hands which can do this, and this, and this.
82
I swear I'm sober right now.
83
Stan mentioned this too.
84
How many fingers do you see right now?
85
Like, two or three?
86
Like, it doesn't know there's five, because sometimes there's two, sometimes there's three, sometimes four, sometimes five.
87
You can see these problems with AI hands, but the jankiness is all over AI art.
88
Just look at horses.
89
You can often have, like, three legs, five legs, six legs.
90
The model does not learn to explain this because there's too much diversity, and it doesn't have as much bias as we do.
91
Okay, did you hear that last part he said?
92
Good, because it's really important.
93
It doesn't have as much bias as we do.
94
We care a lot about hands and need them to be perfect.
95
There is a low margin for error.
96
But because the model doesn't understand hands, hasn't seen many, and because hands act weird,
97
it makes pictures that are like hands it's seen in the museum, but not an exact hand.
98
That's good enough for a ton of stuff, but not hands.
99
Here, let me give you some examples.
100
Come over here.
101
So I typed, make me a person with exactly five freckles.
102
So this one's from Dolly 2, this one is from Stable Diffusion, and this one is from Mid -Journey.
103
So it's like, you know, great job.
104
You've got, you know, a red -haired person, they're more likely to have freckles, but there are not exactly five freckles here.
105
Here that doesn't really matter, because we see a freckly face, but hands require higher standards.
106
Look at our apple holding man again.
107
I made three other variations.
108
The hands are all weird, but don't look at them right now.
109
It changed the shirt stripes, the buttons, the apple style.
110
None of that matters because it's stripe -like and button -like and apple -like, but hand -like isn't good enough.
111
I came away from this thinking a couple of things.
112
AI art is basically bad at art.
113
We're just able to see it with hands.
114
And B, it's never going to get any better.
115
But both of those things are a bit wrong.
116
I will say
117
that the newest AI art generator to come out at the time of this video is Mid -Journey version 5.
118
And they made some progress with hands, for sure.
119
But it's not totally fixed yet.
120
Don't tell the AI to hold an umbrella.
121
I think they're, like, spending lots of time on, like, some things that you appreciate, which is why you like the images, and a lot of stuff that you don't actually even notice.
122
I think, like, for a lot of natural scenery or something like that, I feel like the model might be getting at that people.
123
And they are working on two things.
124
First, they have the AI look at a ton more pictures, which requires more computing power.
125
They're trying to solve that on a big scale, because if you want to train on more than a handful of images
126
if you want to train
127
or more than 100 images this would take tremendous resources from
128
you to retrain the model itself the other solution might be
129
to invite more people into the museum there's an interesting analog
130
so like have you heard of like chat gbt the big difference was
131
that it basically used human feedback so like they generated many many sentences
132
and asked people to rate which ones are good and
133
which ones are not good they basically fine -tune the model
134
so that it would generate sentences
135
that are convincing to people I guess it would require a lot of engineering to get people to label
136
so much data but i think
137
if we could just get like people to rank how good
138
the images are generated by these models then like a lot of these issues will go away actually
139
because they're just training the models to do what people like it's not just the hand teeth
140
and abs anything where there's like a pattern a large amount
141
of something it doesn't know the rule of there are this many because it's trained on different amounts

Vocabulário e dicas de fala para esta lição

Este vídeo tem 140 frases e 1669 palavras para praticar shadowing. A fala dura 9:47. O falante fala em ritmo natural, cerca de 171 palavras por minuto, próximo de uma conversa do dia a dia. 85% das palavras estão entre as 3.000 mais comuns do inglês; vale a pena estudar o restante antes de começar.

Vocabulário principal deste vídeo

15 palavras do vídeo que vale a pena aprender, com pronúncia e significado:

PalavraPronúnciaSignificado
generate verbo/ˈd͡ʒɛn.ə.ɹeɪt/gerar
finger substantivo/ˈfɪŋɡɚ/dedo (da mão)
umbrella substantivo/ʌmˈbɹɛl.ə/guarda-chuva, guarda-sol
pixel substantivo/ˈpɪk.səl/pixel, píxel
recognition substantivo/ˌɹɛkəɡˈnɪʃ(ə)n/reconhecimento
delicious adjetivo/dɪˈlɪ.ʃəs/delicioso, saboroso
trap substantivo/tɹæp/armadilha, arapuca
margin substantivo/ˈmɑɹ.d͡ʒɪn/margem
bias substantivo/ˈbaɪ.əs/predisposição, inclinação
thumb substantivo/ˈθʌm/dedo polegar, polegar
thickness substantivo/ˈθɪknəs/grossura, espessura
stripe substantivo/stɹaɪp/listra
guitar substantivo/ɡɪˈtɑɹ/guitarra, violão
complicated adjetivo/ˈkɑm.plɪˌkeɪ.tɪd/complicado
compare verbo/kəmˈpɛɚ/comparar

Phrasal verbs que você vai ouvir

PalavraSignificado
figure out verbodescobrir, deduzir
go away verbosair, partir
grow up verbocrescer

Gramática neste vídeo

As estruturas que o falante mais usa, com as palavras exatas do vídeo:

EstruturaNo vídeo
Present perfect have/has + particípio passado — uma ação passada que ainda importa agorahas taught · has allowed · has done
Voz passiva be + particípio passado — o foco está no que acontece, não em quem fazare curled · is made · are generated
Orações relativas who / which + oração — informação extra sobre uma pessoa ou coisapeople who have · hands which can · appreciate, which is

Pronúncia para ficar de olho

O falante usa 32 contrações e formas reduzidas, como doesn't, don't, they're. Diga-as na forma curta, do jeito que você ouve.

  • Os sons de “th”: thumb /ˈθʌm/, thickness /ˈθɪknəs/
  • Os sons de “sh” e “zh”: recognition /ˌɹɛkəɡˈnɪʃ(ə)n/, delicious /dɪˈlɪ.ʃəs/, imagination /ɪˌmæd͡ʒəˈneɪʃən/, variation /ˌvɛəɹiˈeɪʃn̩/
  • Palavras longas — acerte a sílaba tônica: recognition /ˌɹɛkəɡˈnɪʃ(ə)n/, complicated /ˈkɑm.plɪˌkeɪ.tɪd/, diversity /daɪˈvɜː(ɹ)sɪti/, celebrity /səˈlɛb.ɹɪ.ɾi/, differently /ˈdɪf.ə.ɹənt.li/

Como praticar com este vídeo

  1. Ouça o vídeo inteiro uma vez sem falar e anote as palavras que você não conhece.
  2. Comece na velocidade 0,75×, faça shadowing frase por frase e volte à velocidade normal quando ficar fácil.
  3. Grave a sua voz e compare com o original, prestando atenção a palavras como generate, finger, umbrella.

O que é a Técnica de Shadowing?

Shadowing é uma técnica de aprendizado de idiomas com base científica, originalmente desenvolvida para o treinamento de intérpretes profissionais. O método é simples, mas poderoso: você ouve áudio em inglês nativo e repete imediatamente em voz alta — como uma sombra seguindo o falante com 1-2 segundos de atraso. Pesquisas mostram melhora significativa na precisão da pronúncia, entonação, ritmo, sons conectados, compreensão auditiva e fluência na fala.

Técnica de shadowing: leia o guia completo passo a passo →