Prática de Shadowing: How AI is Changing Data Scientist's Workflow - Aprenda a falar inglês com vídeo

Criando lição...
1
Data scientist is the top job title referencing generative AI.
2
So that means that companies don't just need data talent.
3
They need data talent that knows AI.
4
So what does know AI actually mean?
5
Well, I found that it's actually twofold.
6
Firstly, you know how to leverage AI to speed up your existing workflow, not just in a way of prompting ChatGBT to write Python code kind of way,
7
but it's more about rethinking the way you get things done from generating or collecting data, exploratory data analysis, modeling, reporting, and communication.
8
Secondly, you know how to build AI systems.
9
For example, creating and deploying tools and agents to automate part of or an entire process or workflow for your team.
10
Now, let's first talk about different ways to leverage AI to speed up a data science workflow.
11
I'll talk about a few specific tools in this video, but I want to be clear that tools isn't the point because tools can change.
12
What matters is knowing what's now possible so you can think about how to adapt and rethink the way you're working.
13
So the first major way data scientists can leverage AI is in EDA and data pre-processing.
14
If you've worked in a real data science project before, you probably already experienced this.
15
Too much time spent preparing data, not enough time analyzing it.
16
You got the job thinking you spent all day finding data patterns, coming up with hypothesis and interesting analysis.
17
In reality, you spent 60 to 70% of your time cleaning data, writing repetitive code to check data types,
18
handling missing values, or creating standard visualizations.
19
Here are two open source tools I've come across that can make your work a little bit easier and more fun.
20
The first one is Pandas AI.
21
It's a Python library that basically allows you to ask questions to your data frame in natural language.
22
For example, what's the total sales per coffee type?
23
Depending on your question, it can return different kinds of responses like a string, a data frame, a chart, or a number.
24
It's totally open source and you can hook it up with any LLM you like.
25
And I'm connecting Pandas AI with an LLM from OpenAI here.
26
Now, if you're wondering if your whole data set would be sent to an LLM or not, the answer is no. When you ask Pandas AI a question, it actually sends just enough context.
27
So your question along with the column names, a few sample rows, and some metadata.
28
The LLM reads that context
29
and writes a small Python script to answer your question pandas AI then checks
30
that the script is safe to run
31
and then it runs the script locally on your actual data using pandas
32
so if you want to get away with writing pandas
33
and matplotlib code yourself you can use this library to do quick analysis data cleaning
34
and data manipulation but for more complex tasks like joining data frames
35
or creating correlation matrix you should still use pandas the second tool
36
that I think is pretty cool is data formulator this one comes out of microsoft research
37
and the easiest way to think about it is the combination of excel formulas power query
38
and modern ai it also allows you to interact with your data in natural language
39
but the nice part is the tool shows you exactly how the data is being transformed
40
and visualized step by step nothing hidden which means the workflow stays transparent and reusable.
41
Give it a broad prompt like show me interesting trends in this data set and we'll start transforming the data,
42
generating visualizations and suggesting directions you can explore further.
43
This tool is also open source and you can install it locally.
44
So in your workflow if you happen to spend a lot of time in the Microsoft ecosystem like Excel, Power Query, Power BI, it's definitely worth experimenting with this tool.
45
Now with these kind of tools you still need to understand your data
46
but they can drastically speed up the exploration phase where you're
47
just trying to understand what you're working with before you get into building the actual pipeline
48
and the real modeling work now as a top one percent data scientist you don't just limit yourself to tools
49
that are already available you also know how to build ai tools
50
and systems to support your team
51
so i want to quickly shout out to datacamp who sponsors this part of the video I used Datacam back
52
when I was learning Python for work.
53
What I really like is that it's hands-on.
54
You actually write code, solve exercises, and build projects right in your browser.
55
They're offering some solid tracks for data scientists to level up your skills.
56
The first one I'd recommend is the Associate AI Engineer for Data Scientist track that covers scikit-learn, PyTorch, Hugging Face, building LLM apps with Lanchain,
57
and the MLObs basics to help you take models into production.
58
If you're more on application building side, their associate AI engineer for developers track covers the more practical stack so OpenAI API,
59
prompt engineering, embeddings, PyCon and LangChain.
60
They also offer an AI engineer for data scientist certification if you want a credential you can put on LinkedIn.
61
If you want to level up your skills in AI with these tracks check the links in the description.
62
Okay the second way data scientists can leverage AI is creating synthetic data sets for prototyping.
63
Getting access to real and high quality data is the problem every data scientist has faced.
64
Probably the data is locked behind privacy constraints, doesn't exist yet, or you just want to test an idea quickly.
65
So instead of waiting, many data scientists use AI to create synthetic data sets, whether they are structured or unstructured data.
66
LLMs have become so good that they can actually generate artificial data that mimics real world patterns. In fact,
67
synthetic data is widely used in privacy sensitive industries like healthcare
68
finance where real data can't always be shared freely for example
69
i used claw to generate this whole e-commerce data set of
70
5 000 rows with customers products order items you just describe what you need the structure the number of rows
71
and it gives you something you can actually work with in
72
minutes similarly you can also generate unstructured text data for testing
73
a model training pipeline now here's another interesting way you can use ai
74
that a lot of data scientists are sleeping on
75
that is using foundation models foundation models those are trained on massive
76
and varied data sets they are different from traditional machine learning models
77
that are custom built for specific tasks today foundation models are no longer just for chatbots
78
and image generation they're starting to show up in core data science work like time series forecasting tabular prediction
79
and even recommendation systems let's say you're building a time series forecasting model you usually collect your data
80
and clean it and engineer your features and then pick a model
81
and train it and fine-tune it and evaluate it now the problem is
82
if your company wants forecasts across 200 product lines you would normally have to repeat
83
that process over and over but nowadays you can use foundation models like chronos by amazon
84
and time gpt
85
so here's an example where i tried one of those models you just feed your data
86
and get a forecast no feature engineering no model selection no
87
training loop this approach is really changing the whole modeling workflow
88
netflix recently replaced a whole stack of separate recommendation models with one foundation model trained on billions of user interactions
89
and now every team just fine-tunes on top of it the
90
data scientists pulling ahead right now are the ones who know
91
when to build from scratch and when to build on top of something
92
that already works next let's talk about how a data scientist
93
can automate workflows with mcp model context protocol think about your
94
typical day you're jumping between five different apps you're querying databases
95
pushing code to github sharing updates on slack pulling some data
96
that your colleague shared on google drive that constant back
97
and forth burns your energy so the idea of using mcp is
98
that it lets you connect all of those tools to an ai assistant like cloud
99
so instead of switching between apps you're doing everything from one
100
place for example i connected cloud desktop with my postgres sql database through mcp
101
and now instead of opening a sql client writing a query
102
running it copying results i just ask cloud for example show me the top 10 customers my revenue this month
103
and i get my answer right there.
104
Now, quick heads up, not every employer allows tools like Claude Desktop or Claude Code due to security policies.
105
But even if you can't use it at work yet, it's useful, I think, to exploit personally because I think sooner or later,
106
this kind of integrated tools will become the standard because it removes a lot of friction in the day-to-day work.
107
Now, beyond just using AI tools to speed up your work
108
the data scientists who are really standing out right now are
109
the ones who can actually build things with AI systems
110
that solve real business problems for their teams
111
and companies from my conversations with all of my data scientist friends their jobs are less
112
and less about prototyping a model in a notebook it's more
113
and more about delivering something people can actually use for example
114
imagine your finance team processes hundreds of invoices every week manually you could build an AI system
115
that read those invoices extracts the key information
116
and saves it straight to your database
117
or say your company wants to understand patterns across thousands of customer interactions
118
or research documents you could build a knowledge graph
119
that maps out how everything connects that way you uncover insights
120
that are impossible to find with traditional analysis methods i also
121
have a tutorial on this i'll link here somewhere in the screen these are the kinds of projects
122
that make you invaluable at a company these days also there's a real gap
123
that most data scientists need to close right now myself included
124
a lot of data scientists can build a great prototype
125
but they can't get into production they can't deploy it monitor it
126
or make it reliable enough for other people to depend on
127
and to close
128
that gap you need to start learning things like how to containerize your application with docker
129
so it runs anywhere how to deploy and serve your models on cloud platforms
130
and how to monitor
131
and maintain ai systems once they're alive these skills are more engineering skills ai engineering specifically
132
but they're becoming essential for data scientists as well
133
if you want to help your team build ai systems now we've talked about a lot of different tools
134
and hard skills but i think eventually data professionals are increasingly about judgment
135
and influence understanding how business works how your company makes money domain knowledge stakeholder management, trust building, communication, and data storytelling,
136
these soft skills in the long term are becoming the core of what you do.
137
Right now, I think one thing that's going to pay off dividends for you is to stay open to trying out new tools, learning them, using them, questioning them, adapting them for your own needs.
138
I run a free newsletter where I share my latest insights and experiments in data science and AI.
139
So if you're interested, check it out in the description below.
140
Thank you for watching.
141
Bye-bye.

Vocabulário e dicas de fala para esta lição

Esta aula de conversação de nível C1 usa o vídeo “How AI is Changing Data Scientist's Workflow”. As palavras que mais se repetem: tool, scientist, model, build. Este vídeo tem 141 frases e 1957 palavras para praticar shadowing. A fala dura 10:46. O falante fala rápido, cerca de 182 palavras por minuto, então espere sons ligados e reduzidos. Apenas 81% das palavras estão entre as 3.000 mais comuns do inglês, por isso o vocabulário é exigente.

Vocabulário principal deste vídeo

As 15 palavras mais avançadas do vídeo, com pronúncia e significado:

PalavraPronúnciaSignificado
query substantivo/ˈkwɪɹ.i/pergunta
leverage substantivo/ˈlɛv.(ə.)ɹɪd͡ʒ/alavancagem
forecast substantivo/ˈfɔːkɑːst/previsão
deploy verbo/dɪˈplɔɪ/posicionar
prototype substantivo/ˈpɹəʊtətaɪp/protótipo
prompt verbo/pɹɑmpt/incitar, impelir
synthetic adjetivo/sɪnˈθɛtɪk/sintético
automate verbo/ˈɔ.təˌmeɪt/automatizar
invoice substantivo/ˈɪnˌvɔɪs/fatura, guia de remessa
rethink verbo/ɹiːˈθɪŋk/repensar
visualization substantivo/ˌvɪʒ.ʊ.ə.laɪˈzeɪ.ʃən/visualização
excel verbo/ɪkˈsɛl/superar, ultrapassar
desktop substantivo/ˈdɛsktɒp/desktop, computador de mesa
adapt verbo/əˈdæpt/adaptar
transform verbo/tɹænsˈfɔɹm/transformar

Phrasal verbs que você vai ouvir

PalavraSignificado
come up with verbopensar (em), alcançar
get away with verbopassar batido
pay off verbocompensar
try out verboprovar, experimentar

Gramática neste vídeo

As estruturas que o falante mais usa, com as palavras exatas do vídeo:

EstruturaNo vídeo
Voz passiva be + particípio passado — o foco está no que acontece, não em quem fazwould be sent · being transformed · is locked
Present perfect have/has + particípio passado — uma ação passada que ainda importa agorayou've worked · I've come · has faced
Orações relativas who / which + oração — informação extra sobre uma pessoa ou coisahidden which means · ones who know · ones who can

Pronúncia para ficar de olho

O falante usa 30 contrações e formas reduzidas, como you're, can't, they're. Diga-as na forma curta, do jeito que você ouve.

  • Os sons de “th”: synthetic /sɪnˈθɛtɪk/, rethink /ɹiːˈθɪŋk/, hypothesis /haɪˈpɒθɪsɪs/
  • Os sons de “sh” e “zh”: visualization /ˌvɪʒ.ʊ.ə.laɪˈzeɪ.ʃən/, recommendation /ˌɹɛkəmɛnˈdeɪʃən/, credential /kɹɪˈdɛnʃəl/, visualize /ˈvɪʒuəˌlaɪz/, friction /ˈfɹɪkʃən/
  • Palavras longas — acerte a sílaba tônica: visualization /ˌvɪʒ.ʊ.ə.laɪˈzeɪ.ʃən/, recommendation /ˌɹɛkəmɛnˈdeɪʃən/, exploratory /ɛkˈsplɒɹ.ə.tə.ɹi/, generative /ˈd͡ʒɛnəɹətɪv/, metadata /ˈmɛt.əˌdeɪ.tə/

Sons difíceis para falantes de português:

  • Consoante final — sem acrescentar um “i” depois: forecast /ˈfɔːkɑːst/, prototype /ˈpɹəʊtətaɪp/, prompt /pɹɑmpt/, synthetic /sɪnˈθɛtɪk/, automate /ˈɔ.təˌmeɪt/
  • /l/ final — a língua toca o céu da boca, não vira “u”: excel /ɪkˈsɛl/, credential /kɹɪˈdɛnʃəl/, tutorial /ˌtjuːˈtɔːɹɪəl/
  • /h/ — um sopro suave, diferente do “r”: stakeholder /ˈsteɪkˌhəʊl.də/, hypothesis /haɪˈpɒθɪsɪs/

Como praticar com este vídeo

  1. Ouça o vídeo inteiro uma vez sem falar e anote as palavras que você não conhece.
  2. Comece na velocidade 0,75×, faça shadowing frase por frase e volte à velocidade normal quando ficar fácil.
  3. Grave a sua voz e compare com o original, prestando atenção a palavras como query, leverage, forecast.

O que é a Técnica de Shadowing?

Shadowing é uma técnica de aprendizado de idiomas com base científica, originalmente desenvolvida para o treinamento de intérpretes profissionais. O método é simples, mas poderoso: você ouve áudio em inglês nativo e repete imediatamente em voz alta — como uma sombra seguindo o falante com 1-2 segundos de atraso. Pesquisas mostram melhora significativa na precisão da pronúncia, entonação, ritmo, sons conectados, compreensão auditiva e fluência na fala.

Técnica de shadowing: leia o guia completo passo a passo →