シャドーイング練習: How AI is Changing Data Scientist's Workflow - 動画で英語スピーキングを学ぶ

レッスンを作成中...
1
Data scientist is the top job title referencing generative AI.
2
So that means that companies don't just need data talent.
3
They need data talent that knows AI.
4
So what does know AI actually mean?
5
Well, I found that it's actually twofold.
6
Firstly, you know how to leverage AI to speed up your existing workflow, not just in a way of prompting ChatGBT to write Python code kind of way,
7
but it's more about rethinking the way you get things done from generating or collecting data, exploratory data analysis, modeling, reporting, and communication.
8
Secondly, you know how to build AI systems.
9
For example, creating and deploying tools and agents to automate part of or an entire process or workflow for your team.
10
Now, let's first talk about different ways to leverage AI to speed up a data science workflow.
11
I'll talk about a few specific tools in this video, but I want to be clear that tools isn't the point because tools can change.
12
What matters is knowing what's now possible so you can think about how to adapt and rethink the way you're working.
13
So the first major way data scientists can leverage AI is in EDA and data pre-processing.
14
If you've worked in a real data science project before, you probably already experienced this.
15
Too much time spent preparing data, not enough time analyzing it.
16
You got the job thinking you spent all day finding data patterns, coming up with hypothesis and interesting analysis.
17
In reality, you spent 60 to 70% of your time cleaning data, writing repetitive code to check data types,
18
handling missing values, or creating standard visualizations.
19
Here are two open source tools I've come across that can make your work a little bit easier and more fun.
20
The first one is Pandas AI.
21
It's a Python library that basically allows you to ask questions to your data frame in natural language.
22
For example, what's the total sales per coffee type?
23
Depending on your question, it can return different kinds of responses like a string, a data frame, a chart, or a number.
24
It's totally open source and you can hook it up with any LLM you like.
25
And I'm connecting Pandas AI with an LLM from OpenAI here.
26
Now, if you're wondering if your whole data set would be sent to an LLM or not, the answer is no. When you ask Pandas AI a question, it actually sends just enough context.
27
So your question along with the column names, a few sample rows, and some metadata.
28
The LLM reads that context
29
and writes a small Python script to answer your question pandas AI then checks
30
that the script is safe to run
31
and then it runs the script locally on your actual data using pandas
32
so if you want to get away with writing pandas
33
and matplotlib code yourself you can use this library to do quick analysis data cleaning
34
and data manipulation but for more complex tasks like joining data frames
35
or creating correlation matrix you should still use pandas the second tool
36
that I think is pretty cool is data formulator this one comes out of microsoft research
37
and the easiest way to think about it is the combination of excel formulas power query
38
and modern ai it also allows you to interact with your data in natural language
39
but the nice part is the tool shows you exactly how the data is being transformed
40
and visualized step by step nothing hidden which means the workflow stays transparent and reusable.
41
Give it a broad prompt like show me interesting trends in this data set and we'll start transforming the data,
42
generating visualizations and suggesting directions you can explore further.
43
This tool is also open source and you can install it locally.
44
So in your workflow if you happen to spend a lot of time in the Microsoft ecosystem like Excel, Power Query, Power BI, it's definitely worth experimenting with this tool.
45
Now with these kind of tools you still need to understand your data
46
but they can drastically speed up the exploration phase where you're
47
just trying to understand what you're working with before you get into building the actual pipeline
48
and the real modeling work now as a top one percent data scientist you don't just limit yourself to tools
49
that are already available you also know how to build ai tools
50
and systems to support your team
51
so i want to quickly shout out to datacamp who sponsors this part of the video I used Datacam back
52
when I was learning Python for work.
53
What I really like is that it's hands-on.
54
You actually write code, solve exercises, and build projects right in your browser.
55
They're offering some solid tracks for data scientists to level up your skills.
56
The first one I'd recommend is the Associate AI Engineer for Data Scientist track that covers scikit-learn, PyTorch, Hugging Face, building LLM apps with Lanchain,
57
and the MLObs basics to help you take models into production.
58
If you're more on application building side, their associate AI engineer for developers track covers the more practical stack so OpenAI API,
59
prompt engineering, embeddings, PyCon and LangChain.
60
They also offer an AI engineer for data scientist certification if you want a credential you can put on LinkedIn.
61
If you want to level up your skills in AI with these tracks check the links in the description.
62
Okay the second way data scientists can leverage AI is creating synthetic data sets for prototyping.
63
Getting access to real and high quality data is the problem every data scientist has faced.
64
Probably the data is locked behind privacy constraints, doesn't exist yet, or you just want to test an idea quickly.
65
So instead of waiting, many data scientists use AI to create synthetic data sets, whether they are structured or unstructured data.
66
LLMs have become so good that they can actually generate artificial data that mimics real world patterns. In fact,
67
synthetic data is widely used in privacy sensitive industries like healthcare
68
finance where real data can't always be shared freely for example
69
i used claw to generate this whole e-commerce data set of
70
5 000 rows with customers products order items you just describe what you need the structure the number of rows
71
and it gives you something you can actually work with in
72
minutes similarly you can also generate unstructured text data for testing
73
a model training pipeline now here's another interesting way you can use ai
74
that a lot of data scientists are sleeping on
75
that is using foundation models foundation models those are trained on massive
76
and varied data sets they are different from traditional machine learning models
77
that are custom built for specific tasks today foundation models are no longer just for chatbots
78
and image generation they're starting to show up in core data science work like time series forecasting tabular prediction
79
and even recommendation systems let's say you're building a time series forecasting model you usually collect your data
80
and clean it and engineer your features and then pick a model
81
and train it and fine-tune it and evaluate it now the problem is
82
if your company wants forecasts across 200 product lines you would normally have to repeat
83
that process over and over but nowadays you can use foundation models like chronos by amazon
84
and time gpt
85
so here's an example where i tried one of those models you just feed your data
86
and get a forecast no feature engineering no model selection no
87
training loop this approach is really changing the whole modeling workflow
88
netflix recently replaced a whole stack of separate recommendation models with one foundation model trained on billions of user interactions
89
and now every team just fine-tunes on top of it the
90
data scientists pulling ahead right now are the ones who know
91
when to build from scratch and when to build on top of something
92
that already works next let's talk about how a data scientist
93
can automate workflows with mcp model context protocol think about your
94
typical day you're jumping between five different apps you're querying databases
95
pushing code to github sharing updates on slack pulling some data
96
that your colleague shared on google drive that constant back
97
and forth burns your energy so the idea of using mcp is
98
that it lets you connect all of those tools to an ai assistant like cloud
99
so instead of switching between apps you're doing everything from one
100
place for example i connected cloud desktop with my postgres sql database through mcp
101
and now instead of opening a sql client writing a query
102
running it copying results i just ask cloud for example show me the top 10 customers my revenue this month
103
and i get my answer right there.
104
Now, quick heads up, not every employer allows tools like Claude Desktop or Claude Code due to security policies.
105
But even if you can't use it at work yet, it's useful, I think, to exploit personally because I think sooner or later,
106
this kind of integrated tools will become the standard because it removes a lot of friction in the day-to-day work.
107
Now, beyond just using AI tools to speed up your work
108
the data scientists who are really standing out right now are
109
the ones who can actually build things with AI systems
110
that solve real business problems for their teams
111
and companies from my conversations with all of my data scientist friends their jobs are less
112
and less about prototyping a model in a notebook it's more
113
and more about delivering something people can actually use for example
114
imagine your finance team processes hundreds of invoices every week manually you could build an AI system
115
that read those invoices extracts the key information
116
and saves it straight to your database
117
or say your company wants to understand patterns across thousands of customer interactions
118
or research documents you could build a knowledge graph
119
that maps out how everything connects that way you uncover insights
120
that are impossible to find with traditional analysis methods i also
121
have a tutorial on this i'll link here somewhere in the screen these are the kinds of projects
122
that make you invaluable at a company these days also there's a real gap
123
that most data scientists need to close right now myself included
124
a lot of data scientists can build a great prototype
125
but they can't get into production they can't deploy it monitor it
126
or make it reliable enough for other people to depend on
127
and to close
128
that gap you need to start learning things like how to containerize your application with docker
129
so it runs anywhere how to deploy and serve your models on cloud platforms
130
and how to monitor
131
and maintain ai systems once they're alive these skills are more engineering skills ai engineering specifically
132
but they're becoming essential for data scientists as well
133
if you want to help your team build ai systems now we've talked about a lot of different tools
134
and hard skills but i think eventually data professionals are increasingly about judgment
135
and influence understanding how business works how your company makes money domain knowledge stakeholder management, trust building, communication, and data storytelling,
136
these soft skills in the long term are becoming the core of what you do.
137
Right now, I think one thing that's going to pay off dividends for you is to stay open to trying out new tools, learning them, using them, questioning them, adapting them for your own needs.
138
I run a free newsletter where I share my latest insights and experiments in data science and AI.
139
So if you're interested, check it out in the description below.
140
Thank you for watching.
141
Bye-bye.

このレッスンの語彙とスピーキングのポイント

このC1レベルのスピーキングレッスンは、動画「How AI is Changing Data Scientist's Workflow」を教材にしています。 繰り返し出てくる語は次のとおりです:tool, scientist, model, build。 この動画には、シャドーイング用の文が141文、単語が1957語あります。 音声の長さは10:46です。 話す速さは速く、1分あたり約182語です。音のつながりや弱く発音される音が多くなります。 英語の頻出3,000語に含まれる単語は81%だけなので、語彙は難しめです。

この動画の重要語彙

動画の中で特に難しい単語15語を、発音と意味つきで紹介します。

単語発音意味
workflow 名詞/ˈwɝkfloʊ/ワークフロー
query 名詞/ˈkwɪɹ.i/質問
leverage 名詞/ˈlɛv.(ə.)ɹɪd͡ʒ/レバレッジ, レバ
forecast 名詞/ˈfɔːkɑːst/予想, 予報
deploy 動詞/dɪˈplɔɪ/配置する
prototype 名詞/ˈpɹəʊtətaɪp/プロトタイプ, 原型
automate 動詞/ˈɔ.təˌmeɪt/自動化する
invoice 名詞/ˈɪnˌvɔɪs/送り状, インボイス
rethink 動詞/ɹiːˈθɪŋk/再考する
visualization 名詞/ˌvɪʒ.ʊ.ə.laɪˈzeɪ.ʃən/映像化, 可視化
excel 動詞/ɪkˈsɛl/超える, 越える
desktop 名詞/ˈdɛsktɒp/机の上
adapt 動詞/əˈdæpt/適応する
transform 動詞/tɹænsˈfɔɹm/変形する, 変換する
pipeline 名詞/ˈpaɪpˌlaɪn/パイプライン

動画に出てくる句動詞

単語発音意味
speed up 動詞/spiːdˈʌp/スピードを出す
pay off 動詞報われる
try out 動詞試す, 試してみる

この動画の文法

話し手がよく使っている文型を、動画の実際の表現とともに紹介します。

文型動画での表現
受動態 be + 過去分詞 — 誰がするかより、何が起きるかに焦点を当てるwould be sent · being transformed · is locked
現在完了形 have/has + 過去分詞 — 過去の出来事が今も関係しているyou've worked · I've come · has faced
関係詞節 who / which + 節 — 人や物について情報を加えるhidden which means · ones who know · ones who can

注意したい発音

話し手はyou're, can't, they'reなど、短縮形や弱形を30回使っています。聞こえたとおりの短い形で発音しましょう。

  • 「th」の音: synthetic /sɪnˈθɛtɪk/, rethink /ɹiːˈθɪŋk/, hypothesis /haɪˈpɒθɪsɪs/
  • 「sh」と「zh」の音: visualization /ˌvɪʒ.ʊ.ə.laɪˈzeɪ.ʃən/, recommendation /ˌɹɛkəmɛnˈdeɪʃən/, credential /kɹɪˈdɛnʃəl/, visualize /ˈvɪʒuəˌlaɪz/, friction /ˈfɹɪkʃən/
  • 長い単語(アクセントの位置に注意): visualization /ˌvɪʒ.ʊ.ə.laɪˈzeɪ.ʃən/, recommendation /ˌɹɛkəmɛnˈdeɪʃən/, exploratory /ɛkˈsplɒɹ.ə.tə.ɹi/, generative /ˈd͡ʒɛnəɹətɪv/, metadata /ˈmɛt.əˌdeɪ.tə/

日本語話者が苦手な音:

  • /r/ と /l/ の区別 — /l/ は舌先を歯茎につけ、/r/ はどこにもつけない: leverage /ˈlɛv.(ə.)ɹɪd͡ʒ/, credential /kɹɪˈdɛnʃəl/, exploratory /ɛkˈsplɒɹ.ə.tə.ɹi/, firstly /ˈfɜɹstli/, tutorial /ˌtjuːˈtɔːɹɪəl/
  • /v/ — /b/ にならないように、上の歯を下唇に当てる: leverage /ˈlɛv.(ə.)ɹɪd͡ʒ/, invoice /ˈɪnˌvɔɪs/, visualization /ˌvɪʒ.ʊ.ə.laɪˈzeɪ.ʃən/, generative /ˈd͡ʒɛnəɹətɪv/, uncover /ʌnˈkʌvɚ/
  • /f/ — 「フ」ではなく、上の歯と下唇で出す: workflow /ˈwɝkfloʊ/, forecast /ˈfɔːkɑːst/, transform /tɹænsˈfɔɹm/, twofold /ˈtuːfəʊld/, friction /ˈfɹɪkʃən/

この動画での練習方法

  1. まず声を出さずに動画を最後まで聞き、知らない単語をメモします。
  2. まず0.75倍速で一文ずつシャドーイングし、慣れてきたら通常の速度に戻します。
  3. 自分の声を録音して元の音声と比べます。workflow, query, leverageなどの単語に特に注意しましょう。

シャドーイングとは?英語上達に効果的な理由

シャドーイング(Shadowing)は、もともとプロの通訳者養成プログラムで開発された言語学習法で、多言語習得者として知られるDr. Alexander Arguelles によって広く普及されました。方法はシンプルですが非常に効果的:ネイティブスピーカーの英語を聞きながら、1〜2秒の遅延で声に出してすぐに繰り返す——まるで「影(shadow)」のように話者を追いかけます。文法ドリルや受動的なリスニングと異なり、シャドーイングは脳と口の筋肉が同時にリアルタイムで英語を処理・再現することを強制します。研究により、発音精度、抑揚、リズム、連音、リスニング力、そして会話の流暢さが大幅に向上することが確認されています。IELTSスピーキング対策や自然な英語コミュニケーションを目指す方に特におすすめです。

シャドーイングのやり方: ステップ別の完全ガイドを読む →