跟读练习: Learn Databricks for FREE (Step-by-Step Guide) - 通过视频学习英语口语
正在创建课程...
1
hey friends so now i'm gonna show you exactly how to create a free databricks account
2
and after that i'm gonna guide you quickly into the interface of databricks
3
because if it is your first time it can feel a bit overwhelming
4
so now let's get started i guess it's starting now before
5
we start creating any account i would like you to understand something important
6
so that you understand what you are actually creating
7
so now the thing is databricks is not a tool
8
that you're gonna go and download
9
and install locally at your pc instead it is a platform
10
that you're gonna simply open in your browser and start working with it
11
so that means it doesn't matter whether you are a windows user
12
or a mac user you can work with databricks
13
and of course you don't need any super computer to use databricks
14
because we are not running anything locally at your pc now
15
of course you're gonna ask okay how we process actually the data we need things like cpu memory storage
16
so where we're gonna get those
17
if we are not using your local pc now the thing
18
is as well in data bricks itself it doesn't go
19
and build like data centers
20
and servers in order to process the data the idea is very simple it is a platform
21
that is built on top of other cloud platforms like azure aws
22
and google clouds so
23
that means it is not like stand alone platform it will be using other cloud service providers
24
so now because of all of
25
that Databricks is not like fixed price we pay my friends on the usage
26
so the more compute and storage
27
that we use from the cloud the more it gonna cost us
28
and now of course I might get you scared I'm gonna think okay
29
if I need to learn Databricks I have to go
30
and pay for all those compute and storage
31
and the cloud I have to configure the whole setup well
32
I have for you good news only recently Databricks now offers
33
a free edition where everything is prepared for you to learn databricks
34
so that means you're gonna get an access to databricks the same platform
35
that we use in real projects with 90 of all features
36
and this is more than enough actually to learn databricks
37
and currently behind the scenes it is connected to aws cloud platform
38
so everything is prepared with only few clicks you're gonna get the whole setup
39
and yes my friends for free databricks is covering all the compute
40
and storage that we're going to use in the aws
41
but of course there will be some limitations about creating high -end fast clusters
42
because they are very expensive and they don't allow you to do
43
that so this can be for personal use
44
and as well for learning purposes
45
but on the other hand we still have the professional setup
46
if you are working in a company then it really depends
47
on the cloud platform your company is using like here in europe mostly we are using the azure platform
48
so you can connect databricks to any of those cloud providers
49
and run everything there but of course you have to pay for it
50
so let's go and create that
51
so let's go to the home page of databricks .com
52
and then we have here on the right side try databricks let's go
53
and click on it okay so the first thing
54
that we have to do
55
if you don't have already an account with databricks you have to go
56
and sign up so let's go and do that and now after
57
that we clearly have two options either we can use it for personal use
58
or for work for personal use it's going to be for free
59
so if you are just learning databricks i'm gonna say go with this edition but
60
if you're gonna use it for work here actually you have two options either you can use the express edition
61
so everything is already prepared with the aws you don't have to go
62
and set up your own cloud so everything is there
63
and by the way they are currently giving free for 14 days
64
and some credits or you say you know what i have my own cloud
65
so i have already an azure account or aws
66
and i would like to go and connect it to databricks
67
and with this actually you can unlock the full professional version of databricks the enterprise edition
68
that we actually use in companies actually here we have three options the free one the express edition
69
and the full enterprise edition pick the one that suits you and go
70
and create the accounts all right after creating the accounts finally we're gonna land inside a databricks interface
71
and as you can see we are just using the web
72
browser without any installations now i know the interface might be overwhelming
73
because we have a lot of tools inside databricks we're gonna do it step by step
74
so don't worry about it now the first thing
75
that we might notice is to go to the right side over here
76
so we can see the workspaces now since you are using
77
a new free account you might see only one workspace
78
and that's okay
79
but once you start working in re -projects you might get an access to multiple databricks workspaces
80
and in order to switch between them you will be using here the workspace now the next thing
81
that we can check is actually the settings so
82
if you go to your icon
83
and then to the settings you're gonna get here in middle
84
a new interface where you're gonna get the usual stuff about
85
configuring the profile the appearance for example you can configure the date
86
and time formats you can go and manage the access to your workspace
87
so you can go and make groups and add users to it
88
and many other options about the security the compute for now
89
we will not change anything we're gonna leave it as it is the only thing
90
that i might say it is interesting for this phase is to go to the linked accounts
91
And here what we can do, we can go and connect one of our GitHub repositories the databricks
92
so that anything that you do inside databricks will be stored
93
and saved inside your github repo and as well you have the version control
94
so if you have one already you can go
95
and connect it to your databricks account
96
so as you can see we don't have any crazy things
97
inside the settings it is the classical usual things now we're
98
gonna go to something really interesting in the interface by the way this is completely new
99
if you go over here to this icon it is the genie code
100
so let's go and open it so basically the genie code is the ai assistant
101
that is inside databricks that's gonna help you work with data and code
102
so here we have a basic interface where you're gonna start writing your prompts like for example what is databricks
103
and what is very interesting you can decide between those two
104
so either an agent or a shot a shot it's just simple like question
105
and answer like we do in chat gpt
106
but an agent it is way more powerful it can go
107
and handle for you a task so
108
that means you can give it a goal and it's gonna go
109
and complete it step by step
110
and this could be anything like creating pipelines creating for you a script a notebook exploring the data building visualizations
111
and charts so this is very powerful things
112
and as well very new in data breaks
113
so you can write your prompts by asking questions
114
or giving it a task as you can see we got an answer for our question
115
and later we can start giving it command in order to do things for us in databricks
116
so this is where you can interact with the genie code
117
now let's go to the landing page of databricks now my
118
friends the most important part in this interface is actually the left navigation bar
119
because you're gonna see a lot of tools and things in databricks
120
and since we have many things databricks is gonna go
121
and group them into like sections
122
so the first set of tools they are the main section then after
123
that we have a lot of tools
124
that are dedicated only for the sql a whole section for sql
125
and another one for data engineering and the last one for ai
126
and machine learning so with that you're now you are getting the feeling
127
that i can use this platform actually for many things for
128
data engineering data analytics data science machine learning even ai engineering
129
now before we start talking about all those tools usually
130
if you are learning any new data platforms you have to
131
think about three things where is the data where is the compute
132
and what tools do i have in order to use my data
133
and the compute to do my projects
134
so let's start with the most important question where is the data in databricks
135
and you can find that in the main section in the catalog
136
so this is the most important component in the interface the catalog of databricks now
137
if you open it you're gonna have like two sections on
138
the left side it's gonna look like a traditional database
139
so you have like hierarchy and you can start browsing this catalog by just expanding things
140
so at the start we have the my organization
141
and then we have different workspaces and
142
if you start like expanding you're gonna find inside them schemas
143
and inside the schemas guess what we have tables
144
so we are just browsing like the structure the hierarchy of the catalog
145
and by the way the data that we see underneath samples it is something
146
that prepare from databricks but course you can go
147
and bring your data to the catalog now
148
if you want to see more details about each of those objects you can go
149
and click for example on the schema
150
and then on the right side you're gonna see more details about it like what are the tables inside it
151
and then here more details about the storage the permissions the policies
152
so you're gonna have the details on the right side same things
153
if you go for example to a table you're gonna see
154
more details like the history the lineage the insights quality
155
so we are seeing here all the details same thing
156
if i go for example to the workspace we're going to
157
see as well what is inside the workspaces with more details
158
and permissions so that's it this is the catalog where the data lives in
159
and of course we have to deep dive into this later
160
because this is the fundamentals that we have to learn how to organize
161
and structure our data inside Databricks
162
so this is about the data now the second most important
163
thing to check is where is the compute where are the resources the powers
164
that we're use in order to process our data
165
and here in databricks we have like let's say two tools
166
for it in the interface in the main section you can see here we have compute
167
and as well in the sql we have the sql warehouses now let's go inside the warehouses
168
so now whether you are using the free edition
169
or the express like me you're gonna find that
170
that are already created for us like a warehouse
171
so this is our compute you can see the status is offline
172
so it stops we see the name we see the size it is very small
173
and the type the serverless and at the end you can go
174
and start the warehouse now my friends the compute is usually the most expensive things
175
that we could use in the platform that's why databricks has a lock in you cannot go
176
and change the size of your warehouse or you cannot go
177
and create other computes
178
because then the costs of the free edition is going to be very high
179
but now for me since i'm using the express edition i can go
180
and create a new warehouse or change the size of it.
181
And of course, don't worry about it.
182
It is more than enough to practice in Databricks.
183
You are not processing terabytes of data you're gonna use some medium data sets in order to explore the platform
184
and learn the features inside it so that's why it is for you frozen
185
if you are using free edition but for me I can go
186
and create here like for example new warehouse now look at
187
this this is exactly why those data platforms are speeding up a lot of projects
188
so we are just deciding on very simple things like for example the name of the warehouse
189
and as well the cluster size that's it everything else is
190
going to be done behind the scenes like configuring the servers
191
the infrastructure the sizes of the workers the scaling their resource management
192
and monitoring everything going to be done behind the scenes
193
and we don't have to configure anything only the size of our compute
194
and for the free edition you cannot do that of course
195
and with this you can have some understanding how easy is things in such platforms
196
so i will not go
197
and create anything i'm going to stick with the default one
198
that's it as you can see the status is already online
199
so the server is ready to be used
200
and i can use it everywhere i can use it in
201
order to execute my notebooks my pipelines my sql queries the dashboards the ai
202
so all the tools
203
that i have over here on the list side it's going
204
to be powered by this sql warehouse of course in the enterprise edition you can go
205
and create multiple different types of clusters you can see here we are using the serverless
206
but there are other types for different purposes
207
but currently this is one of the best types in order to use both sql
208
and python by spark on all other tools
209
so as you can see things are simple and easy it's not that hard
210
but don't forget behind the scenes we still have this big data engine the spark with all data distribution
211
and parallel processing and stuff that is running as you start the warehouse okay
212
so with this you know where is the data where is the compute
213
and of course the question is
214
which tools i'm gonna use in order to do my projects
215
my job well of course this depends on what you plan to do
216
if you want to do data analytics data engineering science ai
217
engineering you're gonna have different tools to use now let's say i am data analyst
218
and i would like to analyze the data inside the data bricks then of course beside the catalog
219
and the warehouse the most important section for me gonna be the sql section
220
and everything else is not that important or let's say you can go
221
and ignore so what do we have here the first one is the sql editor
222
so here again you're gonna see our catalog in order to navigate through your data
223
and in the middle you can go and write sql query
224
so it looks like any traditional sql databases you will just go
225
and write here an sql query from table and then you're gonna go
226
and execute it and see the result in the output so
227
that you are like exploring and analyzing your data
228
that you have inside the catalog
229
and i can say this can be your main tool in
230
order to interact with the data to do data exploration
231
and data analyzes so here it's like a place where you're gonna go
232
and store all your sql queries as you can see i just created one
233
and we have now a query for
234
that it just gives you an overview of all the queries
235
that is used inside databricks for now it is empty
236
because we didn't do yet anything so i think but not
237
that important now we go to something really nice
238
that dashboards my friends as a data analyst you can actually go and build dashboards like i'm not
239
very strong in order to do data explorations
240
and let's say to create the first drafts for the analyzers
241
and here we have some examples I'm gonna go
242
and just click on one of those samples it might take some time in order to query the data
243
and present the result in the dashboard look at this we
244
have a dashboard where you have a title for it you can go
245
and filter the data you have big number some details dot plot
246
and analyzes over the time but again not pixel perfect like power bi
247
or tableau but still you can build the first visuals
248
and dashboards directly inside the same platform where you are working
249
with the data without having you to jump to a new tool
250
so this is just a dummy dashboard
251
that we get from databricks now let's go to more interesting
252
stuff for you we have something called genie now of course at the start everything gonna be empty
253
but what you actually do you can go
254
and create an ai bi genie by just selecting any of your data
255
and now look at this this already looks like shad gpt
256
so this is another ai assistance where you go
257
and define a subset of your data that you want to analyze
258
but this time without writing any code you will just go
259
and use the natural language in order to ask any question about your data
260
so what is the total number of rows inside my table
261
so look at this we have already answered the total number of rows in your table is 750
262
and if you go and expand this you can see it as a table
263
and as well you can see the code
264
that is generated behind the scenes and it is sql
265
so the genie gonna generate from our prompts and sql query
266
that gonna use our warehouse in order to interact with the data
267
and you can see the code is totally correct crazy stuff right
268
so now let's keep going the other things is not
269
that important so the alerts over here now we don't see anything yet
270
but what you can do we can go and create an sql query
271
and based on the thresholds it can go and give us notifications
272
and alerts so again based on sql query
273
and the last one you can see the query history
274
so here you can see a full log and protocol of all queries
275
that got executed in your databricks it's like we are monitoring let's say what is going on
276
so can see the query when it started the duration on
277
which data sets what is the resource that is used to compute
278
and who did that so again as i said some kind of monitoring it is not thing
279
and the last one we have the sql warehouse
280
so my friends for you as a data analyst we have three main tools in order to explore
281
and analyze our data so either we're going to go
282
and use the sql editor in order to write simple queries
283
or we can go and build a dashboard in order to analyze our data
284
or a third cool option is that you're gonna go
285
and use the ai bi genie where you're gonna go
286
and connect your data to the ai and start having simple chats
287
so that's it those are the tools
288
that you have inside data bricks for you as a data analyst in order to work with your data okay
289
so now let's go and switch and let's say
290
that you are a data engineer
291
and you would like to use data bricks in order to build your projects
292
so which tools you have to focus on again the catalog is the most important thing
293
because your main job is actually to build the catalog for the others
294
and the compute we talked about it
295
so either you can use the compute here
296
or the sql warehouse now the third component
297
that is really important for you to understand how to deal with it is the workspace
298
so now what is our main job as data engineers we go
299
and write codes so we need a place for that
300
and this is gonna happen inside the workspace so here we can go
301
and create notebooks write codes and organize the whole projects
302
so we can go and create folders and put inside it our python
303
or notebook codes and by the way all other interactions
304
that we are doing like for example writing an SQL query
305
building a dashboard creating a new genie you're gonna find all those stuff inside the workspace again
306
so that you can go and organize it
307
and reuse it inside your projects
308
so look at this we have the genie space inside the workspace
309
that we just created same thing goes for the dashboards and
310
if i do any queries it's going to be landing inside
311
this workspace now of course at this level at the home anything
312
that i'm creating in notebook code it's going to be stored inside my account in the databricks
313
but of course usually we don't do that we usually organize
314
and store our code inside github repository and for
315
that we have a section over here called workspace and then repos
316
so if you go inside it now you can go
317
and connect your github repository inside databricks
318
and inside it you will be creating your code and project
319
and with that you have version control on your code
320
and your work will not stay only inside your account in data breaks
321
and there are many other like not important things like for example here you can see all the notebooks
322
that are shared with you and so on
323
so again workspace is the place where you organize your projects
324
and to be honest as a data engineer i spend most
325
of my time just working in this component inside the workspace now the second thing
326
that we usually do is we go and create jobs
327
and orchestrate our work so in this place the automation of my work can happen
328
so i can go and create a pipeline
329
and then schedule it to run automatically instead of me doing everything manually
330
and here in the interface like we have different types on
331
how to do this the most like important one i'm gonna say the jobs this is the way
332
that i use in order to orchestrate the notebooks and codes
333
that i'm creating so now like any data pipelining
334
and orchestration i'm gonna go and connect all the notebooks
335
that have created at the workspace and start connecting them one after one
336
and at the ends what we're gonna do we're gonna go
337
and add a trigger for it to run automatically like for
338
example here we have different options the scheduled one i'm gonna say okay run it once every day
339
so i'm gonna say this is very simple for data engineering
340
to create an automated pipeline now for you it can be empty
341
but previously i have created something in order like just to
342
test now we can see like more details about my pipeline
343
and in the task you can see here i have created
344
like like three notebooks that's gonna build for me a data leak house
345
so I'm gonna say this is the easiest thing
346
that we do in Databricks
347
and data engineers in order to create a visual data pipeline
348
now what else we do as data engineers after we write the code
349
and build the pipelines we have to monitor every day whether everything is correct
350
and for this in Databricks we have a dedicated section in
351
that data engineering so here we have something called runs
352
and this kind of gives you a nice
353
and easy overview of the health of your jobs like for example here i tried to run it twice
354
and as you can see the status is failed of course you can see the start time
355
which notebook the type who run this and so on
356
so again here is the place where we can monitor our jobs
357
and do the operations and if something is wrong
358
so it's not green you have to go
359
and debug where exactly the issue happens
360
so actually that's it those are the main components
361
that are specially used by data engineers and data bricks
362
so the workspace where we go and write our logic our code
363
and organize the whole project and then we go
364
and connect everything to automate our code by building jobs
365
and pipelines and at the end so we go
366
and do the operations by monitoring the health of the jobs okay
367
so now let's move to the AI
368
and machine learning part this is mainly used by data scientists ML engineers
369
and AI engineers each role is slightly different
370
but they all work on building models
371
and intelligent system on top of our data
372
so let's go through the components that are relevant for them
373
so again they will be using workspaces because they're gonna go
374
and write codes so they have to write notebooks in order to explore data train models
375
and write code in python and of course they are interacting with data
376
so the catalog is again important the compute in order to execute things
377
but at the end over here we have a whole dedicated sections for ai
378
and machine learning
379
so what do you have here the first one is the
380
playground this is where you experiment with the ai models especially the language models
381
so you can test prompts you can try a few ideas quickly see how the model responds
382
so it is mainly used for exploration
383
and testing now the next one is actually new one the agents
384
so again agent it is an advanced ai system
385
that actually gonna go and do something you give it a task
386
and it's gonna go
387
and complete it step by tip it is more than just conversation with the ai
388
and here we can see that databricks is already like suggesting for us ideas
389
and some common use cases like the document parsing where you can go
390
and give the ai your documents like pdfs and images
391
and you can use it in order to start extracting the data from them another one we have information extraction
392
so again the same thing you're gonna give it your documents
393
and extract some key features and classifications another one
394
which is very famous one where you can go
395
and build a shot bot based on your documents
396
so again you can upload your things
397
and start having conversation with the ai about the contents of your files
398
and then the last one we already talked about it the ai bi genie we're gonna go
399
and convert your text to sql and many other things
400
so i'm gonna say
401
if you want to work with ai in databricks this component is gonna be the most interesting one
402
so what else do we have we have the ai gateway way.
403
It's all about managing the AI models that are used in your projects.
404
It's like the other monitoring tools that we saw inside Databricks but this time for the AI models.
405
Now let's go to the next one.
406
The experiments.
407
This is mainly used by data scientists in order to track the training of the models.
408
Go and compare the different trends
409
and parameters of the models in order to see at the end which model performs the best.
410
Of course we don't see anything yet because we didn't train any model
411
so we cannot see the content now let's go to the features
412
so this is something classical for machine learning engineers where you're gonna go
413
and store the features that you extracted in order to reuse it later
414
so that you don't have to go and rebuild them every time
415
and the next one we have the models
416
so again it's all about the ai models
417
so here you're gonna see where the models are stored
418
and versions you can go and manage the different versions
419
that are used inside your projects track the changes and reuse models when needed
420
and finally we have the serving this is where you deploy your model once the model is ready you go
421
and expose it as an api so the applications can go
422
and use it in real time
423
so as you can see the full section is all about ai
424
and machine learning it is the full life cycle from experimenting to building managing
425
and finally deploying the ai and machine learning solutions and
426
if you are not data scientist
427
or ai engineer i can tell you the most interesting one gonna be the agents all right my friends
428
so that was a quick tour into the data breaks interface as you can see there are many things
429
that we can do inside this platform
430
and i totally understand at the beginning it might looks a lot
431
and you might get overwhelmed with all those things
432
but now you have at least a clear idea on what each component is doing
433
and who uses what and now my friends there is something
434
that i really wanted you to understand that we don't just go
435
and learn everything that we see inside data bricks we focus on the parts
436
and the components that actually match our role
437
so now what i want you to do next based on your goal
438
and your role whether data analyst engineering scientist ai engineering go
439
and explore the platform and the interface on your own
440
so i want you to take few minutes by clicking around
441
and trying all those components
442
that i just showed you in order to get comfortable with the interface
443
and don't worry about it you will not break anything at the end
444
if you are using the free edition you are not paying for anything
445
so this is what you're gonna do next but before
446
that there is something important if you enjoy this type of content
447
and you would like to support my work then subscribe like
448
and comment this really helps in order to reach nice people like you thank you
449
so much for watching and i will see you in the next video bye
📺 同频道
✨ 推荐视频
关于本课
您正在使用跟读技巧通过视频"Learn Databricks for FREE (Step-by-Step Guide)"练习英语口语和发音。
每天练习15到30分钟,将显著提高您的英语流利度和发音准确度。
什么是跟读法?
跟读法 (Shadowing) 是一种有科学依据的语言学习技巧,最初开发用于专业口译员的培训,并由多语言者Alexander Arguelles博士普及。这个方法简单而强大:您在听英语母语原声的同时立即大声重复——就像是一个延迟1-2秒紧跟说话者的影子。与被动听力或语法练习不同,跟读法强迫您的大脑和口腔肌肉同时处理并模仿真实的讲话模式。研究表明它能显着提高发音准确性,语调,节奏,连读,听力理解和口语流利度——使其成为雅思口语备考和真实英语交流最有效的方法之一。












