Shadowing-Übung: Learn Data Modeling in 8 minutes: Dimensional Data Modeling, Data Vault, and One Big Table - Englisch Sprechen Lernen mit Video

Lektion wird erstellt...
1
Data modeling is one of the most important skills that data engineers have to learn.
2
So how is it so important?
3
Well, at the end of the day, data modeling is the structure of the data that we create for businesses.
4
Because the asset that we produce for businesses is data to make better decisions.
5
And data modeling has a bunch of different components to it.
6
And we're going to talk about each of them.
7
There's the highest level, it's called conceptual data modeling.
8
Then we're going to be going over the next layer down, which is logical data modeling.
9
And then the final layer that we're going to talk about is physical data modeling.
10
In this video, we should be covering all of those.
11
So at the end of this, you should know all the ins and outs of all three types of these data modeling.
12
So let's first talk about the highest level conceptual data modeling, which is the highest level way of thinking about it is like, what data do we need?
13
What data does the business desire?
14
What data do we have access to?
15
What are the sources?
16
Where can we find this data?
17
Can we like think about all the relationships of this data, like at a very high level? because sometimes you might want data in a business, but it's actually unfeasible to get it,
18
or it's going to take months and months and months.
19
And it's not ROI positive
20
because the time it would take to get the data is more than the value you can get out of it.
21
So this area of data modeling is very, very, very important because if you get this wrong,
22
you can end up spending a lot of extra time on the data model for things that don't matter that much.
23
Like there might be an extra column or an extra table
24
that like some data scientists ask for that they're only going to use one time.
25
And if you push back on that, it makes your job as a data engineer a lot easier because you remove that requirement because it's a low value requirement.
26
So you need to be able to like walk through all the requirements of the data model at this highest level.
27
And really having a deep understanding of the business is where you can really, really shine as a data engineer with conceptual data modeling.
28
So after you have mastered the conceptual data model, you skip into the next layer, which is the logical data model.
29
So the logical data model is where you start talking about like facts and dimensions and how they're related.
30
Like, so facts are events.
31
You can think of that as like I click the login button.
32
I purchase a product.
33
I send a message to a friend.
34
Like all of these are like events and actions that you can take on a website.
35
And these are going to be your facts because they can't really change.
36
After I've logged in on Facebook at 6.03 p.m., like I can't go back into a time machine and change that.
37
So then you have dimensions, which are going to be like your nouns or like your actors in this space, right?
38
So you have things like users or listings or posts or, you know, devices.
39
These are all like things that bring context.
40
Cause like I could be like, I'm a user in China who clicks an ad, right?
41
And then the user in China bit, that part of my user object and then clicking the ad is part of the fact.
42
And so those are the ways to be thinking about all of these things.
43
So the logical data model is all about how do you find all of these entities and how are they related.
44
So like, for example, when you click on an ad on Facebook, the entities that are involved, the user who clicked the ad, you also have the device that clicked the ad,
45
you also have the browser potentially that clicked the ad.
46
And so those are three dimensions
47
that you can bring in to your fact data to make your fact data richer and more powerful.
48
So after you have gotten the logical data layer nailed down, then you move into the final layer, which is the physical layer.
49
And this is the layer where people oftentimes, like they think the physical layer is the only part of data modeling when it's actually not,
50
it's actually like probably the least important in some regards, but it's also where it becomes real.
51
It becomes reality.
52
It's physical data now.
53
So this is where you start looking at the schema.
54
What are the columns of your data?
55
Like what are the data types of your data?
56
How are you storing this data?
57
How can you compress this data to make it smaller?
58
So in this space, you have a couple different data modeling techniques
59
that you can look into
60
that can really help understand all the way up from like physical to logical to conceptual
61
and you can like understand there's a couple different paradigms here
62
so one is called one big table
63
so one big table essentially says you don't need facts
64
and dimensions you can essentially have a fact
65
that has all of its context all the context is already in the fact
66
so you don't need to do joins you can just do aggregations you don't need to join at all
67
because when data gets very big joins become very expensive like actually joining in the the nouns into your data.
68
But the trade here, right, is that you're duplicating data, right?
69
Because every now my user data is now in every single fact, right?
70
And say I do 300 events on a data set.
71
This is what's called the compute storage trade off, right?
72
Because it is duplicated in storage.
73
But why is this potentially worth it is
74
because you can now do aggregations without doing join
75
and aggregations are faster than join because you don't have to like pull data from all these different locations, you can just aggregate right down and group very quickly.
76
And so this is what is the advantage of one big table, even though one big table also duplicates a lot of data and is super painful.
77
So then you have the more traditional approach of facts and dimensions, where like you separate your nouns and you separate your verbs and your events, and then you join them together.
78
That's called dimensional data modeling.
79
Dimensional data modeling is the oldest technique of all the techniques that we have talked about.
80
And it's the one that most data engineers should probably pick.
81
It's the most foundational and the most fundamental.
82
So definitely, if you haven't learned about dimensional data modeling, I'll put a link in the description below.
83
I have many hours of content on dimensional data modeling on my YouTube channel.
84
And then last but not least here is called a data vault.
85
So data vault is very interesting because data vault is all about preserving the rawness of the data.
86
So you want to bring in data in its most untransformed state and kind of stack it onto things.
87
And this allows you to essentially know exactly what the data was in its rawest form.
88
So you don't end up filtering something out that you didn't mean to or transform something that you made a mistake on.
89
And Data Vault allows you to back up into like the most raw form of the data.
90
It's kind of related to ELT, you know, extract load transform instead of extract transform load.
91
That's it's kind of related to that.
92
But it's also about how you actually model the data.
93
And it has a very strict kind of thing in it and like a strict kind of ideology to it.
94
That's why I don't really I don't really subscribe to Data Vault as much.
95
I really like dimensional data modeling and one big table data modeling.
96
But at the end of the day, like it's still a very useful technique.
97
And you can combine all of these techniques.
98
I'm going to give an example.
99
So when I worked at Airbnb, I actually built a data model that was a mix of one big table and data vault and dimensional data modeling.
100
So how it worked was we use data vault to get all of the inputs.
101
So the inputs in this case were for availability and price
102
because the host sets the price rules like what is the daily price?
103
Are there any discounts?
104
Are there any like any length of stay requirements?
105
Maybe you have to book for three days or seven days or you know, you can only book 48 hours in advanced, there's a lot of these different rules that hosts can set in Airbnb.
106
And so I took all those rules in their raw form and combine them into one data set called inputs.
107
And that was the data vault kind of technique that I use.
108
And then from the inputs table, we process it through to a new table called listing pricing.
109
So now we actually know what the actual prices would have been given those rules.
110
And that table was one big table because it had all of the nights in one column.
111
So one of the giveaways of like a one big table data set is that it uses complex data types, like it uses things like struct and array and like map
112
and those kind of like more complicated data types more than just like, you know, string and integer and decimal.
113
It's like one big table very often is a table within a table.
114
And that was a big thing I did here at Airbnb.
115
Keeping in mind that like you don't have to adhere to any one of these three religions
116
when you are working with with data modeling, you can also kind of do whatever you want, right?
117
And mix and match.
118
It's more of an art than a science.
119
And it's really about matching business requirements to the physical data modeling needs that you, or the physical data modeling tools that you have.
120
So yeah, I hope you found this content interesting and make sure to like, follow, and subscribe.

Über diese Lektion

Was ist die Shadowing-Technik?

Shadowing ist eine wissenschaftlich fundierte Sprachlerntechnik, die ursprünglich für die professionelle Dolmetscherausbildung entwickelt und durch den Polyglotten Dr. Alexander Arguelles populär gemacht wurde. Die Methode ist einfach aber wirkungsvoll: Du hörst englisches Audio von Muttersprachlern und wiederholst es sofort laut — wie ein Schatten, der dem Sprecher mit nur 1–2 Sekunden Verzögerung folgt. Anders als passives Hören oder Grammatikübungen zwingt Shadowing dein Gehirn und deine Mundmuskulatur, gleichzeitig echte Sprachmuster zu verarbeiten und zu reproduzieren. Studien zeigen, dass es Aussprachegenauigkeit, Intonation, Rhythmus, verbundene Sprache, Hörverständnis und Sprechflüssigkeit signifikant verbessert — was es zu einer der effektivsten Methoden für die IELTS Speaking-Vorbereitung und reale englische Kommunikation macht.

Shadowing-Technik: die vollständige Schritt-für-Schritt-Anleitung lesen →