ฝึกพูดภาษาอังกฤษด้วยเทคนิค Shadowing จากวิดีโอ: What is Master Data Management? MDM Explained for Beginners

กำลังสร้างบทเรียน...
1
Up on the screen here, I have an example of one customer, Priya Sharma, as three systems know her, where your CRM has her at 14 Elm Road in Leeds,
2
billing calls her P Sharma and sends her invoices to 3 Canal Wharf, and support has her in her email that no other system has seen.
3
And each record was correct the day that someone typed it in, so which one is true?
4
That's where a whole discipline exists to answer that question, and that's what we're going to be talking about today so
5
that you'll be able to decide whether records like this are one person then build the one record
6
that everyone should trust field by field
7
and this discipline is called master data management
8
so the records here are all made up right nine of them across three systems built
9
so every count on screen can be checked by hand
10
so just you know illustrative for the purpose of this
11
and not going to compare vendors or you know what MDM program to choose
12
but I want to give you the definition of what is master data management
13
and how can you start implementing it in your day -to -day
14
so the dictionary definition of master data management is just the
15
practice of keeping one trusted record for every important business entity
16
and then mapping every system's copy back to that trusted record
17
and the trusted record is you know something like the golden record right uh
18
and say mapping every copy back to it though is the piece
19
that a lot of people will forget about um you know
20
everyone has knows the concept of hey you have one gold standard but with, you know, programmatic controls, you actually can map everything back to that gold standard.
21
And that matters because it means that, you know, no system is going to delete its own customer table.
22
So what counts as master data?
23
There's a textbook, the DAMA Data Management Body of Knowledge,
24
which defines it as data about the business entities that provide context for business transactions.
25
So customers, products, suppliers, employees, locations, the nouns, the business.
26
And then they have another term, which is the transactional data, which are the verbs, the orders, payments, tickets, page views.
27
Each one of those belongs to the system that recorded it.
28
And then all of those, you know, each point at master data.
29
So reference data is also the small agreed list like country codes, which DAMA groups with master data because both are shared.
30
And so a quick test is, is it shared by more than one system
31
and do transactions point it right and so an order lives in one system
32
and the customer in that order that lives in all of them
33
and that's where the trouble starts so you know
34
if it is only in one system
35
and it's just referencing the data there then you know it's probably just a transactional uh piece of data but
36
if it's across a bunch of different systems then that's where
37
you're going to have the problem of dealing with master data
38
so then today nobody plans to have you know something like nine records for four people, right?
39
A merger brings in a second CRM.
40
Maybe, you know, every team bought a SaaS tool with its own customer table
41
or customers signed up once with a work email and then again with a personal one.
42
And so in this example, right, you count the rows, you get nine, you got distinct email addresses, which is a typical trick, you know, people look for for uniqueness.
43
I mean, even after you lowercase all of them, you're only getting six unique email addresses.
44
And the true answer is that there are just four people contained within this data set of nine rows.
45
Priya and Tom appear three times each and two different people called Okafor share one address.
46
So which number goes in the board report?
47
Marketing emails, Priya two addresses and sales
48
and finance can never agree on a customer account
49
and support can't see that the caller has an unpaid invoice because billing filed under P.
50
Sharma.
51
And none of these systems are wrong.
52
The problem exists only between them, right?
53
We need to reconcile data from all those different teams
54
because no single team owns it and so the first job is describing
55
and deciding which job records describe the same person and then
56
that turns out to be a search problem
57
and that's where we start with the first piece of MDM
58
so MDM works to solve this problem in three moves
59
and matching is the first one and it comes in two flavors
60
and a lot of systems will use both of them
61
and the first one is deterministic matching which means exact rules
62
so the same email address lowercase and trimmed means the same person, it's fast, it's really wrong when it fires, and Priya's CRM record and her billing record share an email,
63
so that pair is settled, but the trouble is everything those exact rules misses.
64
And so her support record actually is a different email, and Tom Fletcher's billing record calls him Thomas and uses another email entirely, and to an exact rule, those are strangers.
65
So that's where probabilistic matching scores the evidence instead, field by field, and And here, Priya is billing record against her support record with illustrative weights.
66
So surname agrees, that's plus four.
67
First name is only an initial, plus one.
68
Same street, plus three.
69
Same postcode, plus two.
70
Different email, minus one, because people have more than one typically.
71
And that's nine, a combined score of nine, and let's say we have a rule that says eight or more is a match, below four is not, and anything in between that gets manually reviewed.
72
So now you have a case where that makes scoring worth the effort.
73
for Sarah Okafor and Sam Okafor, who share a surname and a home address, postcode included, but not a first name or an email.
74
So the different first name costs three points and they score five.
75
And a rule like same surname and same address would have merged them without asking.
76
Instead, the score sends them to a data steward because they're in that gray area four to eight, which is the person who accountable for the data, who knows a household when they see one, right?
77
This is pretty obviously a married couple.
78
And Realtools don't handpick those points.
79
Splink, an open source library, uses a model from a 1969 paper that learns each weight from your data.
80
How often a field agrees between true matches against how often it agrees by coincidence.
81
And it also adjusts for how common a value is.
82
So sharing a rare surname counts for more than sharing Smith.
83
And then there's the cost as well, right?
84
Comparing every record with every other one grows with the square of the row count.
85
And a million records gives you about 500 billion pairs.
86
So you block.
87
So you only compare records that share something cheap to catch, like a postcode or an email.
88
And the catch is that a pair sharing no blocking key is never compared at all, which is also why you run more than one blocking rule.
89
So Priya's three records are now one person carrying two spellings of her name
90
and two home addresses and something has to pick now.
91
So we have to decide, hey, what is going to be the true home address?
92
So now the second piece of MDM is merging.
93
Its rules are called survivorship rules because they're going to decide
94
which value survives into the golden record
95
and survivorship works one field at a time
96
so you keep one whole record instead
97
and you lose something whichever one you pick right
98
and the crm record has no phone number
99
and the newest record carries a support desk email
100
that nothing else uses so here each field gets its own rule
101
and here are four of the common ones source trust makes one system the authority for a field
102
and here the crm owns name
103
so the golden name is priya sharma rather than p sharma frequency takes the value most records agree on, and priya .sharmatexample .com appears in two of the three.
104
Regency takes the newest value.
105
Her support record changed on 15th of September, 2025.
106
So the address is 3 Canal Wharf.
107
Completeness takes whatever value actually exists.
108
Only billing has the phone number.
109
So 07700900461 goes in.
110
And then that's the golden record, right?
111
One row that no single source had held that we can bind from all the different sources of truth.
112
And so if we look at the email again, right, frequency would have picked the address that two systems share and recency would have picked the one only the support system knows.
113
And neither answer is technically wrong.
114
Instead, which rule governs which field is just a business decision with an owner, right?
115
It's just something you have to decide and, you know, hey, there are the trade -offs there.
116
And that's where data governance comes in.
117
And then the third, which is the third move, right?
118
So every source ID gets written into a cross -reference table that points at a master ID.
119
So C1042, B77810, and S3391 all point at M0001, and that's that golden PRIA record.
120
And so that table is then what lets one report join all three systems on a single customer, pulling from all those different source systems,
121
which raises the question of where does that golden record actually live?
122
And how do people like billing actually see it when it comes time to check?
123
So to answer that question, there's really four different main implementation styles.
124
And that's really what differs between MDM systems is where that result lives.
125
And so the vendor guides, Informatica and Stevo systems among them, describe four styles.
126
And two questions really separate them.
127
Is the golden record stored?
128
And is it written back into the source systems?
129
And that gives you a matrix of four different options here.
130
So registry is the lightest where the hub stores the cross reference table
131
and the match results and assembles the golden record on request
132
and nothing is written back so billing still says psharma and
133
because no source system gets touched it's a low risk place
134
to start then consolidation adds one thing where it stores the
135
golden records in a hub usually to feed reporting in the warehouse
136
and the sources stay untouched so your board report can agree with itself
137
while the support desk still can't see that invoice and coexistence goes even further and publishes the golden record back.
138
So billing gets corrected to Priya Sharma.
139
The sources can still edit their own copies.
140
So changes flow in both directions and the hub has to reconcile them, which is where the complexity lives here.
141
And then centralized is the far end where the hub becomes the system of record.
142
And there, new customers get created there.
143
Every other system subscribes to it.
144
And this buys the strongest consistency and also demands the biggest change
145
because every application that used to create customers now has to stop.
146
And so the styles really run from read -only to fully in charge.
147
so now let's run through the all three moves over all those nine records
148
so now i'm going to run through kind of the entire example um from start to finish right
149
so let's say those same nine records right that's going to give us 36 possible pairs
150
and that we're then going to block on two rules same postcode or same email.
151
The postcode rule will yield five pairs, and the email rule yields three.
152
And Sarah's two records turn up under both, so seven pairs get compared, the other 29 never need to.
153
Send three of the seven, share an email, and settle on the spot.
154
And then of the four that get scored, Priya's pair and Tom's pair match at nine and ten.
155
And both Sarah against Sam pairs score five and go to review, where the steward will then keep them apart.
156
And that leaves five matched pairs.
157
And the next step is where matches get grouped in the clusters by following the links.
158
Priya's CRM record and support records were never compared because they shared neither a postcode nor an email, and they still open the same cluster because each one matches her billing record.
159
So then those five matches collapse into four clusters, which become four golden records.
160
M01 for Priya, M02 for Tom, 3 for Sarah, M4 for Sam, and the cross -reference table then gets nine rows, one per source record.
161
However, chaining will cut both ways.
162
Had Sam wrongly scored eight or more against either of Sarah's records, he would have been pulled into her golden record and every report downstream would count two people as one.
163
So that's the entire mechanism.
164
So where is it going to show up in the wild, when not just in a whiteboard in that conceptual session?
165
So you'll mostly see MDM take two forms in the wild.
166
The first is a commercial MDM suite, Informatica, Relto, Stebo Systems, Prophecy all sell one with matching and survivorship built in,
167
plus review queues for stewards.
168
And that's usually what someone senior means by we need MDM.
169
The second is MDM Lite, which is built by the data team itself.
170
So source tables land in the warehouse, DBT models clean them up, an entity resolution library like Splink does the matching, and more DBT models apply survivorship and build a cross -reference table.
171
And Splink comes from the UK Ministry of Justice under an MIT license and runs on DuckDB or Spark.
172
And then MDM also sits next to two other topics I've covered.
173
Governance, you know, who owns the customer entity and who signs off a survivorship rule and data catalog records, which system is the source of truth for which field.
174
So now what are people going to commonly get wrong though when they try to implement MDM on their own?
175
So the first wrong thing to think is that the MDM is just a tool you buy.
176
The suite will do the matching, but someone's told to decide that the CRM owns names and
177
that Sam and say are two people and there's no license that's going to make those decisions for you.
178
The second is that MDM is the data warehouse.
179
A warehouse stores history for analysis, while MDM decides which records are the same thing.
180
Consolidation can feed a warehouse, but coexistence and centralize right back into operational systems, which a warehouse never does.
181
The third is that MDM and data governance are the same thing.
182
Governance sets the rules and names the owners, and MDM is one of the places those rules get executed.
183
A survivorship rule is a governance decision just running as code.
184
The last one is that once a golden record exists, you're done.
185
New signups arrive every day and Priya will move again.
186
So matching and survivorship keeps running and the stewards review queue is going to keep filling.
187
So it's really an ongoing system.
188
So that's everything I have for you today.
189
Just to kind of recap, master data management keeps one trusted record for every important business entity and maps every system's copy back to it.
190
And both halves of that should mean something to you matching decides
191
which records describe the same thing with exact rules where they work
192
and scored evidence where they don't blocking keeps that affordable
193
and survivorship builds the golden record one field at a time
194
and the cross -reference table points every source id adding master id the implementation style only decides where
195
that record lives and whether it gets written back
196
so for priya the answer to the opening question is m1
197
the name her crm holds the email two systems agree on her newest address
198
and the only phone number anyone ever had and Obviously, there are limits to this.
199
Matching is going to make mistakes in both directions, so merging a household or missing someone who moved, and the threshold only chooses which mistake you'd rather make.
200
Survivorship rules are policy, and that needs an owner, and those nine records are also built to be checkable.
201
Real data is a lot messier, and real numbers are only going to come from real runs, so try this out for yourself.
202
Maybe do a little pilot for MDM and see if it actually helps you out.
203
But I hope you enjoyed this video.
204
I hope you learned something.
205
Have a great rest of your day.
206
Daddy I out.

เกี่ยวกับบทเรียนนี้

คุณกำลังฝึกภาษาอังกฤษกับ "What is Master Data Management? MDM Explained for Beginners" ด้วยเทคนิค Shadowing — วิธีที่พัฒนาขึ้นสำหรับการฝึกนักแปลมืออาชีพ

ฟังทีละประโยค สังเกตการเน้นเสียงและการเชื่อมเสียง แล้วพูดตามดังๆ อย่างมั่นใจ ฝึกวันละ 15–30 นาทีจะเห็นผลลัพธ์ที่ชัดเจน

เทคนิค Shadowing คืออะไร?

Shadowing เป็นเทคนิคการเรียนรู้ภาษาที่ได้รับการรับรองทางวิทยาศาสตร์ พัฒนาขึ้นสำหรับการฝึกนักแปลมืออาชีพ วิธีการนี้เรียบง่ายแต่ทรงพลัง: คุณฟังเสียงภาษาอังกฤษจากเจ้าของภาษาและพูดตามทันที — เหมือนเงาที่ตามผู้พูดด้วยช่วงเวลาห่าง 1-2 วินาที การวิจัยแสดงว่าเทคนิคนี้ปรับปรุงความแม่นยำในการออกเสียง ทำนองเสียง จังหวะ การเชื่อมเสียง การฟังเข้าใจ และความคล่องแคล่วในการพูดได้อย่างมีนัยสำคัญ

เทคนิค shadowing: อ่านคู่มือฉบับเต็มทีละขั้นตอน →