Shadowing Practice: A/B Testing in Data Science Interviews by a Google Data Scientist | DataInterview - Learn English Speaking with Video

Les maken...
1
If you are preparing for a data science interview, A-B testing is a must-know concept.
2
Whether it's for Google, Meta, and Uber, and so forth, A-B testing is a very popular topic in interview questions
3
because data scientists in those companies will use A-B tests as a way to figure out whether the change
4
that they have made on those platforms are due to random chance or because of the actual change that they have implemented. And
5
so what we're going to do in this video is I'll
6
provide a walkthrough of an A-B test based on real life example
7
and along the way I'll pepper in some hints that can be helpful in terms of acing interview questions on A-B tests.
8
Hey everyone, I'm Dan, the founder of DataAntiv.com, ex-Google and PayPal data scientists.
9
In this video, we're going to do a deep dive on A-B testing based on a real life example
10
and what I will do is I'll walk through the procedure of setting up an A-B test
11
and I'll talk about a couple of things that you should definitely mention whenever you're walking through an A-B testing case.
12
Now let's get started.
13
The first thing that you need to realize is that when you are walking through the A-B testing procedure, there are essentially seven steps that you need to consider.
14
The first thing to do is basically understanding the problem, the problem statement.
15
This is where you try to make sense of the case problem
16
that you need to solve by asking clarifying questions to the interviewer
17
and also figuring out what is the success metric and what is the user journey.
18
And we will definitely do a deep dive on this topic in a second.
19
The second thing is that you want to define your hypothesis testing.
20
And what this basically means is that you set up what your null hypothesis and alternate hypothesis is, and you're going to set up some parameter values for your experiment,
21
such as the significance level and statistical power.
22
The third step is designing the experiment itself.
23
And so this is where you talk about what is the randomization unit
24
and which user type you're going to actually target for this experiment
25
and various other things that you definitely need to consider when you're designing the experiment.
26
The next step is to run the experiment itself.
27
And this is where you need to think about the instrumentation
28
that is required to actually collect the data and analyze the result.
29
Now, once you've collected the data, the next thing that you need to do, even before you actually interpret the result and decide to launch,
30
is basically do some sanity check or validity checks because
31
if your experiment design was flawed or if there was some bias that was implemented into the data collection itself,
32
then you have flawed result and so you might end up making poor decisions.
33
So this is where doing sanity check even before you think about the interpretation and launch decision is very crucial.
34
Once you have done the sanity check, the next step is to basically interpret the result in terms of what is the lift that you saw,
35
the p-value encompassed interval, And lastly, now that you have the statistical result along with the business context,
36
this is where you make a decision in terms of whether you're going to launch the change or not.
37
Now with all of the steps covered, now let's do a deep dive on an actual case problem.
38
Suppose
39
that an online clothing store called Fashion Web Store wants to
40
test a new ranking algorithm to provide products more relevant to customers.
41
How would you design an experiment?
42
So in order to tackle this problem that requires A-B testing, we want to first of all understand the nature of this product.
43
So we know that this is an e-commerce store that sells goods such as basically clothing goods such as clothes, shoes, bags, and other type of merchandise.
44
And what this store uses are basically a product recommendation system, an algorithm, where once the user searches some keyword,
45
let's just say clothing for instance, it's going to generate a result with products that could be relevant for the customer based on, let's just say, their profile,
46
transaction history, and so forth.
47
And so what we want to test in this example is that if we change the recommendation system, maybe we're providing more relevant products to the users.
48
Thereby, it should boost the revenue sales of this e-commerce store.
49
So we have some general sense about what this problem involves.
50
And one thing I want to mention is that I've worked with a number of clients on A-B testing interview questions.
51
One thing I've noticed is that clients will often skip this problem statement part and jump right into basically the statistical methodology,
52
proposing what the experiment design is.
53
You don't want to engage the interview question that way.
54
You always want to start with the business context and then segue over to the statistical methodology.
55
Now, as you're fleshing out the business goal of this problem one thing
56
that you want to do is basically clarify what is the user journey
57
or the product experience that you're trying to change
58
so this e-commerce store has a following user funnel a user will visit meaning they land on this landing site
59
and then they will search an item
60
which basically produces a result based on the recommendation
61
and then their a user is going to browse a couple items
62
and then eventually they might click an item and then they'll purchase it
63
which is the ultimate success that we seek to achieve through the change.
64
Now thinking about the user journey is important because later down the road
65
when you're thinking about what is the success metric
66
and what is the target user population you want to think
67
about at what stage do you want to consider the user to be participant for the experiment
68
and this is something we're going to talk about very soon.
69
Now, one pro tip I definitely want to mention is that when you have an A-B testing interview round, whether it's for meta, Google, go through the individual product,
70
basically the core product, the core features of that platform, and then create an outline of what that user journey is.
71
Once you have done this, this is going to be really helpful
72
when you are in the actual interview setting and you're asked to design an A-B test based on a particular product.
73
Once you have establish the user journey the next thing to do is basically define the success metric
74
so what is it that you need to move in order to basically be certain
75
that the change that you're applying is actually better for the platform overall now
76
when you think about the success metric you want to think about a couple qualities
77
so the first thing that you need to think about is
78
that is it measurable
79
so is it the type of user behavior you can actually collect through the instrumentation
80
or the platform and the next quality you want to think about is is your metric attributable
81
and what this basically means is that can you establish a clear linkage
82
that the cause the treatment
83
that you applied to the platform is has led to the effect
84
that you saw the change in the metric in this case the next quality
85
that you want to think about is is your metric sensitive
86
and you essentially have metric that serves as a proxy
87
that yes the there is a genuine difference in terms of
88
the user experience in terms of the old algorithm versus a new algorithm
89
and so statistically speaking you want to find a metric
90
that has low variability
91
and a bad metric for in this case let's just say time spent on the website
92
so the total time spent on the on the website for a given user might have high variability and
93
so you cannot clearly tell whether there is an actual difference
94
in terms of how the user is engaging the e-commerce given
95
that there are some underlying changes to the ranking algorithm the fourth thing is
96
that a b experiments needs to be very quick it's very
97
it's a very iterative process as a way to improve a product very quickly
98
and so you want to ensure that what you're measuring is timely.
99
You don't want to wait weeks and months to observe a user behavior
100
and then eventually make the change because that's a very costly way of running an experiment.
101
So you want to think about what is that short-term behavior that can serve as a proxy for long-term desired behavior.
102
So based on these four qualities,
103
ultimately the success metric that we want to use for this case is the revenue per day per user.
104
Now once you've established the problem statement clearly, the next thing to do is establish your hypothesis testing.
105
So this is where you state your null hypothesis and alternative hypothesis.
106
So the null hypothesis in this case is going to be
107
the average revenue per day per user between the baseline and the variant ranking algorithms are the same.
108
And the alternative hypothesis in this case is going to be that the
109
average revenue per day per user between the baseline and the variant ranking algorithms are different.
110
Once you've state the hypothesis statements, the next thing that you want to do is you want to set the significance level, or your alpha in this case.
111
So the significance level is basically the decision threshold.
112
If the probability of observing a particular event is very low, then it is deemed statistically significant.
113
And so in this example case, the significance level that we want to set is 0.05.
114
And that's the usual value that is often set in online experiments the next value
115
that we want to set is this physical power the statistical power usually is set as 0.80
116
and what 0.80 basically means is is the 80 probability of detecting effect given
117
that the alternative hypothesis is true
118
and lastly what you want to set is your practical significance basically the minimum detectable effect
119
and typically for a large online platform with millions of users
120
the mde is one percent lived once you've set up the hypothesis testing the next thing
121
that you want to do is design the experiment itself
122
so the first thing
123
that you want to consider as you design the experiment is
124
you want to consider what is the randomization unit in this
125
case we're going to randomly assign at the user level basically randomly assign them in terms of the control group
126
or the treatment group once you have defined the randomization unit you you need to think about
127
which population of the users you want to target.
128
And this is something that I talked about earlier in terms of the user funnel, right?
129
So you have a user that visits, you have a user that searches, they browse, they will visit an item and then they'll eventually purchase
130
so at what level do you want to allow the user to participate in the experiment
131
so in this case we actually want to target users who have actually started searching something why
132
because this is where the algorithm actually kicks in
133
and then they're actually exposed to the treatment in this case
134
which is either the the old algorithm or the new algorithm the third thing
135
that you want to define is your sample size.
136
The general rule of thumb is the following formula, which is that n is approximately equal to 16 times the variance divided by the delta square,
137
where delta essentially represents the difference of the key metric between the treatment and the control.
138
And this formula is based on the assumption that your significance level is 0.05 and your statistical power is 0.80.
139
Once you have determined your sample size, the next thing that you want to do is determine the duration of your experiment.
140
And the typical duration of the experiment is going to be one to two weeks.
141
You don't want to run the experiment less than one week
142
because you want to account for the day of the week effect, meaning
143
that there could be some underlying difference in terms of how the user engages the website during the weekdays versus the weekends.
144
Once you have designed the experiment, the next thing that you want to do is you want to run the experiment
145
and this is where you use instrumentation some experiment platforms as a way to collect the data
146
and track your result now it is very important that
147
while you're running the experiment you do not peak at the
148
p-value meaning you don't make any decision in terms of whether you're going to launch
149
or not while the experiment hasn't been completed yet the reason is
150
that when you're peaking there's a chance that
151
when you have low sample size for instance there's a lot of variability in terms of where
152
that lift goes and so you might falsely conclude that there is an underlying difference when there isn't.
153
So once you have determined what the experiment time is, given your statistical power and your sample size, you have to ensure that you wait it out.
154
Otherwise, you're going to increase the chance that you falsely reject the null hypothesis given that it is actually true.
155
After you have run the experiment, the next thing that you want to do is you want to perform validity checks.
156
So this is where you conduct sanity checks including
157
instrumentation effect or are there any bugs or glitches
158
that could potentially affect the experiment result another potential issues
159
that you want to look at are external factors
160
so maybe you run the experiment during the holiday or
161
when competition launched something very important
162
or it could be some general economic conditions like covid
163
or recessions and when you run the experiment when some external disruptions happen this could potentially impact your experiment result.
164
So ideally you want to run an experiment that avoids periods like this.
165
The next thing that you want to do is you want to check for selection bias.
166
And what this basically means is that you want to assume that the underlying distribution between the control and the treatment group,
167
even before they're exposed to the treatment condition in this case, is that they're homogeneous.
168
And so one way to confirm that the distributions are the same is to basically run an AA test.
169
The next thing that you want to check is a sample ratio mismatch.
170
What it basically means is that if you're randomly assigning a user into the control or the treatment group, then out of all the participants of the experiment,
171
50% of them should be in the control and 50% of them should be in the treatment.
172
But there are cases where because of some flaws with the randomization algorithm, that the ratio is actually not 50 to 50%.
173
It might be 49 to 51%.
174
So in order to ensure whether this could pose a potential issue later down the road,
175
you want to use a chi-square test as a way to ensure that the ratio between the two samples is sound.
176
The next item you want to check is a novelty effect.
177
What this basically means is that if you made some change to the website itself,
178
a user might have reacted simply because there's a novelty behind being exposed to something new.
179
One way to detect a novelty effect is basically look at it by user segment.
180
Look at the underlying difference of the success metric in terms of new visitors versus the recurrent visitors.
181
And if you see that there is a change between the two, then there is a presence of novelty effect.
182
So you may actually want to run the experiment where you
183
segment it by the new visitor group versus the recurrent group first.
184
You've conducted the validating check and there's no issue with the experiment experiment, now you can actually interpret the result.
185
And when you interpret the result, you want to look at the direction of the success metric, is a lift negative or positive?
186
And you want to consider the p-value because
187
that helps you establish whether the lift that you saw is statistically significant or not.
188
And you also want to consider the confidence interval in this case.
189
So based on the experiment that we ran for this example, what we see is
190
that the average revenue per day per user in the control group is $25 whereas in the treatment group is $26.10.
191
This produced a falling lift.
192
So in terms of the absolute difference is $1.10.
193
In terms of relative lift, it is an increase of 4.4%.
194
And the lift we saw is statistically significant because we see
195
that the p-value 0.001 is less than the statistical significance at 0.05.
196
And we also see that the confidence interval at the significance level of 0.05 is between 3.4 and 5.4% lift.
197
So the initial interpretation based on what we see is the following.
198
There is statistical significance to reject the null hypothesis and conclude
199
that the average revenue per day per user between the baseline and the variant ranking algorithms are different.
200
With this result, we can now consider whether we want to launch or not.
201
Now when you think about whether you want to launch it or not, and in this step there are three factors that you want to consider.
202
The first factor is the metric trade-off.
203
So you might have a case where the success metric might have improved, but the guardrail metrics or the secondary metrics might have declined.
204
And so you have to think about what are the pros
205
and cons of launching it considering that the guardrail metrics might have declined.
206
The next factor you want to consider is the cost of launching.
207
If you see that the cost of building this out, basically rolling out to all of the users, and the cost of maintaining this change is highly costly,
208
then maybe this isn't something that you want to actually launch.
209
The last factor you want to consider is the risk of committing false positives, or your type 1 error rate.
210
For instance, if you falsely conclude that there is an effect when there isn't, and you've made a change then it might have a negative
211
consequence to the user you might end up providing poor experience to users
212
and they might turn and ultimately
213
that is going to negatively affect the bottom line of the product
214
so these are three important business factors
215
that you want to consider now you want to kind of go back to the interpretation of the result
216
that you got from the experiment and along with the business context
217
and the statistical result ultimately you want to decide whether you're going to launch it or not
218
so what we're going to do is we're going to look
219
at a couple example cases where we'll look at a possible range of lift along with the confidence intervals
220
and then think about what is a sound decision based on the result that we have seen.
221
So in this first case what we see is
222
that the lift is placed at a positive value
223
but it's still less than practical significance and you also see
224
that the lower bound
225
and the upper bound of the confidence interval are less than the practical significance in this case
226
which is going to be
227
positive one percent so in this case what you want to consider is
228
that maybe you want to change the algorithm
229
or scrap the idea all of it together in the next case what we see is
230
that the lift and the bounds the confidence interval is uh practically significant
231
so this provides a strong support
232
that we should make a launch in the third case we see
233
that the entire interval is less than zero
234
so it's in the negative territory in this case you want to consider perhaps iterating the idea or scrapping it all together.
235
And the next example is you have a positive direction in the expected lift, but then the bounds are in the negative territory and the positive directory.
236
And you can see that it's a very wide bound.
237
Now, what is very important to note is that the upper bound as interval is practically significant.
238
It's greater than 1%.
239
So there is still some likelihood that you might see a lift with practical significance.
240
So what you want to consider in this case is actually to rerun the experiment with increased statistical power
241
and this is going to help improve the precision of that lift that you're seeing.
242
And in the last case you see that the lift itself you see
243
that the expected lift is practically significant but you see
244
that the lower bound of this is not practically significant but it's still in the positive territory.
245
And the best thing to do in this case is to
246
rerun the experiment with increased statistical power just to be absolutely sure that the underlying change is practically significant.
247
So based on all of these considerations in terms of the business context, along with various statistical outcomes.
248
Ultimately, the decision you want to make is
249
that you want to launch this new algorithm as a way to provide a more relevant product recommendations to the users, thereby improving revenue overall.
250
So there you have it, guys.
251
This is the end-to-end process of how to walk through an A-B test, how to actually address an A-B testing interview question.
252
I hope you found this video really helpful
253
if you need any help in terms of mock interview coaching
254
or courses that comes with a b testing courses
255
and various business case problems and slack community access definitely check out data ntv.com
256
and if you have any questions along the way feel free to drop a comment down below
257
or feel free to send me an email at dan at data ntv.com i'll see you in the next video

Context & Background

The video features a former Google data scientist breaking down A/B testing for data science interviews, using real-world examples like optimizing an e-commerce platform. The dialogue is clear, analytical, and packed with technical terms, making it ideal for English learners looking to practice understanding and mimicking professional communication. Its structured approach—with step-by-step explanations—also helps build confidence in following complex, logical conversations, a key skill for IELTS speaking practice and everyday professional interactions.

Top 5 Phrases for Daily Communication

  • "must-know concept": Useful for emphasizing essential skills (e.g., "Time management is a must-know concept for students").
  • "walk through the procedure": Perfect for explaining steps (e.g., "Let me walk through the procedure for setting up the meeting").
  • "sanity check": A casual term for verifying accuracy (e.g., "Let's do a sanity check on the numbers before submitting").
  • "pepper in some hints": To add small, useful tips (e.g., "I'll pepper in some hints to improve your essay structure").
  • "deep dive": For exploring a topic thoroughly (e.g., "We'll take a deep dive into climate change in tomorrow's class").

Step-by-Step Shadowing Guide

To master the video's dialogue using the shadowing technique, follow these steps. First, listen to 10-15 second clips and repeat immediately, focusing on matching the speaker's pace and intonation—this improves English pronunciation and fluency. Next, identify technical terms like "statistical power" and practice saying them aloud until they feel natural; this builds confidence for IELTS speaking practice. Then, use a shadowing app to record yourself and compare it to the original, noting gaps in clarity or rhythm. Finally, mimic the speaker's analytical tone by emphasizing key words like "crucial" or "definitely"—this helps convey confidence in professional settings. By breaking the video into manageable parts, you'll gradually internalize its structure and vocabulary, making complex conversations easier to follow and replicate.

Wat is de Shadowing-techniek?

Shadowing is een wetenschappelijk onderbouwde taalleermethode die oorspronkelijk is ontwikkeld voor professionele tolkentraining en gepopulariseerd door polyglot Dr. Alexander Arguelles. De methode is eenvoudig maar krachtig: je luistert naar native Engelse audio en herhaalt het onmiddellijk hardop — als een schaduw die de spreker volgt met slechts 1–2 seconden vertraging. In tegenstelling tot passief luisteren of grammaticadrills, dwingt shadowing je hersenen en mondspieren om echte spraakpatronen tegelijkertijd te verwerken en te reproduceren. Onderzoek toont aan dat het de uitspraaknauwkeurigheid, intonatie, ritme, verbonden spraak, luisterbegrip en spreekvaardigheid aanzienlijk verbetert — waardoor het een van de meest effectieve methoden is voor IELTS Speaking-voorbereiding en echte Engelse communicatie.

Shadowing-techniek: lees de volledige stap-voor-stap-gids →