Shadowing Practice: FULL Data Science Mock Interview | SQL, Statistics, Business - Learn English Speaking with Video

Creando lección...
1
In this video, you're going to watch a data scientist work through a technical on-site round that includes SQL, stats, and a business case.
2
Johnny, right?
3
Yes, Johnny.
4
Nice to meet you.
5
How's it going?
6
Nice to meet you.
7
You're not related to Johnny Bravo at all, right?
8
Not at all, though.
9
Yeah, I can't say I haven't watched a show before.
10
I did watch it, definitely.
11
You got the same haircut and everything.
12
It was like the first thing I noticed.
13
Yeah, there's definitely some similarities.
14
Cool.
15
Well, thanks for taking the time to do this interview.
16
Just wanted to start off with a little bit of logistics, just to give you a groundwork for kind of what we're going to go through during this time.
17
We're going to be doing the interview for 45 minutes.
18
It's going to be a technical round and there's going to be three sections, a SQL section, a stat section and a business use case.
19
that we're going to walk through together.
20
Most of it will be more conversational in nature, though in the beginning, the first section is going to be the SQL section,
21
and we will go through a little bit of coding.
22
So what I'm going to do is for that section, just give you a prompt and if you can share your screen
23
and just pull up a simple notepad and show me kind of how you would write things.
24
It doesn't have to be a specific SQL syntax or anything, but just pseudo-SQL, if that's a thing, like that would be good enough.
25
So any questions before we get started?
26
No, I'm guessing from the pseudo-SQL part, we're not going to like,
27
obviously no AI, no chat GPT on the side and also no reference docs.
28
Is that correct?
29
Yeah, that's a great question.
30
I think
31
if you need like a specific function that's pretty like you
32
know um you know like we want to simulate the work
33
environment as much as possible uh barring the chat gpt ai stuff
34
so um i think within reason you can definitely
35
but um feel free to let me know if you're going to look something up that would be helpful sounds good
36
okay sounds good all right
37
so we'll let's start with the first one you're working with the right share company
38
and you have access to the following tables and here's a snippet just to to be able to look at each table.
39
So that is the first table.
40
OK.
41
Let me know if that looks too weird.
42
That should be OK.
43
I'll go ahead and share the screen with the notepad.
44
Sounds great.
45
And here is the second table.
46
OK, cool.
47
So we've got two.
48
All right, let's see if we can format it a little bit.
49
It's like we're going to be fighting this a little bit.
50
But in general, so I see we've got this.
51
Okay, so trip, user ID, driver ID, city, distance.
52
And the prompt over here, we are a rideshare company.
53
That's right.
54
Cool.
55
Okay.
56
This...
57
Cool.
58
Okay.
59
And then the second table, shorter, thank goodness.
60
just three columns
61
cool okay user id sign update um user type and
62
so just understanding here so i know trip is your the trip driver is the the one who's working for Uber.
63
And so user here is really our customer.
64
That's right.
65
Okay, cool.
66
Yeah.
67
So yeah, you kind of nailed it on the dot.
68
There's two tables here.
69
One that shows the trip data and some of the metadata around it.
70
And then one is the users.
71
And actually, I want to clarify a little bit.
72
Let me give you another row of that user table.
73
Okay.
74
because the user type is relevant here.
75
It could be the writer, but it could also be the driver in that user data as well.
76
So you can feel free to paste in a second row.
77
Yeah, and thanks for doing the formatting.
78
I think it's actually helpful for me as well.
79
So I appreciate it.
80
Cool.
81
Great.
82
So here's the first question.
83
I'll paste it in as well so that you can have it while we're going through it.
84
So the first one is write a SQL query to calculate the average trip distance per city.
85
So obviously you don't have the whole data set here that you're looking at, but if you can just write a query that gets at,
86
you know, how you would calculate that.
87
Yeah, cool.
88
Okay, so when I write SQL, I immediately go, I know there's a select,
89
there's my from, and then we'll have some other group buys or wears.
90
So we're talking about calculate average trip distance.
91
So that's immediately like telling me, all right, so I want the average trip distance.
92
It looks like it's provided here on the trip data.
93
I'm going to make an assumption that the table name here for that is trips and the table name here is users.
94
Is that okay?
95
that yes that was a copy and paste there yes very good cool
96
so um all right so i want average trip distance
97
so we've got distance kilometers uh which is going to be from our trips table and
98
per city so per city if we're doing an average we're gonna need a group by we can do this
99
group by city here.
100
So I believe this should give us the average distance of trips by city.
101
And I guess we can add the city up top so that we know which distance connects to what city.
102
Okay, great.
103
Yeah, that looks good to me.
104
All right, let's move on to the next one.
105
Find the top five users who are writers.
106
We're talking about writers in this situation who spent the most in July of 2023, and be sure to include their total spend as well.
107
Okay.
108
All right, so we're thinking top five users who spent the most.
109
Okay, so explicitly riders here.
110
In July of 2023, it's like our time frame, include total spend.
111
So our select statement, so we eventually want rider,
112
riders, which is probably going to be users.
113
Yes, I just want to clarify that because it wasn't clear in the schema there.
114
User ID, it could also be rider ID.
115
That's what should connect to that user ID column in the user data table.
116
Gotcha.
117
Okay.
118
So why don't maybe let's start a little simpler.
119
Let's go with just select writers who spent money in July 23.
120
So if I want writers, so let me just select,
121
let me find all users who had a trip in July 2023.
122
Now this July 2023 data is in our trips,
123
so I'm going to need to join here into trips on maybe like trips.usrID.
124
and I should probably just make this a little bit more easier.
125
Sorry, that was a pretty quick jump down.
126
All right.
127
So from users, left join trips on trips user ID equals, so we are the user table, so users user ID.
128
So now we've got users trips.
129
So we can now say where – so I wanted to do July 2023.
130
So we've got trip dates here.
131
So we should be able to do trips.tripdate maybe less than – we are at 2023.
132
2023, so I want to be within 2023.071.
133
Pinterest.tripdate.lesson.081.
134
If I knew how many days were in July, I would use 731, but I don't remember if it's 30 or 31 days.
135
No, that seems perfect.
136
Okay, cool.
137
And one more thing here.
138
So users could be riders or drivers.
139
So we should probably make sure that we're always a driver.
140
Oh, sorry, rider.
141
Cool.
142
so so far looks like we've got users trips in july um
143
so we still need total spend so we just want top five um so that's that's a pretty easy limit five there
144
and we want total spend
145
so spend is probably price trips up price USD we want the sum of that so trip stock
146
cool so we're going to sum this up but in order to do an aggregation we'll need a group by. So let's group by.
147
We want user ID at the end of it.
148
I don't know if it'll get confused between the two tables user ID, so I'm just going to go with users.
149
All right, let's read through this one more time
150
so find the top five so that's limit five users users
151
or writers who spent the most we've got prices in july 2023
152
all right i think i think that's my final answer okay
153
yeah it looks great uh one question i had is is what you think those first five rows would look like
154
if you were to execute this at the moment.
155
What they would look like?
156
Well, I would have a user ID and the total spend on the,
157
for that user yeah for that writer oh
158
but we're looking for top five yeah well top five good good call
159
so that would be i guess we need an order by
160
and we want top and we want to do it by this column
161
so let's do ads maybe trip uh let's do total spend and then we can order by total spend sending
162
awesome all right yeah
163
that looks really good awesome thank you um we'll quickly work through um one more question
164
and this one just like let's um maybe we can talk through how we would do this
165
and you can fill in and write as much code as you can
166
but we have about like a minute more for this section
167
so we'll try to burn through this um for each city
168
rank drivers by total earnings in august 2023 uh okay
169
so you said we had a minute so i probably won't focus on coding too much and more on
170
explaining what i'm thinking so you want drivers by total earnings for each city so i'm thinking
171
we still have drivers we're gonna do a left join on trips to get our
172
city data and then we'll want to group by probably the cities and the drivers Yeah,
173
because this is kind of our tuple right here for each city drivers by total earnings.
174
And I noticed this rank here.
175
So I'm not too familiar with this,
176
but I'm thinking maybe like a rank by rank over partition with city.
177
Yeah.
178
And it looks like we're doing it over the total earnings.
179
So maybe I'll wear like some sum of total earnings.
180
Yeah.
181
Okay.
182
Yeah.
183
That sounds good.
184
That sounds like a very sensible approach to me.
185
So yeah, sounds great.
186
Awesome.
187
Okay.
188
Thanks for running through this portion.
189
I think writing code in an interview setting is never fun.
190
So thanks for humoring me through that.
191
We're going to move on to the next section.
192
And this, I will continue to paste questions to you in the, so that you can have access to it,
193
but you won't code as much.
194
So we're going into more of the stats section here.
195
So here is the prompt just so you started started um
196
and i'll i'll walk through this um so
197
that we can have it together
198
so did you uh right before you start should i will i need the notepad should i just like close
199
that yeah you can feel free to close it yeah okay cool yeah
200
okay so um here's the scenario um you're running an a b test for an online ad platform to evaluate
201
a new bidding algorithm so there's already one um in place
202
and you're evaluating this new one version b against the current one version a
203
and comparing on the two rates so this is ads so um just the rate at which you know
204
a user is looking at something and then whether or not they've clicked
205
so after running the test for a few days the results
206
are as follows group a is the control group it shows 100
207
and 1200 clicks out of 60 000 impressions uh a two percent click-through rate
208
and then group b is the test group had 1350 clicks out of the same number of impression 60k
209
at a 2.25 click-through rate so um we see that just by you know uh just by absolute comparison
210
there's a 0.25% increase.
211
So the marketing team is eager to roll out the new algorithm immediately.
212
And so here are the questions that we want to walk through regarding this.
213
This is mostly a discussion meant to kind of talk through some of the statistical concepts.
214
Okay.
215
First one, I wanted to focus on the central limit theorem here.
216
Why would the central limit theorem
217
be essential like how how does the central limit theorem play into and now analyzing these results
218
yeah so um so the central limit theorem uh as i understand it allows us to
219
take results from a sample and uh say and it kind of gives us this distribution
220
or it allows us to take the results of a sample and apply it on a normal distribution, even if the total population might be skewed.
221
So taking that knowledge and for these particular results, it looks like the test was run on 60,000 impressions,
222
which is a pretty good sample size, I think, unless you're like Amazon or something with millions.
223
but hopefully this is actually a big enough sample size
224
that we can we can now apply a normal distribution to this
225
and therefore make statistical decisions or run statistical metrics such as the z-score standard error
226
or etc okay yeah great awesome um and so regarding those statistical um
227
tests um that kind of segues well into the next question how would you determine whether this difference
228
that we're seeing is meaningful enough to justify rollout like how
229
would you just talk me through like what kinds of things you would look into
230
and um how
231
that would play into you know how the results of those things would play into informing the recommendation you might make
232
yeah i think um well we want to uh identify the statistical significance um so 0.25
233
is i guess our absolute lift in this case
234
and what we'd want to do is we're trying to define figure out the p value
235
and given some alpha such as like 0.05 if our p value given these inputs is much smaller than 0.05,
236
that would tell us that, hey, this test is actually statistically significant.
237
So that would be one sort of check mark I'd want to make sure of.
238
the other thing is on the cost side so even
239
if we've established that this is statistically significant we'd also want to make sure
240
that if we do deploy our new model that
241
it kind of the profit is going to be greater than the actual cost of deploying the model
242
and there are some metrics that we can use to compare okay yes
243
that sounds great um
244
so going back to the statistically significant um portion um how
245
would you help a marketer like let's say marketer has very
246
very like thin statistical background how would you help them understand like what
247
that means like statistically significant and the p-value and the 0.05 like
248
could you could you help create could you help provide like an intuitive understanding of why
249
that leads to um like hey this difference actually is meaningful That might be tough.
250
My understanding is that although we're seeing changes,
251
how would I frame it?
252
So what we're seeing is changes in a very small sample,
253
And we want to make sure that these changes will scale to the larger population.
254
And there are some times where we may see like a 0.25% in a small population.
255
But when we actually deploy it into the larger population,
256
that number could actually drop and be closer to 0.1 or 0.01 and we want to make sure by using the statistics,
257
using these central limit theorem,
258
these confidence interval tests, we're doing our due diligence to make sure that this 0.25% that we see in the sample,
259
it will be the same 0.25% or better in production.
260
Okay, yeah.
261
So if I'm understanding correctly, like the statistical significance, it points to some indicator that
262
that difference is real when we deploy it to the greater population
263
is that how is
264
that a good way of understanding kind of what you're saying there yes yeah okay okay yeah
265
that sounds great um
266
and then going into the cost side um how would you like how how would you um put
267
that into metrics um if you were to
268
so i think it was a really good call out
269
that you have to consider the rollout costs relative to the
270
potential lift it would bring um what what numbers would you
271
consider bringing to this marketing stakeholder group um as you're bringing
272
this analysis to them as you're talking about the rollout plan to them yeah um something
273
things that I might start thinking about, hey, what is our ROI?
274
So right now we have 2% CTR, which is click-through rate.
275
How does that translate into profits?
276
Right now it's just a number.
277
How would we translate that into a dollar amount?
278
So we might have some data about,
279
well uh maybe one percent click-through rate gives us some net profit of like a thousand dollars and
280
so an extra dot to five percent is like all right
281
we're getting i don't know the math like two hundred dollars let's just say uh
282
or it's probably 250 that we're getting so using that we might calculate our roi our return on investment
283
Another, that's definitely one metric.
284
The other metric I want to say is MDE,
285
which I'm forgetting the meaning.
286
Minimum detection effect?
287
Is that the one you're referring to?
288
i believe so yeah minimum detection um the definition is escaping me right now
289
apologies no no worries no worries yeah i think the roi
290
is a great call out i think it's a good place to start
291
and it's it's something that the marketing stakeholders will definitely understand
292
so i think it's a really good place to kind of like um um bridge the gap yeah so Okay, that sounds great.
293
Okay, so let's move on to the third question in this section, which is, what would break our ability to do this analysis?
294
So this is pointing back to the central limit theorem concepts that we had just talked about.
295
As you had mentioned, it kind of underpins our ability to do an analysis like this to begin with.
296
So if you could help me to understand, in what circumstances would we not be able to use those assumptions that we laid out?
297
Yeah.
298
I think there's a few items I'm thinking about.
299
One of them is the segmentation.
300
if we were to find out that, hey, we didn't have a good distribution of segments.
301
So for example, maybe we only got PC users on our A side and device users on our B side.
302
That's a clear skew, which would break that central limit theorem.
303
That's the one my mind goes to.
304
Sambo.
305
There's another one early peaking.
306
So this idea that the more times, if this is maybe the third or fourth time we're running this test,
307
this actually increases the likelihood that maybe this actually increases the likelihood that there was a false positive.
308
And the false positive being that out of the four times we ran it
309
and in the example the fourth time showed a significant like a statistical significance
310
and we want to avoid that because the other three were normal
311
so I'd probably start with those two yeah yeah those are really good considering ad as the industry here.
312
Are there any things that you can think of
313
that would create some potential weirdness or things that are maybe normal about the ad industry,
314
but would introduce gotchas and things to watch out for as we look at this data?
315
Yeah.
316
Okay.
317
So ad bidding algorithms. So probably
318
i think of ads and i just immediately think politics like
319
so there are certain very charged subjects uh that would um that could like skew the results one way or the other
320
versus a more kind of stable topic such as i don't know chicken eggs like who doesn't who doesn't hate chicken eggs.
321
So that could be one example where like now we've got this, people already have their own biases about the ads.
322
And so they might choose A or B based on their own personal biases instead of the actual test.
323
Yeah.
324
And the last question I wanted to follow up about this one is, What questions might you want to ask about how the data was sampled?
325
When you think about the control in the test here, there conveniently hasn't been given to you many details about how the data was sampled.
326
So what questions might you want to ask just to make sure you're covering your bases?
327
Yeah.
328
So in terms of sampling, there's a fancy metric, I think, called SRM.
329
But try not to dive into that.
330
Probably I'd want to ask, like, hey, so it's ad bidding on top of this.
331
I'd want to know kind of how they decided.
332
Was it a random sample?
333
Were these segments or were these divided on a particular customer segment?
334
How was the sample actually recorded?
335
Like, did we actually, when we were recording it, was there a network outage in between?
336
Was there some kind of external economic indicator or something that could have possibly
337
skewed the results during the case of the test.
338
Yeah.
339
So what I'm hearing from there is like, it sounds like you want to really make sure that the two samples weren't taken differently because of some situation,
340
some aberration that could have happened.
341
Yeah, I think those are really good ways to think about that. So awesome.
342
Okay, great.
343
We're making good progress on time here.
344
So we're going to move on to the next, the last section, the business case analysis.
345
So we're going to switch gears.
346
Yeah, we've been kind of switching gears each time.
347
But this one, we're going to go to food delivery.
348
So now you're a data scientist at a food delivery company.
349
And over the past quarter, you noticed a 15% drop in order volume in your largest market.
350
so leadership understandably wants someone to investigate and they they've asked you
351
and they've asked you to investigate and to recommend actions
352
so you have access to customer orders delivery times ratings marketing spend data you may
353
or may not use all this data
354
but that's what you have at your disposal at the moment
355
and so this this one is fairly open-ended
356
and again like would love to just have discussions around it and feel free to ask any questions for clarification.
357
And so we'll start with a very open-ended question, where would you start and why?
358
And yeah, feel free to just, I guess, put to words and articulate as much of your thought process as possible.
359
Yeah.
360
So, I mean, I hear 15% drop and I'm thinking about, am I going to lose my job soon?
361
but we've got the uh obviously that is a huge marker something
362
that i'm thinking about did this happen all at once was it a sudden event
363
or did this happen gradually over time i'd want to
364
probably focus on given that time nature looking at temporal data so for example
365
an easy one might be hey what does my delivery times look like over time what is my
366
looking at marketing spend data
367
so maybe did we have changes in marketing spend has it
368
been relatively similar I'd also might want to compare it not just
369
against I might also want to compare these trends against the same time period last year to see if like,
370
hey, maybe this is kind of expected for a business.
371
Yeah.
372
Yeah.
373
Okay.
374
That's great.
375
So let's, let's focus in on the temporal data.
376
So what hypotheses would you want to form as you think about this, this particular area?
377
yeah so uh well we're thinking food delivery um
378
so i i almost uh want to say has there been a change in uh drivers maybe
379
so the order history uh
380
so have the ratings kind been going down are we seeing a correlation between orders
381
and ratings that would be one of them marketing spend could be huge i mean
382
that is exactly how we get customers onto our platform are we losing like conversion rate there
383
and the third one is i kind of want to stick with seasonality
384
or in general some sort of external factor
385
and maybe we could look at like i would say like maybe customer order trends by year and like compare them
386
and see are we seeing the same sort of trends okay yeah um okay
387
so So let's just deep dive a little bit into one of these areas.
388
So the first one you mentioned, the change in drivers, looking at order history and ratings.
389
What are some of the, like, if you were to form like a testable hypothesis, where would you go with that?
390
you know, trying to figure out some, trying to detect some change in drivers using order history and ratings.
391
Yeah.
392
Drivers order history rating.
393
So, and we're trying to track order volume drop.
394
So driver history.
395
Something I think about is maybe there's a, subsection of drivers that is causing orders
396
or users who have ordered from them to to not order again
397
so because of because of an order
398
that came in from a particular driver will the user order again
399
or will they not so
400
that could be one sort of testable hypothesis system okay okay yeah um
401
and then depend uh so how would your strategy
402
or how would your thoughts on these things change
403
if you saw a sudden drop in 15 percent um in
404
this largest market versus is like a gradual drop over a quarter, let's say.
405
Yeah.
406
If it was sudden, I'm thinking there was probably a flip that was switched.
407
I think more like a flip that was switched,
408
like maybe there was an update that happened right on that time because we're thinking about like, hey, it's a food.
409
I'm presuming an app here maybe a website i'd track it to a singular event versus if it's something gradual
410
um i'm thinking more uh long larger trends maybe there was a tick tock someone made
411
and it's slowly blowing up and as these more and more users are seeing it it's affecting their ordering habits
412
or maybe there's an economic factor like uh i don't know maybe the stimulus checks have stopped
413
and so people just aren't ordering as much food um
414
so yeah okay great great um and yeah
415
so when it comes to like the gradual kind of seasonal trend changes that are happening,
416
let's say we take the hypothesis
417
that you mentioned of maybe there's a subsection of drivers that are causing orders to not happen again because of something.
418
Maybe it's a bad experience.
419
Maybe it's...
420
Delivery times.
421
delivery times maybe yeah maybe it's there's something there
422
and you see a correlation between that and uh ratings um so
423
that what is the metric like let's get really granular let's
424
try to get as granular as we can what is the actual metric you would want to calculate
425
when you are trying to test this hypothesis right so we're thinking like some sort of success metric um
426
So the user, so subsection of drivers and our success is whether a user actually orders again.
427
So I think, I don't know what the name of it would be, but we'd specifically be testing,
428
hey, does the percentage of users that order again, given a particular driver?
429
okay and that would be all time so if a user how like how much time can elapse for a particular
430
user to order again uh well we should probably uh define like some baseline
431
so before this period what was the average time between orders and then use that as the interval in our metric.
432
Yeah.
433
Yeah, that makes sense.
434
And when you say given a particular driver, so percent of users that order again given a particular driver,
435
so is this one metric for each driver or how would you aggregate if you would?
436
Yeah, so I'd want to be on the per driver level,
437
but obviously there's a number of drivers in the database.
438
So we would need a way to kind of compile this metric so we can get per average driver.
439
um i want to say average as well
440
if we average the value well
441
so we have average there's also uh median is what i'm thinking of uh
442
and i'm trying to see like hey what would make more
443
sense i guess i'd have to know a little bit more about the data to make a decision about median or average.
444
Yeah.
445
So between median or average, do you have something like a criteria in mind for when?
446
So how would the data need to look for you to use average versus median, would you say?
447
Yeah.
448
So if the distribution of the per driver metric is closer to normal,
449
normal, I'd want to use average versus if it was skewed.
450
And especially if we have large outliers, I'd want to use median instead.
451
Yeah.
452
Yeah.
453
That sounds good.
454
And then just maybe one last thing there is what
455
if the distribution looked like a really long tail where you have a couple of
456
drivers who are taking up like 80% of the orders
457
and then like the average driver is taking like very few orders okay
458
so in the case that in the case that there's a subsection
459
so there's we could rely on median uh we could also decide to
460
chop off some of those outliers just outright in order to get a better performance.
461
We'd have to be careful about what type of error.
462
We don't want to cut off too much, but we might have a threshold
463
so that we're not polluting our data and our statistics with this group that maybe only happens 5% of the time.
464
Yeah.
465
Okay.
466
Yeah.
467
And then last question I want to ask here is, so that seems like a good metric overall to use.
468
Are there any like guardrail metrics or any like counter metrics you'd want to measure
469
while looking at this just so that you're not,
470
you're kind of being cognizant of any potential tradeoffs that the results of this analysis can cause?
471
Yeah, as we go through this data, I'd want to make sure that we're not affecting latency, for example.
472
As we're doing these tests, we want to make sure we're not adversely affecting the number of orders coming in.
473
We don't want to start the test and orders don't dive even further.
474
So those are like two examples of guard worlds where we don't want to affect our main product
475
while we conduct this test.
476
Yeah.
477
Yeah.
478
So it sounds like the experience of the product
479
and then making sure that the business is staying stable while we're looking through.
480
okay yeah sounds good okay yeah um so those are all the questions
481
that i had for you today um so thank you
482
so much for taking the time to do this interview
483
and uh yeah it was really nice to meet you yeah no i appreciate it thank you that's that's the
484
that was the mock interview um uh uh
485
Khalid what did you think about the candidate how was their performance to you uh i was okay It was definitely,
486
I think, they had a strong start in the beginning.
487
Kind of as we started getting into more of the statistics side, I felt there was a bit more of a general,
488
like a high-level view.
489
And maybe there was a little bit of struggle.
490
I know he wasn't able to go into MDE a little bit.
491
But something to note there, I don't know how that'll play in the larger role.
492
Yeah.
493
Johnny, what do you think?
494
Does that track what Khalid's saying?
495
Yeah, I would totally agree that the candidate was really strong in the beginning with the sequel.
496
I think he did a really great job of different aspects of that section overall.
497
I think so that I can kind of dig into it section by section since.
498
Yeah, that'd be great.
499
Yeah, I think so for the first section, the first two questions, the candidate nailed it out of the park.
500
The only thing there was that he used a left join where I think we could have used an inner join, which is more efficient generally.
501
The other thing high level to note is that with these tables,
502
I generally like it's helpful when the candidate pushes the interviewer to ask about whether
503
or not the rows for each table are unique.
504
Because that's how the SQL should be written out.
505
I think in the understanding the table section of that section, I guess, like just general tip,
506
pushing the interviewer to give you as much information about the tables, you know, like, and unique, whether or not the row, the ID is unique, you know,
507
making these, making it really explicit what assumptions you're making there.
508
is really helpful because at the end of the day, if you're a SQL practitioner, you know that like all of the bugs that you have,
509
not all of it, like 80% is probably in the joins and these kinds of places.
510
So just knowing to ask those questions is a good, in my experience, I think that's a good way to conduct those interviews.
511
The third question was not very well worded in my opinion.
512
So that's a doc on the interviewer, I would say.
513
But it was meant to get at the window function, like the rank window function.
514
And I think the candidate did a good job of articulating how that should be written out.
515
But we ran out of time to be able to get to that.
516
So part of that is I'm not sure if there were too many questions for the amount of time we've allotted.
517
So, yeah, I'm not sure about that.
518
Did it feel realistic?
519
Khalid, did the interview, just in terms of surface area covered, did it feel realistic for you?
520
Yeah, it did.
521
And something that you mentioned, I think, on the first one, I was also hesitating to, I didn't know if I should bring up,
522
I think the candidate was focused a lot more on solving the problem.
523
And it's hard sometimes to like, hey, what sort of how do you engage the interviewer
524
while you're building the SQL query and you're kind of like thinking your thoughts aloud
525
when you mentioned unique values another question
526
that popped into my head was should I bring up missing values
527
or null values how are we going to track
528
that in these queries yeah so that's probably another question that could have been asked Yeah, that's really good.
529
I think that would, yeah, those are generally, in my experience, pretty helpful.
530
Yeah, moving on to the stat section, I thought overall it was pretty good.
531
Khalid said that this was the weakest section.
532
Was this Khalid's weakest section?
533
Was this the candidate's weakest section?
534
Yeah.
535
I would say if I had to choose, yes, because I think the first section was strong and the last section was strong.
536
Okay that um here i think um the candidate did a great job of explaining the central limit theorem um
537
at the first question um and
538
and i think also like i i had strong signal
539
or i guess the interviewer had a strong signal um on um whether
540
or not the candidate would be able to implement like a test to be able to compare and analyze the results.
541
So I didn't really have any questions there.
542
I think where I found a little bit of a gap was
543
when we started to talk about maybe the more intuitive explanation for statistical significance.
544
And that's relevant because generally,
545
like, I think more experienced data scientists will have to interface with non-technical stakeholders quite a bit and so
546
that actually is is a pretty relevant skill set to test for in an interview setting like this
547
and so um being able to articulate like hey it's
548
because intuitively you know there's this distribution we assume about these two groups
549
and if they're both coming from the same group then the difference would be like not statistically significant
550
but if they're actually coming from two different distributions then it is statistically
551
and that's what this test is measuring so that they can see that, oh yeah, like this difference is a real difference.
552
You know, that, that kind of explanation is what marketers would be able to, you know, understand a little better.
553
So that, that I wouldn't necessarily dock a candidate for that one, but I think, I think it's like a seniority level,
554
like experience level signal, I would say.
555
Just as your own experience as an interviewer Khalid, you probably have done these kinds of of stats like questions
556
before do you can you can you um can you speak a little bit to this point
557
that johnny just made about like that kind of technical
558
that kind of communication to non-technical stakeholders
559
that can make it make sense for them can you speak to that is
560
that in your mind a seniority signal as well uh yeah
561
i i do want to like be up front here
562
that i am i have done data science kind of early on in my career
563
and then have swapped into full stack and so I'm kind of now coming back in.
564
Cool.
565
Which is why that maybe you might understand in the interview that I kind of
566
still remember the high level ideas but I was at that level of time and the position that I was at
567
I was doing a lot more implementation.
568
And so it shows even in the interview, as you can see, where I get what I need to use and how to do it,
569
but not so much.
570
I seem to be missing some of that kind of like foundational stuff of, okay, how do you break down this idea for other stakeholders?
571
It's okay.
572
I mean, honestly, Johnny was like super accurate.
573
He's like, yeah, I can tell he knows how to do it, but I'm not sure that he's done it before.
574
Yeah, but everyone can relate to this.
575
That's doing a career pivot or is getting back into data science.
576
Do you two have any recommendations for people to strengthen that skill, whether they're rusty and getting back into DS or just like pivoting into DS as a career?
577
I feel like there's such a difficult balance right now.
578
And the one thing I'd love to talk about is like the roles, right? as an applied ai
579
or applied ml you don't really need to know some of this foundational level stuff i would argue
580
that you at all you need to know is
581
that i'm making an api call to chat gbt or an agent
582
and that'll help you get the job done whereas something like
583
a data science role is um maybe johnny can speak more
584
to this as a data science role you really have to understand some of the statistics
585
and the math part in order to make a better decisions
586
and really be able to do the job properly would you would you agree with
587
that johnny yeah i totally would
588
and um yeah i i think um for the most part
589
like it depends on the kind of data science role as
590
well um data scientist is like kind of a broad term in my experience like company to company
591
and it's usually some mix of
592
the analysis slash almost like consultant like product manager like kind of type of person
593
and then the engineer type person who's like implementing stuff
594
and like actually coding or whatnot
595
and the the scales of like what's more like prominent in
596
the role it's kind of differs company to company um where
597
i'm at right now at um at meta it's it's very much like these communication skills are very important
598
and it um they expect the data scientists to kind of
599
stand on their own feet in terms of being able to drive conversations
600
and influence and things things like
601
that um other roles it's more like your your engineering work is kind of where where um they're focused on
602
so i i really think it depends there um i think i just had a like having said
603
that i I think being able to communicate to non-stakeholders generally
604
is like a seniority experience level kind of skill set to have.
605
So it's always good.
606
But like when you need that, like how important is that for the hire specifically?
607
Like I don't that's why I said it's more like a seniority signal than a hire versus not hire, I think.
608
Awesome.
609
Awesome.
610
Really, really cool.
611
What about that last section?
612
The business case?
613
Johnny, what are your what are your notes here?
614
yeah um i thought it was really good um on on the whole
615
if i had to stack rank just for clarity i would say the school section was the best
616
and then this was the second and then the stats section on
617
i think for this section what i generally look for is structured thinking
618
and ability to start out with kind of like an exploration level
619
and then hone into hypotheses and as we go down be able to actually develop the metrics and maybe even a test,
620
design a test to validate those hypotheses.
621
So on the whole, I think that the candidate did it really well, I think.
622
But here also, like a measure of deeper experience would be
623
if the candidate can go end to end with that on their own and kind of drive the conversation forward.
624
I think the interviewer had to do a little bit of
625
kind of through the questions kind of shaping the answers along the way.
626
So I guess what I mean by that is, so I think up till like what hypothesis do you have for this drop,
627
like that question, like it's okay.
628
But then from there on being able to move to hypotheses that make sense.
629
And then like kind of going deeper and kind of driving the whole flow of the analysis in the conversation,
630
that's something I would look for as an interviewer, I guess.
631
And so I think my experience was that when I asked a question, there were answers that were pretty solid.
632
But I think
633
if there was a little bit more of being able to present the whole picture in a fuller way
634
and to kind of drive and kind of, you know, maybe stop at checkpoints to check with the interviewer to see if this is the right direction to go into.
635
That's something that generally shows like comfort level with the structured thinking
636
and experience having done an analysis of these kinds of things, of these types.
637
Amazing.
638
Amazing analysis.
639
Before we get Johnny's decision, any other things that either of you wanted to say point out about what was interesting or cool
640
or unexpected about this interview uh i i didn't realize how much stats would again this is like coming from someone who's
641
getting back into data science i didn't realize the stats portion
642
was very intriguing to me it's like wow this is a
643
lot of statistics um kind of my thoughts sort of my
644
expectation coming in was like oh this was going to be a design a recommendation system for google
645
and you go start from the beginning of the pipeline you do the some of the high level building blocks
646
and then you dive deeper but the reality is is
647
that it seems to be a little bit more targeted on the statistics especially and then the metrics portion
648
that was I think one of my biggest notice the other
649
one really interesting to hear Johnny say like I should be the one driving
650
or the candidate rather should be the one driving the conversation
651
right that's definitely a tip that I'm going to keep in the back of my head
652
because I guess my mindset was definitely here are the options it's up to you as the stakeholder to decide
653
which one of these paths you think we should go down.
654
So that's definitely a tweak that I'll probably take with me.
655
Yeah.
656
And I just wanted to jump in to say, I think like to your first point, that yeah, it might be
657
that this is more of a certain type of data scientist interview
658
that we're doing right now relative to the maybe more AI engineer type roles,
659
it sounds like, that you might be maybe perhaps aiming for as well.
660
So yeah, I would take that with a grain of salt, like in that, like it might not be as stats heavy as,
661
you know, some of those AI engineering roles, for example, might not be as stats heavy as like a data scientist role.
662
So I don't really know, like a lot of those roles are like brand new too.
663
So yeah, that's something I just want to look at.
664
All right.
665
And Johnny, your decision for this candidate?
666
Yeah, I think for like an entry to mid-level, definitely a higher.
667
I think at the senior level, I think there were some things that I saw as gaps.
668
But overall, I would have been happy to have somebody like this on the team.
669
All right.
670
Thank you.
671
That's all we got.
672
Well done.
673
And thank you for watching.
674
you

Why Practice Speaking with This Video?

Practicing English speaking with real-world dialogues like the one in this video is a game-changer for learners. The conversational flow—mixing casual small talk, technical instructions, and clarifying questions—mimics real-life interactions, making it perfect for IELTS speaking practice and everyday fluency. Unlike scripted lessons, the natural pauses, follow-ups, and adaptions in the conversation train your brain to think on your feet, a critical skill for exams and professional settings. Plus, using videos to learn English helps you pick up context clues and intonation, which are key to sounding confident. Try the shadowing technique here: repeat lines immediately after the speakers, matching their pace and tone. It’s a proven way to boost pronunciation and fluency fast.

Grammar & Expressions in Context

The video is packed with practical grammar and phrases you can use today. Here are 3 key structures:

  • Conditional Clauses for Clarity: "If you need a specific function... you can definitely [look it up]." This "if + present, can" structure is ideal for polite requests or problem-solving, common in both IELTS and workplace conversations.
  • Casual Small Talk Starters: "How's it going?" and "Nice to meet you" are simple but essential for breaking the ice. Notice how the speaker uses humor ("You're not related to Johnny Bravo?") to build rapport—this softens technical interactions, a skill worth practicing.
  • Logistical Phrasing: "We're going to be doing the interview for 45 minutes. It's going to be a technical round..." Using "going to" for plans and "it's going to be" for descriptions helps organize information clearly, a must for explaining processes or answering IELTS cue cards.

Common Pronunciation Traps

Even advanced learners stumble over certain sounds in this dialogue. Watch out for:

  • "Logistics": Pronounced /ləˈdʒɪstɪks/, not /loʊˈdʒɪstɪks/. The first syllable is short, like "luh," not "low."
  • "Pseudo-SQL": The "pseudo" part is /ˈsuːdoʊ/, with a long "u" sound (like "sue-doh"), not "pseh-doh."
  • "Conversational": Stress the third syllable: /ˌkɑːnvərˈseɪʃənl/, not /kɑːnˈvɜːrseɪʃənl/. Misplacing stress can make the word hard to understand.

Practice these with the shadowing technique—a shadowing site or app can help you compare your pronunciation to the speakers. Mastering these details will make your English sound more natural, whether you're prepping for IELTS or a professional interview.

¿Qué es la Técnica de Shadowing?

Shadowing es una técnica de aprendizaje de idiomas respaldada por la ciencia, desarrollada originalmente para la formación de intérpretes profesionales y popularizada por el políglota Dr. Alexander Arguelles. El método es simple pero poderoso: escuchas audio en inglés nativo y lo repites en voz alta de inmediato, como una sombra que sigue al hablante con solo 1-2 segundos de retraso. A diferencia de la escucha pasiva o los ejercicios de gramática, el shadowing obliga a tu cerebro y músculos de la boca a procesar y reproducir simultáneamente patrones de habla reales. Las investigaciones muestran que mejora significativamente la precisión de la pronunciación, la entonación, el ritmo, el habla conectada, la comprensión auditiva y la fluidez al hablar, convirtiéndola en una de las metodologías más efectivas para la preparación del IELTS Speaking y la comunicación en inglés en el mundo real.

Técnica de shadowing: lee la guía completa paso a paso →