Shadowing Practice: Why Shopify Moved Inventory Reservations from Redis to MySQL - Learn English Speaking with Video

Creating lesson...
1
Hey everyone, today I want to break down one of Shopify engineering's latest blog posts, where they talk about how they replaced Redis with MySQL in their inventory reservation system.
2
I thought this post when I read it was super interesting, not only because it overlaps with a lot of the patterns that we teach, it's similar to the Ticketmaster breakdown that we have, but because it moves in the opposite direction of maybe what's most common,
3
which is that teams start on MySQL or Postgres, they realize that there's some table or some subsection of their system which doesn't scale to the throughput that they're after,
4
and they move that part out into Redis.
5
In this case, Shopify ended up taking the opposite approach.
6
And it's good evidence of how much MySQL can really handle scale, MySQL or Postgres, of course.
7
And they did some interesting things in the process.
8
So without further ado, let's break it down.
9
Let me start up by quickly defining the problem that Shopify needed to solve, at least from a product perspective.
10
So when a buyer goes to purchase an item on Shopify, Shopify has to guarantee that that item is actually available for purchase.
11
If they get this wrong, it can fail in one of two different directions.
12
The first failure scenario is that two buyers could think that they both purchased the same last unit.
13
And so now both of them think that they have an item coming, but one of them is going to have to have a cancellation and a refund put back on their card later, as well as maybe an apology email.
14
The other direction is that you have a buyer who thinks that an item is sold out, but it actually isn't.
15
And so now the merchant is losing money because they could have sold that item, but Shopify is incorrectly telling you that it's unavailable.
16
And at Shopify scale, while these things don't happen often, they can compound really quickly.
17
Now the product solution here is one that you all know, you've all used it.
18
It's these reservation systems in a transaction at checkout.
19
Basically, you're going to put something in your cart and then Shopify is going to hold that aside.
20
It's going to reserve it for you for some period of time
21
while you continue to type in your credit card information and try to make your payment.
22
If you complete your payment, then that unit is going to come off the merchant's inventory for good.
23
It's yours now.
24
But if the payment fails or if the clock runs out, you use too much time, then the inventory is going to go back into the pool so that somebody else can have it.
25
That whole thing is a reservation system.
26
And from a technical perspective, for years, Shopify had been using Redis to do this.
27
They had a Redis instance for which every single Redis key was a single piece of inventory.
28
So the blue hoodie.
29
And then the count or the value of that key was how much was currently available to sell.
30
So if a user needed to reserve an item, you would decrement that count.
31
And then if they didn't end up purchasing, you would need to increment that count back.
32
The whole time, the source of truth still lives in MySQL.
33
And so let me walk you through what this would look like.
34
Imagine there are 10 hoodies available.
35
So both Redis and MySQL agree on 10.
36
Then a user wants to check out, so they put a hoodie in their cart, which triggers a reservation, and it decrements Redis down to nine.
37
But MySQL stays at 10.
38
Once they complete their purchase, MySQL also decrements down to nine, and the two are in sync again.
39
That part's pretty straightforward.
40
Now, what I found interesting when reading it is
41
that the blog doesn't actually touch on what I think is the most interesting part, which is how do you handle the case where the user doesn't complete their payment.
42
Maybe time expired, maybe they shut their laptop, walked away, you know, whatever it may be.
43
In these cases, you need something that's gonna trigger that increment back to Redis, basically to put this thing back in the pool.
44
Like I said, the blog doesn't cover how, so this is just my guess.
45
But let me walk you through how I think this would happen and what is most likely happening under the hood.
46
And that's that in Redis, they probably have two data structures.
47
They've got the raw current of the remaining inventory that we're incrementing and decrementing.
48
But then alongside it, they probably also keep a sorted set, which is basically just a heap for those who aren't familiar with the sorted set.
49
Now, every outstanding reservation, two things are gonna happen.
50
It's going to decrement the count, but it's also going to add that reservation to the sorted set where the score is the time that it expires.
51
In doing that, now whatever is sitting at the top of that heap, the top of that sorted set, is always the reservation that's due to expire next.
52
It can pull off the top of the heap, make sure that that thing is already expired.
53
If it is, get rid of it off the heap
54
and then go and increment that count of how many things are available to be reserved.
55
Now, this cron job is going to work great for the majority of cases.
56
But if you've read any of our contents, then something might be coming into your mind, which is that if something's selling really, really, really fast, then you don't have to wait for that cron job to add things back to the pool.
57
Even if the cron job was running every 10 seconds, you would have a user who's told that there are no hoodies available, when in reality, there should be a hoodie that's back in the pool.
58
The cron job just hasn't run yet, right?
59
There's a 10 second lag there.
60
So what you could do to fix this, and this is probably what I would do, is that you could tie this all together in a single atomic transaction.
61
Basically, before you take a unit out of the reserve pool, you would look at the head of the heap for anything that's already expired, remove those from the heap, increment them back, and then run your own decrement.
62
So say the user comes in and they reserve a hoodie.
63
The count would then go from 10 to nine and they would add their reservation to that heap.
64
Now imagine that some time goes by and their time expires.
65
Then another user comes to check out and what they do is that they go first check the head of that heap.
66
They see that there is a reservation that's already expired.
67
So they remove it from the heap and they run the increment.
68
So we're back up to 10.
69
And then they run their own decrement.
70
And all of this can be part of what's called a Lua script in Redis.
71
So that full chunk happens atomically, just like a transaction in a database.
72
If you do it this way, then you never have the problem where you're turning away a buyer because of a reservation that's already dead, right?
73
The expired stuff gets cleaned up right at the moment somebody actually wants it.
74
And the cron job is just there maybe as a backstop to clean things up periodically in the background.
75
Okay, so back to Shopify.
76
This is what was running in production, or at least our best guess of what was running in production.
77
and it's a totally reasonable design.
78
The reason that you reach for Redis in the first place is that it can handle really high
79
read and write throughput on a single key.
80
If that count were just a row in MySQL, then every checkout for that item would be fighting over the same row.
81
And at scale, that's where things start to really slow down and back up.
82
But the problem with this introduced, and this was the motivation for why they wanted to evolve into this new system we'll go into in a second, is that you now have two systems that have to be kept in sync, Redis and MySQL.
83
And so when a payment succeeds, you deduct that inventory from MySQL, but then you also have to go over to Redis and you have to pull that reservation off of the heap
84
because if you were to leave it sitting there, they would eventually expire and then the cron job
85
or whoever came next would increment it back to the pool of units that are still available, right?
86
So that's two writes in two different system and there's no easy way to wrap those in a single atomic step.
87
And so if there's ever a crash or a failure between those two things, then you can be left in an inconsistent state where the two disagree on how many things can be sold.
88
So in an ideal world, you would love for all of this to exist in just a single data store as part of a single transaction.
89
And in fact, this is exactly what Shopify had originally tried.
90
The post is a bit terse about this, but it says that the earlier attempts would fail and that the single row with a quantity column couldn't handle the contention.
91
So if you could imagine that in MySQL, they have two different tables or two different rows, one that has the reservation count and one that has the actual total inventory count.
92
Take what was in Redis and put it in MySQL.
93
Now every single checkout for that hoodie is an update competing for the same row.
94
And MySQL hands out a row lock to one transaction at a time.
95
So your buyers just end up in a single file line
96
waiting for whoever is in front of them to end up committing.
97
So the next natural thing that people tend to do is that you try to take that single row, maybe it's a row of quantity 10, and you would turn it into 10 distinct rows.
98
This is a really logical thing to do because it's breaking up the thing that you're competing for.
99
Now not everybody is competing for a single row.
100
They're competing against the 10 rows.
101
But the issue is what happens when you actually go to write that query.
102
So you end up doing a select for update statement to lock the row yourself.
103
But then every one of those queries is still grabbing just the first available row.
104
And so the second buyer locks up behind the first one, the third one locks up behind them, and you haven't actually solved anything.
105
There's all these rows under that they could go grab, but they don't know that, right?
106
Now the big breakthrough here, and this is maybe you know, one of the top three most interesting things from the Shopify blog is
107
that this all changed in 2018 when MySQL 8 shipped a new feature called Skip Locked.
108
Now, when your query is scanning for a row to lock and it hits one that somebody already has, instead of just waiting there like it did before for the transaction to finish,
109
it just steps right over it and it keeps looking for the next free one.
110
So now two buyers who are going for the same hoodie, they can land on those two different rows like we wanted and both of them can commit at the same time.
111
Now, the new simple flow for five hoodies is going to look something like this.
112
Instead of one row with five on it, you've got five rows sitting in a reservation table and one row for each hoodie.
113
Now when a buyer goes to checkout, they're going to open a transaction.
114
They're going to grab one of these rows with four update skip locked this time.
115
They're going to pull it out of that pool.
116
They're going to write a hold saying that that unit belongs to them at checkout and they're going to commit.
117
So now you only have four left.
118
If the payment does go through, then you deduct that hoodie from the merchant's inventory and you can retire the hold.
119
And both of those two things can happen as one transaction in the same database.
120
If the payment fails, you delete the hold and you put the row back into the pool and you're back to five.
121
But very interestingly, this introduces a new problem, right?
122
When you're talking about millions of merchants who themselves have thousands of items, each of which have hundreds of thousands or tens of thousands of units of inventory per item,
123
then the row count just got huge.
124
We used to have a single row per item, even if there were 10,000 units.
125
Now that's 10,000 rows.
126
Multiply that by all these different merchants and you're looking at millions of millions of millions of rows.
127
What's worse is that almost all of these rows are just sitting there doing nothing.
128
A merchant might have 10,000 hoodies in the warehouse, but they're never going to have 10,000 people checking out at the exact same moment.
129
And so you're carrying all these millions of idle rows per merchant to serve maybe a few dozen reservations
130
that could be competing at the exact same time, which just means that you have a bigger database, which needs to be replicated, backed up, maintained, and of course paid for.
131
So what Shopify ended up doing was something I thought was pretty clever.
132
They capped it.
133
So now no matter how much inventory a merchant actually has, there are never more than a thousand rows in that pool at any given time for that given item, right?
134
And then it just gets refilled out of the merchant's real inventory as it starts to drain.
135
So you have a thousand there.
136
If a bunch of people reserve and buy things and it ends up going down to 500, whatever, then we can come in, fill it back up to a thousand.
137
That thousand in effect just ends up being how many hoodies can be reserved at once.
138
That's like your contention number.
139
We'll be able to survive a blast of a thousand people trying to buy all at once.
140
But the real count still lives in the inventory table where it always has.
141
And the pool is just sized for a concurrency that's reasonable as opposed to millions, right?
142
As for the refill, that ends up just happening on the reserve path.
143
And so during a big enough flash sale, you could drain that pool completely, but then a checkout goes to grab a row and there's nothing there.
144
And that request does the refill itself before then retrying to pull something from the pool, right?
145
That thousand number seems a little arbitrary, but they chose it just by looking at peak reservation rates.
146
So they saw the most that ever comes in was maybe something like a couple hundred.
147
Let's add a buffer, make it a thousand.
148
That's how they ended up at that number.
149
So what happens now when a reservation expires though?
150
The post doesn't say this in detail, but it gives us enough to be able to reverse engineer probably what's happening
151
because it does tell us that the replenishment process refills the pool from the inventory ledger.
152
So when that replenishment runs, it has to answer one question.
153
How many units should be reservable right now?
154
And that number is just the merchant's inventory
155
that hasn't sold yet minus whatever is still held by a reservation that hasn't expired.
156
And then of course, cap it at a thousand.
157
So let's put some numbers on it again with an example.
158
Say a merchant had 800 hoodies that haven't been sold and 200 buyers are sitting in checkout with live reservations.
159
Well, 800 minus 200 is 600 units that could be reserved right now, which is under the cap of a thousand.
160
So 600 rows is what belongs in that pool.
161
If there are 400 rows, already in the pool at that moment, then the replenishment would just insert 200 more rolls.
162
Now imagine that 50 of those buyers just shut their laptop and their reservations end up expiring.
163
Well, the nice thing with the system is
164
that nothing has to go clean that up because the next time the replenishment runs, only 150 reservations are still live.
165
And so the number it lands on is 650 and those 50 units are back in the pool.
166
And that's basically just the same inline trick that we were talking about in the Redis scenario, where the buyer who needs the inventory is the one who goes and reclaims it,
167
who refills the pool, as opposed to waiting for some external cron job to do it on a schedule.
168
Okay, so just to wrap up, a couple things here stood out to me that I want to make sure I emphasize as your takeaways.
169
First, most teams assume that MySQL and Postgres can't handle their scale, so they pull the hot paths out into an in-memory store to shed that load.
170
And I want to be clear, that is a completely reasonable thing to do.
171
Redis was the right call for Shopify for years, and their post says outright that it handled their throughput great.
172
What ended up pushing Shopify over the edge was not the scale.
173
It was at the reservation and the inventory lived in two different systems.
174
And so whenever the payment succeeded, they had to do two writes that couldn't commit together.
175
One lesson for you to take away is that if two things have to be atomic, then they really have to live in the same data store.
176
The other big takeaway is that MySQL and Postgres can handle way more throughput than people think nowadays.
177
The chances are your site, whatever you're building, is not doing the same volume as Shopify.
178
And so if they can handle this sort of volume, there's a really high chance you can too.
179
And it's a totally reasonable place for you to start.
180
The last thing I'll emphasize is that Shopify first ruled out MySQL and they were reasonable or right to do so.
181
Skip locked didn't exist yet.
182
The database then changed, not their reasoning.
183
And so if you're a team that made decisions a long time ago, it's worth coming back and revisiting the technologies to see how much they've evolved
184
and whether some of the complexity that you added to a system could now be reversed into something much more simple.
185
And so if you enjoyed this, check out the other blog posts that we've written about so far, and we'll have a lot more coming soon.
186
Take care.

About This Lesson

You're practicing English with "Why Shopify Moved Inventory Reservations from Redis to MySQL" using the Shadowing technique — a method originally developed for professional interpreter training.

Focus on sounding like the speaker — not just repeating words. With 15–30 minutes of daily practice, you'll build real-world speaking confidence.

What is the Shadowing Technique?

Shadowing is a science-backed language learning technique originally developed for professional interpreter training and popularized by polyglot Dr. Alexander Arguelles. The method is simple but powerful: you listen to native English audio and immediately repeat it out loud — like a shadow following the speaker with just a 1–2 second delay. Unlike passive listening or grammar drills, shadowing forces your brain and mouth muscles to simultaneously process and reproduce real speech patterns. Research shows it significantly improves pronunciation accuracy, intonation, rhythm, connected speech, listening comprehension, and speaking fluency — making it one of the most effective methods for IELTS Speaking preparation and real-world English communication.

Shadowing technique: read the full step-by-step guide →