跟读练习: How Does a URL Shortener Work? - 通过视频学习英语口语

正在创建课程...
1
Everyone wonder what happens when you click a tiny URL?
2
You know, those short links like bit.ly and tinyurl, they somehow know exactly where to take you.
3
Today, we're going to reverse engineer how URL shorteners actually work.
4
And trust me, there's way more going on behind the scenes than you might think.
5
So what exactly does a URL shortener do?
6
Simple, it takes a massive URL like this Amazon link that goes on forever and turns it into something short and clean.
7
Click the short version and boom, you end up at the same place.
8
But here's where it gets interesting.
9
Let's say you're building the next bit.ly.
10
How many URLs do you think you need to handle?
11
Try 100 million new URLs every day.
12
That's over a thousand new short links created every single second.
13
And since people click links way more than they create them, you're looking at over 10,000 clicks per second.
14
Now think about storage.
15
Over 10 years, that's 365 billion URLs you need to keep track of.
16
Just storing the URLs themselves would take 36 terabytes.
17
So the real question becomes, how do you even generate these short URLs?
18
This is where the math gets fun.
19
Short URLs can use numbers and letters.
20
That's 62 possible characters, 0 through 9, lowercase A through Z, and uppercase A through Z.
21
But how short can we make them?
22
Let's work backwards.
23
We need 365 billion unique combinations.
24
With one character, you get 62 possibilities.
25
With two characters, 62 squared.
26
That's about 3000.
27
With three characters, 62 cubed. About 238,000.
28
Keep going you get to 7 characters, 62 to the 7th power.
29
That's 3.5 trillion possible combinations, way more than we need.
30
So 7 characters it is, but how do we actually create them?
31
There are two approaches, and they couldn't be more different.
32
The first approach, just hash the long URL.
33
Take any hash function like md5, run the long URL through it, and you get back a long string of random characters.
34
The problem?
35
That string is way too long.
36
Even the shortest hash gives you 32 characters when you only want 7.
37
So you take the first 7 characters and call it a day.
38
But wait, what happens when two different URLs give you the same first 7 characters?
39
You've got a collision.
40
Now you're stuck.
41
You have to try again with some variation of the original URL until you find 7 characters that nobody else is using.
42
Every time you want to create a short URL, you have to check if those 7 characters are already taken.
43
That's a lot of database lookups.
44
The second approach is way more elegant.
45
Instead of hashing, you just count.
46
Every time someone wants to shorten a URL, give you the next number in sequence.
47
URL number 1, number 2, number 3, and so on.
48
Then convert that number to what's called base 62.
49
Here's how that works.
50
Let's say you're on URL number 11157.
51
To convert to base 62, you divide by 62 over and over, keeping track of the remainders.
52
11157 divided by 62 is 179, remainder 59.
53
179 divided by 62 is 2, remainder 55.
54
2 divided by 62 is 0, remainder 2.
55
Now read the remainders backwards.
56
2, 55, 59.
57
In base 62, 2 stays as 2.
58
55 becomes T.
59
59 becomes X.
60
So URL number 11157 becomes 2TX.
61
Your short URL is tinyurl.com slash 2tx.
62
No collisions, no database lookups to check if it's taken, just clean math.
63
To trade off, you need a way to generate unique numbers across multiple servers.
64
That's its own engineering challenge.
65
And there's a security issue.
66
If someone figures out your pattern, they can guess the next short URL.
67
But for most cases, this approach is much cleaner.
68
Now, generating the short URL is just half of the problem.
69
The other half is what happens when someone clicks it.
70
When you click a short URL, the system needs to look up the original URL and redirect you there.
71
And this happens a lot more often than creating new short URLs.
72
So speed matters.
73
First, check the cache.
74
If the mapping is there, redirect immediately.
75
If not, hit the database, cache the result for next time, then redirect.
76
The redirect itself uses what's called a 301 status code.
77
That tells your browser this URL has permanently moved to the other location.
78
Your browser remembers this, so it might skip the URL shortener entirely next time.
79
But here's what makes this really interesting a scale.
80
A single database can handle 10,000 lookups per second, so you need multiple database replicas to spread the load.
81
Eventually, you will need to split the data across multiple databases entirely.
82
This is called sharding, and it's another whole engineering problem.
83
You need to figure out how to distribute the data evenly, how to route requests to the right database, what happens when one database goes down, how to rebalance when you add more servers.
84
And that's just the beginning.
85
In the real world, you also need to think about rate limiting so people can't spam your service.
86
You need analytics to track how many people click each link.
87
You need security to block malicious URLs.
88
What started as make this URL shorter becomes a lesson in distributed systems, caching, database scaling, and performance optimization.
89
Every major tech company has built some version of this.
90
Twitter shortens URL in tweets.
91
Slack does it in messages.
92
Even your company's internal tools probably do this.
93
And the techniques we talk about, unique ID generation, caching strategies, database sharding, these patterns show up everywhere.
94
Instagram uses similar ID generation for photos.
95
Netflix uses similar caching for video metadata.
96
Uber uses similar database splitting for trip data.
97
The next time you click a shortened link, you will know there's a host system working in milliseconds to get you where you're going.
98
And you will start recognizing these same patterns in every app you use.
99
Ready to ace your next technical interview?
100
Join our community where we offer comprehensive courses on system design, coding, behavioral questions, machine learning, and object-oriented design.
101
Learn more at bytebytego.com.

本课的词汇与口语要点

这段视频共有 100 个句子、978 个单词可供跟读。 讲话部分时长为 6:41。 说话人语速平稳,每分钟约 146 个词,很适合跟读。 只有 81% 的单词属于英语最常用的 3,000 词,词汇难度较高。

视频中的重点词汇

视频中 15 个值得学习的单词,附发音和释义:

单词发音释义
database 名词/ˈdeɪtəˌbeɪs/數據庫 /数据库
remainder 名词/ɹəˈmeɪndɚ/其餘 /其余
cache 名词/kæʃ/緩存 /缓存, 快取
divide 动词/dɪˈvaɪd/分, 分裂
hash 名词/ˈhæʃ/井號 /井号, 井字
server 名词/ˈsɝvɚ/服務器 /服务器, 伺服器
string 名词/stɹɪŋ/線 /线
mathematics 名词/mæθ(.ə)ˈmæt.ɪks/數學 /数学, 算學 /算学
collision 名词/kəˈlɪʒn̩/碰撞
distribute 动词/dɪˈstɹɪbjuːt/分配
backward 形容词/ˈbækwɚd/逆向
lesson 名词/ˈlɛs.ən/課業 /课业, -課 /-课
technique 名词/tɛkˈniːk/技術 /技术, 技巧
comprehensive 形容词/ˌkɑm.pɹəˈhɛn.sɪv/全面的, 綜合的 /综合的
sequence 名词/ˈsiː.kwəns/序列, 順序 /顺序

视频中出现的短语动词

单词发音释义
figure out 动词弄清楚
go down 动词下降, 下去
look up 动词/ˌlʊk ˈʌp/仰, 上看
show up 动词露面

需要注意的发音

说话人用了 9 次缩略和弱读形式,例如 you're, can't, couldn't。请按听到的简短形式来说。

  • “sh” 和 “zh” 音: cache /kæʃ/, hash /ˈhæʃ/, collision /kəˈlɪʒn̩/, combinations /kɑmbɪˈneɪʃənz/, variation /ˌvɛəɹiˈeɪʃn̩/
  • 长单词——注意重音位置: mathematics /mæθ(.ə)ˈmæt.ɪks/, combinations /kɑmbɪˈneɪʃənz/, comprehensive /ˌkɑm.pɹəˈhɛn.sɪv/, permanently /ˈpɜː.mə.nənt.li/, behavioral /bɪˈheɪvjəɹəl/

如何用这段视频练习

  1. 先完整听一遍视频,不要开口,记下不认识的单词。
  2. 用正常速度逐句跟读,每句重复到你的节奏与说话人一致为止。
  3. 录下自己的声音并与原声对比,特别注意 database, remainder, cache 这类单词。

什么是跟读法?

跟读法 (Shadowing) 是一种有科学依据的语言学习技巧,最初开发用于专业口译员的培训,并由多语言者Alexander Arguelles博士普及。这个方法简单而强大:您在听英语母语原声的同时立即大声重复——就像是一个延迟1-2秒紧跟说话者的影子。与被动听力或语法练习不同,跟读法强迫您的大脑和口腔肌肉同时处理并模仿真实的讲话模式。研究表明它能显着提高发音准确性,语调,节奏,连读,听力理解和口语流利度——使其成为雅思口语备考和真实英语交流最有效的方法之一。

影子跟读法: 阅读完整分步指南 →