Load Assumptions
Before choosing a design, we need an estimate of the load. These numbers are my assumptions, not the companies’ real figures:
t.co |
Drive | O’Reilly | |
|---|---|---|---|
| Links created per day | 200 million | 50 million | 200 |
| Redirects per day | 20 billion | 500 million | 50,000 |
We assume each link takes about 500 bytes to store. That covers the code, a long URL of a few hundred characters, the creation time, the status, and, for Drive and O’Reilly, the document or the book.
Per second and per year
A day has 86,400 seconds. For a quick estimate, we round that up to 100,000 and divide each daily number by it. For t.co, 200 million links a day is about 2,000 links per second.
For storage, we multiply the links per day by 365 days and by 500 bytes. For t.co, that is 200 million × 365 × 500 bytes. That comes to about 36 TB a year.
t.co |
Drive | O’Reilly | |
|---|---|---|---|
| Writes per second | About 2,000 | About 500 | Under 1 |
| Reads per second | About 200,000 | About 5,000 | Under 1 |
| Reads per write | 100 | 10 | 250 |
| Storage per year | About 36 TB | About 9 TB | About 36 MB |
Here, a write means creating a link, and a read means following one. These numbers are averages. A viral post can make the peak rate several times higher than the average.
Compared with Shopend
Earlier, we assumed Shopend’s workload was 10,000 reads and 1,000 writes per second, and one database server could not keep up with it. Compared with that:
t.coneeds far more than one server for its reads, its writes, and its storage.- Drive’s reads and writes fit on one primary with replicas. But its storage grows by about 9 TB every year.
- O’Reilly fits on one small server.
With Shopend, we used a load test to find out that the database was the bottleneck. Here we do not need one. X and Google already have this load, so we design for it from the start.