Load Assumptions

Before choosing a design, we need an estimate of the load. These numbers are my assumptions, not the companies’ real figures:

t.co Drive O’Reilly
Links created per day 200 million 50 million 200
Redirects per day 20 billion 500 million 50,000

We assume each link takes about 500 bytes to store. That covers the code, a long URL of a few hundred characters, the creation time, the status, and, for Drive and O’Reilly, the document or the book.

Per second and per year

A day has 86,400 seconds. For a quick estimate, we round that up to 100,000 and divide each daily number by it. For t.co, 200 million links a day is about 2,000 links per second.

For storage, we multiply the links per day by 365 days and by 500 bytes. For t.co, that is 200 million × 365 × 500 bytes. That comes to about 36 TB a year.

t.co Drive O’Reilly
Writes per second About 2,000 About 500 Under 1
Reads per second About 200,000 About 5,000 Under 1
Reads per write 100 10 250
Storage per year About 36 TB About 9 TB About 36 MB

Here, a write means creating a link, and a read means following one. These numbers are averages. A viral post can make the peak rate several times higher than the average.

Compared with Shopend

Earlier, we assumed Shopend’s workload was 10,000 reads and 1,000 writes per second, and one database server could not keep up with it. Compared with that:

  • t.co needs far more than one server for its reads, its writes, and its storage.
  • Drive’s reads and writes fit on one primary with replicas. But its storage grows by about 9 TB every year.
  • O’Reilly fits on one small server.

With Shopend, we used a load test to find out that the database was the bottleneck. Here we do not need one. X and Google already have this load, so we design for it from the start.