The API

We have designed REST APIs and GraphQL APIs. Which one fits a URL shortener?

REST or GraphQL

A browser follows a link by sending a GET request to the link’s URL. When a reader clicks https://short.example/x7Kp2Qa, the browser sends GET /x7Kp2Qa to our server. That is a REST request to a resource identified by its URL.

A GraphQL request is a POST to one endpoint, with the query in the body. A browser cannot follow a link that way. So follow must be a REST endpoint.

Shorten could use GraphQL. But each operation returns one fixed representation, with no nested data, so GraphQL would not save us anything. I keep the whole API on REST.

Endpoints

Operation Request
Shorten POST /links
Follow GET /{code}
Block, turn off, change target PATCH /links/{code}
List a book’s links GET /links?bookId={id}

The last two rows are for the variants that need them. In t.co, PATCH blocks a link. In Drive, it turns a link off and on. In O’Reilly, it changes a link’s target. Only O’Reilly lists a book’s links.

In all three variants, the company’s own software sends the shorten requests: X’s servers when someone posts, Drive’s share dialog, and O’Reilly’s production tool. Only readers’ browsers send follow requests.

Shorten

Here is a request to shorten a link:

POST /links HTTP/1.1
Host: short.example
Content-Type: application/json

{ "url": "https://docs.google.com/document/d/1aB2.../edit" }

The server creates the code and sends this response:

HTTP/1.1 201 Created
Location: /links/x7Kp2Qa
Content-Type: application/json

{ "code": "x7Kp2Qa", "shortUrl": "https://short.example/x7Kp2Qa" }

Suppose X’s servers send this request, and the request times out. X’s servers retry it. If the first request did reach our server, the server creates two codes for one link. The second code is wasted, but both codes work. An idempotency key, which we saw in the chapter on designing a write API, would prevent the duplicate.

Follow

Here is the request a browser sends when a reader clicks the short link:

GET /x7Kp2Qa HTTP/1.1
Host: short.example

The server looks up the code and sends this response:

HTTP/1.1 302 Found
Location: https://docs.google.com/document/d/1aB2.../edit

The browser then requests the URL in the Location header.

Redirects

The Public Courses API uses one status code that starts with 3: 304 Not Modified, which tells the client to reuse its cached copy. That code is not a redirect. Most of the other codes that start with 3 are redirects. That API does not need them. A redirect tells the client that the resource is at another URL. The new URL is in the Location header.

Two redirect codes matter here:

  • 301 Moved Permanently: the browser caches the redirect. The next time the reader follows the same short link, the browser goes straight to the long URL. It does not send a request to our server.
  • 302 Found: the browser sends a request to our server every time, unless the response says the redirect may be cached.

A browser caches a 301 with no expiry, unless the response says otherwise. It caches a 302 only when the response gives it a freshness period, such as Cache-Control: max-age=300.

With a 301, later clicks never reach our server. So we also cannot count them. If we needed click analytics, we would have to think about those missing counts when we choose a redirect code. We left click analytics out.

Aside: There are two more redirect codes, 307 and 308. They work like 302 and 301, but the browser keeps the method of the original request. For a GET, 307 behaves like 302, and 308 behaves like 301.

Choosing a redirect

We use 302 in all three variants. A browser that cached a 301 would keep following:

  • a t.co link after X blocks it,
  • a Drive link after the owner turns it off,
  • an O’Reilly link’s old target after staff change it.

The variants differ in whether the browser may cache the 302.

For t.co, we add Cache-Control: max-age=300 to the redirect. A browser reuses the redirect for up to 5 minutes. A reader who clicks the same link twice in a few minutes sends our server only one request. We still meet the requirement that a block takes effect “within minutes”.

For Drive and O’Reilly, we add Cache-Control: no-store to the redirect. This header tells the browser not to cache the redirect. Drive needs a link that is turned off to stop working within seconds. O’Reilly receives so few requests that caching would save nothing.