Skip to content
Shailu-sPublic

About

URL shortener in Go which can handle 110 million user per day. 37,893 request/sec measured, and the bugs that only showed up under load.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

27 Commits

Folders and files

Repository files navigation

shortn

A URL shortener in Go, built to find out where a read-heavy service actually breaks.

Everything below was measured with k6 and the service's own /metrics, on one laptop running Postgres, Redis and the API together. No estimates, no scaled-up projections.

1. Roughly 110 million daily users. Redirects measured 37,893/sec at a p95 of 2.32ms with zero errors. Holding 30% back as headroom gives a usable peak of ~26,500/sec. Real traffic is not flat, so at a 4x peak-to-average ratio that is ~6,600/sec sustained, or ~570 million redirects a day. At five clicks per user per day, ~110M DAU. The measured number is the throughput; the user count is arithmetic on top of three stated assumptions.

2. One billion links is not the constraint. Writes measured 12,141/sec with zero errors. A service holding 1B links creates roughly 100k new ones a day, which is 1.2 writes/sec, so the write path has four orders of magnitude of headroom. Storage is about 240GB at 1B rows including both indexes, which fits one disk. The code space runs out at 3.5 trillion, not 1 billion, because codes lengthen as the counter grows.

What actually binds is neither of those. It is Redis memory, now capped and evicting, and the fact that this is a single instance. Both are covered below, along with the bugs that only appeared under load.

Architecture

flowchart LR
    client([client])

    nginx["<b>nginx</b><br/>route by path<br/>round-robin per pool"]

    write["<b>write instance</b><br/>ROLE=write<br/>next id from local block<br/>base62 encode, INSERT"]
    read["<b>read instance</b> ×2<br/>ROLE=read<br/>Redis first<br/>302, or 404"]

    counter[("<b>Redis</b><br/>shortn:counter<br/><i>no TTL, never evicted</i>")]
    cache[("<b>Redis</b><br/>code → url<br/><i>volatile-lfu, 256MB</i>")]
    pg[("<b>PostgreSQL</b><br/>links<br/><i>source of truth</i>")]

    client --> nginx
    nginx -->|"POST /shorten"| write
    nginx -->|"GET /{code}"| read

    write -.->|"claim 1000 ids<br/>1 write in 1000"| counter
    write -->|INSERT| pg

    read -->|"GET, then SET"| cache
    read -.->|"on cache miss<br/>0.14% of reads"| pg

    classDef svc fill:#e8f0fe,stroke:#4285f4,stroke-width:2px
    classDef store fill:#fff4e5,stroke:#f9ab00,stroke-width:2px
    classDef edge fill:#f1f3f4,stroke:#5f6368,stroke-width:2px
    class write,read svc
    class counter,cache,pg store
    class nginx,client edge
Loading

Read path. Redis first. On a miss, Postgres, then populate the cache and redirect. Under load Redis serves 99.86% of reads.

Write path. Never touches the cache. It takes an id from an in-process block, encodes it, and inserts. Redis is contacted once per 1000 writes to claim the next block, so writes keep working even with Redis stopped entirely.

Measured

Single instance, no load balancer, CODE_COUNT=1000:

Workload Throughput Median p95 p99 Errors
Read 37,893/s 1.06ms 2.32ms 3.85ms 0%
Write 12,141/s 3.77ms 6.41ms 9.70ms 0%

/metrics after the read run:

requests_total          770619
redirect_cache_hits     769563
redirect_cache_misses     1056
redirect_db_reads         1056
status_302              770619
status_404                   0
status_500                   0
writes_failed                0
redirect_expired             0

1,056 database reads out of 770,619 requests. The misses are the 1,000-code pool being read for the first time; everything after that is served by Redis.

What it does

POST /shorten {"url": "https://example.com"}                    -> {"code": "4c92"}
POST /shorten {"url": "...", "protected": true}                 -> unguessable code
POST /shorten {"url": "...", "expires_at": "2027-01-01T00:00:00Z"}
GET  /{code}                                                    -> 302, or 404
GET  /metrics

Codes come from a counter, not randomness. Redis holds one counter and each instance claims a block of 1000 ids with a single INCRBY, then serves the rest from memory. Collisions are impossible by construction rather than merely unlikely, and 999 of every 1000 writes make no Redis call.

protected swaps in an 11-character crypto/rand code. Counter codes are sequential, so anyone holding one can walk to its neighbours, and the gap between two codes leaks how many links were created in between.

expires_at is enforced on read and the row is kept, so a link can be audited or extended later.

Things that only showed up under load

The counter does not survive a Redis restart. Redis is configured as a cache with no persistence, so a restart sets the counter to zero and every insert then fails on the primary key. ensureCounterFloor reseeds it from MAX(id) in Postgres at startup. Verified by running FLUSHALL and restarting.

maxmemory-policy noeviction fails silently. Redis shipped with the default, which refuses writes at full memory while still serving reads. Cache population would fail with its error ignored, the hit rate would decay from 99.9% toward zero, and Postgres would take the full read load with nothing logged anywhere. volatile-lfu with a memory cap fixes it. It has to be volatile-, not allkeys-: the id counter is touched once per 1000 writes, which makes it one of the coldest keys in the keyspace and an early eviction candidate under allkeys. Only cached URLs carry a TTL.

An in-process cache in front of Redis made things worse. It served 0 of 1.2M reads while costing 40% write throughput, had no size bound, and could not be invalidated once more than one instance existed. Removed, and reads went from 31,629/s to 38,233/s.

The connection pool is sized per process. SetMaxOpenConns(50) is correct for one binary and becomes 150 connections across three instances, against a Postgres default of 100. 93.8% of writes failed with "too many clients already" the first time the split ran.

GET /shorten returned 404, not 405. GET /{code} matched the literal path, so the router looked up a link called "shorten" and cached the miss in Redis.

Destinations were unvalidated. A stored javascript: URL is handed back in a Location header. Only absolute http and https URLs are accepted now.

Scaling reads and writes separately

Reads outnumber writes 3:1 here and have the opposite resource profile: reads are 98.9% Redis, writes are 100% Postgres. One process cannot scale them independently, so ROLE splits the same binary into read and write instances behind nginx.

On a single machine the split is slower, which is the expected result and the reason it was measured rather than assumed:

Combined :8080 Split :8090
Read 37,125/s, p95 2.29ms 23,183/s, p95 3.53ms
Write 12,610/s, p95 6.14ms 12,241/s, p95 6.70ms

An extra network hop and four processes sharing one CPU cost 38% of read throughput. The split buys independent scaling across machines, not local latency. Both shapes run at once against the same data, so the comparison is direct.

Running it

./run.sh                                    # postgres, redis, combined API on :8080
./try.sh                                    # walk every endpoint, metrics before and after
./bench.sh                                  # the three k6 benchmarks

./split.sh                                  # nginx + 2 read + 1 write on :8090
BASE_URL=http://localhost:8090 ./bench.sh   # benchmark the split shape
./run.sh stop

run.sh polls Postgres with a real query rather than pg_isready, which reports OK while the container is still running its init scripts.

Deliberately not built

No user accounts, no click analytics, no custom aliases, no multi-region. Postgres is a single instance: at 1B rows and roughly 1 write per second there is nothing yet to shard, and read replicas only become interesting once cache misses rather than cache hits are the bottleneck.

Layout

cmd/
  main.go      startup, role selection, routing
  redirect.go  read path
  shorten.go   write path
  idgen.go     block-allocated ids, counter recovery
  base62.go    code generation
  url.go       destination validation
  metrics.go   counters, flushed to Redis
db/schema.sql
bench/         k6 scripts
nginx.conf     load balancer for the split shape

RUNBOOK.md has the operational detail and every benchmark run. IMPROVEMENTS.md is the ordered backlog.

About

URL shortener in Go which can handle 110 million user per day. 37,893 request/sec measured, and the bugs that only showed up under load.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages