A URL shortener in Go, built to find out where a read-heavy service actually breaks.
Everything below was measured with k6 and the service's own /metrics, on one laptop
running Postgres, Redis and the API together. No estimates, no scaled-up projections.
1. Roughly 110 million daily users. Redirects measured 37,893/sec at a p95 of 2.32ms with zero errors. Holding 30% back as headroom gives a usable peak of ~26,500/sec. Real traffic is not flat, so at a 4x peak-to-average ratio that is ~6,600/sec sustained, or ~570 million redirects a day. At five clicks per user per day, ~110M DAU. The measured number is the throughput; the user count is arithmetic on top of three stated assumptions.
2. One billion links is not the constraint. Writes measured 12,141/sec with zero errors. A service holding 1B links creates roughly 100k new ones a day, which is 1.2 writes/sec, so the write path has four orders of magnitude of headroom. Storage is about 240GB at 1B rows including both indexes, which fits one disk. The code space runs out at 3.5 trillion, not 1 billion, because codes lengthen as the counter grows.
What actually binds is neither of those. It is Redis memory, now capped and evicting, and the fact that this is a single instance. Both are covered below, along with the bugs that only appeared under load.
flowchart LR
client([client])
nginx["<b>nginx</b><br/>route by path<br/>round-robin per pool"]
write["<b>write instance</b><br/>ROLE=write<br/>next id from local block<br/>base62 encode, INSERT"]
read["<b>read instance</b> ×2<br/>ROLE=read<br/>Redis first<br/>302, or 404"]
counter[("<b>Redis</b><br/>shortn:counter<br/><i>no TTL, never evicted</i>")]
cache[("<b>Redis</b><br/>code → url<br/><i>volatile-lfu, 256MB</i>")]
pg[("<b>PostgreSQL</b><br/>links<br/><i>source of truth</i>")]
client --> nginx
nginx -->|"POST /shorten"| write
nginx -->|"GET /{code}"| read
write -.->|"claim 1000 ids<br/>1 write in 1000"| counter
write -->|INSERT| pg
read -->|"GET, then SET"| cache
read -.->|"on cache miss<br/>0.14% of reads"| pg
classDef svc fill:#e8f0fe,stroke:#4285f4,stroke-width:2px
classDef store fill:#fff4e5,stroke:#f9ab00,stroke-width:2px
classDef edge fill:#f1f3f4,stroke:#5f6368,stroke-width:2px
class write,read svc
class counter,cache,pg store
class nginx,client edge
Read path. Redis first. On a miss, Postgres, then populate the cache and redirect. Under load Redis serves 99.86% of reads.
Write path. Never touches the cache. It takes an id from an in-process block, encodes it, and inserts. Redis is contacted once per 1000 writes to claim the next block, so writes keep working even with Redis stopped entirely.
Single instance, no load balancer, CODE_COUNT=1000:
| Workload | Throughput | Median | p95 | p99 | Errors |
|---|---|---|---|---|---|
| Read | 37,893/s |
1.06ms |
2.32ms |
3.85ms |
0% |
| Write | 12,141/s |
3.77ms |
6.41ms |
9.70ms |
0% |
/metrics after the read run:
requests_total 770619
redirect_cache_hits 769563
redirect_cache_misses 1056
redirect_db_reads 1056
status_302 770619
status_404 0
status_500 0
writes_failed 0
redirect_expired 0
1,056 database reads out of 770,619 requests. The misses are the 1,000-code pool being read for the first time; everything after that is served by Redis.
POST /shorten {"url": "https://example.com"} -> {"code": "4c92"}
POST /shorten {"url": "...", "protected": true} -> unguessable code
POST /shorten {"url": "...", "expires_at": "2027-01-01T00:00:00Z"}
GET /{code} -> 302, or 404
GET /metricsCodes come from a counter, not randomness. Redis holds one counter and each instance
claims a block of 1000 ids with a single INCRBY, then serves the rest from memory.
Collisions are impossible by construction rather than merely unlikely, and 999 of every
1000 writes make no Redis call.
protected swaps in an 11-character crypto/rand code. Counter codes are sequential,
so anyone holding one can walk to its neighbours, and the gap between two codes leaks how
many links were created in between.
expires_at is enforced on read and the row is kept, so a link can be audited or
extended later.
The counter does not survive a Redis restart. Redis is configured as a cache with no
persistence, so a restart sets the counter to zero and every insert then fails on the
primary key. ensureCounterFloor reseeds it from MAX(id) in Postgres at startup.
Verified by running FLUSHALL and restarting.
maxmemory-policy noeviction fails silently. Redis shipped with the default, which
refuses writes at full memory while still serving reads. Cache population would fail with
its error ignored, the hit rate would decay from 99.9% toward zero, and Postgres would
take the full read load with nothing logged anywhere. volatile-lfu with a memory cap
fixes it. It has to be volatile-, not allkeys-: the id counter is touched once per
1000 writes, which makes it one of the coldest keys in the keyspace and an early eviction
candidate under allkeys. Only cached URLs carry a TTL.
An in-process cache in front of Redis made things worse. It served 0 of 1.2M reads while costing 40% write throughput, had no size bound, and could not be invalidated once more than one instance existed. Removed, and reads went from 31,629/s to 38,233/s.
The connection pool is sized per process. SetMaxOpenConns(50) is correct for one
binary and becomes 150 connections across three instances, against a Postgres default of
100. 93.8% of writes failed with "too many clients already" the first time the split ran.
GET /shorten returned 404, not 405. GET /{code} matched the literal path, so the
router looked up a link called "shorten" and cached the miss in Redis.
Destinations were unvalidated. A stored javascript: URL is handed back in a
Location header. Only absolute http and https URLs are accepted now.
Reads outnumber writes 3:1 here and have the opposite resource profile: reads are 98.9%
Redis, writes are 100% Postgres. One process cannot scale them independently, so ROLE
splits the same binary into read and write instances behind nginx.
On a single machine the split is slower, which is the expected result and the reason it was measured rather than assumed:
Combined :8080 |
Split :8090 |
|
|---|---|---|
| Read | 37,125/s, p95 2.29ms |
23,183/s, p95 3.53ms |
| Write | 12,610/s, p95 6.14ms |
12,241/s, p95 6.70ms |
An extra network hop and four processes sharing one CPU cost 38% of read throughput. The split buys independent scaling across machines, not local latency. Both shapes run at once against the same data, so the comparison is direct.
./run.sh # postgres, redis, combined API on :8080
./try.sh # walk every endpoint, metrics before and after
./bench.sh # the three k6 benchmarks
./split.sh # nginx + 2 read + 1 write on :8090
BASE_URL=http://localhost:8090 ./bench.sh # benchmark the split shape
./run.sh stoprun.sh polls Postgres with a real query rather than pg_isready, which reports OK while
the container is still running its init scripts.
No user accounts, no click analytics, no custom aliases, no multi-region. Postgres is a single instance: at 1B rows and roughly 1 write per second there is nothing yet to shard, and read replicas only become interesting once cache misses rather than cache hits are the bottleneck.
cmd/
main.go startup, role selection, routing
redirect.go read path
shorten.go write path
idgen.go block-allocated ids, counter recovery
base62.go code generation
url.go destination validation
metrics.go counters, flushed to Redis
db/schema.sql
bench/ k6 scripts
nginx.conf load balancer for the split shape
RUNBOOK.md has the operational detail and every benchmark run. IMPROVEMENTS.md is the
ordered backlog.