Case study — Stripe-style payments API

Design a Global Rate Limiter

PayFlow: 2M req/s, 3 regions, one merchant key. Throttle it fairly, globally, in under 5ms — live demo on the right →

live_demo.shARMED
merchant_482
capacity10 tok
refill2 tok/s
left10
result—
00 Problem Statement

Six constraints, one system

1Per-key — 100 req/s, enforced no matter which region it hits
2Per-endpoint — /v1/charges tighter than /v1/customers
3Tiered — Enterprise ceiling > Starter ceiling
4Global — can't cheat the limit by spraying across regions
5Fast — under 5ms added to every request
6Resilient — degrades predictably when its own store is down
2M/s
global peak traffic
500K
merchant API keys
3
active-active regions
±10%
acceptable overshoot
01 Functional Requirements

Must-have vs. nice-to-have

Per-key throttling
MUST
Per-endpoint policy
MUST
Tiered plans
MUST
Global enforcement
MUST
429 + Retry-After
MUST
Dynamic limit updates
NICE
Usage dashboards
NICE
Temporary overrides
NICE
02 Non-Functional Requirements

What it must be, in numbers

<5ms
p99 check latency
99.99%
limiter availability
±10%
global accuracy band
<2s
policy propagation

Fail behavior isn't uniform — split by endpoint cost, see §09.

03 Capacity Estimation

2M req/s, traced to hardware

2,000,000 req / sec global ~700K/s US-East ~700K/s EU-West (peak) ~600K/s AP-South ÷100K/shard 12–14 Redis shards / region ~100K ops/s per shard, headroom for failover ~500 MB RAM / region 500K merchants × 5 endpoint classes × ~200B per bucket key trivial — fits fully in-memory
04 API Design

Same request, two paths

POST /v1/charges
key: mk_482
→
rate limiter
→
200 OK
X-RateLimit-Remaining: 87X-RateLimit-Reset: 42
429 Throttled
Retry-After: 3policy: endpoint_limit

◆ Control plane

PUT/internal/limits/{id}
GET/internal/limits/{id}
GET/internal/usage/{id}
05 High-Level Architecture

Build it piece by piece, then fire a request

One region's request path. Add each component, then send a request and watch the rate limiter make the call. (Global sync across 3 regions is covered in §06.)

0 / 5 components
Merchant Client sends POST /v1/charges API Gateway auth + routing Rate Limiter token bucket check tokens: 10.0 ✕ 429 THROTTLED Redis token store, atomic Lua Charges Service business logic Config Plane pushes tier + limits
request path token check forwarded on allow policy push

Zoom out: the full global picture

The block you just built lives inside each of these three regions.

Merchant Client Anycast Edge (GeoDNS) US-EAST API Gateway Rate Limiter Sidecar Redis Cluster (local) Charges Service usage → Kafka EU-WEST AP-SOUTH Global Usage Aggregator — syncs every ~500ms–1s Control Plane Policy Config DB Config API + Pub/Sub pushes limits to all sidecars
request path hot-path check allowed / policy push async best-effort
06 Deep Dive

Why token bucket wins

Fixed Window
2× burst at boundary
Sliding Log
accurate, memory-heavy
Token Bucket
burst then steady, O(1) mem
✓ CHOSEN
-- one atomic round trip to local Redis tokens = min(cap, tokens + elapsed*rate) if tokens >= 1 then decrement(); return ALLOW else return THROTTLE end

① regional slice

100 req/s global → ~35/s per region

② aggregator sync

streams usage every 500ms–1s

③ bounded worst case

overshoot capped at (regions−1)×slice

① write

PUT /internal/limits/{id}

② publish

pub/sub broadcasts new policy

③ swap

sidecars hot-swap cache, no restart
07 Data Model / Schema

Two stores, two jobs

⚡ Redis — bucket state

keyrl:{merchant}:{endpoint}
tokensfloat
tsint64 ms
capacityint
refill_ratefloat
TTL~120s

▤ Postgres — merchants

iduuid PK
api_key_hashtext
tierenum
region_affinitytext[]

▤ Postgres — policies

tierenum
endpoint_classtext
limit_per_secint
burst_capacityint
daily_quotaint
updated_attimestamptz
08 Scalability

500K → 5M merchants

shard A shard B shard C shard D shard E shard F shard G merchant_482 Consistent hash ring Whale merchant → counter sharding rl:482:0 rl:482:1 rl:482:2 round-robin routed, summed for reporting Adding a 4th region US EU AP +NEW joins aggregation topology — additive, no redesign
09 Bottlenecks & Failures

Failure → mitigation, at a glance

Redis node dies
replica failover, sub-second
whole regional store down
reads fail open, spends fail toward a strict local cap
cross-region partition
falls back to static regional slice
hot key / whale merchant
counter sharded across sub-keys
window-boundary burst
solved by design — token bucket has no edge
clock skew
Lua script uses Redis's own clock
10 Tradeoffs

Bounded imprecision over unbounded latency

Accuracy vs. Latency
chose local-first + async sync
accuracylatency
Fail-open vs. Fail-closed
split by endpoint cost
openclosed
Algorithm
token bucket over sliding log
precisionmemory
Global state model
regional slices + reconciliation
single truthno cross-hop