Case study — Stripe-style payments API
Design a Global Rate Limiter
PayFlow: 2M req/s, 3 regions, one merchant key. Throttle it fairly, globally, in under 5ms — live demo on the right →
live_demo.shARMED
merchant_482
capacity10 tok
refill2 tok/s
left10
result—
00 Problem Statement
Six constraints, one system
1Per-key — 100 req/s, enforced no matter which region it hits
2Per-endpoint —
/v1/charges tighter than /v1/customers3Tiered — Enterprise ceiling > Starter ceiling
4Global — can't cheat the limit by spraying across regions
5Fast — under 5ms added to every request
6Resilient — degrades predictably when its own store is down
2M/s
global peak traffic
500K
merchant API keys
3
active-active regions
±10%
acceptable overshoot
01 Functional Requirements
Must-have vs. nice-to-have
Per-key throttling
MUSTPer-endpoint policy
MUSTTiered plans
MUSTGlobal enforcement
MUST429 + Retry-After
MUSTDynamic limit updates
NICEUsage dashboards
NICETemporary overrides
NICE02 Non-Functional Requirements
What it must be, in numbers
<5ms
p99 check latency
99.99%
limiter availability
±10%
global accuracy band
<2s
policy propagation
Fail behavior isn't uniform — split by endpoint cost, see §09.
03 Capacity Estimation
2M req/s, traced to hardware
04 API Design
Same request, two paths
POST /v1/charges
key: mk_482
→
rate limiter
→
200 OK
X-RateLimit-Remaining: 87X-RateLimit-Reset: 42
429 Throttled
Retry-After: 3policy: endpoint_limit
◆ Control plane
PUT/internal/limits/{id}
GET/internal/limits/{id}
GET/internal/usage/{id}
05 High-Level Architecture
Build it piece by piece, then fire a request
One region's request path. Add each component, then send a request and watch the rate limiter make the call. (Global sync across 3 regions is covered in §06.)
0 / 5 components
request path
token check
forwarded on allow
policy push
Zoom out: the full global picture
The block you just built lives inside each of these three regions.
request path
hot-path check
allowed / policy push
async best-effort
06 Deep Dive
Why token bucket wins
Fixed Window
2× burst at boundary
Sliding Log
accurate, memory-heavy
Token Bucket
burst then steady, O(1) mem
✓ CHOSEN
-- one atomic round trip to local Redis
tokens = min(cap, tokens + elapsed*rate)
if tokens >= 1 then decrement(); return ALLOW
else return THROTTLE end
① regional slice
100 req/s global → ~35/s per region
② aggregator sync
streams usage every 500ms–1s
③ bounded worst case
overshoot capped at (regions−1)×slice
① write
PUT /internal/limits/{id}
② publish
pub/sub broadcasts new policy
③ swap
sidecars hot-swap cache, no restart
07 Data Model / Schema
Two stores, two jobs
⚡ Redis — bucket state
keyrl:{merchant}:{endpoint}
tokensfloat
tsint64 ms
capacityint
refill_ratefloat
TTL~120s
▤ Postgres — merchants
iduuid PK
api_key_hashtext
tierenum
region_affinitytext[]
▤ Postgres — policies
tierenum
endpoint_classtext
limit_per_secint
burst_capacityint
daily_quotaint
updated_attimestamptz
08 Scalability
500K → 5M merchants
09 Bottlenecks & Failures
Failure → mitigation, at a glance
Redis node dies
replica failover, sub-second
whole regional store down
reads fail open, spends fail toward a strict local cap
cross-region partition
falls back to static regional slice
hot key / whale merchant
counter sharded across sub-keys
window-boundary burst
solved by design — token bucket has no edge
clock skew
Lua script uses Redis's own clock
10 Tradeoffs
Bounded imprecision over unbounded latency
Accuracy vs. Latency
chose local-first + async sync
Fail-open vs. Fail-closed
split by endpoint cost
Algorithm
token bucket over sliding log
Global state model
regional slices + reconciliation