Rate Limiting You Can Watch Work
What rate limiting is, with a bucket analogy anyone can follow — then how I built a distributed token-bucket limiter in NestJS + Redis with a live dashboard that shows requests getting throttled in real time.
Rate Limiting You Can Watch Work
On this page
Most explanations of rate limiting open with jargon — tokens, windows, 429s. I wanted the opposite: something you can watch. So I built a rate limiter with a live dashboard where you drag a slider and see requests turn red the moment you push past the limit. Here is the idea, and how it is built.
What is rate limiting?
Rate limiting caps how many requests a server will handle in a window, so a flood of traffic cannot overwhelm it. Requests over the limit are turned away politely (HTTP 429 Too Many Requests) instead of taking the whole system down.
The bucket analogy
Picture a bucket that holds 10 tokens and refills 5 every second. Every request spends one token. Send requests slowly and the bucket keeps up — everything succeeds. Send a flood and the bucket empties; until it refills, extra requests get rejected. That is a token bucket: it allows a short burst (the bucket size) while enforcing a steady long-run rate (the refill).
Why not a simpler counter?
The naive approach counts requests per fixed window — "100 per minute". It works until the window edge: a client can send 100 at 0:59 and another 100 at 1:01, doubling the intended rate. A token bucket avoids that by thinking in terms of a continuously refilling balance rather than hard window boundaries.
Making it correct under load
Here is the subtle part. A rate-limit decision is read-compute-write: read the token count, refill based on elapsed time, then check and decrement. If two requests do that at the same instant, both can read the same count and both succeed — spending one token twice. Across multiple server instances sharing one Redis, that is a real bug.
The fix is to make the whole decision atomic. I put it in a single Redis Lua script, which Redis runs as one uninterruptible operation — no interleaving, one round-trip, no retry loop.
-- one atomic decision: refill, check, decrement
tokens = math.min(capacity, tokens + elapsed * refill)
if tokens >= 1 then tokens = tokens - 1; allowed = 1 endAn integration test fires 100 concurrent requests at a bucket of 20 and asserts that exactly 20 are allowed. No double-spend.
Making it visible
The part I enjoyed most: every limiter decision is published over Server-Sent Events, and a Next.js dashboard animates it — a token gauge that drains and refills, a stream of green/red squares for allowed vs throttled requests, and a throughput chart. Two one-click scenarios ("calm" and "spike") let anyone see both behaviours without knowing what a token is.
Proven three ways
- Unit tests drive the pure token-bucket function with a fake clock.
- An integration test (Testcontainers + real Redis) proves no double-spend under concurrency.
- A k6 load test sends 25 req/s at a 5/s limit and confirms most requests get throttled.
The whole thing runs with one docker compose up, deploys free, and ships with CI on GitHub Actions.
Takeaways
- A token bucket models "burst up to N, then steady rate" in two numbers per key.
- Rate limiting is read-compute-write — make it atomic or it is wrong under load.
- The clearest way to explain a system is often to let people watch it run.
Comments
Loading comments…