The first version of our public API had no rate limiting at all. It survived for about four months, right up until a customer's integration developer wrote a retry loop with no backoff and pointed it at our sandbox on a Friday afternoon. By the time the on-call engineer noticed, that one misbehaving script was generating roughly 40,000 requests a minute and our payment status endpoint was returning timeouts to everyone else. Nobody was attacking us. A junior developer had just written a while-loop that never slept.
That incident is why I now treat rate limiting as a first-class feature of any public API, not a bit of infrastructure plumbing you bolt on later. It protects you from your own customers, and it protects your customers from each other. Here is how I think about building one, and the specific decisions I would not make differently a second time.
Rate limiting is a product feature, not defense
Most people frame rate limiting as protection against abuse, and it is that, but if you stop there you build the wrong thing. The far more common case is not a malicious actor. It is a well-meaning customer who deployed a bug, or a batch job that used to run nightly and now runs every fifteen minutes because someone changed a cron entry. Your limits exist to keep one tenant's mistake from becoming everyone's outage.
Once you see it as a product feature, the design questions change. You stop asking "how do I block bad traffic" and start asking "what is a fair, predictable envelope I can promise every customer, and how do I make it obvious when they hit it." Predictability matters more than the exact number. A developer who knows they get 100 requests per second can design around it. A developer facing a mysterious ceiling that moves around will file support tickets and distrust you forever.
The algorithm choice is mostly settled
There is a lot of writing comparing fixed windows, sliding windows, token buckets, and leaky buckets, and most of it overstates how much the choice matters. But the differences are real at the edges, so here is my blunt take on the four you will actually consider:
- Fixed window: count requests per calendar minute, reset at the boundary. Trivial to build, but it allows a burst of double your limit across a window edge. A client can fire 100 requests at 10:00:59 and another 100 at 10:01:00.
- Sliding window log: store a timestamp for every request and count the ones in the trailing window. Accurate to the millisecond, and memory-hungry in exactly the way you do not want at scale.
- Sliding window counter: a weighted blend of the current and previous fixed windows. Cheap, and it smooths out the boundary-burst problem well enough for almost everyone.
- Token bucket: a bucket refills at a steady rate up to a cap, each request spends a token. This is the one I reach for, because it models real usage. Customers do bursty work, and a bucket lets them spend saved-up capacity for a short spike without letting them sustain it.
For a payments API we settled on token bucket for the per-tenant limits and a plain fixed window for a few crude global guards. Token bucket maps cleanly onto how integrators actually behave: quiet for minutes, then a flurry when a batch of invoices settles. Letting them burst to, say, 200 requests when their steady rate is 50 keeps legitimate workloads happy without opening the door to sustained hammering.
Count in one place, and make it fast
The state has to live somewhere both shared and fast, because every node behind your load balancer needs the same view of a tenant's budget. In-process counters are a trap. They feel great in a single-instance test and then quietly let a customer do N times their limit once you scale to N instances. I have watched a team ship exactly that and spend a week confused about why their limits "didn't work in production."
Redis is the boring, correct answer for the shared store. The important detail is that the read-check-decrement has to be atomic, or two concurrent requests both read "1 token left" and both proceed. Do it in a Lua script so the whole operation runs server-side as a single step. Here is the token bucket we run, trimmed down:
-- KEYS[1] = bucket key, ARGV = capacity, refill_per_sec, now_ms, requested
local capacity = tonumber(ARGV[1])
local refill = tonumber(ARGV[2])
local now = tonumber(ARGV[3])
local requested = tonumber(ARGV[4])
local state = redis.call('HMGET', KEYS[1], 'tokens', 'ts')
local tokens = tonumber(state[1]) or capacity
local last = tonumber(state[2]) or now
-- refill based on elapsed time since last touch
local elapsed = math.max(0, now - last) / 1000.0
tokens = math.min(capacity, tokens + elapsed * refill)
local allowed = 0
if tokens >= requested then
tokens = tokens - requested
allowed = 1
end
redis.call('HMSET', KEYS[1], 'tokens', tokens, 'ts', now)
redis.call('PEXPIRE', KEYS[1], 60000)
return { allowed, math.floor(tokens) }
The expiry matters more than it looks. Setting a TTL means idle tenants' keys clean themselves up, so you are not carrying a live Redis key for every customer who called you once in 2023. On the .NET side, this runs as a piece of middleware that fails open if Redis is unreachable. That last part is a deliberate choice and a slightly uncomfortable one: if my limiter's backing store is down, I would rather serve traffic than reject everyone. A rate limiter that becomes a single point of total failure has defeated its own purpose.
Enjoying this article?
Get more like it in your inbox — practical engineering leadership, fintech, and AI. No spam, unsubscribe anytime.
What you key on defines fairness
The identifier you count against is the real design decision, and it is easy to get wrong. Keying purely on IP address is almost always a mistake for a business API, because half your customers sit behind the same three cloud NAT gateways and you will punish them for each other's traffic. Key on the authenticated tenant, the API key, or ideally both in a hierarchy.
We run limits at two levels: a per-API-key limit so one leaked or runaway key cannot exhaust a customer's whole allowance, and a per-tenant ceiling above it so a customer with fifty keys still cannot collectively swamp us. Then a small set of expensive endpoints, like bulk reconciliation exports, get their own tighter per-tenant bucket regardless of the general limit, because a single one of those calls costs us far more than a status check.
Tell the client exactly what happened
When you reject a request, the response is part of your API's contract, and a bare 429 with an empty body is hostile. Return the standard headers so a competent client can self-regulate without guessing. Three of them belong on every response, success or failure, not only the ones you reject:
A rate limit the client cannot see is indistinguishable from a random outage. If your developers have to reverse-engineer your ceiling from timing experiments, you have shipped a puzzle, not an API.
Send RateLimit-Limit, RateLimit-Remaining, and RateLimit-Reset on every successful response so clients can throttle proactively. On a 429, add a Retry-After header with the number of seconds to wait, and put a machine-readable error code plus a short human message in the body. The single highest-leverage thing you can do for developer experience here is make Retry-After honest. If you tell them to wait two seconds, the request must actually succeed after two seconds.
Setting the actual numbers is guesswork you can improve
Do not agonize over the first numbers. You genuinely cannot know the right limit before you have real traffic, and any figure you pick in a planning meeting is a guess dressed up as a decision. What you can do is start conservative, log everything, and watch the distribution.
When we launched, I set the default to 100 requests per second per tenant with a burst of 200, entirely on gut feel. Three weeks of data told the real story: our 99th percentile customer never exceeded 12 requests per second, and the two tenants brushing the ceiling were both running inefficient polling loops we could help them fix with webhooks instead. The lesson is that your first job is not to enforce, it is to observe. Run the limiter in a log-only "shadow" mode first, where it records what it would have rejected but lets everything through. You will find your real limits in that data, not in a spreadsheet.
Tiers, and the temptation to over-engineer
Sooner or later someone in a commercial meeting will ask for per-plan rate tiers, where the enterprise customers get higher limits than the starter plan. This is reasonable, and it is also where rate limiting quietly turns into billing logic, so tread carefully. Keep the mechanism dumb and the policy in configuration.
Store the limit values as data attached to the tenant, not as branches in code. The middleware should look up "what is this tenant's bucket size" and apply it mechanically, so that changing a customer's tier is a config write, not a deploy. I have seen a team hard-code three tiers as an enum and then need an engineering release every time sales negotiated a custom limit for a big account. That is a self-inflicted wound. Numbers that a salesperson might negotiate belong in a table, never in a switch statement.
Test the limiter like it is load-bearing, because it is
A rate limiter is one of those components where the failure modes only show up under concurrency, which means normal unit tests will happily pass while the thing is broken in production. You have to test it the way it actually gets used: many simultaneous requests racing for the same bucket.
Write a test that fires several hundred concurrent requests against a single key and asserts that the number allowed through is within one or two of the configured limit, not wildly over. That test is what catches the non-atomic decrement. Also test the clock edges deliberately, because refill logic that quietly assumes time only moves forward will do strange things when a server's clock drifts or a container's monotonic time resets. And test the fail-open path by killing Redis mid-run and confirming traffic still flows. If you never test the degraded state, you do not actually know how the system behaves when it matters most.

Conclusion
If I had to compress all of this into one instruction for a team building their first public API, it would be this: ship the limiter in shadow mode on day one, before you have a single external customer, and just watch. The engineering of a token bucket is a solved problem you can get right in an afternoon. The hard part, the part that earns trust or destroys it, is choosing humane limits, telling developers the truth about them in every response header, and never letting your own safety mechanism become the outage. Get those three right and your rate limiter stops being a wall your customers hit and becomes something they barely notice, which is exactly what good infrastructure is supposed to be.
Get new posts in your inbox
Occasional, practical notes on engineering leadership, fintech, and building with AI. No spam, unsubscribe anytime.
Comments (0)
Leave a Comment
No comments yet. Be the first to comment!
