Skip to content

Article

Your Rate Limiter Will Fail Open or Closed: Choose on Purpose

When a shared counter fails, decide which requests may continue, which must stop, and how local limits, timeouts, and load shedding prevent another outage.

Published 4 Aug 2026Updated 1 Oct 202613 min read
Rate Limiting · Redis · Reliability · API Design · Production
Requests reaching four API instances while a failed shared counter splits the path into open, closed, and local-share fallbacks.
On this page (16)

A customer sends 120 requests per second to an API that allows 50. Four application instances ask Redis whether each request fits. Then Redis stops answering for 30 seconds. If the instances allow every request, the noisy customer can overload the backend. If they refuse every request, two quiet customers lose service because the limiter's database failed.

A rate limit is a rule that gives one caller a measured amount of work over time. A limiter shared by several servers needs shared counters in something such as Redis, an in-memory data store that can update those counters quickly. When that dependency is unreachable, fail open means allow the request and fail closed means refuse it. The code always makes one of these choices; when the product has not chosen, a library default chooses for it.

Most discussions of rate limiting spend their time on algorithms. The algorithm matters less than this decision, and less than two others that sit beside it: whether the check is atomic, and what the limiter tells clients when it says no.

For a limiter that keeps customers from crowding each other out, a local share is a useful degraded mode: each server enforces the global limit divided by an estimate of the live server count until the store returns. Start with a short store timeout chosen from measured network latency, put a circuit breaker beside it, and alert on every fallback. Fail closed where allowing traffic is itself the larger harm, such as credential guessing, one-time-code abuse, or a hard spend boundary. Scope that decision to the protected route instead of applying it to the whole service.

How to read the diagrams

  • The limiter and its counters
  • The shared store and its failures
  • Requests let through
  • Requests refused, and outages
  • Degraded mode: timeouts and fallbacks

One request through a limiter

A load balancer spreads requests across application instances. Each instance derives a caller key, such as an authenticated customer ID, and asks the limiter for one decision. An allowed request reaches the API. A caller over its policy receives HTTP 429 Too Many Requests; Retry-After: 1 asks it to wait one second before trying again.

The normal path while Redis is healthy

  1. 0 msClient → Load balancer

    POST /orders

    customer = acme

  2. 1 msLoad balancer → API instance

    route request

  3. 2 msAPI instance → Redis

    check and increment

    key rl:{acme}

  4. 3 msRedis → API instance

    allowed · 12 remain · reset in 1 s

  5. 4 msAPI instance → Client

    continue to the API

The decision and counter update must happen as one indivisible operation, or two instances can both consume the final allowance.

Two common algorithms decide how allowance changes over time. A token bucket refills tokens at a steady rate; a request spends one, and saved tokens permit a bounded burst. A sliding window counter estimates how many requests fell within the immediately preceding interval, smoothing the hard edge between fixed windows. The interactive token-bucket and failure-mode simulator lets you change traffic, rate, burst capacity, and outage policy.

A race condition occurs when the outcome depends on which concurrent operation happens first. Redis prevents it here by executing the read, comparison and increment atomically, as one operation no other request can interleave with.

Your framework already chose

Start by checking what the current stack does. Defaults disagree across products, and sometimes within one product family:

what happens when the store or limit service failsdefaults checked against each project's current docs or source · September 2026
tool
Envoy global rate limit
default
open
detail
failure_mode_deny defaults to false. The call to the limit service times out after 20 ms. If you set it to deny, the response is a 500.
tool
Envoy Gateway global rate limit
default
open
detail
The BackendTrafficPolicy global-rate-limit call lets traffic pass unless failClosed is set. External authorization is a separate feature with a different default.
tool
Istio (documentation example)
default
closed
detail
The official EnvoyFilter example sets failure_mode_deny: true and a 10 s rate-limit-service timeout.
tool
Kong rate-limiting
default
open
detail
fault_tolerant defaults to true: requests are proxied anyway, "effectively disabling the rate-limiting function until the data store is working again."
tool
Kong Rate Limiting Advanced
default
local
detail
With a local strategy, each node limits independently. Its sync modes periodically reconcile counters, so accuracy depends on the configured sync interval; a store outage leaves nodes enforcing local knowledge until synchronization resumes.
tool
Spring Cloud Gateway
default
open
detail
The Redis limiter allows the request on any Redis error. The code comment: "We don't want a hard dependency on Redis to allow traffic."
tool
express-rate-limit
default
closed
detail
passOnStoreError defaults to false: "the default is to 'fail closed' and block all requests if the datastore becomes unavailable."
tool
NestJS throttler
default
closed, in effect
detail
Not documented; from the source, the guard doesn't catch storage errors, so a store failure becomes a 500.
tool
rate-limiter-flexible
default
your choice
detail
An optional "insurance" limiter takes over on store errors. Without one, the error reaches your code.
tool
NGINX limit_req
default
no shared store
detail
State lives in each instance's shared memory, so there is nothing remote to lose. Refusals are 503 by default, not 429.

The same configuration idea gets opposite defaults: Envoy and Kong allow while express-rate-limit refuses. Switching a library can therefore flip the failure mode without changing the product requirement. Several fail-closed paths also answer with a 500 rather than a 429. Many clients retry 5xx responses, so the refusal can create more traffic during the dependency failure.

A running example

To make the options concrete, take an illustrative setup. An orders API runs on four instances behind a load balancer. Each customer may make 50 requests per second, and the counters live in one Redis. One customer, acme, has a sync job stuck in a loop and is sending 120 requests a second. Two others, globex and initech, send 20 and 5. Then Redis fails over and is unreachable for 30 seconds.

Here is what each policy lets through for those 30 seconds:

orders API · per-customer limit 50/s · Redis unreachable for 30 s
customersendinglet throughrefused
Redis up
acme120/s50/s70/s, as 429
globex20/s20/snone
initech5/s5/snone
Fail open
acme120/s120/snone
globex20/s20/snone
initech5/s5/snone
Fail closed
acme120/s0120/s, as 500
globex20/s020/s, as 500
initech5/s05/s, as 500
Local share
acme120/s50/s70/s, as 429
globex20/s20/snone
initech5/s5/snone

Failing open costs 70 extra requests a second from one customer for 30 seconds. Whether that matters depends entirely on what sits behind the limiter. Failing closed turns a Redis failover into a 30-second outage for every customer, including the two who did nothing wrong. The local share gets almost everything right: each instance enforces 12.5 requests a second per customer, so acme's evenly spread traffic is held to 50 in total. "Almost" hides two assumptions, that acme's requests are spread evenly and that there are still four instances; both come up below. The rate limiter simulator runs these same failure modes against generated traffic.

Choose by what the limiter protects

There is no single right answer, because limiters protect different things, and the cost of a wrong answer differs by an order of magnitude between them:

failure policy by purpose
the limiter protects
Fairness between customers of a shared API
when the store fails
Fail open to a local share
why
A few minutes of an approximate limit costs little. Refusing every customer is an outage.
the limiter protects
A backend that falls over under load
when the store fails
Local share, plus load shedding that needs no store
why
The limiter shouldn't be the only thing between a spike and the database. Concurrency limits and client-side throttling keep working when Redis doesn't.
the limiter protects
Credentials: logins, password resets, one-time codes
when the store fails
Fail closed or fail static, on those routes only
why
Here an unlimited minute is the attack. Keep the refusal to the routes that need it.
the limiter protects
Spend on a paid upstream, such as model APIs or SMS
when the store fails
A strict local cap
why
A best-effort limiter isn't a spend control. AWS says the same of its own usage plans.
the limiter protects
Batch work sharing a resource
when the store fails
Fail closed: pause
why
A batch job can wait. Google's Doorman design calls this the right mode for offline work.

Google's Building Secure and Reliable Systems frames the trade-off in general terms: "you must balance between optimizing for reliability by failing open (safe) and optimizing for security by failing closed (secure)," and "each organization must first determine its minimal nonnegotiable security posture." The practical consequence is that one service often needs both policies, one per route.

Who fails open, and why

The teams that have written about running limiters at scale mostly chose to fail open, and they gave the same reason: a limiter that can take the service down is a bigger risk than the traffic it would have stopped.

2017

Stripe

"Make sure that if there were bugs in the rate limiting code (or if Redis were to go down), requests wouldn't be affected. This means catching exceptions at all levels so that any coding or operational errors would fail open." The accompanying code adds: "Make sure to set an alert so you know if this is happening too much."

2026

Uber

"If the control plane became unavailable, clients failed open, allowing traffic to continue rather than risk self-inflicted outages." Uber moved away from Redis counters because every request would have paid a remote round trip; the new system runs at around 80 million requests per second across more than 1,100 services.

2025

Databricks

"Our single Redis instance was our single point of failure." The replacement is optimistic: requests are let through by default "unless we already knew we wanted to reject those particular requests," because backends could tolerate some traffic over the limit.

2017

Cloudflare

Cloudflare described an asynchronous counting path: "the increment jobs are run asynchronously without slowing down the requests." That design avoids adding a synchronous counter round trip to each request.

Stripe's post also lists the two tools that make failing open safe to operate: a kill switch to disable any limiter that misbehaves, and dark launches, where a new limiter only logs what it would have blocked before it blocks anything.

Exact limits are rarer than you think

Failing open to an approximate limit sounds like a compromise. It helps to know that the large managed limiters are approximate all the time, by design:

how exact the managed limiters claim to be
service
AWS API Gateway
in their own words
Throttles and quotas are applied "on a best-effort basis and should be thought of as targets rather than guaranteed request ceilings." And: "Don't rely on usage plan quotas or throttling to control costs or block access to an API."
service
AWS WAF rate-based rules
in their own words
"It's not intended for precise request-rate limiting." Detection can lag "for up to several minutes… Usually, this delay is below 30 seconds."
service
Cloudflare rate limiting rules
in their own words
"Each data center maintains its own counters." There is no global counter across the network.
service
Cloudflare Workers rate limiting
in their own words
"Permissive, eventually consistent, and intentionally designed to not be used as an accurate accounting system."
service
NGINX Plus, synced zones
in their own words
Nodes decide without consulting each other; "In average, the cluster limit is properly imposed." NGINX calls it an AP system.

Cloudflare measured its own approximation on 400 million requests: "0.003% of requests have been wrongly allowed or rate limited," with an average difference of 6% between the real and the estimated rate. A local fallback that is off by some percentage for the duration of a Redis failover is well within what these systems accept every day.

The local share, and the multiplier hiding in it

The local fallback gives each instance a fraction of the global limit. The fraction is the part that goes wrong, because it depends on how many instances are running, and that number changes.

incident.io described both halves of the problem. Their limiter used to fail open: "If Valkey went down, we would fail open, meaning we wouldn't rate limit at all." They changed it so that when the store is unreachable, "a pod could now grant up to its own budget and then start denying, instead of failing open." That introduced the next problem: "the effective limit across the whole system ended up being limit × pod count." The fix was to divide the local limit by the live pod count, read from Kubernetes.

The other failure is a divisor that goes stale. In March 2026, part of Discord's voice outage came from a per-host limit that had not been updated after a migration: "As we were running a higher pod count than the previous VM architecture, the per-host rate limit allowed more capacity through" (post).

In the running example, the effective limit on acme during the outage depends on the divisor:

acme's effective limit while Redis is down (requests per second)

Effective limit across all instances

A fixed divisor is right only until the next scale-up. Dividing by the live count keeps the total near the limit.

y: req/s

From left: the shared store; limit ÷ 4 with 4 instances; limit ÷ 4 after scaling to 8; limit ÷ live count with 8; and the full limit on each of 8 instances.

The other assumption is even spread. AWS's Builders' Library notes that dividing a quota by the number of servers "assumes that requests are relatively uniformly distributed across servers." If acme's sync job holds one keep-alive connection that always lands on the same instance, its local share is 12.5 requests a second, a quarter of what it's entitled to. For a fallback that runs for seconds or minutes, that is usually acceptable: it errs on the strict side, and only for the callers whose traffic is concentrated.

The request path combines those decisions as follows:

Sketch: the store first, then this instance's share
const STORE_TIMEOUT_MS = 20;
 
async function allow(caller: string): Promise<Decision> {
  if (!breaker.isOpen()) {
    try {
      const decision = await withTimeout(store.check(caller), STORE_TIMEOUT_MS);
      breaker.recordSuccess();
      return decision;
    } catch (err) {
      breaker.recordFailure();
      metrics.increment("ratelimit.fallback", { reason: kindOf(err) });
    }
  }
  // The store is slow or down: enforce this instance's share of the limit.
  const share = limitFor(caller) / Math.max(1, liveInstanceCount());
  return localBuckets.take(caller, share);
}

The timeout is short because every request pays it while the store is slow. The circuit breaker stops calling a store that is already failing, so later requests go straight to the local bucket instead of each waiting 20 milliseconds. Every fallback increments a metric; otherwise a limiter can remain degraded for days without an obvious symptom.

A slow store is worse than a dead one

A store that refuses connections fails fast. A store that accepts connections and answers slowly holds every request that asks it something. That is the case timeouts exist for, and the defaults are worth reading together.

Envoy and its reference rate limit service when Redis stalls

  1. 0 msClient → Envoy

    POST /orders

  2. 0 msEnvoy → Limit service

    should this be limited?

    descriptor: customer = acme

  3. 1 msLimit service → Redis

    increment acme's counter

    Redis is mid-failover: no reply

  4. 20 msEnvoy

    timeout (default 20 ms): allowed, counted in failure_mode_allowed

  5. 20 msEnvoy → Client

    request forwarded to the orders service

  6. ≤ 10 sLimit service

    still waiting on Redis (REDIS_TIMEOUT default 10 s)

Read together, the defaults mean the proxy gives up after 20 ms and lets the request through, while the service behind it can wait up to ten seconds on Redis. Nothing is refused. The only trace is a counter.

Timeouts on the store call have a cost even when they work. Shopify's write-up on circuit breakers works through a worker whose Redis calls start timing out: "A worker which had a request processing rate of 5 requests per second now can only process half a request per second. That's a tenfold decrease in throughput!" That is why the timeout needs a circuit breaker next to it. And the right timeout depends on the network: Databricks reported a P99 network latency of 10 to 20 milliseconds in one cloud provider, so a timeout of a few milliseconds there would send ordinary slow requests to the fallback.

When failing closed is right, and what it costs

Failing closed is the right call where letting requests through is the harm itself: guessing passwords, requesting one-time codes, triggering password-reset emails. Even there, it should be scoped to those routes. A limiter that fails closed across a whole service makes every store outage a service outage, and the incidents where that happened were rarely store outages at all.

On 12 June 2025, a new quota policy check in Google Cloud's Service Control "did not have appropriate error handling nor was it feature flag protected." A null pointer crashed the binary everywhere, and Google Cloud APIs failed across regions for hours. Among the remediations: "We will modularize Service Control's architecture, so the functionality is isolated and fails open." Recovery then hit a second problem, a herd of restarting tasks overloading the Spanner table they depended on, because they "did not have the appropriate randomized exponential backoff" (incident report).

When you do fail closed on purpose, choose the status code on purpose too. Envoy returns a 500 by default (configurable with status_on_error), Envoy Gateway returns a 500, and Kong's documentation says that with fault tolerance off, "the clients will see 500 errors." A 503 with a Retry-After header tells clients the service is temporarily refusing work, which is what is actually happening.

Fail static: keep the decisions you already have

Google's security book describes a third option, from its denial-of-service defenses: "If the central controller fails, we don't want to either fail closed (as that would block all traffic, leading to an outage) or fail open (as that would let an ongoing attack through). Instead, we fail static, meaning the policy does not change." The payoff is economic: "Because we fail static, the DoS engine doesn't have to be as highly available as the frontend infrastructure, thus lowering the costs."

For a rate limiter, failing static means continuing to enforce the decisions each instance already holds. Some systems keep those decisions locally for speed. The reference Envoy rate limit service has an optional local cache of keys already over their limit, which it answers without reading Redis again. Cloudflare's servers, once they start mitigating a client, "will not even run another query for the subsequent requests" from it. A cache like that, kept through a store outage, gives you a limiter that stays closed for the callers it had already caught and open for everyone else.

Google's Doorman, an open-source capacity system, names the whole range for clients that can't renew their share. Pessimistic mode stops using the resource, "the right course of action for an off-line batch job." Optimistic mode continues, right "for on-line components that interact with users and with resources that are protected against overloading by other means." And safe mode uses a safe capacity the server calculates: "the available capacity divided by the number of clients."

Layers that need no store at all

The best protection against a store outage is not needing the store for the protection that matters most. Two layers work entirely locally.

The first is load shedding by concurrency. Google's SRE book suggests returning "an HTTP 503 (service unavailable) to any incoming request when there are more than a given number of client requests in flight." It protects the instance itself, whatever any counter says.

The second is client-side throttling. Each client tracks, over the last two minutes, how many requests it sent and how many the backend accepted, and rejects some of its own requests locally with probability max(0, (requests − K × accepts) / (requests + 1)). With K = 2, which the book says Google generally prefers, a client keeps sending until it is sending twice what gets accepted, and "requests above the cap fail locally without even reaching the network." It has, in the book's words, "no additional dependencies or latency penalties."

Envoy combines local and global limits for a related reason: "a local token bucket rate limit can absorb very large bursts in load that might otherwise overwhelm a global rate limit service." The global limit handles fairness; the local one makes sure the global one is never the thing that falls over.

The check itself: atomic, and on one clock

A shared limiter's check is read, compare, write. If two instances read the counter at the same moment, both see room, and both increment, the limit has been exceeded. Figma's write-up names it: "the 'read-and-then-write' behavior creates a race condition, which means the rate limiter can at times be too lenient." In Redis, the fix is a Lua script, because "Redis guarantees the script's atomic execution."

Two further rules from the Redis documentation shape the script. Every key it touches must be passed in: "Scripts should never access keys with programmatically-generated names," which is also what makes a script work on Redis Cluster. And the current time should come from Redis, not from the application, so that instances with slightly different clocks don't disagree about which window they're in. Since Redis 5, scripts replicate their effects rather than the script itself, so calling TIME inside one is safe. Spring Cloud Gateway's limiter does exactly this.

A sliding window counter fits both rules in one key per caller:

Sketch: a sliding window counter, one key per caller, time from Redis
-- KEYS[1] = the caller's counter, e.g. "rl:{acme}" (a hash: w, c, p)
-- ARGV[1] = limit per window, ARGV[2] = window length in seconds
local limit, window = tonumber(ARGV[1]), tonumber(ARGV[2])
 
local t   = redis.call('TIME')                     -- {seconds, microseconds}
local now = tonumber(t[1]) + tonumber(t[2]) / 1e6
local w   = math.floor(now / window)               -- which window we're in
local pos = (now - w * window) / window            -- how far into it, 0 to 1
 
local s = redis.call('HMGET', KEYS[1], 'w', 'c', 'p')
local sw, c, p = tonumber(s[1]), tonumber(s[2]) or 0, tonumber(s[3]) or 0
if sw == nil or sw < w - 1 then c, p = 0, 0         -- idle for two windows or more
elseif sw == w - 1 then p, c = c, 0 end             -- a new window: current becomes previous
 
local estimate = p * (1 - pos) + c                 -- previous window, weighted by overlap
local reset = math.ceil(window - (now - w * window))
local allowed = estimate + 1 <= limit
if allowed then c = c + 1 end
 
redis.call('HSET', KEYS[1], 'w', w, 'c', c, 'p', p)
redis.call('EXPIRE', KEYS[1], window * 2)
return { allowed and 1 or 0, math.max(0, math.floor(limit - estimate - 1)), reset }

The hash stores three numbers: the window it last saw (w), that window's count (c), and the previous window's count (p). Work through a ten-second window with a limit of ten. At 19.0 seconds the request is 90% through window 1. At 20.5 seconds it has crossed into window 2, so the old current count moves into p.

Redis hash rl:{acme} · limit 10 per 10 s
19.0 s · before request 1
moment
window 1 · pos 0.90
w
1
c
8
p
4
estimate before request
4 × (1 − 0.90) + 8 = 8.40
result
8.40 + 1 ≤ 10 · allow
19.0 s · after request 1
moment
same window
w
1
c
9
p
4
estimate before request
stored count increased
result
return allowed = 1 · remaining = 0
20.5 s · request 2 crosses the boundary
moment
window 2 · pos 0.05
w
2
c
0 → 1
p
9
estimate before request
9 × (1 − 0.05) + 0 = 8.55
result
8.55 + 1 ≤ 10 · allow
Crossing a boundary does not erase the previous window. Ninety-five per cent of its nine requests still contributes at 20.5 seconds.

The estimate weights the previous window by how much of it still overlaps the trailing interval. Thirty per cent of the way into a window, seventy per cent of the previous count remains. The braces in rl:{acme} are a Redis Cluster hash tag; if the script later uses more keys for this caller, the tag can keep those keys in one cluster slot.

Shared state has smaller traps too. GitHub's rate limiter migration found that Redis replicas "don't expire data until they receive instructions to do so from their primaries," so reads from a replica could see a window that had already ended. Their earlier design, on a shared Memcached, "would sometimes evict rate limiter data, even when it was still active." A limiter's store should be its own, not a cache that something else can fill up.

Telling clients what happened

A refusal is useful when the client can act on it. The current IETF document was draft-ietf-httpapi-ratelimit-headers-11, last updated 23 May 2026 and still an active Internet-Draft when checked in September 2026. It is work in progress rather than an RFC. Its current syntax uses one header for policies and another for changing service state.

Sketch: a refusal that says which limit was hit
HTTP/1.1 429 Too Many Requests
RateLimit-Policy: "per-second";q=50;w=1, "per-day";q=1000000;w=86400
RateLimit: "per-second";r=0;t=1
Retry-After: 1
Content-Type: application/problem+json
 
{"title": "Rate limit exceeded", "policy": "per-second", "retry_after_seconds": 1}

q is the quota and w the policy window in seconds; r is what remains. In draft 11, t is the effective window associated with the reported service state, not a countdown to reset. Retry-After carries the server's retry advice. Naming the policy matters when there are several: a client that hit the per-second limit should slow down, while one that hit the per-day limit may need to stop until the longer window changes.

Status codes carry meaning too. 429 says this caller is over its limit. 503 says the service is overloaded for everyone. NGINX's limit_req defaults to 503, which is worth changing if the limit is per caller, so clients can tell the two apart.

When the limiter is the outage

Limiters are code on every request's path, and the incidents they cause are rarely store failures. More often the limiter enforces the wrong thing, confidently:

Published incidents where the limiter itself was the fault

  1. 2025-06

    Google Cloud

    A new quota check without error handling crashed Service Control globally. The fix is to isolate the check so that it fails open.

    status.cloud.google.com

  2. 2025-10

    PostHog

    A per-IP limiter behind a load balancer "saw all traffic as coming from a single IP (our load balancer)." About 97% of feature flag evaluation requests got 429s worldwide for 72 minutes. Nothing alerted on 429s, so detection took 62 minutes, and the limiter had only been tested directly, never behind the production load balancers.

    posthog.com/handbook post-mortem

  3. 2026-01

    GitHub

    Blocking rules added during past incidents stayed in place and kept catching legitimate users, about 3 to 4 requests in every 100,000. GitHub now treats "incident mitigations as temporary by default."

    github.blog engineering

  4. 2026-03

    Discord

    A per-host limit that wasn't re-tuned after a migration to more pods let more traffic through than intended, and a second limiter "was outdated or unset."

    discord.com/blog

  5. 2026-04

    GitHub

    A bug "incorrectly applied a rate limit globally across all users, rather than scoping it to the individual installation." About 84% of new agent session requests were delayed. In a related incident, a caching bug kept callers limited after their window had ended.

    github.blog availability report

PostHog's incident is worth reading twice, because every part of it is common. The key was the client IP, and the client IP was the load balancer. The failure was invisible because nobody alerted on 429s. The test that would have caught it needed production's network shape. The SRE book makes the same point about rarely exercised fallback paths: "the code path you never use is the code path that (often) doesn't work."

A checklist

Decide

  • For each limiter: what it protects, and what it does when the store fails
  • Check your library's default rather than assuming it
  • Fail closed only on routes where an unlimited minute is the attack
  • Choose the refusal status: 429 for a caller, 503 for overload

Fall back

  • Local share = limit ÷ live instance count
  • Store timeout in tens of milliseconds, with a circuit breaker
  • Keep over-limit decisions locally, so callers already caught stay caught
  • Load shedding and client throttling that need no store

The check

  • One atomic script, every key passed in KEYS
  • Time from the store, not from each instance
  • Key by the real caller, never the load balancer's address
  • A store of its own, not a shared cache

Operate

  • Alert on fallbacks and on 429 rates
  • Dark-launch new limits before enforcing them
  • A kill switch for every limiter
  • Test the fallback by taking the store away
  • Expire limits added during incidents

Sources