backoff-lab

Retries are load

Every retry is another request the server has to handle at the worst possible moment. Below, the same seeded clients run through five backoff schedules fighting over one database row, then through four retry policies while a server loses capacity for a few seconds. Everything runs in this tab.

1. Contention: who wakes up together collides together

N clients update one row with optimistic concurrency. A write only commits if nobody else committed since that client's read, so losers back off and retry. Network hops take 10ms ± 2ms, as in Marc Brooker's AWS article this reproduces.

2. Overload: the storm that outlives its cause

A server handles 1,000 requests a second against steady traffic of . Clients time out after , but the server still spends capacity on requests nobody is waiting for. Capacity drops to 20% for a few seconds. Whether it comes back depends only on the retry policy.

3. In code

retryFetch applies all of it to fetch: only idempotent methods or requests with an Idempotency-Key are retried, Retry-After is respected, every retry spends from a shared budget, and an open breaker refuses at once.

  1. Brooker, Exponential Backoff And Jitter, AWS 2015
  2. Google SRE book, ch. 21 Handling Overload
  3. Huang et al., Metastable Failures in the Wild, OSDI 2022
  4. RFC 9110 9.2.2 idempotent methods, 10.2.3 Retry-After
  5. Fowler, CircuitBreaker
import { retryFetch } from "./src/retry.js";
import { createRetryBudget, createCircuitBreaker } from "./src/budget.js";

const budget = createRetryBudget({ ratio: 0.1 });
const breaker = createCircuitBreaker({ failureThreshold: 5 });

const res = await retryFetch("/api/orders/42", {}, {
  strategy: "full", base: 100, cap: 10_000,
  maxAttempts: 3, budget, breaker,
});