Systems / 2026 / Live in the browser
Tail Latency Lab
Fan-out tail amplification, hedged requests on idle and loaded replicas, and power-of-two-choices load balancing, each simulated on seeded runs and checked against its closed form.
01
The problem
A page that waits on 100 backends is only as fast as the slowest. If each backend has a 1% chance of a one-second hiccup, 63% of page loads hit one, so a leaf's p99 becomes the request's median. The usual fixes, sending duplicate requests or spreading load, can make things worse when the cluster is busy.
02
How I approached it
Dependency-free modules on a mulberry32 RNG and a lognormal service-time model with a rare slow mode. Fan-out is checked against 1 - (1 - p)^n. Hedged requests follow Dean and Barroso (CACM 2013): the backup fires past the p95, calibrated on a separate seed, and the simulated tail is checked against the two-replica survival product. A second simulation runs hedges on loaded FIFO replicas, where a backup queues and holds another server until one copy answers. Load balancing is Mitzenmacher's supermarket model (IEEE TPDS 2001), with random and best-of-d runs matched to the fixed-point formula for expected time in system.
03
The outcome
On idle replicas a p95 hedge cuts p99.9 from 1,555 ms to 34 ms for 4.7% extra requests and 1.5% extra work. On loaded replicas it costs about 3% of capacity at every load, while duplicating every request wins at 20% load and pushes the median up 26x at 70%. At 90% load, best of two servers cuts mean time from 10.0 to 2.63 against a predicted 2.61. Zero dependencies, 19 tests.