Post

Retry Storms: How Client Retries Turn a Short Dip into a Long Outage

A two-second capacity dip, four retry policies and one retry budget, simulated in about a hundred lines of Go: how retries multiply load, why jitter helps, and what a budget changes.

Retry Storms: How Client Retries Turn a Short Dip into a Long Outage

Retries are the cheapest reliability feature there is, and the only one that can make a small failure bigger. A request fails, the client sends it again, and usually that works. The trouble starts when the failure is not random bad luck but a service that is out of capacity. Every retry is then extra work aimed at the one component that has none to spare.

That is not a theoretical worry. In its postmortem for the August 17, 2026 outage, GitHub wrote that errors in a set of services "triggered a client-side retry loop that increased traffic during recovery," and that the team "had to mitigate that behavior before we could safely restore traffic." The same document says they are now applying "consistent retry limits, retry budgets, and variable timeouts across service-to-service interactions to prevent retry storms and cascading load." The outage lasted 7 hours and 47 minutes, and the retries were not its root cause. They were what made recovery slow.

This post builds the mechanism from the bottom: the arithmetic of amplification, what backoff and jitter each fix, what a retry budget adds, and a small simulation you can run and change.

Where the amplification comes from

One client retrying up to three times is a 4x multiplier in the worst case. That alone is survivable. The dangerous version is layered:

flowchart LR
  U[User] --> A[Frontend]
  A -->|up to 3 tries| B[Service]
  B -->|up to 3 tries| C[Database]

If every layer makes three attempts and the database is rejecting everything, the frontend's one request becomes 3 calls to the service, each of which makes 3 calls to the database: 9 database attempts. Add a third layer in between and it is 27. The multiplication happens in exactly the place that is already failing.

Google's SRE book addresses this with two limits and one rule. A per-request limit caps a single request at three attempts. A per-client budget lets a client retry only "as long as" the ratio of retries to requests is below 10%. The book says layering that 10% budget on top of the per-request limit "reduces the growth to just 1.1x in the general case," where the per-request limit alone could triple the traffic. The rule: retry at "the layer immediately above the layer that is rejecting them," and nowhere else.

What backoff fixes, and what it does not

Exponential backoff spaces attempts out: wait 1 unit, then 2, then 4. It reduces how much load retries add per unit of time, but it has a flaw that the AWS Architecture Blog's analysis of exponential backoff and jitter makes explicit: clients that failed together also retry together. If a thousand clients are rejected in the same instant, they all wake up after 1 unit, collide, all fail, and all wake up after 2. The retries arrive in synchronised waves.

Jitter breaks the synchronisation by randomising the wait. The AWS post compares three variants, and the simplest one, full jitter, is:

1
sleep = random(0, min(cap, base * 2 ^ attempt))

In their simulation of clients competing for a contended resource, full jitter and decorrelated jitter did substantially less total work than plain exponential backoff or the "equal jitter" variant, which keeps half of the deterministic delay. Full jitter needed the least work overall. The cost is that an individual client can occasionally retry almost immediately, which is fine, because the point is the aggregate arrival rate, not any one client.

Backoff and jitter shape when retries arrive. Neither limits how many there are. Only a budget does that.

A simulation to see it

The model is deliberately small. Time advances in ticks of 100 ms. A service completes up to 100 requests per tick and rejects the rest instantly. Users send 95 new requests per tick, so the service runs at 95% utilisation, which is high but normal for a busy system. For 20 ticks (two seconds), capacity drops to 20. Each rejected request is retried up to three times according to the policy. A request that exhausts its retries counts as a failed user.

The budget policy uses one shared token bucket: every new request adds 0.1 token (capped at 10) and every retry costs 1. This is a single-bucket approximation of many independent per-client ratios, and it is a simplification, not a faithful model of any production client.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
package main

import (
	"fmt"
	"math/rand"
)

const (
	ticks     = 400 // one tick = 100 ms
	capacity  = 100 // requests the service completes per tick when healthy
	arrivals  = 95  // new user requests per tick (95% utilisation)
	dipStart  = 100 // capacity dips for 20 ticks (two seconds)
	dipEnd    = 120
	dipCap    = 20
	maxRetry  = 3
	baseDelay = 1 // ticks
	capDelay  = 64
)

type policy struct {
	name   string
	delay  func(attempt int, r *rand.Rand) int // ticks to wait before retry number attempt (1-based)
	budget bool
}

func expo(attempt int) int {
	d := baseDelay << attempt
	if d > capDelay {
		d = capDelay
	}
	return d
}

func main() {
	policies := []policy{
		{"no retries", nil, false},
		{"immediate retry", func(a int, r *rand.Rand) int { return 1 }, false},
		{"exponential, no jitter", func(a int, r *rand.Rand) int { return expo(a) }, false},
		{"exponential, full jitter", func(a int, r *rand.Rand) int { return 1 + r.Intn(expo(a)) }, false},
		{"full jitter + 10% budget", func(a int, r *rand.Rand) int { return 1 + r.Intn(expo(a)) }, true},
	}
	fmt.Printf("%-28s %10s %12s %10s %10s\n", "policy", "peak load", "total sent", "failed", "recovery")
	for _, p := range policies {
		run(p)
	}
}

func run(p policy) {
	r := rand.New(rand.NewSource(1))
	// pending[t][a] = number of requests that will make attempt a (retry number) at tick t
	pending := make([][maxRetry + 1]int, ticks+capDelay+2)
	tokens := 0.0
	peak, sent, failed, recovery := 0, 0, 0, -1
	for t := 0; t < ticks; t++ {
		var offered [maxRetry + 1]int
		offered = pending[t]
		offered[0] += arrivals
		total := 0
		for _, n := range offered {
			total += n
		}
		sent += total
		if total > peak {
			peak = total
		}
		cp := capacity
		if t >= dipStart && t < dipEnd {
			cp = dipCap
		}
		served := cp
		if served > total {
			served = total
		}
		// which attempts get served is random: spread the rejections proportionally
		rejected := total - served
		tokens += 0.1 * float64(offered[0])
		if tokens > 10 {
			tokens = 10
		}
		for a := 0; a <= maxRetry; a++ {
			rej := 0
			for i := 0; i < offered[a]; i++ {
				if r.Intn(total) < rejected && rejected > 0 {
					rej++
				}
			}
			// exact proportional rounding is not needed; the seed keeps it reproducible
			for i := 0; i < rej; i++ {
				if p.delay == nil || a == maxRetry {
					failed++
					continue
				}
				if p.budget {
					if tokens < 1 {
						failed++
						continue
					}
					tokens--
				}
				pending[t+p.delay(a+1, r)][a+1]++
			}
		}
		if t >= dipEnd && recovery < 0 && total <= capacity {
			recovery = t - dipEnd
		}
	}
	fmt.Printf("%-28s %10d %12d %10d %8d\n", p.name, peak, sent, failed, recovery)
}

"Recovery" is the number of ticks after the dip ends until the offered load first fits within capacity again. With the fixed seed above, one run prints:

1
2
3
4
5
6
policy                        peak load   total sent     failed   recovery
no retries                           95        38000       1508        0
immediate retry                     356        44038       1413       19
exponential, no jitter              345        46649       1213       65
exponential, full jitter            361        45638       1299       48
full jitter + 10% budget            108        38197       1495        1

This is one seed of a toy model, so read the shape and not the third digit.

What the numbers say

Three things stand out.

First, retries buy very little here. Without any, 1508 users fail. With the unbudgeted policies, between 1213 and 1413 do. That is a saving of 6 to 20 percent of failures, paid for with 16 to 23 percent more total traffic and a peak load of 345 to 361 requests per tick against a capacity of 100. The retries rescue some users, but mostly they re-request work that the service was never going to finish during the dip.

Second, the damage outlasts the cause. The capacity dip is over after 20 ticks, but with unbudgeted retries the offered load stays above capacity for another 19 to 65 ticks, because the backlog of waiting retries keeps arriving on top of the new traffic. At 95% utilisation there is only 5 requests per tick of spare capacity to drain it with, so the backlog clears slowly. This is the shape of a retry storm: the trigger has gone, and the system is still overloaded by its own echo.

Third, jitter helps but is not the fix. Plain exponential backoff is the slowest to recover in this run (65 ticks) because its retries stay synchronised in waves; full jitter recovers faster (48), but it is still far from the budgeted policy (1). The budget does not rescue more users than "no retries" (1495 against 1508 failed, essentially the same) but it adds only 0.5 percent more traffic, keeps the peak at 108 and returns to normal one tick after the dip ends.

Rules that follow

  • Retry only errors that can plausibly succeed on a second try, and treat "I am overloaded" as a signal to back off, not as an error to retry.
  • Retry at one layer. If the layer above already retries, the layer below should fail fast.
  • Add backoff with full jitter to every retry that exists. It is a few lines and it removes the synchronised waves.
  • Cap retries by ratio, not only by count. A per-request limit of three still allows a 4x amplification; a budget of roughly 10% bounds the aggregate near 1.1x regardless of how many requests fail.
  • Honour deadlines. A retry whose caller has already given up is pure waste.
  • Test the recovery, not just the failure. Inject a short capacity dip into a staging system at high utilisation and measure how long after it ends the system takes to return to normal. If that time is many multiples of the dip, the retries are the problem.

Limits of this model

Real services queue rather than reject instantly, clients time out instead of getting fast errors, and load shedding changes who is served first. The simulation also uses one shared budget where production clients each keep their own, and it never models users who give up and leave. Those details change the numbers. They do not change the mechanism: a retry is new load, it arrives when the service is least able to take it, and only a limit on the ratio keeps it from growing with the failure.

Sources

This post is licensed under CC BY 4.0 by the author.