What the simulation models
It is small, but the parts that make real systems fall over are there. Each tier is a pool with a fixed capacity. Latency stays flat until about 80% utilisation and then bends sharply upward, the way queueing delay does. Past 100% a tier builds a backlog, and once the wait passes the one-second client timeout the excess is dropped. Overload doesn't just fail the extra requests: it makes every request slow.
That is why the rate limiter is worth a dollar an hour. Shedding the traffic you can't serve keeps the requests you can serve fast. Rejected requests still count against the SLO, but it is a few percent instead of everything.
- Cache. Absorbs up to 90% of reads once warm. A cold cache is the same as no cache, which the 21:00 deploy will remind you of.
- Read replicas. Split read load with the primary. They take longest to provision because they copy data first.
- Queue and workers. Writes behind a queue are acknowledged at once and drained by workers. The workers back off when the primary has no headroom, so a burst becomes a backlog instead of an outage.
The p99 shown is derived rather than sampled. Each path a request can take (cache hit, primary read, replica read, sync write, queued write) has a mean latency and a backlog wait. Its tail is treated as exponential, and p99 is the latency that 1% of all requests exceed. The error budget is 1% of the day's projected traffic.