Lab 05 of 09Progressive delivery simulator
Canary Release
Roll v2 out to 1% → 5% → 25% → 100% of traffic, slip a bug into it, and watch automated burn-rate analysis halt the rollout before most users notice.
- 1%
- 5%
- 25%
- 50%
- 100%
- Traffic on v2
- 0%
- v2 error rate
- —
- Burn rate
- —
- Failed on v2
- 0
Simulation · 600 req/s · SLO 99.5% · multi-window burn-rate alerts (fast 10 s > 10×, slow > 2×) · p95 guard 1.25× · time is compressed
Built by Melih Kızmaz · runs entirely in your browser
What you are looking at
Six hundred requests a second flow from users through a weighted router. When you deploy, v2 starts on a single pod and gets 1% of the traffic; every step bakes for a few seconds while an analysis run watches it, then the weight goes up — 5%, 25%, 50%, 100% — and the stable pods scale down as the canary scales up. This is the shape of an Argo Rollouts or Flagger canary, compressed from minutes into seconds.
Burn rate, not raw error rate
The analysis speaks in error budget. With a 99.5% success SLO you may fail 0.5% of requests; a burn rate of 10× means v2 is spending that budget ten times too fast. The rollout halts on a fast burn (>10× over the last 10 s) or a slow burn (>2× over everything v2 has served), or when v2’s p95 is more than 1.25× the stable version’s. It also refuses to judge with too little data: at 1% of traffic the first verdict is often “inconclusive”, which is honest — thirty requests cannot tell 0.1% from 1.5%.
Why the subtle bug gets further
Inject 500s and the fast-burn alert trips on the first or second step; only a handful of users ever see an error. Inject the subtle bug and it usually sails through 1% — the sample is too small — and is caught a step or two later by the slow-burn window. That is the real trade-off of canaries: small first steps limit the blast radius, but they also limit how much you can learn, so good rollouts pair them with long-window checks.
Blue/green for contrast
Switch the strategy to blue/green: the new version comes up on an idle environment, passes its smoke tests (they almost always do), and then the router flips 100% of traffic at once. Rollback is instant — flip back — but by the time the analysis has a verdict, every user has been exposed. Compare the Failed on v2 counter after the same bug under both strategies. All numbers here are simulated; the thresholds mirror common SRE multi-window burn-rate alerting.