So you'd think this would be the case. I certainly did when I first joined. In fact, I actually proposed something even more extreme than your suggestion (use thompson sampling to control the rate of task restarts and rollbacks). But in practice such a continual canarying process isn't actually any better than a staged canary (like .1% -> 1% -> 10%).
Consider the types of issues you run into, there are
- Things that are painful, but not destructive (minor performance regressions)
- Things that are highly destructive (data gets deleted, major performance regressions, etc.)
The first you can handle being deployed to many people, so if you detect at 7.8% of your users instead of 10% doesn't much matter, you can run it at 10% forever without issue.
The second you'll detect on a smaller population because the changes are catastrophic.
>In reality, neither continual-canarying or google-canarying work well for detecting anomalous metrics because the rollout process itself causes applications to restart, which in itself changes their performance characteristics (cold caches, empty queues, first-use delay, jit optimization, etc.).
The problem here is that most of the time, you don't care about startup behavior, but steady-state behavior. Imagine that a new version introduces a memory leak, so the old behavior was to linearly increase to 1GB of memory over 1 hour and then level off, and the new behavior is to increase unbounded until eventually the system OOMs and restarts or whatever.
A "google style" canary would roll out to 1% of tasks or something, wait a few hours, and notice the difference and roll back. You might even experience 1% of tasks restating, but it's likely that the system can sustain that.
With the continuous canary, you'll release to 1 task per minute or whatever, and only be able to notice any change after you've pushed out to 60 tasks, and the change will likely only be detectable with any confidence once you have a notable difference in 120 or so. At that point, you've released to more than 1% of your tasks (either that, or it takes you 8 days to release a new version).
You can fix that by slowing the rate of releases, but now it takes you 8 weeks to release a new version.
Plus, even worse, you're now in a much more fail-open environment. With a 1% release, you can sustain in the kind of bad state if all of your qualification tooling fails. If you're continuously canarying though, you have to be much more careful to make sure that your tooling won't continue to push if the tooling itself is broken or getting unusual results. It's a more risky set up.