Blog / Engineering
Patch ring design in depth: what has to be true before a ring is allowed to promote
Ring-based rollout is one of those practices nobody argues with and almost nobody specifies. "Canary, then pilot, then everyone" is a diagram, not a design. The design is in the rules: how membership is decided, what freezes and when, what counts as evidence that a ring passed, and what happens the moment it did not.
Those rules are where a ring stops being a safety mechanism and becomes a delay you route around at 4pm.
Membership has to be decided once, then frozen
The most common failure is a ring defined by a live query — "all workstations at site A." It reads well and it is wrong, because the population changes underneath the rollout. A device enrolled after the canary passed inherits a pass it was never part of. A device that fails is removed from the group by an unrelated cleanup and the failure disappears with it.
A deployment should snapshot its target list at creation and use that frozen list from then on. What you validated is then a specific set of machines, not a description that happened to match a different set of machines yesterday. It also makes the rollout auditable after the fact, which a live query never is.
A ring defined by a live query cannot tell you what it actually tested, only what it would match if you asked again now.
Ring one should be small enough to be boring
Our own split is a single canary machine, then a pilot ring of roughly a fifth of the population, then everything else. The canary being exactly one is deliberate: it is the smallest population that can produce a real failure, and a one-machine failure is a conversation rather than an incident.
The pilot ring is where proportion matters. Too small and it is a second canary that tells you nothing new. Too large and you have skipped straight to broad deployment with a longer name. A fifth is enough to surface the environmental variation — different hardware, different line-of-business software, different users doing unusual things — that a single machine cannot.
Promotion needs evidence, not absence of complaints
This is the rule that does the most work: a ring may only promote when every device in it has reached a terminal state, and every one of those terminal states is a success. Not "no failures reported." Not "most of them finished." Every device, terminal, successful.
- A device that has not finished blocks promotion — an incomplete ring is not a passing ring, it is an unfinished one, and the difference matters most when a machine is off because it broke.
- Any non-success terminal result halts the deployment rather than logging a warning. The failure is the signal; a rollout that continues past it has converted its own safety mechanism into telemetry.
- Promotion reads the canonical, verified job records, not a status the dispatcher reported about itself. A record that does not match on tenant, job, and device is a hard failure, not a row to skip.
- Dispatch is idempotent per device, so a retried rollout does not double-issue against a machine that already has a job.
The "verified records only" rule is the least obvious and the one worth stealing. If a ring can advance on the deployment system's own optimistic bookkeeping, then a bug in that bookkeeping promotes a broken patch to your whole fleet — and it does it quickly, because optimistic bookkeeping is fast.
The halt condition is the feature
Rings are usually sold on their first property: fewer machines exposed to a bad update. That is real, but it is the smaller half. The larger half is that a staged rollout gives you a defined moment to stop, with a defined trigger, before the blast radius grows.
A rollout with rings and no halt condition just distributes the same failure over a longer afternoon, and gives everyone involved the impression that something careful happened.
In Nexus, deployment rings are persisted rather than reassembled per rollout, the target snapshot is frozen at creation, and promotion enforces the conditions above at the server rather than in the console. Patch deployment is live end-to-end on real hardware today; the ring mechanics described here are the same ones any capability rollout uses, not a patch-only special case.