Skip to content
Sections
All notes

All notes · Patching

Ring Deployment and Why It Is Worth It

Staged rollout costs a little setup and prevents the failure that damages a provider most. How to structure it.

Patching · Procedure

A ring is a group of machines that receives updates before the rest. The structure is simple and the discipline is what makes it work.

The control described in “Ring Deployment and Why It Is Worth It” also consumes technician time before, during and after each maintenance window. A provider evaluating the official explanation can record that operational effort by client and work item, making verification and follow-up visible without confusing a timesheet with proof that a patch succeeded.

For an independent operational benchmark, compare the local practice with CISA guidance on patches and updates; the important test is whether the control remains proportionate, documented and recoverable when the usual technician is unavailable.

The structure

Ring zero: your own machines. You take the risk first.

Ring one: a small test group per client — five to ten machines, representative, belonging to people who will tell you if something breaks.

Ring two: the bulk of workstations.

Ring three: servers, with their own windows.

Four rings, a few days apart, and most bad updates are caught before they reach anybody important.

Choosing ring one

Representative, not convenient.

It should include the applications that matter: the line-of-business system, the design software, whatever the client actually depends on.

A ring of five identical machines running nothing tells you nothing.

And the users should know they are in it, because their cooperation is the mechanism.

The gap between rings

Long enough for a problem to appear, short enough to matter for security.

Three to seven days between rings is typical.

Shorter for anything being actively exploited, which the policy should state explicitly.

Verification, not just delay

The ring only works if somebody checks.

A short list per client: does the main application open, does printing work, does the line-of-business system connect.

Five minutes, performed after ring one.

Without it the ring is a delay mechanism, which is better than nothing and much less than it appears.

Servers

Separate rings, separate windows, separate verification.

Dependencies matter: patching a database server before the application server that depends on it has an order.

Document the order per client, because it is exactly the knowledge that lives in one technician's head.

What to do when ring one breaks

Stop the rollout. This is the whole point and it requires that somebody is watching.

Record what broke, on what configuration.

Add the exclusion with a date and a reason.

And tell the other clients, if they run the same software, which is a genuine advantage of serving many estates.

The cost

Setup: a day across your client base.

Ongoing: the verification checks, a few minutes per client per cycle.

Against: one bad update reaching three hundred machines, which costs a week and some of the relationship.

What to check

Do you have rings, and is your own estate ring zero?

Is ring one representative or convenient?

Does anybody verify after ring one, with a checklist?

And has a rollout ever actually been stopped?

The point

A ring only works if somebody verifies afterwards.

Without the check it is a delay mechanism dressed as a control.

Underlying all of this

Everything in this collection reduces to four habits: tune until every alert is read, verify rather than assume at every stage from ring one to script execution, treat the console as the privileged system it is, and know what each client costs you. None needs a better platform, and a provider doing all four runs a quieter service than one twice its size.

The recurring pattern

The recurring pattern across every section here is the same: the appearance of control substituting for control. An unread alert queue looks like monitoring. A compliance percentage that excludes pending reboots looks like protection. A script that reports success looks like automation. In each case the provider believes a risk is handled and it is not, which is worse than knowing it is open.