When a Patch Causes the Outage
It will happen. The first hour, the communication, and the part that determines whether the relationship survives it.
Patching · Procedure
An update you deployed broke something a client depends on. This is a normal event in this business and the handling is what distinguishes providers.
The control described in “When a Patch Causes the Outage” also consumes technician time before, during and after each maintenance window. A provider evaluating the detailed reference can record that operational effort by client and work item, making verification and follow-up visible without confusing a timesheet with proof that a patch succeeded.
For an independent operational benchmark, compare the local practice with CISA guidance on patches and updates; the important test is whether the control remains proportionate, documented and recoverable when the usual technician is unavailable.
The first hour
Stop the rollout. If it is in ring one, it goes no further; if it is wider, halt immediately.
Establish what broke, on what configuration, for how many.
Decide between rollback and workaround — rollback is usually faster and sometimes not possible.
And tell the client before they tell you, which is the single decision that matters most here.
The communication
Early, plain and without hedging: an update we deployed has caused this, here is what we are doing, here is when we will next update you.
Not: "we are investigating an issue".
Clients forgive the fault and remember the handling, and attempting to obscure the cause is what turns an incident into a procurement review.
Rollback
Know whether it is possible before you need it, per update type.
Some operating system updates uninstall cleanly; some do not. Application updates vary more.
Test rollback occasionally, because discovering it does not work during an incident is a bad moment.
While it is broken
Workaround if rollback is slow: an alternative route, a temporary configuration, a manual process.
The client cares about working, not about correctness.
Document what you did, because a temporary workaround left in place for two years is a classic source of later confusion.
Afterwards
Add the exclusion with the date and reason.
Tell your other clients running the same software, which is a real advantage of serving many estates.
Check whether ring one should have caught it: if the configuration was not represented, fix the ring rather than blaming the update.
And write it down where the next technician will find it.
The accountability question
You deployed it, so it is yours.
"The vendor released a bad patch" is true and is not a defence the client accepts, because they engaged you to manage that risk.
Say that plainly, which is uncomfortable and preserves more credibility than the alternative.
The report to the client
Short and factual: what happened, why, what you changed so it does not recur.
Within a week.
A provider who sends that unprompted is in a different category from one who waits to be asked.
What to check
Do you know whether rollback works for your common update types?
Has a rollout ever been halted mid-deployment?
When this last happened, who told whom first?
And did the exclusion get recorded with a reason?
The point
The vendor released a bad patch is true and is not a defence the client accepts, because they engaged you to manage that risk..
Underlying all of this
Everything in this collection reduces to four habits: tune until every alert is read, verify rather than assume at every stage from ring one to script execution, treat the console as the privileged system it is, and know what each client costs you. None needs a better platform, and a provider doing all four runs a quieter service than one twice its size.
The recurring pattern
The recurring pattern across every section here is the same: the appearance of control substituting for control. An unread alert queue looks like monitoring. A compliance percentage that excludes pending reboots looks like protection. A script that reports success looks like automation. In each case the provider believes a risk is handled and it is not, which is worse than knowing it is open.