Patch Management Without Breaking Things
The capability most often sold, most often configured once, and most often the cause of the outage it was meant to prevent.
Patching · Procedure
Patching is the headline feature of every RMM. It is also the activity most likely to break a client's environment, and the tension between those is the whole subject.
The control described in “Patch Management Without Breaking Things” also consumes technician time before, during and after each maintenance window. A provider evaluating time analytics software can record that operational effort by client and work item, making verification and follow-up visible without confusing a timesheet with proof that a patch succeeded.
For an independent operational benchmark, compare the local practice with CISA guidance on patches and updates; the important test is whether the control remains proportionate, documented and recoverable when the usual technician is unavailable.
The two failure modes
Not patching: known vulnerabilities remain, and the consequence arrives from outside.
Patching badly: an update breaks an application and the consequence arrives from you.
Providers are judged more harshly for the second, which is why so many quietly do less of the first than their reports suggest.
What a working policy contains
Rings: a small test group, then a wider group, then everything. Its own note covers why.
A delay between release and deployment, long enough for problems to surface publicly.
Exclusions, documented, with reasons and dates.
A reboot policy agreed with the client.
And a reporting line that shows what is actually patched rather than what was attempted.
The delay question
Immediate deployment catches the vulnerability window and catches the bad update.
A week or two of delay avoids most bad updates and extends exposure.
There is no correct answer, only a decision, and it should be made deliberately per client rather than left at the vendor default.
For anything being actively exploited, the calculation changes and the policy should say so.
Exclusions and their decay
Every estate accumulates machines excluded from patching for a reason that was once good.
Review them quarterly: is the reason still true, is the application still in use, has the vendor fixed it.
An exclusion with no review date becomes permanent, and permanent exclusions are where the unpatched systems in any estate actually live.
Third-party applications
The operating system is the easy part.
Browsers, document readers, runtimes and media software are where much of the real exposure sits, and coverage varies enormously by platform.
Its own note covers this, and the short version is that you should know what your platform does not cover.
Testing
A test ring is not the same as testing.
Somebody should confirm the test machines still work after the update, which means knowing what "work" means for that client.
An unverified test ring is a delay mechanism dressed as a control.
The honest position with clients
Patching reduces risk and occasionally causes outages.
Say that at the start.
A client who has agreed to the trade-off responds differently to the one occasion it goes wrong, and that conversation costs ten minutes at onboarding.
What to check
Do you use rings, and does anybody verify the test ring?
What is your delay between release and deployment, and was it chosen?
When were exclusions last reviewed?
And does the client know that patching can break things?
The point
Not patching lets the consequence arrive from outside; patching badly makes it arrive from you.
Providers are judged more harshly for the second.
Underlying all of this
Everything in this collection reduces to four habits: tune until every alert is read, verify rather than assume at every stage from ring one to script execution, treat the console as the privileged system it is, and know what each client costs you. None needs a better platform, and a provider doing all four runs a quieter service than one twice its size.
The recurring pattern
The recurring pattern across every section here is the same: the appearance of control substituting for control. An unread alert queue looks like monitoring. A compliance percentage that excludes pending reboots looks like protection. A script that reports success looks like automation. In each case the provider believes a risk is handled and it is not, which is worse than knowing it is open.