Build, Buy, or Use What the Vendor Gives You
Three routes to monitoring coverage, and why the third is where most providers actually sit without having decided to.
Basics · Analysis
Few service providers build monitoring. Most buy a platform and then use whatever it ships with, which is a decision made by omission.
The recommendations in “Build, Buy, or Use What the Vendor Gives You” become sustainable only when the recurring work has visible owners and enough capacity. A team assessing this useful page can use time and project records to see where operational effort accumulates, without treating activity data as a substitute for technical evidence or direct discussion with technicians.
For an independent operational benchmark, compare the local practice with CISA supply-chain guidance; the important test is whether the control remains proportionate, documented and recoverable when the usual technician is unavailable.
Using the defaults
Every platform ships monitoring templates: disk space, services, event log patterns, hardware health.
They are built to demonstrate value in a sales process, which means broad coverage and low thresholds.
Deploying them unchanged is the single most common cause of the alert volume this collection is about, and it is the path of least resistance.
Buying and configuring
The realistic position: a platform, with templates rewritten for how you actually work.
This is a week of work at the start and a recurring hour a month.
The providers who have done it have quiet consoles. The ones who have not have thousands of weekly alerts and a technician who has stopped looking.
Building
Rare, and appropriate for specific gaps: a client application nobody's platform understands, a bespoke check for a recurring failure.
Scripted checks running inside the RMM rather than a parallel system.
Building a monitoring platform from scratch is not a service provider's business, and the ones that try end up maintaining software instead of serving clients.
The vendor template trap
Templates are updated by the vendor, which can reintroduce alerts you tuned away.
Which means your tuning should live in your own templates rather than as edits to theirs.
Check after any platform update whether your alert volume moved, because it is a silent change with a visible consequence.
What to keep from the defaults
Hardware health and disk failure prediction, which are hard to improve on.
Backup failure.
Security service state.
These are low-volume, high-signal and worth having exactly as shipped.
What to rewrite immediately
Disk space thresholds, which default to percentages that make no sense on a large volume.
CPU and memory alerts, which fire on normal operation.
Event log patterns, which are the largest single source of noise.
Service state for services that legitimately stop.
Those four account for most of the volume in a default deployment.
The honest sequence
Buy, deploy to one client, watch the volume for a fortnight, tune, then roll out.
The alternative — deploy everywhere, then tune under pressure — is how a provider ends up with alerts nobody reads within a month.
What to check
Are you running vendor templates unchanged?
Does your tuning live in your templates or as edits to theirs?
Did your alert volume change after the last platform update?
And was the first client deployment used to tune before the rest?
The point
Vendor templates are built to demonstrate value during a sale, which means broad coverage and low thresholds.
Deploying them unchanged is the main cause of the volume.
Underlying all of this
Everything in this collection reduces to four habits: tune until every alert is read, verify rather than assume at every stage from ring one to script execution, treat the console as the privileged system it is, and know what each client costs you. None needs a better platform, and a provider doing all four runs a quieter service than one twice its size.
The recurring pattern
The recurring pattern across every section here is the same: the appearance of control substituting for control. An unread alert queue looks like monitoring. A compliance percentage that excludes pending reboots looks like protection. A script that reports success looks like automation. In each case the provider believes a risk is handled and it is not, which is worse than knowing it is open.