Skip to content
Sections
All notes

Remote monitoring and management

An unread alert queue is worse than none

Fifty notes on running an RMM: why the shipped configuration is unusable, how to tune it per client, patching without causing the outage, and treating a tool that reaches every client at once as the privileged system it is.

Practical guidance, vendor comparisons and no substitute for contractual or security advice.

Alerts per week, after tuning
Illustrative chart of weekly alert volume falling after tuning
Actionable 6 a day
50notes
4patch rings
12common failures listed
0vendors named

The queue nobody reads

A platform deployed with its shipped templates to a hundred machines will produce alerts in the thousands per week. This is not a fault. The defaults are designed to demonstrate capability during an evaluation, and a platform that showed nothing in the first week would look inert.

Tuning the process in “An unread alert queue is worse than none” is easier to defend when the team can compare alert volume with the human effort required to investigate it. Teams evaluating automatic time tracking can record time against monitoring and remediation work, while the RMM remains the authoritative source for device state, thresholds and event history.

For an independent operational benchmark, compare the local practice with CIS Critical Security Controls; the important test is whether the control remains proportionate, documented and recoverable when the usual technician is unavailable.

What happens next is consistent. The queue is read for a week, skimmed for a month, and ignored thereafter. At that point monitoring has stopped while the dashboard continues to show that it is working.

Then a genuine disk failure arrives in a queue with four hundred other alerts from that night, nobody sees it, and the client discovers the outage before you do. That specific sequence has happened to most providers in this industry, and it is the argument for everything that follows.

The objection is that tuning might mean missing something. You are already missing things, inside a queue nobody reads. Tuning does not reduce what you detect; it reduces what you are told about, and the distinction is the whole discipline.

The estate is not yours

Internal IT monitors machines its own organisation owns, configured to its own standard, used by colleagues subject to its own policies. A service provider has none of that.

You do not control what is installed, who has local administrator rights, when machines are on, how old the hardware is, or whether a user will accept an interruption. You are accountable for uptime, patch compliance and response times anyway.

That mismatch between control and accountability is the structural problem of this business, and it belongs in the contract rather than in the tooling. A service agreement promising patch compliance without a reboot policy is promising something the client can prevent.

Four things settle most of the friction, agreed at onboarding and written down: who can approve a reboot and in what window, which machines are out of scope, what happens to a machine that is never on during the patch window, and who to contact when a user refuses an action.

Thresholds that mean something

Every check should answer one question: if this fires, what will the technician do? If the answer is "look at it and close it", the threshold is wrong. A check with no action attached should not exist.

Disk space in percentages is the wrong unit — ninety per cent of a two-terabyte volume is two hundred gigabytes free, which is not an emergency. Use absolute space, and consider rate of change: a volume losing a gigabyte an hour matters more than one sitting steadily at ninety per cent for two years.

CPU and memory spikes are normal; set a duration before a level. Alert on services that should never stop, and on repeated restarts rather than each one. And for event logs, alert on specific identifiers you have decided matter — "all errors" is not a monitoring strategy, and it is the default in many templates.

Patching, and the gap nobody reports

Not patching lets the consequence arrive from outside. Patching badly makes it arrive from you, and providers are judged more harshly for the second — which is why many quietly do less of the first than their reports suggest.

Rings help: your own machines first, then a small representative test group at each client, then the bulk, then servers. But a ring only works if somebody verifies afterwards. A five-minute checklist — does the main application open, does printing work, does the line-of-business system connect — is what turns a delay mechanism into a control.

And then the gap that most reports conceal. A patch installed and pending reboot is not protecting anything. Machines run for weeks between restarts, especially laptops that are only ever suspended. A compliance figure counting installed patches overstates protection by whatever that gap is, and most providers do not measure it.

An honest report is six lines rather than one percentage: patched and rebooted, patched and pending, excluded with reasons, unreachable, machines without agents, and third-party coverage separately.

Automation, and the fault it conceals

Automation is where an RMM pays for itself. It is also where a working mechanism can leave an organisation less informed than before.

The easiest win is the one nobody reaches for: automatic diagnostic collection. When an alert fires, run a script that gathers the obvious context — logs, versions, disk state, recent changes. The technician opens a ticket that already contains what they would have spent ten minutes collecting. No risk, immediate return, and it works even where you would not trust automation to act.

Beyond that, choose from ticket data rather than from ideas, and do the thing manually ten times before scripting it. Automating something you have not done by hand produces a script that works on the happy path and fails in ways nobody anticipated.

Then the concealment. A service restarting itself five times a day is a resolved alert and an unresolved fault. The ticket count falls, the cause persists, and because the automation works nobody investigates. A disk cleanup running daily on a machine filling up from a runaway log is the same pattern.

The remedy is counting. Every automated fix should be recorded even though no ticket is raised, and any remediation firing repeatedly on the same machine should raise one regardless of whether it succeeded. That single rule converts the whole category from a blind spot into an early warning.

And the worst case in any script library is the one that reports success and does nothing — a cleanup whose target path changed, a check whose condition can no longer be true. It appears in the schedule, runs daily, and the thing it was protecting against goes unmanaged. Every script should report what it did, not that it ran.

The commercial side nobody measures

In most providers, two or three clients account for a large share of support effort while paying the same as everybody else. A client paying the average and consuming three times the average is subsidised by the others, out of your margin and out of the service they receive.

Finding them needs time tracking that actually happens, or tickets per user per month compared across clients. Before the commercial conversation, confirm the noise is theirs rather than yours — a client generating thousands of alerts from an untuned profile is your problem, and starting that discussion without checking is an embarrassing position.

Scope creep arrives the same way: one small favour at a time, each reasonable, nobody counting. The response that is not a refusal is to do the work, record it as out of scope, and show it in the monthly summary. The accumulation becomes visible to the client rather than invisible to you, and the annual conversation happens with evidence instead of frustration.

And the service agreement needs the clause most of them lack: a client dependency provision, suspending obligations where the client has not approved a reboot, replaced unsupported hardware, or funded a fix you recommended. Without it a client can block the work and then hold you to the target.

The tool that fixes can also break

Described neutrally, an RMM is execution of arbitrary code at the highest available privilege on every managed machine across every client, from one console, usually without the user noticing. No other tool a provider operates concentrates that much reach.

Which makes the asymmetry worth stating plainly: an internal IT compromise affects one organisation. A provider compromise affects everybody they serve, and those clients did not choose your security posture.

The controls are not exotic. Multi-factor on every console account without exceptions. Technicians scoped to the clients they actually serve, with a named few holding multi-client execution. Credentials in a secrets system rather than embedded in scripts — which is the recurring finding in reviews of this industry, and they persist in version history long after the script is corrected. Logs exported where the console cannot alter them. And alerts on the platform itself: new administrator accounts, wide-scope scripts, unusual logins.

What is deliberately absent here

No vendors or products are named. The category consolidates, the names date, and the practices do not.

No breach statistics. The available figures about how many providers have been compromised come from surveys commissioned by companies selling controls.

And no claim that automation solves this. The automation section ends with the case where it hides the problem it keeps fixing, which is the part the marketing omits.

What this covers

From the queue nobody reads to the console that reaches every client

Eight sections, in the order the work actually runs: what the platform does, getting the noise down, onboarding an estate, patching, automation, securing the tool itself, and the commercial side that decides whether any of it is sustainable.

What it does

The machines belong to somebody else, and the gap between what you control and what you are accountable for shapes everything.

6 notes →

Getting the noise down

A platform at default settings produces more alerts than anybody will read, and an unread queue looks like monitoring.

7 notes →

Onboarding an estate

The machine count a client gives you is wrong. Finding the real one before the contract is better than after.

7 notes →

Patching

A patch installed and pending reboot protects nothing. That gap is where compliance quietly fails.

6 notes →

Automation

A script runs against every client at once. There is no partial failure and no time to notice.

6 notes →

Securing the tool

An internal compromise affects one organisation. A provider compromise affects everybody they serve.

7 notes →

The commercial side

Two or three clients usually consume most of the capacity, and the pricing model decides whether that is visible.

7 notes →

Reference

The end state as a description, the twelve failures, and the order to do things in.

4 notes →

If the queue is unreadable

Three things that cost a week and change the service

Count, then tune the top three

Alerts raised, opened, actioned, and arriving out of hours. A small number of check types produce most of the volume, and fixing those usually removes most of it.

Reconcile agents against the directory

Nobody looks for absence. A machine without an agent is invisible from the console and looks identical to a healthy one, which is how estates acquire unmonitored machines.

Scope script execution by client

Running a script across estates is code execution with high privilege on machines you do not own. Most incidents in this category are a correct script with a wrong scope.

All 50 notes

All fifty notes, by subject

The short version

Measure what the client notices

Not alert volume, but how often a client tells you about an outage before your monitoring does. Tune until every alert is read, verify rather than assume, and run the console as the privileged system it is.

Managed IT tool guides

Detailed comparisons for RMM, service-desk and endpoint operations.