Skip to content
Sections
All notes

All notes · Reference

Common Failures, Listed

Twelve ways RMM deployments go wrong, each with the signal and the fix. Most providers have four or five.

Reference · Reference

In the configuration

Vendor templates unchanged. Signal: thousands of alerts a week. Fix: tune on one client, then roll out.

The recommendations in “Common Failures, Listed” become sustainable only when the recurring work has visible owners and enough capacity. A team assessing time management software can use time and project records to see where operational effort accumulates, without treating activity data as a substitute for technical evidence or direct discussion with technicians.

For an independent operational benchmark, compare the local practice with NIST Cybersecurity Framework; the important test is whether the control remains proportionate, documented and recoverable when the usual technician is unavailable.

Alerting on things nobody can act on. Signal: checks with no action attached. Fix: the three-in-the-morning test.

No dependency configuration. Signal: one server down raises ten alerts. Fix: an afternoon in the platform settings.

In the coverage

Machines without agents. Signal: nobody has compared agent count against the directory. Fix: monthly reconciliation.

Agents that stopped reporting. Signal: no alert for prolonged silence. Fix: one rule.

Unsupported third-party software. Signal: compliance reports cover the operating system only. Fix: compare installed software against your patching catalogue.

In the operation

Pending reboots counted as patched. Signal: compliance above ninety per cent with uptime in months. Fix: report both.

Scripts that report success and do nothing. Signal: nobody checks what automation actually did. Fix: every script reports its effect.

Self-healing masking a fault. Signal: a remediation firing daily on the same machine. Fix: escalate on repetition.

In the platform security

Everybody can run scripts everywhere. Signal: one role for all technicians. Fix: scope by client; a named few for multi-client.

No multi-factor, or partial. Signal: an exception for convenience. Fix: no exceptions.

No plan for compromise. Signal: nobody could suspend execution in five minutes. Fix: write it, rehearse once.

The pattern

Seven of the twelve are configuration and fixed in under a week.

Three are habits: reconciliation, verification, review.

And two are structural — the permission model and the compromise plan — which need a decision rather than an afternoon.

The order to fix them

If you do one thing: count your alerts and tune the top three sources.

If two: reconcile agents against the directory.

If three: scope script execution by client.

Those three cost about a week and address most of the twelve.

What to check

Which of the twelve do you have?

Which is cheapest to fix?

Which would be worst if it mattered tomorrow?

And has anything from your last review actually changed?

The point

Seven of the twelve failures are configuration and fixed in a week.

Two are structural and need a decision rather than an afternoon.

Underlying all of this

Everything in this collection reduces to four habits: tune until every alert is read, verify rather than assume at every stage from ring one to script execution, treat the console as the privileged system it is, and know what each client costs you. None needs a better platform, and a provider doing all four runs a quieter service than one twice its size.

The recurring pattern

The recurring pattern across every section here is the same: the appearance of control substituting for control. An unread alert queue looks like monitoring. A compliance percentage that excludes pending reboots looks like protection. A script that reports success looks like automation. In each case the provider believes a risk is handled and it is not, which is worse than knowing it is open.