Skip to content
Sections
All notes

All notes · Noise

Thresholds That Mean Something

Most alert conditions are set to a number somebody guessed. What to set them to instead, and the question that settles each.

Noise · Procedure

A threshold should mark the point where action is needed and possible. Most are set where the condition becomes visible, which is much earlier.

Tuning the process in “Thresholds That Mean Something” is easier to defend when the team can compare alert volume with the human effort required to investigate it. Teams evaluating the complete guide can record time against monitoring and remediation work, while the RMM remains the authoritative source for device state, thresholds and event history.

For an independent operational benchmark, compare the local practice with CIS Critical Security Controls; the important test is whether the control remains proportionate, documented and recoverable when the usual technician is unavailable.

The question to ask of each

If this fires, what will the technician do?

If the answer is "look at it and close it", the threshold is wrong.

If the answer is "nothing can be done yet", the threshold is early.

Every check should survive that question, and most do not.

Disk space

Percentages are the wrong unit on modern volumes.

Use absolute free space, set to how much you need to act: enough for an update, a log rotation, a day of growth.

And consider rate of change rather than level — a volume losing a gigabyte an hour matters more than one sitting steadily at ninety per cent for two years.

CPU and memory

Instantaneous spikes are normal and should not alert at all.

Sustained saturation over a period is a finding.

Set the duration first, then the level, which removes most of this category immediately.

Services

Alert on services that should never stop, not on everything.

And check whether the service restarts itself, because an alert for something that recovered in ten seconds is noise by construction.

Where restart is automatic, alert on repeated restarts rather than on each one.

Event logs

The largest noise source and the one needing the most restraint.

Alert on specific event identifiers you have decided matter, never on severity levels.

"All errors" is not a monitoring strategy, and it is the default in many templates.

Hardware

Predictive disk failure, temperature, fan and power supply faults.

Low volume, high signal, leave as shipped.

This is the category where the platform earns its keep.

Per-client variation

The same threshold is wrong for a design studio and a warehouse.

Which is why the next note argues tuning belongs at client level rather than at product level.

Writing them down

Each check: the condition, the threshold, why that number, and what the technician does.

The last column is the one that keeps the set honest, because a check with an empty action column should not exist.

What to check

Pick three checks: can you say what a technician does when each fires?

Are your disk alerts in percentages or absolute space?

Do you alert on event log severity rather than specific identifiers?

And does any check have no action attached?

The point

Every check should answer one question: if this fires, what will the technician do? A check with no action attached should not exist..

Underlying all of this

Everything in this collection reduces to four habits: tune until every alert is read, verify rather than assume at every stage from ring one to script execution, treat the console as the privileged system it is, and know what each client costs you. None needs a better platform, and a provider doing all four runs a quieter service than one twice its size.

The recurring pattern

The recurring pattern across every section here is the same: the appearance of control substituting for control. An unread alert queue looks like monitoring. A compliance percentage that excludes pending reboots looks like protection. A script that reports success looks like automation. In each case the provider believes a risk is handled and it is not, which is worse than knowing it is open.