Skip to content
Sections
All notes

All notes · Noise

Counting Your Alerts Honestly

The measurement that precedes any tuning, and the four numbers that tell you where the noise is.

Noise · Procedure

You cannot reduce what you have not counted. The exercise takes an hour and most providers have never done it.

Tuning the process in “Counting Your Alerts Honestly” is easier to defend when the team can compare alert volume with the human effort required to investigate it. Teams evaluating the official page can record time against monitoring and remediation work, while the RMM remains the authoritative source for device state, thresholds and event history.

For an independent operational benchmark, compare the local practice with CIS Critical Security Controls; the important test is whether the control remains proportionate, documented and recoverable when the usual technician is unavailable.

The four numbers

Total alerts raised, per week.

How many were opened by a human.

How many resulted in action.

And how many arrived outside the hours anybody was working.

The gap between the first and the second is your real problem.

Where to get them

Most platforms report alert counts; if yours does not, export and count.

Opened and actioned usually come from the PSA, if the integration exists.

And the out-of-hours figure is a timestamp filter.

None of this needs new tooling.

Breaking it down

By check type. Almost always a small number of checks produce most of the volume.

By client. Almost always one or two clients produce a disproportionate share, which has a commercial dimension covered in its own note.

By machine. A single sick machine can generate hundreds of alerts and it is usually known to somebody.

Those three cuts locate the noise within an hour.

The ratio that matters

Actioned divided by raised.

Below about one in twenty, the queue is being skimmed rather than read, whatever anybody says.

Below one in a hundred, monitoring is decorative.

Publish this internally, because it is the number that justifies spending time on tuning.

The duplicate question

One failure raising five alerts counts as five.

Check whether your top sources are genuinely distinct conditions or one condition reported five ways.

Deduplication and correlation settings exist in most platforms and are frequently off.

Setting a target

Pick a number a technician can read in a day alongside other work.

Measure against it monthly.

A published target changes behaviour in a way a general intention to reduce noise does not.

Repeating it

Monthly for the first six months, then quarterly.

And after every platform update, which can reintroduce volume silently.

Keep the series, because the trend is what shows whether tuning is holding.

What to check

Do you know your weekly alert count?

What proportion is opened, and what proportion actioned?

Which three check types produce most of the volume?

And is there a target anybody is measured against?

The point

Count four numbers: alerts raised, opened, actioned, and arriving out of hours.

The gap between the first and second is the real problem.

Underlying all of this

Everything in this collection reduces to four habits: tune until every alert is read, verify rather than assume at every stage from ring one to script execution, treat the console as the privileged system it is, and know what each client costs you. None needs a better platform, and a provider doing all four runs a quieter service than one twice its size.

The recurring pattern

The recurring pattern across every section here is the same: the appearance of control substituting for control. An unread alert queue looks like monitoring. A compliance percentage that excludes pending reboots looks like protection. A script that reports success looks like automation. In each case the provider believes a risk is handled and it is not, which is worse than knowing it is open.