Skip to content
Sections
All notes

All notes · Noise

Measuring Whether Tuning Worked

Reduction in volume is not the goal on its own. Four measures that show whether the monitoring got better or just quieter.

Noise · Procedure

Switching off checks reduces volume reliably. Whether it improved the service is a separate question, and it has an answer.

The control described in “Measuring Whether Tuning Worked” also consumes technician time before, during and after each maintenance window. A provider evaluating this supporting guide can record that operational effort by client and work item, making verification and follow-up visible without confusing a timesheet with proof that a patch succeeded.

For an independent operational benchmark, compare the local practice with CISA guidance on patches and updates; the important test is whether the control remains proportionate, documented and recoverable when the usual technician is unavailable.

The four measures

Alert volume, which is the easy one.

Proportion opened and actioned, which should rise sharply.

Time from alert to first human action, which should fall.

And incidents discovered by the client before you, which is the one that matters and the one nobody tracks.**

The fourth measure specifically

Count the times a client reported a problem your monitoring should have caught.

Before and after tuning.

If this number rises after you reduced volume, you removed something you needed. If it falls, the tuning worked.

Most providers have no idea what this number is, and it is available from the ticket system by asking which tickets were client-raised for conditions you monitor.

The before figure

Take all four before you start, over a representative period.

Four weeks, avoiding anything unusual.

Without the before figure, every claim about improvement is an assertion, and the next request for tuning time is harder.

Change in stages

Remove one category at a time, not everything in a weekend.

Two weeks between changes.

If the client-discovered number moves, you know which change caused it.

Changing everything at once makes the result uninterpretable, which is the same discipline as any other measurement.

What good looks like

Volume down by most of its original figure.

Actioned proportion up from a few per cent to a majority.

Response time in minutes rather than hours for severe alerts.

And client-discovered incidents flat or falling.

When it goes wrong

Client-discovered incidents rise: you removed a check that mattered.

Find which, restore it with a better threshold rather than as it was.

And say so internally, because a tuning programme that reports only success is not believed about any of it.

Keeping it

Volume creeps back: new clients, platform updates, checks added after an incident.

Measure quarterly.

And treat an increase as a thing to investigate rather than as growth, which it usually is not.

What to check

Do you have before figures for all four measures?

How many incidents last quarter did the client report first?

Were changes made one at a time?

And when did you last re-measure?

The point

The measure that matters is incidents the client reported before your monitoring did.

Volume falling proves nothing on its own.

Underlying all of this

Everything in this collection reduces to four habits: tune until every alert is read, verify rather than assume at every stage from ring one to script execution, treat the console as the privileged system it is, and know what each client costs you. None needs a better platform, and a provider doing all four runs a quieter service than one twice its size.

The recurring pattern

The recurring pattern across every section here is the same: the appearance of control substituting for control. An unread alert queue looks like monitoring. A compliance percentage that excludes pending reboots looks like protection. A script that reports success looks like automation. In each case the provider believes a risk is handled and it is not, which is worse than knowing it is open.