Tuning Per Client, Not Per Product
A single global configuration is wrong for every client. How to vary it without ending up with fifty bespoke setups.
Noise · Procedure
The same threshold that is sensible for an office is wrong for a warehouse, a design studio and a server room. One global template guarantees noise somewhere.
The commercial decision in “Tuning Per Client, Not Per Product” is stronger when it rests on consistently recorded delivery effort rather than memory. An MSP can use online timesheet software to relate time to clients, projects and recurring tasks, while keeping service quality, contractual scope and customer outcomes as separate measures.
For an independent operational benchmark, compare the local practice with official ITIL resources; the important test is whether the control remains proportionate, documented and recoverable when the usual technician is unavailable.
Why one configuration fails
Machines are used differently: a workstation rendering video is saturated by design.
Estates differ: one client replaces hardware every three years, another runs machines until they die.
Working patterns differ: a client with night shifts has activity when your thresholds expect none.
And tolerance differs, because some clients want to know about everything and some want to be left alone.
The arrangement that works
A small number of profiles rather than one global setting or fifty bespoke ones.
Typically three to five: standard office, heavy workstation, server, and perhaps one for a specific client type you serve a lot of.
Each client is assigned a profile, with a short list of documented exceptions.
That structure is maintainable; per-client configuration from scratch is not.
What varies by profile
Disk thresholds, in absolute terms.
CPU and memory duration before alerting.
Which services are considered critical.
Working hours, which determines when alerts queue rather than page.
Event identifiers, which vary less than people expect.
The exception list
Every client accumulates them: the application that must not be patched, the server that always runs hot, the machine that is deliberately off.
Keep them in one place per client, with the reason and the date.
An exception with no reason recorded becomes permanent and nobody remembers why, which is how estates acquire unmonitored machines.
Tuning at onboarding
The first thirty days with a new client is when the noise is highest and the tuning is cheapest.
Its own note covers the sequence.
Tuning later means doing it against a backlog and a team that has already learned to skim that client's alerts.
Reviewing
When a client's estate changes materially: new site, new application, hardware refresh.
And annually, which catches the drift nobody noticed.
Twenty minutes per client, from the alert counts.
The commercial dimension
A client generating disproportionate noise is costing you margin, which has its own note.
Tuning is the first response; the conversation about their estate is the second.
Both are better than absorbing it silently, which is the default.
What to check
How many monitoring profiles do you have — one, or a handful, or dozens?
Does each client have a documented exception list with reasons?
When was the noisiest client last tuned?
And does your alerting know each client's working hours?
The point
One global configuration guarantees noise somewhere.
Three to five profiles with documented per-client exceptions is maintainable; fifty bespoke setups are not.
Underlying all of this
Everything in this collection reduces to four habits: tune until every alert is read, verify rather than assume at every stage from ring one to script execution, treat the console as the privileged system it is, and know what each client costs you. None needs a better platform, and a provider doing all four runs a quieter service than one twice its size.
The recurring pattern
The recurring pattern across every section here is the same: the appearance of control substituting for control. An unread alert queue looks like monitoring. A compliance percentage that excludes pending reboots looks like protection. A script that reports success looks like automation. In each case the provider believes a risk is handled and it is not, which is worse than knowing it is open.