Skip to content
Sections
All notes

All notes · Business

What the Service Level Agreement Should Say

Most agreements promise response times and nothing else. What else belongs in them, and the clause that protects both sides.

Business · Analysis

General orientation, not legal advice.

The recommendations in “What the Service Level Agreement Should Say” become sustainable only when the recurring work has visible owners and enough capacity. A team assessing work time tracking tools can use time and project records to see where operational effort accumulates, without treating activity data as a substitute for technical evidence or direct discussion with technicians.

For an independent operational benchmark, compare the local practice with official ITIL resources; the important test is whether the control remains proportionate, documented and recoverable when the usual technician is unavailable.

A service level agreement is where the mismatch between what you control and what you are accountable for gets resolved. Most resolve it badly by not addressing it.

What they usually say

Response times by severity.

Hours of cover.

Perhaps an uptime figure.

All necessary and none of it addresses the structural problem: you do not control the estate.

Response versus resolution

Response time is yours to control and is what should be promised.

Resolution time depends on the fault, the vendor, the hardware availability and the client's own decisions.

Promising resolution times is promising something you cannot deliver, and it is common in agreements written to win a tender.

The client dependency clause

The most important and the most often missing.

It states that certain obligations are suspended where the client has not done something: approved a reboot, replaced unsupported hardware, provided access, funded a fix you recommended.

Without it, a client can block the work and then hold you to the target.

With it, the conversation is about their decision rather than your performance.

What else belongs

Scope: what is covered, and explicitly what is not.

Exclusions: unsupported systems, equipment past end of life, anything out of scope.

Change windows and who approves them.

What happens outside hours.

Reporting: what, how often.

And termination: notice, handover, agent removal, data.

Severity definitions

Everybody has severity one through four and almost nobody defines them in client language.

Define by impact: how many people cannot work, is there a workaround, is revenue affected.

Not by technical category, because a down print server is a severity one for the client who cannot invoice.

And agree who assigns severity, because that is where most disputes actually sit.

Measuring against it

If you promise response times, measure them and report them.

A provider who publishes their own performance against the agreement, including the misses, is in a far stronger position than one who waits to be challenged.

The honest conversation at sale

Say what you cannot promise.

A prospect comparing agreements will see weaker promises and ask why, which is an opportunity rather than a problem — the answer is that the stronger promise is not deliverable and you would rather say so now.

What to check

Does your agreement promise resolution times?

Is there a client dependency clause?

Are severities defined by impact, in client language?

And do you report your own performance against the targets?

The point

Response time is yours to control; resolution time is not.

Promising the second is promising something you cannot deliver.

Underlying all of this

Everything in this collection reduces to four habits: tune until every alert is read, verify rather than assume at every stage from ring one to script execution, treat the console as the privileged system it is, and know what each client costs you. None needs a better platform, and a provider doing all four runs a quieter service than one twice its size.

The recurring pattern

The recurring pattern across every section here is the same: the appearance of control substituting for control. An unread alert queue looks like monitoring. A compliance percentage that excludes pending reboots looks like protection. A script that reports success looks like automation. In each case the provider believes a risk is handled and it is not, which is worse than knowing it is open.