Skip to content
Sections
All notes

All notes · Reference

What a Working Deployment Looks Like

The end state, assembled from everything here, as a description to measure yours against.

Reference · Reference

Not a maturity model. A description of a provider serving a few dozen clients, two years in, whose console is quiet and whose clients stay.

The recommendations in “What a Working Deployment Looks Like” become sustainable only when the recurring work has visible owners and enough capacity. A team assessing the product overview can use time and project records to see where operational effort accumulates, without treating activity data as a substitute for technical evidence or direct discussion with technicians.

For an independent operational benchmark, compare the local practice with NIST Cybersecurity Framework; the important test is whether the control remains proportionate, documented and recoverable when the usual technician is unavailable.

The console

A handful of actionable alerts a day across all clients, every one read.

Three monitoring profiles, with documented exceptions per client.

Dependencies configured, so one server outage raises one alert.

And a target volume somebody is measured against, reviewed quarterly.

What gets monitored

Hardware health and predictive failure, as shipped.

Backup success and failure.

Services that should never stop, alerting on repeated restarts rather than each one.

Disk space in absolute terms, with rate of change.

Event logs by specific identifier, never by severity.

And nothing the provider is not paid to fix.

Patching

Four rings, with ring zero being the provider's own machines.

Verification after ring one, from a per-client checklist.

Exclusions with reasons and review dates, examined quarterly.

Reporting that shows patched, pending reboot, excluded, and unreachable separately.

Automation

A library of thirty scripts everybody trusts, owned by one person, reviewed annually.

Every script reports what it did, not that it ran.

Dry-run mode on anything that changes something.

Remediations counted, with repeated firing raising a ticket.

And multi-client execution requiring a second pair of eyes.

Security of the platform

Multi-factor on every console account, no exceptions.

Technicians scoped to their own clients; two or three full administrators.

Logs exported outside the platform, with alerts on new administrator accounts and wide-scope scripts.

Credentials in a secrets system, never in scripts, scoped per technician.

And a written plan for compromise that somebody has read.

Per client

The six answers: scope, approvals, hours, billing, out-of-scope, exit.

A dated baseline of what was inherited.

Documentation a new technician could work from.

A primary and a secondary technician.

And a monthly report ending with decisions for them.

Commercially

Cost per client known.

Out-of-scope time measured.

Written criteria for declining, and something declined in the last year.

And an annual comparison of what the agreement covers against what is actually done.

The test

How many alerts yesterday, and were they all read?

Could a new technician work your largest client from documentation?

When did a client last tell you about an outage first?

A provider who answers those three has done the work.

The point

Three tests: how many alerts yesterday and were they all read, could a new technician work your largest client from documentation, and when did a client last report an outage first?.

Underlying all of this

Everything in this collection reduces to four habits: tune until every alert is read, verify rather than assume at every stage from ring one to script execution, treat the console as the privileged system it is, and know what each client costs you. None needs a better platform, and a provider doing all four runs a quieter service than one twice its size.

The recurring pattern

The recurring pattern across every section here is the same: the appearance of control substituting for control. An unread alert queue looks like monitoring. A compliance percentage that excludes pending reboots looks like protection. A script that reports success looks like automation. In each case the provider believes a risk is handled and it is not, which is worse than knowing it is open.