A Checklist for the Whole Thing
Everything in this collection as a sequence, from before the contract to the annual review.
Reference · Reference
Before the contract
Discovery: directory, DHCP, network scan with written permission, their records.
The recommendations in “A Checklist for the Whole Thing” become sustainable only when the recurring work has visible owners and enough capacity. A team assessing the reporting guide can use time and project records to see where operational effort accumulates, without treating activity data as a substitute for technical evidence or direct discussion with technicians.
For an independent operational benchmark, compare the local practice with Atlassian knowledge-sharing guidance; the important test is whether the control remains proportionate, documented and recoverable when the usual technician is unavailable.
Count what is actually there, not what you were told.
Name the unsupportable: unsupported systems, equipment you will not touch.
Price the estate you found.
Agree the six: scope, approvals, hours, billing, out-of-scope, exit.
Onboarding
Agents to a representative sample first.
Watch the alert volume for a week.
Tune: assign a profile, record exceptions with reasons.
Then deploy to the rest.
Record a dated baseline: patch levels, hardware age, backup state, security software.
Documentation from day one, written when things are discovered.
Thirty-day review and a written summary to the client.
Monitoring
Alert on what somebody will act on, nothing else.
Disk in absolute terms; CPU with a duration; services that should never stop; event logs by identifier.
Dependencies configured.
Working hours per client, so out-of-hours alerts queue rather than page.
Count weekly volume, opened and actioned proportion, and incidents the client found first.
Patching
Four rings, ring zero your own.
Verification checklist after ring one.
A chosen delay between release and deployment.
Exclusions with reasons and review dates.
Report patched, pending reboot, excluded and unreachable separately.
Automation
Choose from ticket data, not from ideas.
Manual ten times before scripting.
Dry-run mode on anything that changes something.
Every script reports its effect.
Count remediations; escalate on repetition.
Second pair of eyes for multi-client scope.
Platform security
Multi-factor everywhere, no exceptions.
Technicians scoped to their clients.
Credentials in a secrets system, never in scripts.
Logs exported outside the platform.
Alerts on new administrator accounts, wide-scope scripts, unusual logins.
A written compromise plan, rehearsed once.
Monthly
Alert volume and actioned proportion.
Agent reconciliation against the directory.
Out-of-scope time.
Client report ending with decisions.
Quarterly
Exclusion review.
Template and script library review: what is unused, what duplicates.
Permission review against current roles.
Noisiest client re-tuned.
Cost per client.
Annually
Discovery repeated.
Agreement compared against what you actually do.
Vendor security settings reviewed.
Criteria for declining revisited.
The whole thing in one line
Tune until every alert is read, patch in rings with verification, automate what you have done by hand, treat the console as the privileged system it is, and know what each client costs you.
The point
Tune until every alert is read, patch in rings with verification, automate what you have done by hand, and know what each client costs you..
Underlying all of this
Everything in this collection reduces to four habits: tune until every alert is read, verify rather than assume at every stage from ring one to script execution, treat the console as the privileged system it is, and know what each client costs you. None needs a better platform, and a provider doing all four runs a quieter service than one twice its size.
The recurring pattern
The recurring pattern across every section here is the same: the appearance of control substituting for control. An unread alert queue looks like monitoring. A compliance percentage that excludes pending reboots looks like protection. A script that reports success looks like automation. In each case the provider believes a risk is handled and it is not, which is worse than knowing it is open.