Skip to content
Sections
All notes

All notes · Reference

A Checklist for the Whole Thing

Everything in this collection as a sequence, from before the contract to the annual review.

Reference · Reference

Before the contract

Discovery: directory, DHCP, network scan with written permission, their records.

The recommendations in “A Checklist for the Whole Thing” become sustainable only when the recurring work has visible owners and enough capacity. A team assessing the reporting guide can use time and project records to see where operational effort accumulates, without treating activity data as a substitute for technical evidence or direct discussion with technicians.

For an independent operational benchmark, compare the local practice with Atlassian knowledge-sharing guidance; the important test is whether the control remains proportionate, documented and recoverable when the usual technician is unavailable.

Count what is actually there, not what you were told.

Name the unsupportable: unsupported systems, equipment you will not touch.

Price the estate you found.

Agree the six: scope, approvals, hours, billing, out-of-scope, exit.

Onboarding

Agents to a representative sample first.

Watch the alert volume for a week.

Tune: assign a profile, record exceptions with reasons.

Then deploy to the rest.

Record a dated baseline: patch levels, hardware age, backup state, security software.

Documentation from day one, written when things are discovered.

Thirty-day review and a written summary to the client.

Monitoring

Alert on what somebody will act on, nothing else.

Disk in absolute terms; CPU with a duration; services that should never stop; event logs by identifier.

Dependencies configured.

Working hours per client, so out-of-hours alerts queue rather than page.

Count weekly volume, opened and actioned proportion, and incidents the client found first.

Patching

Four rings, ring zero your own.

Verification checklist after ring one.

A chosen delay between release and deployment.

Exclusions with reasons and review dates.

Report patched, pending reboot, excluded and unreachable separately.

Automation

Choose from ticket data, not from ideas.

Manual ten times before scripting.

Dry-run mode on anything that changes something.

Every script reports its effect.

Count remediations; escalate on repetition.

Second pair of eyes for multi-client scope.

Platform security

Multi-factor everywhere, no exceptions.

Technicians scoped to their clients.

Credentials in a secrets system, never in scripts.

Logs exported outside the platform.

Alerts on new administrator accounts, wide-scope scripts, unusual logins.

A written compromise plan, rehearsed once.

Monthly

Alert volume and actioned proportion.

Agent reconciliation against the directory.

Out-of-scope time.

Client report ending with decisions.

Quarterly

Exclusion review.

Template and script library review: what is unused, what duplicates.

Permission review against current roles.

Noisiest client re-tuned.

Cost per client.

Annually

Discovery repeated.

Agreement compared against what you actually do.

Vendor security settings reviewed.

Criteria for declining revisited.

The whole thing in one line

Tune until every alert is read, patch in rings with verification, automate what you have done by hand, treat the console as the privileged system it is, and know what each client costs you.

The point

Tune until every alert is read, patch in rings with verification, automate what you have done by hand, and know what each client costs you..

Underlying all of this

Everything in this collection reduces to four habits: tune until every alert is read, verify rather than assume at every stage from ring one to script execution, treat the console as the privileged system it is, and know what each client costs you. None needs a better platform, and a provider doing all four runs a quieter service than one twice its size.

The recurring pattern

The recurring pattern across every section here is the same: the appearance of control substituting for control. An unread alert queue looks like monitoring. A compliance percentage that excludes pending reboots looks like protection. A script that reports success looks like automation. In each case the provider believes a risk is handled and it is not, which is worse than knowing it is open.