Skip to content
Sections
All notes

All notes · Automation

Self-Healing and Its Limits

Automated remediation closes tickets before they open. It also conceals the problem it keeps fixing.

Automation · Analysis

Self-healing means detecting a condition and correcting it without human involvement. It is genuinely valuable and it has a specific failure mode that takes months to notice.

Automation around “Self-Healing and Its Limits” should reduce repetitive labour while leaving ownership and review visible. Teams considering Monitask can compare the time spent on manual diagnosis, scripted remediation and later investigation, but technical logs must remain the evidence of what the automation actually changed.

For an independent operational benchmark, compare the local practice with Red Hat automation resources; the important test is whether the control remains proportionate, documented and recoverable when the usual technician is unavailable.

Where it works well

Conditions with a known cause and a safe fix: a service that stops, a queue that jams, a temporary file build-up.

High frequency, low consequence, well understood.

These are the cases where automation removes work and adds nothing.

The hidden-problem effect

A service restarting itself five times a day is a resolved alert and an unresolved fault.

The ticket count falls. The underlying cause persists and sometimes worsens.

And because the automation works, nobody investigates.

This is the central risk and it is entirely preventable by counting.

Counting remediations

Every automated fix should be recorded, even though no ticket is raised.

Then review: which remediations fire most, on which machines, trending how.

A remediation firing more often than it used to is a problem developing, and it is invisible unless somebody looks at the count.

The escalation rule

If a remediation fires more than a defined number of times in a period, raise a ticket.

Three times in a day, or five in a week, depending on the condition.

This converts the automation from a mask into a detector, which is the whole difference.

What not to self-heal

Anything where the fix could cause harm if the diagnosis is wrong.

Anything on a server during business hours without notice.

Conditions you do not understand, which is the tempting case — the fix that works for unknown reasons will eventually not work for unknown reasons.

And anything a client has asked to be told about, which is a contractual matter rather than a technical one.

Telling the client

Self-healing performed silently looks like nothing happened.

Which is fine until they ask what they are paying for.

Report remediation counts in the monthly summary: these conditions occurred and were corrected automatically.

It is genuine value and it is invisible unless stated.

The review

Quarterly: which remediations fire, which never, which are masking something.

Remove the ones that never fire.

Investigate the ones that fire constantly.

An hour, and it is the maintenance that keeps automation honest.

What to check

Do you count automated remediations?

Which fires most often, and has anybody investigated why?

Is there a threshold that escalates to a ticket?

And do clients see the remediation count?

The point

A service restarting itself five times a day is a resolved alert and an unresolved fault.

Count remediations or the automation becomes a mask.

Underlying all of this

Everything in this collection reduces to four habits: tune until every alert is read, verify rather than assume at every stage from ring one to script execution, treat the console as the privileged system it is, and know what each client costs you. None needs a better platform, and a provider doing all four runs a quieter service than one twice its size.

The recurring pattern

The recurring pattern across every section here is the same: the appearance of control substituting for control. An unread alert queue looks like monitoring. A compliance percentage that excludes pending reboots looks like protection. A script that reports success looks like automation. In each case the provider believes a risk is handled and it is not, which is worse than knowing it is open.