Blog / Operations
The three outages nobody pages for: certificate, domain, and the second failed check
Uptime monitoring gets designed around the dramatic case: the server falls over at 3am and something wakes somebody up. That case is real and it is also the one every tool handles. The outages that disproportionately hit small businesses are duller and entirely foreseeable — a TLS certificate that lapsed, a domain registration that nobody renewed, and a service that was flapping for a week before it stayed down.
All three are calendar failures. None of them require a monitoring system that is clever; they require one that is looking at the right three things.
Certificate expiry is an outage with a known date
An expired certificate is a total outage for anything that validates it, and it arrives on a date that was knowable months in advance. Automated renewal has made this rarer and, perversely, worse when it happens: the renewal that has worked for two years silently stops working, and nobody is watching because it has always worked.
The check that matters is not "is renewal configured" but "what does a live handshake say the expiry is right now." Configuration describes intent. A handshake describes what a visitor's browser is about to be told.
Domain expiry is worse, and less monitored
A lapsed domain takes down the website, the email, and often the identity provider that everything else authenticates against — simultaneously, and with a recovery path that runs through a registrar's support queue rather than through you. It is the single highest-consequence, lowest-effort item on this list, and it is monitored far less often than certificates because there is no equivalent to the browser warning that trains people to care.
Registration data is publicly queryable, so expiry can be checked directly rather than tracked in a spreadsheet somebody inherited. The useful design detail: an automatically-discovered expiry date should stop auto-refreshing the moment a human edits it by hand, because a manual override usually means the human knows something the public record does not yet reflect.
A domain lapse is the only outage on this list where the fix is a phone call to somebody else's support desk.
The second consecutive failure, not the first
The third case is not an outage detection problem, it is an alert credibility problem. A monitor that alerts on the first failed check will fire on every transient network blip, and a team that has been woken by three false alarms will not react promptly to the fourth alert — which is the real one.
Alerting on the second consecutive failure rather than the first costs you one check interval of detection latency and buys back the thing that actually determines outcomes: whether anyone believes the alert. That trade is almost always correct for small teams, where the on-call rotation is one person and their goodwill is a finite resource.
What this looks like in practice
Nexus carries a built-in monitor set alongside the RMM: admin-defined HTTP, HTTPS, and TCP checks on a per-monitor interval, certificate expiry read from a live TLS handshake, and domain-registration expiry resolved from public registration data with a manual override that sticks. A transition to down raises an alert on the second consecutive failure, and soon-to-expire certificates and domains raise their own — so the two outages with a known date get flagged while there is still time to act, rather than at the moment they take the client offline.
It is a modest feature. It also covers the three failure modes that most often turn into a bad week for a business that has no internal IT, which is a better return than most of what monitoring tooling competes on.