Blog / Operations

Operations
July 21, 2026 · 6 min read · Nexus Team

Runbook-driven support versus tribal knowledge: what actually breaks when the one person who knows leaves

Tribal knowledge is a real asset, not a failure mode by itself — a technician who's worked a client's environment for years genuinely knows things no document captures, and pretending otherwise undervalues real expertise. The failure mode is treating that knowledge as the plan instead of as an input to the plan, because tribal knowledge has a single point of failure built into its name: it lives in one tribe member, and tribe members leave, get sick, go on vacation, or simply aren't the one who happens to be on shift when the problem recurs at 6pm on a Friday.

What actually breaks, specifically, when the person leaves

  • The next technician re-diagnoses a problem from scratch that was already solved once, because the solution existed only as a memory, not as a record anyone else could find
  • A workaround that was never meant to be permanent stays in place indefinitely, because the person who knew it was a workaround — and knew the real fix — is the one who's gone
  • Client-specific quirks — this firewall rule exists for a reason nobody wrote down, this server reboots badly if you don't stop this service first — get rediscovered the hard way, usually during an outage
  • Institutional trust erodes quietly, because the client notices that competence at their account dropped the day a specific person's calendar changed, which is a business risk dressed as a personnel event

What a runbook actually has to be to prevent that

  • Written at the point the knowledge is fresh — during or right after the fix — not reconstructed later from memory when the details have already blurred
  • Specific to the actual environment and its quirks, not a generic vendor procedure that skips the client-specific step that made the last three attempts fail
  • Attached to the asset or the recurring issue itself, so the next person finds it by looking at the thing that's broken, not by remembering it exists and searching for it
  • Treated as a living document that gets corrected when it's wrong, rather than trusted forever the moment it's written once
A runbook isn't a replacement for expertise. It's what expertise looks like after it survives the person who had it.

This is also where an AI drafting layer earns a real, narrow role rather than a dramatic one: because Nexus keeps ticket history, asset records, and resolution notes in one place per tenant, an agent drafting a response to a recurring issue can surface the runbook or the prior resolution automatically instead of relying on a technician to remember it exists — turning tribal knowledge that happened to get written down into knowledge that actually gets found. It doesn't replace the discipline of writing the runbook down in the first place. Nothing does. That part is still on the humans, same as it always was.

Follow the build as it ships.

Nexus is live in our own MSP operations and opening to a limited design-partner cohort. Join the private-preview list.