Phase 13 · Site Reliability Engineering Practices
TopicsRunbooks & Documentation
Part of the DevOps Roadmap.
Summary
Step-by-step guides for handling known operational tasks or incident types — turning tribal knowledge (only in one person's head) into something any on-call engineer can follow at 3am.
How to Learn This
- 1Write a runbook for a common operational task (e.g. restarting a stuck service).
- 2Learn why a runbook should be specific and step-by-step, not a vague general description.
- 3Practice keeping a runbook updated as the underlying system or process changes.
More topics in Site Reliability Engineering Practices
Stuck on this topic? Ask an Insider
Get 1:1 guidance from people who've walked this exact path — free on the InsideEdge app.