How do you run on-call and incident response well, and reduce MTTR?
What they are really testing: Senior/SRE signal: connects observability to operational practice, good dashboards, runbooks, actionable alerts, blameless postmortems, and sustainable on-call, rather than just heroics.
A real interview question
How do you run on-call and incident response well, and reduce MTTR?
What most people say
drag me
“I would make sure someone is on call to fix things quickly when they break.”
Just having someone on call is heroics, not a system. Low MTTR comes from fast detection (good alerts/dashboards), fast diagnosis (runbooks, correlated telemetry), sustainable rotations, and blameless postmortems that prevent recurrence.
The follow-ups they ask next
Why are blameless postmortems important?
They surface honest, systemic causes instead of hiding them out of fear of blame, producing real fixes. Blaming individuals suppresses information and the same failure recurs; blameless culture improves the system.
How does alert fatigue increase MTTR?
Noisy/non-actionable alerts get ignored or muted, so a real incident is detected late (higher MTTD) and responders are desensitized. Trustworthy, actionable alerts are a prerequisite for fast detection.
What the interviewer is listening for
- Connects observability to detection/diagnosis
- Runbooks + correlated telemetry + roles (incident commander)
- Blameless postmortems + tracks MTTR/MTTD + sustainable on-call
What sinks the answer
- "Have someone on call" as the answer
- No runbooks/postmortems
- Ignores alert fatigue and sustainability
If you genuinely do not know
Say this instead of freezing. Reasoning out loud from what you do know beats silence every single time, and a good interviewer is listening for exactly that.
“Lower MTTR by making each phase fast and sustainable: [detect fast with symptom/SLO-burn alerts + golden-signal dashboards, low noise so the pager is trusted], [diagnose fast with runbooks per alert, correlated telemetry, clear escalation, and incident-commander roles], [sustainable on-call: reasonable rotations, manageable volume, clear severities], and [blameless postmortems with action items + track MTTR/MTTD]. Not heroics.”
Keep going with observability
Foundation
What is the difference between monitoring and observability?
Foundation
What are the three pillars of observability, and what is each good for?
Foundation
What should you measure first for a service? Explain the golden signals (or RED/USE).
Junior
What are good logging practices for a distributed system?
Junior
What are the main metric types (counter, gauge, histogram), and when do you use each?
Junior
What is distributed tracing, and how does it work?
Knowing the answer is not the same as recalling it under pressure
Sign in to send the questions you fumble to spaced recall, so they come back right before you would forget them, and learn the concepts behind them with hands-on labs.
Start free