SRE & Incident Management · running systems, not just building them
SRE & Incident Management: the questions that separate "I shipped code" from "I've carried the pager."
Most backend interview prep stops at design and stops before 3am. This guide covers what happens after code ships: SLIs, SLOs, SLAs and the error-budget math that makes them enforceable, SEV1–SEV4 severity and the incident commander role, blameless postmortems, executable runbooks, chaos engineering, and on-call design — with diagrams and 22 interview-depth Q&A.
Why operating a system is a different skill than building one
Almost every interview-prep resource on backend engineering is about design: how you'd shard a database, choose a message queue, or lay out a microservices boundary. Very little of it covers what happens after that system is live and something breaks at 3am on a Sunday. That gap is exactly where interviewers separate candidates who have shipped code from candidates who have operated it — carried a pager, made a call under pressure with incomplete information, and turned a bad night into a system that's actually harder to break the same way twice.
Site Reliability Engineering, as a discipline, is essentially the codification of that operational wisdom into practices you can hire for, interview on, and improve deliberately instead of leaving to whoever happens to be senior and calm under pressure. None of it is exotic: it's a numeric definition of "reliable enough" (SLOs and error budgets), a structured way to run and staff an incident (severity levels and incident command), a way to learn from failure without punishing the people who report it (blameless postmortems), a way to make 3am response fast instead of improvised (runbooks), a way to find weaknesses before they find you (chaos engineering), and a way to keep the humans doing this work sustainable (on-call design).
Jump to a section
SLIs, SLOs, SLAs, and the error budget that makes them enforceable
These three terms get used interchangeably in casual conversation, but they answer three different questions, and the distinction is exactly what most candidates blur together in an interview.
| Term | What it answers | Audience | Example |
|---|---|---|---|
| SLI — service level indicator | What are we actually measuring, right now? | Internal, engineering | % of requests returning <300ms with a 2xx status |
| SLO — service level objective | What target do we hold ourselves to on that measurement? | Internal, product + engineering | 99.9% of requests meet that bar over a rolling 30 days |
| SLA — service level agreement | What have we promised customers, with consequences if we miss? | External, contractual | 99.5% uptime, or the customer receives service credits |
The SLA is deliberately looser than the SLO — if the internal target is 99.9% and the external promise is 99.5%, that gap is intentional margin, so a genuinely bad month still doesn't trigger a customer-facing contractual breach on top of the incident itself.
The error budget: turning "reliable enough" into a number you can spend
An error budget is simply 100% − SLO, expressed as an amount of allowed unreliability over the SLO's time window. It converts a fuzzy cultural argument ("should we slow down and be careful, or ship faster?") into an objective, calculable number both product and engineering can look at without it being a political standoff.
The grey dashed diagonal is a 1x "on-pace" reference — consuming the budget at exactly the rate that empties it right at month end, which is fine and needs no one paged. The rose line is what actually happened: mostly flat, then a steep drop when a real incident burned budget far faster than 1x. That's burn rate — how fast you're consuming budget relative to that on-pace reference. A burn rate of 1x is background noise; a burn rate of 14–20x means the whole month's budget would be gone in under two days if it kept going, which is exactly the signal multi-window, multi-burn-rate alerting is designed to catch quickly, without also paging someone for ordinary, sustainable degradation that a flat threshold alert would flag constantly.
The behavioral payoff is the point of the whole exercise: while budget remains, a team has explicit license to take reliability risk in exchange for velocity — ship faster, run a risky migration, adopt new infrastructure. Once burn rate or remaining budget crosses the freeze threshold, the same team is expected to stop shipping risky changes and redirect effort into reliability work until the trend recovers. The freeze itself isn't the valuable part — it's that the trigger is a number everyone already agreed to, instead of a fresh argument every time between whoever wants to ship and whoever wants to be cautious.
Incident severity levels and the incident commander
Severity classification exists so the intensity of the response matches the actual impact — enough people paged for a real SEV1, not so many that every minor blip trains the org to ignore incident channels entirely.
| Severity | Definition | Response | Example |
|---|---|---|---|
| SEV1 | Full outage or critical data-integrity risk on a core user journey. | Page immediately, all hands, IC declared, exec/status-page comms. | Checkout or login is completely down for all users. |
| SEV2 | Significant degradation or partial outage affecting a meaningful slice of users or a core feature. | Page primary + secondary on-call, IC usually declared. | Checkout succeeds but is 5x slower for 20% of traffic. |
| SEV3 | Minor, contained impact; workaround exists or affects a small user segment. | Ticketed, handled during business hours, no page. | A non-critical admin report is generating stale numbers. |
| SEV4 | Cosmetic or negligible impact, no urgency. | Backlog item. | A typo in an email template. |
The incident commander is a role, not a rank — a single person who owns coordination and decision authority for the duration of the live response. That's explicitly not the same as being the person with the deepest technical context; the IC's job is making sure the team pursues one decision path at a time, triaging competing theories instead of letting every senior engineer chase their own, and shielding the people doing hands-on work from stakeholder interruptions so they can stay heads-down. Supporting roles — a scribe capturing the objective timeline as it happens, a comms lead handling stakeholder and customer-facing updates, an ops lead executing the actual technical fix — can all be delegated by the IC, and often should be on anything beyond a small, single-responder incident.
Why does this matter even among senior engineers? Because seniority doesn't prevent two capable people from independently pursuing conflicting mitigations at the same moment — one rolling back a deploy while another restarts the affected pods can each mask the other's evidence about what's actually happening, and doubles the number of variables changing at once when the whole goal is to reduce uncertainty as fast as possible. A single point of decision authority isn't about ego or hierarchy; it's about not thrashing when the cost of thrashing is measured in minutes of customer impact.
Blameless postmortems: the document structure and why blame kills the data you need
A blame culture doesn't prevent mistakes — it just makes people better at hiding them. If someone believes the honest version of what happened will be used against them personally, they'll unconsciously sanitize the timeline, soften the moment they guessed wrong or hesitated, and steer the story toward whatever sounds least damaging. That omitted moment is almost always exactly where the causal insight lives, which means blame culture doesn't reduce mistakes; it reduces the information available to prevent the next one.
Blameless doesn't mean nobody is accountable. It means accountability is aimed at the system and the conditions that made a mistake easy to make, rather than at the individual who happened to be holding the pager when a pre-existing systemic gap finally got exercised. Practically, that shows up as ground rules: the postmortem describes what a reasonable person would have done given the information and tooling available at the time, not what an idealized engineer would have done with hindsight.
Standard postmortem document structure
Impact leads the document, before the technical detail, because the audience for a postmortem is often wider than the responding team — a leader deciding how much investment this deserves needs to calibrate severity before wading into a technical narrative. Burying "2 hours of checkout downtime, 4% of daily revenue" halfway down the page under stack traces makes the document harder to triage against everything else competing for attention.
The five whys, done well
The five whys technique repeatedly asks "why did that happen" until you reach a mechanism specific enough to actually fix. A worked example: Why did checkout fail? — the order service couldn't reach the database. Why? — the connection pool was exhausted. Why? — a batch reporting job opened hundreds of long-lived connections against the same pool. Why? — the reporting job and the checkout path share a connection pool with no isolation between them. Why? — nobody had flagged that shared infrastructure resource as a single point of failure between an unrelated batch workload and the revenue-critical path. That last answer is concrete and fixable (separate pools, or resource limits per workload) — which is the test for whether you've gone deep enough.
The two failure modes to watch for are opposite directions of the same mistake. Stopping too early — "the engineer ran the wrong command" — blames a person instead of asking why the tooling made that command easy to run by accident. Going too far the other way — "our culture doesn't prioritize reliability" — is too abstract for anyone to act on this quarter. The right stopping point is specific and concrete enough that a named owner could build a fix for it, and general enough that fixing it actually prevents a class of failure, not just this one exact sequence of events.
Root cause versus contributing factor
The root cause is the mechanism that, if it hadn't happened, this specific incident wouldn't have occurred — in the example above, the shared, unisolated connection pool. A contributing factor made the incident more likely, more severe, or harder to detect and recover from, without being sufficient on its own to cause it: no alert on connection pool saturation, a runbook that didn't cover this failure mode, or an on-call engineer new to this service who took longer to orient. Most real incidents have exactly one root cause and several contributing factors, and a postmortem that only fixes the root cause while ignoring the contributing factors still leaves the team slower to detect and recover from the next, different root cause.
Runbooks: executable steps, not documentation
A runbook and a design doc solve different problems, and confusing the two is the most common reason runbooks fail exactly when they're needed most.
A good runbook is a sequence of exact, copy-pasteable commands with the expected output at each step, and explicit branches for what to do next depending on what you see — written so a competent engineer with no prior context on this specific system could follow it correctly under pressure at 3am. A bad runbook is a paragraph of prose explaining how the system works and what might go wrong, written by someone who already understands the architecture, for a reader who's implicitly assumed to understand it too — exactly backwards from who's actually reading it during a live incident.
# GOOD: exact commands, expected output, explicit branches
1. Check current connection pool usage:
kubectl exec -it order-svc-0 -- curl localhost:8080/actuator/metrics/hikaricp.connections.active
Expect: value below 80 (max pool size is 100).
2. If value >= 90, the pool is likely exhausted by the reporting job. Kill it:
kubectl delete job nightly-report --namespace=batch
Then re-check step 1. If it drops below 80 within 60s, go to step 4.
3. If value is normal but checkout is still failing, escalate to #db-oncall now --
do not keep investigating alone past this point.
4. Confirm checkout recovery:
curl -s https://internal/api/health/checkout | jq .status
Expect: "UP". If not "UP" after 3 minutes, escalate per step 3.
Commands beat prose specifically because incident pressure already consumes working memory — reading a paragraph to reconstruct what to actually type costs time and invites misreads that a command you can paste directly doesn't. The explanation of why the fix works belongs in the design doc or the postmortem that produced this runbook step, not in the thing you're executing while the clock is running.
Runbooks decay silently as the system underneath them changes, so keeping them trustworthy takes deliberate effort on two fronts. First, treat "this runbook step was wrong or missing" as its own postmortem action item every time it happens during a real incident — that's the clearest possible signal that it's drifted from reality. Second, exercise runbooks on purpose, during scheduled game days or as part of chaos experiments, so the stale step gets found on a calm Tuesday afternoon instead of in the middle of a real SEV1. A runbook nobody has actually run since it was written should be treated as unverified, not trusted.
Chaos engineering: finding weaknesses on your own schedule
The core principle is simple to state and easy to get wrong in practice: deliberately inject failure into a system, under controlled conditions, to discover weaknesses before they surface as a real incident at a worse time — a time you didn't choose, without the people or tooling in place to handle it well.
A chaos experiment starts from a steady-state hypothesis: a measurable baseline of normal behavior, like "checkout success rate stays above 99.5%," and a prediction that it holds even while the failure is injected. If the hypothesis turns out false, you've found a real weakness on your own schedule, with your own team watching, instead of discovering it during an actual customer-facing outage.
Blast radius is how much of the system and how many real users could be affected if the failure behaves worse than hypothesized. The discipline is starting as narrow as possible — a single canary instance, a capped percentage of traffic, a non-critical environment or off-peak window — and only widening scope in later runs once confidence is earned. Just as important is an automatic abort condition tied to that same steady-state metric: if it crosses a bad threshold, the experiment halts and reverts itself immediately, without waiting for a human to notice. An experiment that can't stop itself the instant it goes wrong isn't a controlled experiment — it's just an outage you scheduled on purpose.
Chaos engineering is also directly connected to the error budget from earlier: running experiments in anything resembling production consumes a small, deliberate slice of that budget, and treating it that way — planned, scoped, accounted for — is what keeps it a legitimate engineering practice rather than reckless improvisation. Common experiments include killing a random instance, injecting artificial latency between two services, dropping network packets on a specific link, exhausting a database connection pool on purpose, and running a full regional failover drill.
On-call design: rotation fairness, alert fatigue, and toil
The people carrying the pager are as much a part of the system's reliability as the code itself, and on-call design that ignores that eventually shows up as attrition, burnout, and slower incident response — not just as an HR concern, but as an engineering risk.
Alert fatigue isn't a personality weakness in whoever's holding the pager that week — it's a leading indicator that the alerting design itself needs to change. Every alert that pages a human should be actionable (there's something a person can actually do about it right now), should link directly to its runbook so the responder isn't starting cold, and should be deduplicated against other alerts already firing for the same underlying issue. An alerting system tuned to maximize sensitivity without regard for signal-to-noise ratio isn't more safe — it's training its own responders to stop trusting it.
Toil reduction earning a place on the actual roadmap, rather than being whatever gets squeezed in after "real" feature work, is the mechanism that keeps a growing system from eventually requiring a growing on-call team just to stay afloat. A team that never budgets time against toil will find that budget consumed anyway — just involuntarily, and at the worst possible moments.
Key design decisions and interview talking points
These are asked as scenario questions far more often than as textbook definitions — an interviewer wants to hear you reason through a real 3am decision, not recite a glossary.
Walk me through your first 10 minutes after being paged for a SEV1.
First, acknowledge the page so the alert stops escalating and other responders know someone's on it. Then spend two minutes confirming it's real and actually customer-impacting before doing anything else — a surprising number of pages are noise, and burning ten minutes mitigating a false alarm is its own kind of incident. If it's real, declare the incident formally (open the incident channel or bridge, not just a Slack DM), and either take the incident commander role explicitly or explicitly hand it to someone else — the worst failure mode is two senior engineers both assuming the other is coordinating. Only after that do you start actually investigating, and even then the first goal is mitigation (stop the bleeding — roll back, fail over, shed load) not root cause, because every additional minute of customer impact matters more than understanding why it happened.
How do you decide whether something is a SEV1 or a SEV2?
Severity should be driven by customer-facing impact and scope, not by how technically scary the underlying bug looks. A full outage of a core user journey (checkout, login) is SEV1 regardless of how simple the eventual fix turns out to be; a degraded but still-functional experience affecting a minority of traffic, or a full outage of a non-critical feature, is usually SEV2. The classification also drives who gets paged and how loudly, so erring toward over-classifying early and downgrading once you understand scope is safer than under-classifying and discovering ten minutes in that you needed the extra hands from the start.
Why does incident response need a single incident commander, even when everyone responding is senior — and should that person also be the one fixing the problem?
Seniority doesn't stop two people from pursuing conflicting mitigations at once — one engineer rolling back a deploy while another restarts pods can each mask the other's evidence and double the number of variables changing simultaneously. The IC's job is coordination and decision authority, not technical superiority: making sure one decision path is pursued at a time, triaging incoming theories instead of letting everyone chase their own, and shielding the responders doing hands-on work from stakeholder interruptions. On a small incident the same person often plays both roles out of necessity, but on anything larger, splitting them is what keeps things fast — someone heads-down in logs loses situational awareness of the bigger picture, while a dedicated IC can track multiple hypotheses and manage communications without breaking the fixer's concentration.
What's the difference between a root cause and a contributing factor?
The root cause is the mechanism that, if it hadn't happened, this specific failure wouldn't have occurred — for example, a database migration that removed an index the checkout query depended on. A contributing factor made the failure more likely, more severe, or harder to detect, without being sufficient on its own to cause it — the missing runbook step that cost the responder ten extra minutes, or the dashboard that didn't have an alert on that particular query's latency. Real incidents usually have one root cause and several contributing factors, and a postmortem that only fixes the root cause while ignoring the contributing factors will still be slower to detect and recover from the next, different root cause.
Why do blameless postmortems produce better outcomes than ones that assign fault to an individual?
If people believe the honest version of events will be used against them, they'll unconsciously sanitize the timeline, omit the moment they hesitated or guessed wrong, and steer the narrative toward whatever sounds least personally damaging — and that omitted moment is usually exactly where the causal insight lives. Blameless doesn't mean nobody is accountable; it means accountability is aimed at the system and the conditions that made the mistake easy to make, not at the individual who happened to be holding the pager when a systemic gap finally got exercised. In practice this produces postmortems people actually want to attend and contribute honestly to, instead of ones people dread and route around.
What's wrong with a postmortem action item that just says "the engineer should be more careful next time"?
It isn't specific, isn't measurable, isn't assigned an owner or a deadline, and most importantly it doesn't change anything about the system — the exact same conditions that produced this incident are still there, waiting for the next tired engineer at 3am to make the same understandable mistake. A real action item changes a system property: add a confirmation step to a destructive command, add an automated check that would have caught this class of error, add an alert that would have surfaced the problem sooner. If the only fix you can think of is "be more careful," that's usually a sign you haven't found the actual root cause yet.
How do you use the five whys without either stopping too early or wandering into irrelevant territory?
Each "why" should stay anchored to this specific incident's actual causal chain, not slide into a general critique of the team or the org — if an answer is "because our culture doesn't prioritize testing," you've gone too abstract to act on, and if you stop at "because the engineer ran the wrong command," you've stopped too early, because you haven't asked why the tooling made that command easy to run by mistake. A good stopping point is when the next "why" would require fixing something outside this incident's actual scope, and you've landed on a mechanism specific and concrete enough that someone could realistically build a fix for it this quarter.
What does a good postmortem document actually contain, and why does impact come before root cause?
The standard shape is: a summary with customer-facing impact and duration first, then an objective timestamped timeline of what happened, then root cause analysis and contributing factors, then action items with named owners and due dates. Impact goes first because the person reading it — often someone outside the team, like a leader deciding how much investment this deserves — needs to calibrate severity before they wade into technical detail; burying "2 hours of checkout downtime, 4% of daily revenue" halfway through a technical narrative makes the document harder to triage and prioritize against everything else competing for attention.
What's the actual difference between an SLI, an SLO, and an SLA?
An SLI (service level indicator) is the raw measurement — the percentage of requests that returned successfully in under 300ms, say. An SLO (service level objective) is the internal target you hold yourself to on that indicator, like 99.9% of requests meeting that bar over a rolling 30 days. An SLA (service level agreement) is the external, usually contractual, promise to customers, and it's deliberately set looser than the internal SLO — if your SLO is 99.9% and your SLA is 99.5%, that gap is margin so that a bad month still doesn't trigger a customer-facing breach with financial penalties attached.
Walk me through the math: what does a 99.9% monthly SLO actually mean in minutes of downtime?
A 30-day month has 43,200 minutes. A 99.9% SLO allows 0.1% of that to be out of compliance, which is 43,200 × 0.001 ≈ 43.2 minutes. That's your error budget for the month — roughly 43 minutes of full downtime, or a proportionally larger amount of partial degradation, whichever combination of incidents actually consumes it. Tightening to 99.95% cuts that to about 21.6 minutes, and 99.99% ("four nines") cuts it to about 4.3 minutes — which is why each additional nine gets disproportionately more expensive to hold, not linearly more expensive.
How does having an error budget actually change what a team is allowed to do day to day?
As long as there's budget left, the team has explicit license to take reliability risk in exchange for velocity — ship faster, run migrations during business hours, adopt new infrastructure — because a bad deploy that costs a few minutes of downtime is an acceptable, budgeted trade. Once the budget is nearly exhausted, the same team is expected to freeze risky changes and redirect effort toward reliability work until the budget recovers. The value isn't the freeze itself — it's that the trigger is an objective number instead of a political argument between whoever wants to ship and whoever wants to be cautious.
What's burn rate, and why would a fast burn get someone paged at 3am when a slow burn wouldn't?
Burn rate is how fast you're consuming error budget relative to the rate that would exhaust it exactly at the end of the SLO window. A burn rate of 1x means you're on pace to use the whole month's budget by month end, which is fine and doesn't need anyone woken up. A burn rate of 20x means you'll exhaust an entire month's error budget in about 36 hours if it continues — that's a real ongoing incident, and multi-window burn-rate alerting is specifically designed to catch that fast-burn case quickly while not paging anyone for the slow, sustainable kind of degradation that a static threshold alert would otherwise flag constantly.
If a team blows through its error budget for the period, what's supposed to happen next?
The team stops shipping features that carry reliability risk and shifts its roadmap toward the reliability work that will prevent a repeat — better testing, more automation, addressing the specific class of failure that burned the budget — until the SLO is back on a healthy trend. This is the mechanism that's supposed to make reliability investment happen automatically instead of only after a bad enough outage forces a reactive scramble; if a team routinely blows its budget and nothing actually changes, the error budget policy has no teeth and is just theater.
What actually separates a good runbook from a bad one?
A good runbook is a sequence of exact, copy-pasteable commands with the expected output at each step and explicit branches — "if you see X, go to step 4; if you see Y, escalate to the database team" — written so a competent engineer unfamiliar with this specific system could follow it under pressure at 3am. A bad runbook is a prose explanation of how the system works, written by someone who already understands the architecture, for a reader who's assumed to understand it too. Commands beat prose specifically because incident pressure already consumes working memory — reading a paragraph to reconstruct what to type costs time and invites misreads that a command you can paste doesn't; the explanation of why it works belongs in the design doc or the postmortem, not in the thing you execute under pressure.
How do you keep runbooks from silently going stale as the system changes?
Treat "this runbook step was wrong or missing" as its own postmortem action item every time it happens during a real incident — that's the signal that it drifted from reality. On top of that, periodically exercise runbooks deliberately, during game days or as part of chaos experiments, so you find the stale step on a Tuesday afternoon instead of discovering it mid-SEV1; a runbook nobody has run since it was written should be treated as unverified, not trusted.
What is chaos engineering actually trying to prove, in one sentence?
That your system can survive a specific, real failure mode you believe it can already handle — proven by actually injecting that failure under controlled conditions, rather than assumed because the architecture diagram implies it should work.
How do you design a chaos experiment so that running it doesn't become the very incident it was meant to prevent?
You start by defining a steady-state hypothesis — a measurable baseline like "checkout success rate stays above 99.5%" — and predicting it will hold under the injected failure. You constrain blast radius as tightly as possible for the first run: one instance, a small percentage of traffic, a non-critical environment or time window. And you define an automatic abort condition tied to the same steady-state metric, so if it crosses a bad threshold the experiment halts and reverts itself immediately without waiting for a human to notice and intervene. An experiment without an automated abort path is just an outage you scheduled on purpose.
What's blast radius, and how do you actually control it?
Blast radius is how much of the system and how many real users could be affected if the injected failure behaves worse than hypothesized. You control it by scoping the experiment as narrowly as the first run allows — a single canary instance, a single shard, a capped percentage of traffic, off-peak hours — and only widening scope in later runs once you've built confidence the hypothesis holds. Combined with an automatic abort condition, tight blast radius is what turns "randomly break production" into a controlled, repeatable engineering practice.
What causes alert fatigue, and why is it more dangerous than just being annoying?
Alert fatigue comes from alerts that don't require action, alerts with no attached context or runbook forcing a cold investigation every time, duplicate alerts firing from multiple monitoring layers for the same root issue, and thresholds tuned too sensitively relative to what actually matters. It's dangerous because humans habituate — once most pages turn out to be noise, the instinctive response becomes acknowledge-and-dismiss, and that reflex doesn't turn itself off for the one page in fifty that's a real SEV1. Alert fatigue isn't a personality problem in the on-call engineer; it's a leading indicator that the alerting design itself needs to change.
Why is reducing toil treated as an engineering priority instead of just an accepted cost of running a service?
Toil — manual, repetitive, automatable work with no lasting value, like re-running a stuck job by hand every week — tends to scale linearly with the system's growth if nobody deliberately fixes it, and left unmanaged it eventually consumes all of a team's available time. That's time taken directly away from the automation and reliability investment that would have prevented the toil in the first place, so the situation only gets worse. Treating toil reduction as budgeted engineering work — a common rule of thumb caps it at roughly half of an SRE's time — is what actually breaks that cycle instead of accepting it as background noise.
How do you design an on-call rotation that stays fair as the team grows or shrinks?
Fairness isn't just an even number of on-call days per person — it has to account for actual page volume and severity, not just calendar time, because a rotation that's fair by the calendar but dumps every 3am page on whoever's on call during a recurring batch-job failure isn't fair in any way that matters. Practical elements include a secondary/backup on-call so no single person is a point of failure, a cap on consecutive on-call stretches, real compensation for the burden, and tracking on-call load itself as a metric that feeds back into whether the team needs to invest in reducing pages rather than just rotating who suffers them.
What's the single biggest mental shift an engineer needs to make going from "I ship code" to "I own this system in production"?
Accepting that reliability is a property you design and budget for deliberately, not a side effect of writing correct code — the best-written service still goes down, and the engineers who handle that well are the ones who've built the muscle of detecting fast, communicating clearly under pressure, and turning every incident into a specific system change instead of a story about someone's bad night. That shift, more than any specific tool, is what interviewers are actually screening for when they ask about incidents.
Related guides
Update these hrefs to your published Blogger post URLs once each page is live.
Post a Comment
Add