SLIs, SLOs, SLAs, and Error Budgets Questions
Defining and operating reliability targets. Covers choosing service level indicators, setting service level objectives and agreements, computing and spending error budgets, and using them to drive engineering decisions. Includes negotiating reliability targets with stakeholders.
Design an SLO-based release gating system that can scale across hundreds of services. Describe the architecture (centralized vs decentralized), how SLIs are ingested and validated, enforcement mechanisms (CI/CD pre-deploy checks, automated gating), handling of flaky metrics and partial outages, and how teams are onboarded or may opt-in/opt-out.
Sample Answer
A release-gating system that scales to hundreds of services needs to be a shared PLATFORM (a decentralized enforcement model plugging into a centralized SLI ingestion and policy layer), not either a fully centralized bottleneck or a fully decentralized free-for-all.
Structured elaboration
Architecture: a central SLI ingestion and validation layer (so every team's SLIs are collected in a consistent, auditable way) with DECENTRALIZED enforcement (each team's CI/CD pipeline calls a shared gating API before promoting a release, rather than a single central system being in the critical path of every deploy across the whole company, which would be both a scaling bottleneck and a single point of failure). Enforcement: a pre-deploy check queries the current error-budget state for the service being deployed; flaky metrics (a brief data gap, a known noisy period) need an explicit "insufficient confidence, default to allow with a logged warning" fallback rather than either blocking indefinitely on missing data or silently treating a gap as healthy. Onboarding: teams should be able to opt IN gradually (dry-run mode, showing what WOULD have been blocked, before the gate becomes enforcing), and an opt-out escape hatch (with mandatory logging and a review trigger) needs to exist for genuine exceptions.
Worked example
On top of core SLO-based gating, richer signals layer in over time: real-user telemetry and even ML-based anomaly detection on top of standard synthetic-test SLO analysis, so the gate can catch a regression that basic SLI thresholds alone might miss (an anomaly in a metric that hasn't yet crossed an absolute threshold but looks statistically unusual relative to that service's own history). Enforcement itself maps burn-rate tiers to specific actions: healthy budget allows normal deploy velocity; moderate burn triggers a monitoring-only warning; severe or sustained burn triggers an automatic freeze requiring a documented runbook step before any override. When such a policy changes (e.g. tightening the gate's default sensitivity), a stakeholder communication plan is needed for both INTERNAL teams (a heads-up and a dry-run period before enforcement) and, where the change could affect commitments made to external customers, an appropriately separate communication.
Trade-offs and pitfalls
A fully centralized gating system that must approve every deploy across hundreds of services becomes both a bottleneck (a single slow dependency check delays every team's release) and an organizational flashpoint (one team's incident blocking every other team's unrelated deploy if the design isn't carefully scoped per-service); the decentralized-enforcement-with-central-data model avoids both. The dry-run-before-enforcing onboarding path is not optional at this scale: forcing every team straight into a hard-enforcing gate on day one, with no chance to see what it WOULD have blocked first, reliably produces resistance and workarounds that undermine the whole system's credibility.
Compare rolling windows and fixed calendar windows for SLO measurement. For each approach list pros and cons, and give an example scenario where you would prefer a rolling 30-day window versus a calendar month window.
Sample Answer
A rolling window slides continuously (every evaluation looks back exactly N days from right now); a fixed calendar window resets at a boundary (the 1st of the month), and that difference in RESET behavior is the whole practical distinction.
Structured elaboration
Rolling windows give a smooth, always-current view of recent reliability and avoid the "reset cliff" problem: with a calendar window, a service that had a terrible incident on the 2nd of the month can look perfect again by the 3rd simply because a new month started, even though nothing was actually fixed. Calendar (fixed) windows are simpler to reason about for REPORTING and BILLING purposes (a customer expects "this month's SLA report," not an arbitrary trailing 30-day snapshot that changes meaning depending on which day you look at it), and they align naturally with typical business reporting cadences.
Worked example
A team debugging day-to-day reliability and deciding release velocity should prefer a rolling 30-day window: it gives a continuously accurate picture of "how much budget is left right now," which is what actually matters for today's release decision, uncontaminated by an arbitrary calendar boundary. A customer-facing SLA report, by contrast, is usually better served by a calendar-month window, since "your July SLA compliance was 99.92%" is a cleanly reportable, auditable, contractually-referenceable statement in a way that "your trailing-30-days-as-of-whenever-you-happen-to-check compliance" is not.
Trade-offs and pitfalls
Using ONLY a calendar window for internal error-budget decisions creates a perverse incentive right around month-end: a team burning budget hard in the last week of the month knows the counter effectively resets in a few days, which can (even subtly, even unintentionally) reduce urgency to fix a real problem versus just waiting it out. Using ONLY a rolling window for external contractual reporting creates the opposite problem: a customer has no fixed, auditable period to check compliance against, since the "current" 30-day figure is different every single day they might look. Many mature setups use both: rolling for internal operational decisions, calendar for external reporting, explicitly choosing the right window for each AUDIENCE and PURPOSE rather than picking one window type company-wide.
Design an enterprise reliability governance model that standardizes SLO ownership, error budget policies, reporting cadence, escalation paths, and incentives across multiple product lines. Include roles (product owner, SRE, engineering manager), required dashboards, and how you would enforce and audit compliance without slowing delivery.
Sample Answer
Enterprise-scale reliability governance needs a genuinely tiered structure, since the same rigid policy applied uniformly across 50 microservices of wildly different criticality either over-constrains the unimportant ones or under-protects the critical ones, and every layer of that structure needs an explicit, named owner and a real audit mechanism.
Structured elaboration
Roles: product owner (owns the business-criticality tier assignment and reporting to leadership), SRE (owns technical SLO measurement and the enforcement mechanism), engineering manager (owns the team's actual remediation work and roadmap trade-offs when budgets burn). Required dashboards: a standard, company-wide reliability dashboard template (current SLO status, burn rate, and tier per service) that every product line adopts, so an executive or auditor can compare services across lines without each team inventing its own bespoke reporting format. Standardized process elements: an explicit SLO-proposal APPROVAL workflow (who can propose a new SLO or change an existing one, and who signs off), a defined escalation and remediation-window procedure for repeated breaches (not just a one-off incident response, but what happens after the SECOND or THIRD consecutive miss), and tiered budget ALLOCATION across services of varying criticality: a concrete 3-tier example assigns critical services a strict policy triggering action at 50% budget consumed, important services at 75%, and best-effort services only at full exhaustion, with corresponding automated or manual actions at each threshold per tier. For a large-scale system (e.g. 50 microservices), budgets need explicit shared-vs-component allocation logic (does a cascading failure across several services draw from one shared pool or each service's own budget) and periodic audits to confirm the allocation still reflects actual current criticality, not a stale assessment from when the service first launched.
Worked example
For a 40-team organization with legal SLAs carrying real financial penalties: the governance model requires any SLO the model designates as SLA-relevant to go through a stricter approval path involving legal, not just engineering sign-off, with a monthly reporting cadence to a cross-functional steering committee (not just the platform team), and an escalation path where two consecutive missed SLOs trigger a mandatory remediation-window review with a defined completion deadline, escalating to executive visibility if the remediation window itself is also missed. Framed as an SLO-driven-development adoption initiative, the model explicitly ties reliability targets back to architecture decisions: a service repeatedly missing its tier's SLO despite remediation attempts should trigger an architecture review, not just another remediation cycle, since repeated tactical fixes failing to hold the target is itself evidence of a structural, not tactical, problem.
Trade-offs and pitfalls
The audit-without-slowing-delivery balance is the hardest part of this design: an overly heavy-handed compliance process (mandatory sign-offs for every minor SLO adjustment) will be quietly worked around by teams under delivery pressure, while an overly light one produces exactly the governance theater problem seen elsewhere, where the framework exists on paper but has no real teeth; the right calibration usually reserves the heaviest process (legal sign-off, executive escalation) for genuinely high-stakes changes (legally-binding SLAs, critical-tier services) while keeping best-effort-tier governance lightweight, so the total compliance burden scales with actual risk rather than being uniform across every tier.
Implement a simplified SLO alert evaluator in Python: given a stream or array of sliding-window error rates (one value per minute) and an SLO threshold, trigger an alert if the error rate exceeds the threshold for N consecutive windows. Define function signature, edge-case handling (missing data), and provide example input/output.
Sample Answer
A consecutive-window alert evaluator is a small state machine over a stream of per-minute error rates, and the interesting design decision is what happens to that state when a minute's data is simply missing.
Structured elaboration
The core logic tracks a running streak of consecutive breaching minutes and fires once the streak reaches N; a non-breaching minute resets the streak to zero, which is the ordinary case. The less obvious case is a MISSING minute (no data point at all, as opposed to a zero error rate): treating it as "good" would let a real ongoing breach silently survive a monitoring gap and then resume counting where it left off once data returns, understating how long the breach has actually persisted; treating it as neutral-but-reset (this implementation's choice) is the conservative middle ground, since claiming the streak continued through a data gap you can't observe would itself be an unverifiable assertion.
Worked example (executed; python3, verified against 5 cases including missing data)
from typing import List, Optional
def evaluate_slo_alert(error_rates: List[Optional[float]], threshold: float, n_consecutive: int) -> bool:
streak = 0
for v in error_rates:
if v is None:
streak = 0
continue
if v > threshold:
streak += 1
if streak >= n_consecutive:
return True
else:
streak = 0
return False
Confirmed: [0.01, 0.02, 0.06, 0.07, 0.08] with threshold 0.05 and N=3 correctly returns True (three consecutive breaching minutes at the end); [0.06, 0.07, 0.01, 0.08, 0.09] returns False because the 0.01 minute in the middle resets the streak before it reaches 3; [0.06, None, 0.07, 0.08] returns False because the missing minute resets the streak, so only two consecutive breaching minutes remain after it; an empty input correctly returns False.
Trade-offs and pitfalls
This function only ever looks BACKWARD at a fixed streak; it does not distinguish "three minutes just over threshold" from "three minutes wildly over threshold," which is exactly the gap a burn-rate-based alert (rather than a fixed error-rate threshold) is designed to close. It is also inherently a synchronous, single-pass evaluator: a real streaming deployment needs to decide how state persists across process restarts (an in-memory streak resets to zero on every deploy of the alerting service itself, which could silently delay a real detection right when a release is happening).
You want to detect regressions from a canary using statistical hypothesis testing rather than fixed thresholds. Describe a test design to detect a significant change in latency or error rate during canary rollout that balances false positives and time-to-detection. Include choice of test, sample sizes, significance level, and how you would action on a positive result.
Sample Answer
Fixed-threshold canary checks either miss a real but modest regression or false-alarm on normal noise; a proper statistical test explicitly balances those two error types and gives you a principled way to decide how much data you need before acting.
Structured elaboration
Test design: a two-sample statistical test comparing the canary's observed metric (error rate, or latency) against the stable control's observed metric over the same time window, using a test appropriate to the metric's actual distribution: a two-proportion z-test (or Fisher's exact test at low sample sizes) for a rate metric like error rate, and a test appropriate for skewed distributions (e.g. a Mann-Whitney U test, since latency data is rarely normally distributed) rather than a naive t-test for a continuous metric like latency. Sample size and significance level: choose a significance level (e.g. alpha = 0.01, a stricter-than-typical bar, since a false positive here triggers an unnecessary rollback with real cost) and compute the minimum sample size needed to detect a MEANINGFUL effect size (e.g. a 20% relative increase in error rate) at that significance level with reasonable statistical power (e.g. 80%), rather than testing continuously on tiny early samples that have no real power to detect anything reliably yet.
Worked example
Balancing false positives against time-to-detection: at a stricter significance level (lower alpha), you need MORE samples before the test can confidently flag a real regression, which means slower detection; a looser significance level detects faster but risks more false-positive rollbacks. The practical compromise many teams use is a SEQUENTIAL testing approach (checking at multiple points as data accumulates, with a significance level adjusted for the repeated testing to avoid inflating the overall false-positive rate from checking too many times), rather than either a single test at a fixed, possibly-too-early sample size or waiting for one enormous sample before ever checking at all.
Trade-offs and pitfalls
The single most common statistical mistake here is repeatedly peeking at a test's p-value as data accumulates and stopping as soon as it first crosses the significance threshold, without correcting for the fact that repeated peeking inflates the TRUE false-positive rate well above the nominal alpha (a well-known problem sometimes called "peeking" or "optional stopping"); a genuinely valid sequential design (like a group sequential test, or a fixed pre-registered set of peek points with an adjusted significance level at each) is needed to avoid this, not an ad hoc "keep checking until it looks significant" approach. Actioning a positive result: a confirmed statistically significant regression should trigger an automatic rollback, but the specific effect size detected (not just "significant," which says nothing about MAGNITUDE) should inform how urgently to act, since a statistically significant but practically tiny effect size may not warrant the same urgency as a large one, even at the same significance level.
Unlock Full Question Bank
Get access to all SLIs, SLOs, SLAs, and Error Budgets interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.