Observability deep dive
Observability at scale: logs, metrics, and traces beyond Micrometer basics.
Why the three pillars each cover a gap the others can't, how a trace id actually survives an HTTP call and a Kafka hop, why a careless metric label turns into an outage of its own, and the alerting and dashboard discipline that separates a system a handful of engineers can operate from one that pages everyone into burnout.
The three pillars — and what each one alone can't tell you
Every team past a handful of services already has Micrometer counters and a log aggregator. The interview-level question isn't "what are logs, metrics, and traces" — it's which one you reach for when the others structurally can't answer the question in front of you, and why operating dozens of services makes that choice matter in ways a single monolith never forced.
Take a concrete case: the p99 latency metric for a payment API jumps from 80ms to 900ms. The metric localizes when it happened and how severe it is — that's genuinely useful, and it's the only pillar cheap enough to watch continuously across every endpoint you own. But it cannot tell you whether the culprit is a downstream wallet-transfer call, lock contention in a ledger table, or an outbox relay stalling behind a backed-up Kafka broker. A metric alone leaves you guessing at the cause with no more evidence than "it got slower." That's the exact moment a trace earns its cost: pull a handful of traces sampled from the spike window and look at where the time actually landed in the call tree. The metric said something is wrong; the trace says what. And once a trace points at a specific downstream call, structured logs filtered to that trace's id fill in the request-level detail — a specific error code, a payload size, a retry count — that neither a span attribute nor an aggregate counter was ever designed to carry.
| Question you're asking | Best pillar | Why the others fall short |
|---|---|---|
| Is error rate or latency trending worse across the fleet? | Metrics | Traces and logs are per-request artifacts; aggregating them over millions of requests to answer a trend question is the expensive, wrong tool for a job metrics do natively. |
| Where did this one slow request actually spend its time? | Traces | A metric only reports aggregate duration; unstructured logs at best approximate ordering from timestamps across services, without a real causal link. |
| What exact input or error caused this one request to fail? | Logs | A trace shows a span failed, not necessarily the specific validation error, stack trace, or payload detail that explains it. |
| Did a change five minutes ago make things worse for everyone? | Metrics | You'd need to sample and eyeball many traces to notice a shift that a single time-series chart shows at a glance. |
Distributed tracing mechanics: spans, context propagation, and OpenTelemetry
A trace is only useful if the trace id actually survives every hop a request takes — and a request in a mature backend rarely stays inside one service or one protocol. It goes over HTTP between synchronous services, and it goes over a message broker for anything async. Each hop is a distinct propagation problem, and each one can silently break.
A span is one timed unit of work: a service and operation name, start and end timestamps, a set of attributes, a status, its own span id, the trace id of the request it belongs to, and — except for the root span — the span id of whatever caused it. A trace is nothing more than the set of spans sharing one trace id, reassembled into a tree purely from those parent-child links, with no span needing to know about any other directly. That reassembly is what makes a trace viewer able to show you a waterfall of exactly which service called which, and for how long, for one specific request.
The mechanism that keeps that tree connected across a network hop is context propagation. Consider the synchronous path this site's payment mini project describes: Payment Service calls Wallet Service over a direct HTTP call to execute a transfer. Before that call goes out, Payment Service's instrumentation injects the current trace id, its own span id, and a sampling flag into an outgoing traceparent header — the W3C Trace Context format that most tracing libraries, including Micrometer Tracing and the OpenTelemetry Java agent, generate and read automatically for common HTTP clients. Wallet Service's instrumentation extracts that same header on the way in, before it creates its own span, and parents that new span to the extracted span id instead of starting a fresh trace. Skip either half of that — no injection on the client side, or no extraction on the server side — and Wallet Service's work shows up in the trace backend as a completely disconnected, orphaned trace with no visible link back to the payment that caused it. From a debugging seat, that's indistinguishable from Wallet Service just not having been touched at all, which is a much worse failure mode than "the trace is incomplete" — it actively lies about what happened.
The Kafka hop is the harder version of the same problem. An HTTP call has a synchronous caller and callee that exist at the same instant, so injecting a header onto the outgoing request and extracting it on the incoming one is straightforward. A Kafka publish and its eventual consume are decoupled in time — the outbox relay might publish an event that a notification or reconciliation consumer doesn't process for seconds, or, if that consumer is down, hours. There's no live call to attach a header to on the way back. Instead, trace context has to be serialized directly into the Kafka record's own headers at publish time, and the consumer has to read those headers back out at poll time, before it creates its own span, parenting it to the producer's span exactly as the HTTP case does. Miss this on either side of the broker and the async fan-out — the exact shape this site's payment mini project uses for notification and reconciliation events — loses its causal link entirely, which is precisely what makes an async failure so much harder to root-cause at 2am than a synchronous one: the two halves of the story sit in two disconnected traces, and nothing tells you they're related.
OpenTelemetry is the vendor-neutral answer to needing this to work consistently across every service, language, and hop rather than reinventing propagation per team. It standardizes the instrumentation API applications code against, the data model for spans, metrics, and logs, the OTLP wire protocol for exporting telemetry, and a Collector that can receive, batch, and forward that data to any backend. The practical payoff is portability: instrument once against the OTel API and common auto-instrumentation libraries, and switching from, say, an open-source backend to a hosted APM vendor is an exporter configuration change in the Collector, not a re-instrumentation project across every service you own.
Metrics cardinality: the trap that turns a cheap counter into an incident
Metrics are supposed to be the cheap pillar. That's only true if the number of distinct time series a metric produces stays bounded — and it's astonishingly easy for one well-intentioned label to break that assumption.
Cardinality is the number of unique label-value combinations a metric name produces, and every combination is a distinct time series the backend has to allocate memory for, index, and store. A request counter tagged with method (a handful of values) and status_class (2xx/4xx/5xx) stays at a few dozen series — trivially cheap regardless of how many requests it counts. Tag that same counter with the raw request path instead of the matched route template, and every order lookup like /orders/9f2a1e40-... mints a brand-new series, because a UUID-bearing path never repeats. What should have been a few hundred series — one per route template times status class — becomes millions, each one a permanent memory allocation in an in-memory time-series store like Prometheus (which doesn't release a series' memory until well past its staleness window) or a separately billed custom-metric line item in a hosted backend. That's the whole failure in one sentence: a metric that was supposed to cost nothing per request now costs real memory and real money per unique value a label happens to take on, and neither the person adding the label nor the dashboard using it usually notices until the backend falls over or the bill arrives.
The damage doesn't scale evenly across label choices. A raw user id is worse than a raw path in practice, because teams tend to reach for it as a debugging shortcut on several metrics at once — request counters, latency histograms, error counters — multiplying the explosion across every metric it's added to instead of confining it to one. The fix isn't "be careful" as a policy; it's a concrete design rule applied at review time: a label's set of possible values should be small, known in advance, and controlled by your own code — a route template, an HTTP method, a status class, a service name, a region — never a value dictated by external, growing input like a user id, a free-text field, or a raw path. If a team genuinely needs per-entity granularity, that's what exemplars are for: a sampled trace id attached to a specific histogram bucket, letting a dashboard viewer click from an aggregate latency bucket into one representative trace, or a structured log query filtered by the entity id directly — the unbounded detail lives in the pillars designed to carry it, not in a metric label.
| Tempting label | Why it explodes | Safer alternative |
|---|---|---|
path="/orders/9f2a..." | Unbounded — a new value per entity id embedded in the URL. | route="/orders/{id}" — the matched template, bounded by how many routes exist. |
user_id="482913" | Unbounded and multiplies across every metric it's added to. | Drop from labels; correlate via trace exemplar or a log query instead. |
error_message="..." (free text) | Effectively unbounded — every stack trace variant is a new series. | error_class="TimeoutException" — a small, known set of categories. |
request_id="..." | Unique per request by definition — the worst possible label. | Never a label; this belongs on a log line and a span attribute only. |
Structured logging and the correlation ID that ties it all together
At one service, grepping a log file for a customer complaint works fine. At dozens of services each writing free-text lines with no shared shape and no shared identifier, the same technique stops working entirely — not because the information isn't there, but because there's no reliable way to find or join it.
An unstructured line like ERROR payment failed for user 4821 is only searchable by whatever exact substring someone thinks to grep for, it can't be reliably parsed into queryable fields by a log aggregator because a slightly different phrasing in another service breaks any regex built against this one, and — the part that actually matters at scale — it carries no explicit link to the request or trace that produced it. Reconstructing "everything that happened across every service for this one failed checkout" from lines like that means correlating by eyeballing timestamps and hoping no other request logged something similar in the same second window. Structured logging — fixed-shape JSON lines with a stable set of fields — turns log data from prose into queryable data: a field can be filtered, aggregated, and joined the same way a database column can.
{"ts":"2026-09-09T14:02:11.408Z","level":"ERROR","service":"payment-service",
"trace_id":"4bf92f3577b34da6a3ce929d0e0e4736","span_id":"00f067aa0ba902b7",
"order_id":"9f2a1e40-...","http_status":502,
"msg":"wallet transfer failed: downstream timeout after 3000ms"}
The single field that unlocks the cross-pillar workflow described earlier is the trace id — and ideally the span id — present on every log line for a request. With it, a query can pivot directly from a metric spike's time window, to the specific traces sampled in it, to every log line across every service those exact traces touched, without three separate ad-hoc searches stitched together by guesswork. This is also where a correlation id becomes a distinct, additional concept rather than a synonym for trace id: a trace id is scoped to one attempt, minted fresh every time a new root span starts, while a correlation id — often a client-supplied idempotency key or a request id assigned at the edge — can be designed to persist across retries. A client that times out and retries produces a second, distinct trace id for that second attempt, but a well-designed system carries the same correlation id across both, so an engineer can still answer "show me every attempt at this logical operation, including the ones that failed and got retried," a question trace ids alone can't answer once a retry has happened.
Alerting design: symptom-based vs cause-based, and why paging on everything backfires
Every team that operates more than a handful of services eventually collects dozens of things they technically could page on: CPU, GC pauses, disk queue depth, connection pool saturation, consumer lag, replica lag, thread pool rejection counts. The interview-level judgment isn't knowing these signals exist — it's knowing which of them should ever wake someone up.
A symptom-based alert fires on something a user actually experiences — checkout error rate crossing a threshold, p99 latency on a payment API breaching its SLO, a spike in legitimate traffic getting rejected by a rate limiter. It should page, because by definition it represents live or imminent user-facing harm, and someone needs to look now. A cause-based alert fires on an internal signal that might explain a symptom — a database connection pool near saturation, elevated GC pause time, a Kafka consumer group's lag climbing — and on its own says nothing about whether users are currently affected. Paging independently on every cause-based signal, regardless of whether it's translating into user-facing harm right now, is the single most common way teams manufacture their own alert fatigue.
The cost of over-paging isn't just annoyance — it's a direct erosion of the response you get when it actually matters. Every page that turns out not to matter trains an on-call engineer, consciously or not, to treat the next one with a little less urgency: acknowledge and go back to sleep, triage a bit more slowly, assume it's probably nothing again. The real outage, when it arrives, gets the same degraded reaction as the noise that preceded it, because there was never a reliable signal in the stream telling the engineer which pages were the ones that mattered. That's the actual mechanism behind alert fatigue, and it's why "just page on everything to be safe" is precisely backwards as a safety strategy.
The practical test is actionability: if nobody can or should do anything about a specific signal right now, it shouldn't interrupt a person. It belongs on a dashboard, in a ticket, or as a linked signal an on-call engineer checks after being paged by an actual symptom — not as an independent page. Elevated disk queue depth on one node, alone, usually fails that test at 3am: it's frequently self-correcting or masked by redundancy, and paging on it regardless of user impact just trains people to distrust pages. It becomes genuinely worth paging on once it's correlated with an actual symptom breach — which is exactly what a well-designed multi-window, multi-burn-rate SLO alert (the pattern popularized by Google's SRE practice) formalizes: page when the rate of error-budget consumption is both fast enough and sustained enough to represent a real, ongoing problem, not a one-minute blip.
Dashboards for different audiences aren't the same graphs resized
A common mistake once a team has good telemetry is building one dashboard and assuming it serves everyone who might look at it. An on-call engineer mid-incident, an executive checking reliability trends, and someone doing next quarter's capacity plan are asking fundamentally different questions of the same underlying data, and a dashboard optimized for one of them actively fails the other two.
The design goal, not the data source, is what should determine a dashboard's shape. An on-call engineer needs density: enough drill-down to go from a service-level anomaly to the specific instance, query, or downstream dependency responsible, with a trace and log link sitting right next to the graph rather than requiring a separate tool switch mid-incident. An executive asking about the payment platform doesn't want a per-pod CPU graph — that's noise relative to the question they're actually asking, which is closer to "is this system trending toward or away from a problem worth their attention," so the right artifact trades drill-down for a small number of trend lines over a much longer window, rendered simply enough to read without context. A capacity-planning dashboard is a third, different artifact again: it needs resource-utilization trends correlated against growth drivers like traffic volume, storage growth, or connection-pool saturation, at a granularity coarse enough (daily or weekly rollups) to be affordable to retain for years and to reveal a trend line without being swamped by second-scale noise that would actively hide the slope you're trying to forecast from.
| Audience | Primary question | Time window | Design goal |
|---|---|---|---|
| On-call engineer | What's broken right now, and where? | Minutes to hours | Speed to root cause — dense, drill-down, linked to traces/logs. |
| Engineering VP / exec | Are we reliable, and is it improving? | Weeks to quarters | Narrative clarity — a handful of trend lines, no per-instance noise. |
| Capacity planner | When do we need to provision more? | Months to years | Forecasting — smooth, trend-fitted rollups correlated to growth. |
Building a single "do everything" dashboard and resizing it per audience is a false economy: it forces every viewer to filter out the majority of what they're looking at to find the slice relevant to their question, and it means a change made to serve one audience — adding drill-down panels for on-call, say — degrades the experience for the others. Treat these as three separate artifacts sharing the same underlying telemetry, not three views of one graph.
Topics
Interview questions and answers
Each answer is aimed at the depth an interviewer probing past "define the three pillars" actually wants: a concrete scenario, the trade-off, and the production reasoning behind the choice.
The three pillars — what each is good at
1. You have a metric showing p99 latency doubled for Payment Service. Walk through what you'd actually do next, pillar by pillar.
The metric localizes when it happened and roughly how much things degraded. Pull traces sampled from the spike window next to see the actual call graph and where the time landed — Wallet Service's response, a database lock, or the Kafka outbox relay stalling. Then pivot to structured logs filtered by the trace ids of those specific slow traces to see request-level detail a span attribute alone might not carry, like payload size or a retry count. Each pillar narrows the search space for the next one.
2. If you could keep only one of the three pillars for debugging a single failed request, which would you pick, and why does that choice not generalize to a systemic issue?
A trace, since it reconstructs the causal path of that one request across every service boundary it touched, which is exactly what single-request debugging needs. But a trace says nothing about frequency or trend — whether this is one unlucky request or forty percent of traffic degrading — which requires the cheap aggregate view only metrics give you across millions of requests, something no per-request artifact can substitute for at volume.
3. Why can't a team just log "enough" instead of also building metrics and tracing?
Cost and query shape. A metric is pre-aggregated at write time, so a counter increment costs a few bytes and a bucket lookup, letting you hold years of trend data cheaply. Recovering "requests per second by status code over ninety days" from raw logs means storing and scanning every individual line, which is orders of magnitude more storage and compute for a question metrics answer almost instantly. Logs are the right tool for what exactly happened on one request, not the shape of traffic over time.
4. A teammate says "we have Micrometer, so we have observability." What's missing from that claim?
Micrometer gives you the metrics pillar, and increasingly a tracing bridge, but instrumenting one service's counters and timers doesn't give you cross-service causality — that needs propagated trace context through every hop, sync and async — structured and correlated logs across those same services, or the operational discipline layered on top: sane cardinality limits, symptom-based alerting, dashboards built for the audience using them. Observability at scale is the combination plus that discipline, not any single library's default setup.
Distributed tracing mechanics
5. Concretely, what is a span, and what does "parent span id" actually encode?
A span is one timed unit of work — a service name, operation name, start and end timestamps, a set of key-value attributes, and a status — tagged with its own span id, the trace id of the request it belongs to, and, except for the root span, the span id of the operation that caused it. The parent id is what lets a trace backend reconstruct the whole call tree purely from a flat set of span records, without any of them needing to know about the others directly.
6. Walk through exactly what has to happen for a trace to survive the synchronous HTTP call from Payment Service to Wallet Service in this site's payment mini project.
Payment Service's instrumentation starts a span for the outbound call and injects the current trace id, its own span id, and the sampling decision into an outgoing traceparent header on the Feign/HTTP request. Wallet Service's instrumentation extracts that header before creating its own span, sets the extracted span id as its new span's parent, and continues the same trace id rather than minting a new one. Skip either half — no injection on the client, or no extraction on the server — and Wallet Service's work shows up as a disconnected, orphaned trace with no visible link back to the payment that triggered it.
// outbound (Payment Service)
GET/POST ...
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
// inbound (Wallet Service) extracts traceparent
// before starting its own span, parented to 00f067aa0ba902b7
7. Why is the Kafka hop between Payment Service's outbox relay and a downstream notification consumer a fundamentally harder propagation problem than the HTTP call?
HTTP context propagation piggybacks on a request that has a synchronous caller and callee at that instant; a Kafka publish and the eventual consume can be seconds, minutes, or hours apart, with no direct connection to inject a header onto in real time. The context has to be serialized into the record's own headers at publish time by the producer and read back out by the consumer at poll time, entirely decoupled from any live connection. Get this wrong and the notification/reconciliation service's processing of that event becomes an unlinked, second trace — exactly the kind of gap that makes an async fan-out failure much harder to root-cause than a synchronous one.
8. What does OpenTelemetry actually standardize, and why does that matter beyond "it's popular"?
OTel standardizes the API surface applications instrument against, the data model for spans/metrics/logs, the wire protocol (OTLP) for exporting that data, and a vendor-neutral Collector that can receive, process, and forward it to any backend. The practical payoff is the same portability argument this site makes about swapping providers behind a stable abstraction: instrument once against the OTel API, and switching from one tracing backend to a hosted vendor is a Collector exporter configuration change, not a re-instrumentation project across every service.
9. What is tail-based sampling, and why would you want it in addition to, not instead of, head-based sampling?
Head-based sampling decides whether to keep a trace at the moment the root span starts, typically a flat percentage — cheap, but it can't know yet whether this particular request will turn out to be slow or fail, so a fixed 1% sample mostly captures boring, fast, successful traces. Tail-based sampling buffers spans until the trace completes and then decides based on the outcome — always keep errors and high-latency outliers, sample the fast/successful bulk at a low rate — which is what actually gets you the traces you need when debugging a p99 spike, at the cost of needing a collector that can buffer and correlate spans across the whole trace before deciding.
Metrics cardinality
10. Define cardinality explosion with a specific example, not the textbook one-liner.
Cardinality is the number of unique label-value combinations a metric produces, and each combination is a separate time series the backend has to store and index. Tag an HTTP request counter with the raw request path instead of the matched route template, and every order lookup like /orders/9f2a1e... or /orders/7c88b2... mints a brand-new series forever, because a UUID-bearing path never repeats — a metric that should have had a few hundred series now has millions, each consuming memory in an in-memory TSDB like Prometheus, or a separate billed custom-metric line item in a SaaS backend.
11. Why does a raw user id as a label do more damage than a raw URL path, even though both are "just one more tag"?
They're both unbounded-cardinality sources, but a user id typically appears across many metric names at once — request counters, latency histograms, error counters, business metrics — if a team is tempted to add it "for debuggability," multiplying the explosion across every metric it's added to rather than just one. It also tends to persist as active, queried series for the lifetime of that user's traffic rather than aging out quickly, so the damage compounds instead of settling.
12. If you genuinely need to see a metric broken down by something as specific as an individual order or user, what's the right way to get that instead of adding it as a label?
That per-entity granularity is what traces and logs are for, not metrics — attach an exemplar (a sampled trace id linked to a specific bucket of a metric's histogram, which Prometheus and OTel both support) so a dashboard viewer can click from an aggregate latency bucket straight into one or two representative traces for that bucket, or query structured logs filtered by the entity id directly. The metric stays cheap and bounded; the per-entity detail lives in the pillar designed to carry unbounded-cardinality detail.
13. How would you review a pull request that adds a new metric label, specifically for cardinality risk?
Ask what the realistic upper bound on distinct values is and whether that bound is fixed by your own design — a route template, a small enum of status classes, a bounded set of region codes — or dictated by external, growing input — a user id, a free-text field, a raw path, a request id. Anything in the second category should be rejected as a label outright regardless of how useful it sounds for one specific debugging session, because the cost is paid by every query and every byte of retention across the whole metric's lifetime, not just the one investigation that motivated the ask.
Structured logging & correlation IDs
14. Why is an unstructured log line like "ERROR payment failed for user 4821" close to useless once you're operating dozens of services?
It's only searchable by whatever exact substrings you think to grep for, it can't be reliably parsed into fields by an aggregator (a slightly different message format in another service breaks any regex built against this one), and it carries no explicit link to the trace or request that produced it — so finding everything related to this one failed request across service boundaries means guessing timestamps and hoping nothing else logged the word "payment" in the same second. At one service that's tolerable; correlating that across dozens of services during an incident is not.
15. What fields does a structured log line need at minimum to be useful at scale, and which one is the most important for cross-pillar correlation?
Timestamp, level, service name, and message are baseline, but the field that actually connects logging to the other two pillars is the trace id (and ideally span id) of the request being logged. With it present on every line, a log aggregator query can pivot directly from a metric spike's time window to the traces in it to every log line across every service those specific traces touched — without the trace id, you're back to correlating by eyeballing timestamps.
{"ts":"...","level":"ERROR","service":"payment-service",
"trace_id":"...","span_id":"...","order_id":"...","http_status":502,
"msg":"wallet transfer failed"}
16. Is a trace id the same thing as a correlation id? When would a team need both?
Not quite — a trace id is scoped to one attempt at a request, minted fresh each time; a correlation id, often the client's idempotency key or a request id assigned at the edge, can be designed to persist across retries, meaning a client retry after a timeout produces a second, distinct trace id but should carry the same correlation id as the first attempt. Teams that need to answer "show me every attempt at this logical operation, including the ones that timed out and got retried" need the correlation id in addition to trace ids; teams that only need "show me what happened during this one attempt" only need the trace id.
Alerting design
17. What's the practical difference between a symptom-based and a cause-based alert, using this site's payment or rate-limiter services as the example?
A symptom-based alert fires on something a user actually experiences — checkout error rate crossing a threshold, p99 latency on the payment API breaching its SLO, the rate limiter rejecting a spike of legitimate traffic — and should page, because it represents live or imminent user-facing harm. A cause-based alert fires on an internal signal that might explain a symptom — Wallet Service's connection pool near saturation, GC pause time elevated, a Kafka consumer group's lag climbing — and on its own says nothing about whether users are currently affected, so it shouldn't independently page; it's useful as a linked signal an on-call engineer checks after being paged by the symptom.
18. Why does paging on every cause-based anomaly actively make incident response worse, not just noisier?
Every page an on-call engineer receives that turns out not to matter trains them, consciously or not, to treat the next page with a little less urgency — acknowledge and go back to sleep, or triage more slowly — which means the one page that represents a real, user-facing outage gets the same degraded response as the false alarms that preceded it. Alert fatigue isn't just an annoyance metric; it's a direct erosion of your actual time-to-acknowledge during the incident that matters.
19. Give the test for whether an alert should page, and apply it to "disk queue depth is elevated on one node."
The test is actionability — if nobody can or should do anything about this specific signal right now, it shouldn't interrupt a person; it should show up on a dashboard, open a ticket, or feed as context into a symptom alert instead. Elevated disk queue depth on one node, alone, usually fails that test at 3am: it's often self-correcting or masked entirely by redundancy, and paging on it independent of whether it's translating into user-facing latency or errors trains on-call to distrust pages. It becomes actionable, and worth paging on, once it's correlated with an actual symptom breach — that combination is what a good multi-signal alert or a linked dashboard captures.
Dashboards for different audiences
20. An on-call engineer and an engineering VP both ask for "a dashboard for the payment system." Why shouldn't they get the same one, just resized?
They're optimizing for opposite things. The on-call engineer, mid-incident, needs second-to-minute granularity, drill-down from service-level to instance-level, and direct links out to traces and logs for the exact window they're looking at — density and speed-to-root-cause matter more than narrative. The VP needs a trend over weeks or months — availability percentage, SLA compliance, cost per transaction, incident count and severity — rendered simply enough to answer "are we reliable and is it improving" in ten seconds, where per-pod CPU graphs are noise, not signal. Building one dashboard for both means it drowns the VP in irrelevant detail or strips the on-call engineer of the drill-down they need under pressure.
21. What makes a capacity-planning dashboard a genuinely different artifact from an on-call dashboard, beyond just "longer time window"?
The design goal is forecasting, not triage, so it needs different granularity (daily or weekly rollups over months to years, not second-level data that would be both noisy and expensive to retain that long), it needs resource-utilization trends correlated against growth drivers (traffic volume, data volume, connection pool or partition saturation) rather than moment-to-moment error rates, and it's read on a planning cadence, not glanced at mid-incident — which means it can afford smoother, trend-fitted visualizations that would actively hide the sharp, second-scale spikes an on-call engineer needs to see.
Related interview guides
Update these hrefs to your published Blogger post URLs once each page is live.
Post a Comment
Add