Performance Engineering & Capacity Planning Interview Questions | JiQuest

add

#

Performance Engineering & Capacity Planning

Advanced module · performance engineering

Performance engineering and capacity planning for backend systems.

The senior-level layer that sits above "write correct code": profiling methodology that separates measuring from guessing, JVM heap and GC tuning across G1, ZGC, and Shenandoah, why JIT warm-up affects your load tests and your rolling deploys, and the queueing-theory math (Little's Law) that explains why latency falls off a cliff near saturation instead of degrading gently.

3Collectors compared
1Core formula: L = λW
20Interview Q&A
Profilemeasure, don't guess Tuneheap, GC, JIT warm-up ModelL = λW, utilization Capacity planpeak + headroom, not averagesized from p99, not mean each stage feeds the next -- skipping profiling means tuning blind

Why this sits above ordinary feature work

Most backend interview prep stops at "does the code work" and "is the design correct." Performance engineering is the layer above that: given a design that's already correct, why is it slow, how do you find out without guessing, and how do you plan capacity so it doesn't become slow under load you haven't seen yet. It gets asked heavily at 5-20 year experience levels precisely because it separates people who've shipped features from people who've owned a service's SLA in production.

Profiling methodologySampling vs instrumentation, CPU vs allocation profiling, and why a laptop benchmark misleads.
JVM heap & GC tuningThe generational hypothesis, and when G1, ZGC, or Shenandoah is the right default.
JIT warm-upWhy a fresh instance is measurably slower, and what that means for load tests and rolling deploys.
Queueing theoryLittle's Law and why latency degrades non-linearly, not gradually, near saturation.
Capacity planningTranslating a growth number into infrastructure sized from p99, with real headroom.
Caching & pool sizingCache-aside vs read-through vs write-behind, and why a bigger connection pool can be slower.

Jump to a section

Profiling methodology: measuring instead of guessing

Every experienced engineer has a war story about a confident guess ("it's probably the JSON serialization") that turned out to be wrong, and a profiler that found the real answer ("it's a connection pool wait") in minutes. The discipline isn't about tools, it's about refusing to change code based on intuition when you could instead measure.

Sampling profilers vs instrumentation profilers

An instrumentation profiler injects timing code at every method's entry and exit, which produces exact call counts and exact per-call durations -- genuinely useful for understanding call graphs precisely. The cost is that the injected code changes the program it's measuring: overhead of 10-100x is common, and worse, it can change which code path looks hot, because a cheap method surrounded by heavy instrumentation overhead now looks expensive relative to code that wasn't instrumented as densely.

A sampling profiler instead periodically interrupts running threads and records their stack trace -- typically hundreds of times per second -- and reconstructs a statistical picture of where time goes from those samples. Overhead is usually under 1-2%, low enough to run continuously in production. The trade-off is statistical noise on infrequently-hit code paths, and (in older tools) safepoint bias, where sampling could only trigger at JVM safepoints, systematically under-representing code that doesn't reach one often. Modern tools like async-profiler solve this with AsyncGetCallTrace and hardware performance counters instead of safepoint-gated sampling, which is why it's become close to the default choice for JVM CPU profiling.

DimensionSampling profilerInstrumentation profiler
OverheadLow (often <2%) -- safe in productionHigh (10-100x) -- rarely safe in production
AccuracyStatistical, converges with more samplesExact call counts and per-call timing
Observer effectMinimal -- measures real behaviorSignificant -- can change which path is hottest
Best forProduction profiling, finding hot paths under real loadPrecise call-graph analysis in a controlled, non-prod environment

CPU profiling vs memory/allocation profiling

CPU profiling answers "where do cycles go." Allocation profiling answers a genuinely different question: "what's creating GC pressure." A method can be cheap per call in CPU terms and still be the dominant cost in the system if it allocates heavily -- those allocations don't show up as CPU time attributed to that method, they show up later as GC pause time attributed to nothing in a naive CPU-only profile. JDK Flight Recorder's allocation events (or async-profiler's -e alloc mode) surface this directly, and the two profiles frequently point at different methods entirely -- both are worth running, not just one.

Why production-representative load, not a laptop microbenchmark A microbenchmark runs a tight, already-warmed-up loop on different hardware (core count, cache sizes, NUMA topology) than production, under GC settings that may not match, and critically without production's concurrent noisy-neighbor load, connection pool contention, or realistic request mix. Any one of those can change which code path actually dominates. A correct JMH benchmark answers "how fast is this method in isolation" -- a legitimately useful but different question from "how does this service behave under real production concurrency," and conflating the two is one of the most common performance-tuning mistakes.

JVM heap and the generational hypothesis

Almost every modern garbage collector's design rests on one empirical observation, borne out repeatedly across decades of real programs: most objects die young. A request-scoped DTO, a temporary string built for logging, an iterator -- these are created, used briefly, and become garbage within microseconds to milliseconds. Long-lived objects (caches, connection pool entries, singleton services) are comparatively rare.

This is why the heap is split into a young generation (further split into Eden and two Survivor spaces) and an old generation. A minor GC only scans the young generation, and because most objects there are already dead, it only has to copy the small surviving fraction to a Survivor space (or promote it to old gen after enough survived collections) -- scanning mostly-dead memory is fast, which is why minor GCs are typically sub-millisecond to a few milliseconds even though they're stop-the-world. An object that survives enough young collections gets promoted to old gen, which is collected far less often because it's expected to hold mostly long-lived, legitimately live data.

The practical consequence: if your young generation is undersized for your allocation rate, objects get promoted to old gen before they'd naturally die (a failure mode called premature promotion), which both increases minor GC frequency (more objects means less room, forcing more frequent minor collections) and pollutes old gen with garbage that now has to wait for a much more expensive major collection to clean up. Sizing the young generation generously enough for your actual allocation rate and object lifetimes is usually the single highest-leverage GC tuning move, well before reaching for a different collector.

G1 vs ZGC vs Shenandoah: picking a default

All three are "low-pause" collectors relative to the old Parallel/CMS world, but they don't target the same pause budget, and picking the wrong one is either wasted complexity or a missed SLA.

CollectorPause targetHow it gets thereRight default when
G1 (default since JDK 9)Tens of milliseconds, soft target, scales with heap sizeRegion-based, mostly concurrent marking, evacuates a subset of regions per pauseThe large majority of services -- no special tuning needed, well-understood behavior, good throughput
ZGCSub-millisecond, largely independent of heap sizeColored pointers and load barriers let marking and relocation happen almost entirely concurrently with the appHard single-digit-ms p99 SLA, or heaps large enough (tens to hundreds of GB) that G1's pauses start showing up in the SLA
ShenandoahSub-millisecond, largely independent of heap sizeConcurrent compaction using a Brooks-pointer forwarding scheme, similar goals to ZGC via a different mechanismSame profile as ZGC -- often a build/platform-availability choice between the two rather than a behavioral one
Typical worst-case pause vs heap size (illustrative, not a benchmark) 0ms 200ms 4GB 32GB 128GB G1 (grows with heap) ZGC (flat, sub-ms) Shenandoah (flat, sub-ms)
What a 200ms pause costs a synchronous web service Every request in flight when a stop-the-world pause begins is frozen for its full duration, regardless of whether it had just started or was about to finish. If your p99 SLA is 150ms, one 200ms pause guarantees every request caught in it breaches SLA at the same moment -- a correlated failure, not a smoothly added average. That's a materially different problem than "average latency went up by a couple of milliseconds," and it's why pause time is a tail-latency and correlated-failure concern first.
Tuning ZGC/Shenandoah as a default is usually a mistake Concurrent marking and relocation aren't free -- they consume CPU and some throughput on every request, all the time, to buy a pause guarantee you only need if you actually have a tight tail-latency SLA or a huge heap. Reaching for them by default trades real, constant cost for a benefit you may never collect. G1 first; switch only with evidence.

JIT compilation and warm-up

The JVM doesn't compile your code to native machine code up front -- it interprets bytecode first, and only compiles the methods that prove themselves hot. That design choice has direct, testable consequences for load testing and deployment.

Why the first minutes are slower

HotSpot uses tiered compilation: every method starts running in the interpreter (correct, but slow), gets promoted to C1 (the client compiler, fast to compile, moderately optimized) once it's been invoked enough times, and eventually to C2 (the server compiler, slower to compile but far more aggressively optimized) once it's proven consistently hot. This promotion is triggered by invocation counts, not wall-clock time -- so a freshly started instance under real traffic genuinely is running much of its code path at interpreter or C1 speed for a real window (often tens of seconds to a few minutes under moderate load) until enough invocations accumulate for C2 to kick in.

What it means for load testing

A load test that starts recording results the instant traffic starts is measuring cold-JVM performance and reporting it as steady-state capacity -- systematically pessimistic, and a common cause of a service being under-provisioned (or worse, rejected in a benchmark bake-off) based on numbers that don't represent how it runs 99% of the time. The fix is straightforward: run a warm-up phase of representative load first (watch JIT compilation log activity taper off, or just use a fixed multi-minute ramp), and only start recording after that.

What it means for rolling deploys

A rolling deploy replaces warmed-up instances with fresh ones, and if the load balancer routes full production traffic to a brand-new instance immediately, that instance is measurably slower right when it's least equipped to prove itself -- which can trip latency-based health checks, false-positive an autoscaler into thinking capacity is short, or in bad cases fail its own readiness probe due to self-inflicted latency. The standard mitigations are a readiness gate that holds new instances out of the load balancer during an explicit warm-up traffic ramp, or a gradual traffic-percentage ramp-up instead of flipping an instance from 0% to 100% of its share at once.

A cheap mitigation worth knowing by name Application Class-Data Sharing (AppCDS) and, on newer JDKs, Project Leyden's ahead-of-time work aim to reduce class-loading and (increasingly) warm-up cost across restarts. They don't eliminate JIT warm-up entirely, but they're worth knowing as the direction the JVM itself is moving on this problem, not just an application-level workaround.

Little's Law: the formula behind queueing intuition

Little's Law is disarmingly simple and explains an enormous amount of production behavior once it clicks: L = λW -- the average number of requests in a system (L) equals the average arrival rate (λ) multiplied by the average time each request spends in the system (W). It holds for any stable system, regardless of arrival distribution or service-time distribution, which is what makes it so broadly useful.

λ = 500 req/sec arriving System (service + queue) W = 100ms average time-in-system requests complete & leave L = λ × W = 500 × 0.1 = 50 in system

That "50 in system" number is the whole point: it's not throughput and it's not latency, it's concurrency -- the number of requests simultaneously holding a thread, a connection, or memory at any instant. Capacity planning that only looks at λ (requests/sec) without deriving L is planning for throughput while silently ignoring the resource that throughput actually consumes. If your thread pool is sized at 40 and Little's Law says you need 50 in flight at normal load, you will queue at the thread pool under completely normal traffic, not just during a spike.

Little's Law also explains a debugging pattern worth recognizing on sight: a service doing 8ms of real work per request but showing a 400ms p99. If every request genuinely took ~8ms, p99 would sit close to the mean. A 50x gap between mean work and tail latency is the signature of queueing -- W in Little's Law includes time spent waiting, not just time spent being served, so a handful of requests queued behind a momentarily saturated resource inflate W (and therefore the percentiles) dramatically without changing the average amount of "real work" at all.

Why latency degrades non-linearly near saturation

This is the part of queueing theory that actually changes how you provision things. For a simple queueing model (M/M/1, a reasonable first approximation for a single saturating resource), the average wait time scales with ρ/(1-ρ), where ρ (rho) is utilization as a fraction of capacity. That denominator is the whole story.

utilization (ρ) latency 0% 50% 80% 99% 1x baseline ~4x ~99x same fixed headroom lost near the top costs vastly more latency

Concretely, using ρ/(1-ρ) as a relative multiplier: at 50% utilization the multiplier is 1x (baseline). At 80% it's 4x. At 95% it's 19x. At 99% it's 99x. The curve is flat and forgiving for most of its range and then turns nearly vertical near the top -- which is exactly why a system that looks perfectly healthy at 70% average utilization can fall over during a 30% traffic spike that pushes it to 91%. You didn't move a proportional amount along a line; you crossed the knee of a curve that gets steep fast.

The practical implication "We're at 70% CPU, we have headroom" is true and also can be dangerously misleading, because the next 20 percentage points of utilization cost far more latency than the last 20 did. Headroom needs to be evaluated against where you are on this curve, not as a flat percentage.

Capacity planning: from a growth number to infrastructure

"We're expecting 40% more traffic next quarter" is a business number. Turning it into "we need N more instances/cores/connections" requires a few deliberate steps, and skipping any of them is how teams end up provisioning for the wrong thing.

Why average load under-provisions

Traffic is diurnal and bursty -- by definition, you're above the daily average roughly half the time, and real-world spikes (a marketing push, a retry storm, a batch job kicking off) push well above that. Combined with the non-linear latency curve above, provisioning to comfortably handle the average is provisioning to be in trouble a meaningful fraction of the time; the only real question is how often and how badly. Capacity planning has to target a peak -- typically a recent observed p95 or p99 of actual traffic, or a load-tested ceiling -- not the mean.

The translation steps

1. Get a real growth numberFrom the business: "40% more orders/day by Q3" -- translate to a peak requests/sec, not an average, using your current peak-to-average ratio.
2. Apply Little's LawNew peak λ × current W gives the concurrency (L) you need to support -- this is what actually sizes thread pools and connection pools.
3. Load test to the new peakConfirm W doesn't itself degrade at the new concurrency (it usually does somewhat) -- feed the measured W back into step 2.
4. Add headroom above that peak30-50% is a common starting range for services with reasonably fast autoscaling and moderately spiky traffic -- more for flash-sale-shaped traffic or slow-to-scale infrastructure.
Headroom isn't a universal constant The right number depends on how spiky your real traffic is, how fast your autoscaling reacts relative to how fast load can grow, and how expensive an overload incident is versus how expensive the extra capacity is. A steady B2B API and a flash-sale retail endpoint should not use the same headroom target, even at identical average load.

Caching strategies for performance

All three patterns below cache the same kind of data -- the difference is where the miss-handling and write-propagation logic lives, and that placement decision has real latency and consistency consequences.

PatternHow it worksTrade-off
Cache-aside (lazy loading)Application checks cache first; on a miss, reads the source itself and populates the cache.Simple, works with any cache/source pairing, application stays in full control -- but every miss is a round trip the caller explicitly pays for and must handle.
Read-throughApplication only ever talks to the cache; the cache itself fetches from the source on a miss.Simplifies the caller, but couples the cache layer to the data source's access pattern and failure modes.
Write-behind (write-back)A write is acknowledged immediately against the cache, then flushed to the source asynchronously.The only pattern of the three that meaningfully reduces write latency -- at real risk of losing the most recent writes if the cache node dies before the flush completes.
Picking one isn't about which is "best" Cache-aside is the safe default for read-heavy data where you can tolerate a cache-miss round trip. Read-through earns its coupling when you want cache-population logic centralized rather than duplicated across callers. Write-behind is reserved specifically for cases where write latency matters enough to justify accepting some durability risk -- a metrics counter, not a financial ledger.

Connection pool sizing: why bigger can be slower

The intuitive assumption is that more connections mean more concurrency mean more throughput. That's true only up to the backing resource's real concurrency ceiling -- past it, more connections actively make things worse, and this is one of the highest-value things to actually understand rather than memorize.

A database like PostgreSQL has a real limit on how many queries it can usefully execute at once, set by its own CPU core count and internal contention (row locks, buffer pool latches, WAL writer contention). HikariCP's own sizing guidance cites a widely-used formula for this ceiling: connections = ((core_count * 2) + effective_spindle_count), which lands around 10 for a typical modern database server. A pool of 100 against that database doesn't get you 10x the throughput of a pool of 10 -- it gets you less, because the database is now spending real cycles on context-switching between ten times as many concurrent transactions and fighting over the same locks, rather than making progress on any of them.

Reconciling this with Little's Law Little's Law describes a relationship, L = λW -- it doesn't say increasing concurrency (L) always increases throughput. It only helps while the added concurrency is queued waiting for genuinely free capacity. Past the backing resource's real ceiling, W (time per unit of work) starts increasing faster than L is increasing, because of the contention effects above, so effective throughput (L divided by W) can actually fall even as you add more connections.

The practical sizing approach: size the pool to roughly the backing resource's real concurrency ceiling (the HikariCP formula above is a reasonable starting point for Postgres/MySQL), and let any additional concurrency queue at the pool level instead of at the database. Queuing at the pool is cheap -- a request waits briefly for a connection to free up. Queuing (via contention) at the database is expensive -- every connection, including the ones that were making progress, gets slower.

Key design decisions and interview talking points

A profiler says a method is "hot" -- what does that actually mean, and can it lie to you?

It means the profiler observed that thread's stack pointing into that method on a large share of its sampling ticks, which correlates with (but isn't identically equal to) CPU time spent there. A sampling profiler can lie in a specific, well-known way: safepoint bias. Older JVM sampling only captures a stack at JVM safepoints, so code that never naturally reaches a safepoint (tight loops without allocation or method calls) can be systematically under-sampled while code near safepoints is over-sampled. Async-profiler fixes this by using AsyncGetCallTrace and hardware perf events instead of safepoint-gated sampling, which is why it's the standard tool now instead of hprof or JFR's older stack-walking path.

Why not just run the JVM with default settings and call it done -- when does GC tuning actually start paying for itself?

G1 with defaults is a genuinely good starting point and for most CRUD services with heaps under a few GB and no tight latency SLA, tuning it further is wasted engineering time. It starts paying off once you have evidence, not suspicion: a service with a p99 latency SLA under roughly 200ms where GC log analysis shows pause times regularly eating 10-20% of that budget, or a heap large enough (16GB+) that G1's pause target becomes hard to hit without help. Tuning before you have that evidence is solving a problem you don't have yet at the cost of one you're creating -- config complexity nobody remembers the reasoning for in two years.

Why does the generational hypothesis matter for how you'd size heap regions, not just as trivia?

Because most objects die young, a young-gen collection can afford to be a stop-the-world copying collector that only scans live (surviving) objects -- and since the vast majority die, that scan is cheap and fast, which is exactly why minor GCs are sub-millisecond to a few milliseconds even though they're stop-the-world. If the hypothesis were false -- if objects survived at random -- young-gen collection would have to scan almost everything, and the entire generational design would stop making sense. Sizing follows from this: an undersized young gen forces objects to be promoted to old gen before they'd naturally die, which is the single most common self-inflicted GC tuning mistake.

A service does 8ms of average work per request but has a 400ms p99 -- where do you even start?

That gap is the signature of queueing, not of any single slow operation -- if every request individually took roughly 8ms, the p99 would sit close to the average, not 50x above it. Start by checking utilization against Little's Law's implication: as a resource (thread pool, DB connection pool, CPU) approaches saturation, queueing time explodes non-linearly, so a handful of requests waiting behind a momentarily saturated resource can produce exactly this shape. Check thread pool queue depth and connection pool wait time under load before touching the 8ms of actual work -- optimizing the fast path further won't move a p99 that's dominated by waiting, not working.

Why does Little's Law (L = λW) matter for capacity planning specifically, in concrete terms?

It gives you the relationship between arrival rate, time-in-system, and how many requests are in flight at once -- and "how many in flight at once" is what actually consumes threads, connections, and memory, not the arrival rate alone. At 500 requests/sec with a 100ms average time-in-system, L = 500 × 0.1 = 50 requests in the system on average, which tells you directly that a thread pool sized at 40 will queue under normal load, not just under a spike. Capacity planning that only looks at requests/sec without applying L=λW to get concurrency is planning for throughput while accidentally ignoring the resource that throughput actually consumes.

Why does latency degrade non-linearly as utilization approaches 100%, instead of just gradually getting worse?

Queueing theory's waiting-time formulas (for an M/M/1 queue, W_q proportional to ρ/(1-ρ) where ρ is utilization) have a (1-ρ) term in the denominator, which approaches zero as ρ approaches 1 -- so wait time approaches infinity, not just a bigger number. At 50% utilization the queueing multiplier is 1x; at 80% it's 4x; at 95% it's 19x; at 99% it's 99x -- the same fixed amount of headroom lost near the top costs vastly more latency than losing it near the bottom. This is precisely why a system that looks fine at 70% average utilization can fall over during a 30% traffic spike that pushes it to 91%: you crossed a knee in a curve, not a straight line.

Why is planning capacity from average load wrong, even if the average is accurate?

Average load hides the peaks that actually determine whether requests queue, and because of the non-linear latency-vs-utilization relationship, a system provisioned to handle the average spends real time above it -- diurnal traffic patterns alone mean you're above the daily average roughly half the time by definition. Providing for the average is providing to be overloaded half the time; the only question is how badly. Capacity planning has to target a peak (typically a recent p95 or p99 of actual traffic, or a load-tested ceiling) with headroom above that peak, not the mean.

What headroom percentage above peak is "enough," and where does that number come from?

There's no universal constant -- it comes from three inputs: how spiky your traffic actually is (a flash-sale retail workload needs more headroom than a steady B2B API), how fast your autoscaling can react relative to how fast load can grow, and how expensive an overload incident is versus how expensive the extra capacity is. A common starting range is 30-50% headroom above the highest expected peak for services that autoscale reasonably fast (a few minutes) with moderately spiky traffic, pushed higher for anything provisioned statically or serving traffic that can spike in seconds faster than autoscaling can respond.

Cache-aside, read-through, and write-behind all cache the same data -- why would you pick one over another?

Cache-aside puts the application in control of when to populate and invalidate the cache, which is simple and works with any cache and any data source, but every cache miss is a round trip the caller explicitly pays for and handles. Read-through moves the miss-fetch logic behind the cache itself, simplifying the caller at the cost of coupling the cache to the data source. Write-behind acknowledges a write immediately and flushes to the source asynchronously, which is the only one of the three that can meaningfully reduce write latency -- at the real risk of losing the most recent writes if the cache node dies before the flush happens, which is why it's reserved for data where that risk is acceptable.

How can a connection pool that's too large actually make a service slower, not just wasteful?

The backing resource -- Postgres, for instance -- has a real ceiling on how many connections it can usefully execute work on concurrently, governed by its own CPU core count and internal lock contention; connections beyond that ceiling don't add throughput, they add context-switching and lock-contention overhead that steals capacity from the connections that are doing useful work. HikariCP's own docs cite the PostgreSQL formula (connections = ((core_count * 2) + effective_spindle_count)) landing around 10 for a typical modern server -- a pool of 100 against that database doesn't process 10x the work, it processes less, because the database is now spending cycles fighting over the same table's row locks and internal latches across ten times as many concurrent transactions.

If Little's Law says more concurrency raises throughput, why would a smaller connection pool ever win?

Little's Law describes the relationship, it doesn't say increasing L always increases throughput -- it only helps as long as the extra concurrency is queued waiting for genuinely free capacity, not competing for capacity that's already saturated. Past the backing resource's real concurrency ceiling, W (time-in-system per unit of work) starts increasing faster than L is increasing, because of the same contention effects covered above, so effective throughput (L/W) can actually fall. The pool should be sized to just cover the backing resource's real concurrency ceiling, with the rest of the concurrency handled by queuing at the pool (which is cheap) rather than at the database (which is expensive).

G1, ZGC, and Shenandoah are all "low-pause" collectors -- what actually differentiates them for a decision?

G1 targets pauses in the tens-of-milliseconds range and scales that target reasonably well up to maybe 16-32GB heaps; it's the right default for the large majority of services because it needs no special tuning and its pause behavior is well understood. ZGC and Shenandoah both target sub-millisecond pauses regardless of heap size (tested into the hundreds of GB) by doing almost all marking and compaction concurrently with the application, which is the right choice specifically when you have a hard, tight latency SLA (single-digit millisecond p99, or a heap large enough that G1's pauses start showing up in your SLA) -- not as a default, because that concurrency has real throughput and CPU cost you're paying on every request even when a pause was never going to hurt you.

What does a 200ms GC pause actually cost a synchronous request-handling web service, concretely?

Every request in flight when the stop-the-world pause begins is frozen for its full duration -- not delayed by 200ms in isolation, but stalled with everything else, so requests that were nearly done and requests that just started both lose the same 200ms. If your p99 SLA is 150ms, a single 200ms pause guarantees every request in flight during it breaches SLA simultaneously, which is why GC pause time isn't just "added average latency" (which a 200ms pause every 30 seconds might barely move) but a tail-latency and correlated-failure problem -- many requests fail together at the same moment, which is exactly the kind of correlated failure that trips alerting thresholds and cascades into retries.

Why would a JVM service be slower in its first minute under load than an hour later, with identical code and traffic?

The JVM starts every method executing in the interpreter, which is correct but slow, and only compiles a method to optimized native code (via C1 then C2, tiered compilation) after it's been invoked enough times to prove it's actually hot -- so the first thousands of invocations of any code path run interpreted or lightly optimized, a function of real invocation counts, not wall-clock time. A freshly started instance under real production load is, for a real window of time, running large portions of its code path at interpreter speed while the JIT profiles and compiles -- this is warm-up, and it's a measurable, reproducible effect, not folklore.

What goes wrong if you load-test a JVM service without accounting for warm-up?

A load test that starts hammering a freshly started instance immediately measures cold-JVM performance and reports it as the service's steady-state capacity, which is systematically pessimistic -- you'll under-provision or reject a perfectly capable service based on numbers that don't represent how it actually runs 99% of the time. The fix is a warm-up phase before measurement: run representative load against the instance for long enough that JIT compilation has stabilized (watch compilation log activity taper off, or just use a fixed multi-minute ramp) and only start recording results after that.

Why does JIT warm-up matter for how you handle a rolling deploy, not just for load testing?

A rolling deploy replaces warmed-up instances (running fully JIT-compiled code) with fresh ones (starting from the interpreter), and if your load balancer sends full production traffic to a brand-new instance immediately, that instance will be measurably slower and can trip latency-based health checks or autoscaling signals that assume steady-state performance -- occasionally even failing its own readiness check because of self-inflicted latency. The standard mitigations are a readiness gate that holds an instance out of rotation during an explicit warm-up traffic ramp, or gradually increasing the share of real traffic a new instance receives, rather than flipping it from zero to full share at once.

Sampling profiler vs instrumentation profiler -- what's the actual trade-off, not just the definition?

An instrumentation profiler injects timing code at every method entry/exit, which gives exact call counts and exact per-call timings but that injected code itself changes the program's performance characteristics -- it can slow execution by 10-100x and, worse, can change which code path is actually hot by making cheap methods look expensive relative to their surrounding code. A sampling profiler periodically captures stack traces (typically far less overhead, often under 2%) and reconstructs hot paths statistically, trading exact counts for a much smaller observer effect. For production profiling, sampling is almost always the right tool because you need the measurement to reflect real behavior, not the behavior of a program that's now running 50x slower.

Why would you profile allocation rate instead of (or alongside) CPU time?

CPU profiling shows where cycles go, but a method can be CPU-cheap per call and still be the dominant cost in the system if it allocates heavily, because those allocations become GC's problem later -- a high allocation rate directly drives more frequent young-gen collections, and GC time doesn't show up as CPU time attributed to the allocating method, it shows up as pause time attributed to nothing in a naive CPU profile. Allocation profiling (JFR's object allocation events, or async-profiler's alloc mode) answers a different question -- not "where does the CPU spend time" but "what's creating GC pressure" -- and the two often point at different methods.

Why is a laptop microbenchmark a bad predictor of production performance for a JVM service?

A microbenchmark typically runs a tight, warmed-up loop on hardware with different core counts, cache sizes, and NUMA topology than production, under GC settings that may not match, with none of production's concurrent noisy-neighbor load, connection pool contention, or realistic request mix -- any one of these can change which code path dominates. JMH exists specifically to control the JVM-level pitfalls (dead code elimination, insufficient warm-up iterations) but even a correct JMH benchmark answers "how fast is this method in isolation," not "how does this service behave under production concurrency and load shape" -- those are different questions with different answers.

If an interviewer asks how you'd approach a vague "the API got slow" ticket, what's the actual process?

Measure before touching anything: pull the latency percentiles (not just the average) over the relevant time window and correlate the onset with a deploy, a traffic change, or an infrastructure event, since "slow" has categorically different causes depending on whether it's every request slightly slower (systemic: GC, CPU contention, a dependency) or a subset dramatically slower (queueing at some saturated resource, or a specific slow query/endpoint). Then profile production-representative load with a low-overhead sampling profiler rather than guessing, check GC logs for pause frequency and duration, and check utilization of the suspected bottleneck resource against Little's Law's implication about non-linear degradation near saturation. Guessing and changing code before that measurement is how you "fix" the wrong thing and the ticket reopens next week.

What's the single most important mental model to carry out of this whole topic into an interview?

That performance work is a measurement discipline before it's a tuning discipline -- profile before you change code, model concurrency (Little's Law) before you size a pool or thread count, and plan capacity from tail behavior (p99, real peaks) rather than averages, because averages hide exactly the non-linear saturation effects that cause outages. Every specific technique in this guide (G1 vs ZGC, cache-aside vs write-behind, pool sizing) is a downstream consequence of that one discipline, not a separate set of facts to memorize independently.

How would you explain to a non-technical stakeholder why "just add more servers" doesn't always fix a slow API?

More servers add capacity for handling more concurrent requests, but if the slowness comes from a shared bottleneck every server depends on -- a database connection pool, a downstream service's own capacity ceiling, GC pauses inherent to each instance's own heap -- adding more servers can leave the real bottleneck untouched or, in the connection-pool case, make it measurably worse by adding more contention against the same shared resource. The useful question isn't "do we have enough servers," it's "what resource is actually saturated," and that requires the same profiling-before-guessing discipline as any other performance problem, just aimed at infrastructure instead of code.

Related guides

Update these hrefs to your published Blogger post URLs once each page is live.

No comments
Leave a Comment