Interview guide · API design & protocols
API design & protocols: the questions every backend interview eventually asks.
The Richardson Maturity Model and what "RESTful" actually means, resource naming and cursor-vs-offset pagination, three ways to version an API and the backward-compatibility rules that decide which changes are safe, idempotency keys, and when a real system reaches for REST, gRPC, or GraphQL — sometimes all three at once.
What "API design" interview questions are actually testing
Every backend role from a 2-year-in engineer to a 20-year staff architect gets asked some version of "design an API for X," and the question is rarely testing whether you can write GET /users/{id}. It's testing whether you understand that a wire contract is a promise to every client that has ever called it, and whether you can reason about what happens when that contract needs to change. This page works through the concrete mechanics: how "RESTful" is measured, how pagination breaks silently under real traffic, how to change an API without breaking the clients already using it, and when REST stops being the right protocol at all.
Jump to a section
The Richardson Maturity Model, and what "truly RESTful" means
Leonard Richardson's model breaks "REST" into four levels of maturity, and interviewers who ask "is this really REST?" are almost always probing whether you know the difference between Level 2, where nearly every API that calls itself RESTful actually lives, and Level 3, the hypermedia-driven design Roy Fielding's dissertation originally described. Knowing the levels turns a vague argument about terminology into a precise, defensible answer.
Reading the levels the way an interviewer wants
Level 0 is one endpoint and one verb doing dispatch by payload — SOAP-style RPC wearing an HTTP costume. Level 1 introduces resource URIs but keeps everything behind a single verb, usually POST for reads and writes alike. Level 2 is where the HTTP verbs and status codes finally carry meaning: GET is safe and cacheable, DELETE is idempotent, a 404 means "not found" instead of a 200 with an error field buried in the body. This is the level almost every API that calls itself "RESTful" actually operates at, and for most systems it's the right stopping point — it gets you cacheability, standard tooling, and predictable semantics without the added complexity of Level 3.
Level 3 adds hypermedia controls: the response itself tells the client what it can legally do next (_links.cancel, _links.pay) instead of the client hardcoding a URL template. In principle this decouples clients from URL structure entirely; in practice, most clients are written against a specific API version by a team that already knows the URL structure, so the payoff rarely justifies the extra payload size and server-side complexity of computing valid state transitions per resource. It survives in domains like payments and travel booking, where third-party clients genuinely benefit from being guided through a multi-step workflow they don't fully control.
Resource naming, URI design, and pagination
Naming conventions are the part of API design that looks like bikeshedding until an inconsistent API actually ships, at which point every inconsistency becomes a support ticket or a client bug. The conventions below aren't arbitrary style preference — each one closes a specific class of ambiguity.
Resource naming and URI conventions
/orders not /getOrders — the HTTP verb already says what to do; the noun says what to do it to./orders/42/items is fine; /customers/9/orders/42/items/3/reviews should usually become /reviews?item=3 once nesting stops mapping to a real containment relationship./orders/42 identifies a resource; ?status=shipped&sort=-created_at filters and shapes a collection of them./order-items not /orderItems or /Order_Items — consistency here is worth more than which convention you pick.Offset vs. cursor pagination
Offset pagination (?page=2&size=20, translating to LIMIT 20 OFFSET 20) is simple and lets a client jump straight to page 47 or show a total page count — but it has a correctness problem, not just a performance one. OFFSET counts rows from the start of the current sort order at query time. If a row is inserted or deleted between the client fetching page 1 and page 2, every row after that point shifts position by one: the client either sees a row twice (it moved from page 2 into page 1's range) or skips one entirely (it moved the other way). On any actively-written table — an order list, a social feed, an audit log — this isn't a rare edge case; it happens routinely under real concurrent traffic.
Cursor pagination fixes this by making each page request relative to a specific row instead of a row count: ?after=eyJpZCI6NDIsImNyZWF0ZWRBdCI6..., an opaque token encoding the sort key and id of the last row the client saw. The next page is "rows after this cursor," which is stable regardless of what got inserted or deleted elsewhere in the table, because it never re-derives a position by counting. The trade-off is real: you lose the ability to jump to an arbitrary page or show a total count cheaply, since neither is a well-defined concept relative to a moving cursor. For an admin screen with page numbers, offset is the right tool; for a live feed, an infinite-scroll list, or any endpoint returning data under concurrent writes, cursor pagination is worth the lost page-jumping.
// Cursor pagination response shape
GET /orders?limit=20&after=eyJpZCI6NDIsImNyZWF0ZWRBdCI6IjIwMjYtMDktMDEifQ==
{
"data": [ /* 20 orders, newest first */ ],
"page": {
"next_cursor": "eyJpZCI6MjIsImNyZWF0ZWRBdCI6IjIwMjYtMDgtMjgifQ==",
"has_more": true
}
}
// The cursor is opaque to the client -- base64 of {id, sortKeyValue} -- so the
// server is free to change its internal pagination implementation later
// without that ever being a breaking change to the API contract.
Filtering and sorting conventions
?status=shipped&customer_id=9 — exact-match filters as flat query params; reserve a dedicated /search endpoint for free-text or complex boolean queries rather than overloading GET params.?sort=-created_at,total — a leading - for descending is a common, compact convention; always document the default sort explicitly since an unspecified order is technically allowed to change between requests.?fields=id,status,total lets bandwidth-sensitive clients (mobile) avoid over-fetching without building a whole GraphQL layer for one use case.{data, page} shape; every error uses the same {error: {code, message}} shape — predictability across endpoints is worth more than any individual shape choice.API versioning strategies
Versioning only becomes a hard problem when a change is genuinely breaking. The three common mechanisms below all solve the same problem — letting old and new clients hit the same server without either breaking — but they disagree on where the version lives and how discoverable it is.
| Strategy | Strength | Weakness |
|---|---|---|
URI versioning (/v2/orders) | Visible in every log line and browser tab; trivially cacheable per version since the URI itself is the cache key. | Technically wrong per REST purists — the URI is supposed to identify a resource, not a protocol revision — and it invites a whole resource tree to fork per version. |
Header versioning (Api-Version: 2) | Keeps URIs stable and clean; natural for internal service-to-service calls between trusted clients. | Invisible to a browser, a shared link, or casual debugging; caching now has to vary by header, which many caches don't do by default. |
Content negotiation (Accept: vnd.api.v2+json) | The most HTTP-correct mechanism — version is genuinely a media-type concern. | Hardest to test with a browser or curl by hand; least tooling support; rarely worth the purity for the debugging cost. |
Backward compatibility: what's safe, what's breaking
"Is this change safe?" has a precise answer once you separate it by direction: a request-body change is judged by whether an old client's existing requests still validate; a response-body change is judged by whether an old client's existing parsing code still works. The two aren't symmetric — a change can be safe on one side and breaking on the other.
| Change | Request body | Response body |
|---|---|---|
| Add a new optional field | Safe | Safe |
| Add a new required field | Breaking | Safe (clients that ignore unknown fields are unaffected) |
| Remove a field | Safe (clients just stop sending it) | Breaking (a client may already depend on it) |
| Rename a field | Breaking | Breaking |
| Change a field's type | Breaking | Breaking, even if all values still fit the new type |
| Add a new enum value | Depends — breaking for strict server-side validation | Depends — breaking for a client with an exhaustive switch/enum |
| Make an existing field nullable | Usually safe if old clients never omitted it | Breaking for clients that assumed non-null |
| Add a new endpoint | Always safe | Always safe |
The type-change row surprises people the most: changing an order id from a number to a string is breaking even though "42" and 42 represent the same value, because statically typed clients — a generated OpenAPI or protobuf client, a Java DTO with a long field — deserialize by declared type, not by inspecting the value. The wire format's type is part of the contract independently of what values ever actually appear in it.
The safest posture for a response body is to design every consumer to ignore fields it doesn't recognize (this is the default in most JSON deserializers, but not all — strict schema validation or @JsonIgnoreProperties(ignoreUnknown = false)-style configurations opt out of it) and to never treat "field exists" as load-bearing without a documented guarantee. That single discipline is what makes additive-only evolution possible without a version bump at all.
Idempotency in API design
Idempotency answers one question: if the same request is executed more than once, does the server end up in the same state as if it ran exactly once? For a network call that can fail after the server did the work but before the client got the response — the classic ambiguous timeout — idempotency is what makes "just retry" a safe default instead of a risk.
Which HTTP methods are naturally idempotent
This isn't a theoretical distinction. A load balancer, an HTTP client library, or a proxy is allowed to automatically retry a GET or PUT on a connection failure without asking the calling code, because idempotency guarantees that's safe. The same infrastructure is never allowed to auto-retry a POST, because it has no way to know whether the first attempt already succeeded server-side. PATCH sits in between and depends entirely on how it's defined: a PATCH that sets a field to an absolute value is idempotent; a PATCH that means "increment the counter by 1" is not.
Idempotency keys for POST
Since POST can't be made idempotent by the method semantics alone, the standard fix is an application-level idempotency key: the client generates a unique key (typically a UUID) per logical operation and sends it as a header, and the server persists which keys it has already processed before doing any real work. This is exactly the pattern this site's payment and wallet mini project builds: Payment Service checks whether a payment with the given Idempotency-Key already exists under a unique database constraint before calling Wallet Service, so a client that times out and retries with the same key gets the original result back instead of triggering a second transfer.
@PostMapping("/payments")
public PaymentResponse createPayment(
@RequestHeader("Idempotency-Key") String idempotencyKey,
@RequestBody PaymentRequest request) {
// unique constraint on idempotency_key makes this check-and-insert race-safe
// even under two near-simultaneous retries of the same request
return paymentRepository.findByIdempotencyKey(idempotencyKey)
.map(existing -> toResponse(existing)) // already processed -- return the same result
.orElseGet(() -> processNewPayment(idempotencyKey, request));
}
How long to remember a key
Idempotency keys need a bounded retention window, not permanent storage — typically 24 hours is enough to cover realistic client retry logic (exponential backoff rarely retries a day later) while keeping the lookup table from growing unbounded. A scheduled job or TTL index expires old keys the same way a webhook receiver's deduplication store does.
REST vs. gRPC vs. GraphQL
These aren't competing answers to the same question — they optimize for different things, and a mature answer to "which would you use" is almost always "it depends which edge of the system we're talking about," not a single universal pick.
| REST | gRPC | GraphQL | |
|---|---|---|---|
| Wire format | JSON, text | Protobuf, binary | JSON, text |
| Transport | HTTP/1.1 (or 2) | HTTP/2, native streaming | HTTP/1.1 (or 2), one endpoint |
| Typing | Convention / OpenAPI (optional) | Enforced by .proto schema, compiled | Enforced by GraphQL schema |
| Caching | Native via HTTP semantics (GET + ETag) | Not designed for HTTP caching | Hard — single endpoint, POST-shaped queries |
| Over/under-fetching | Common — fixed response shape per endpoint | N/A — message shape is explicit and typed | Solved by design — client picks fields |
| Where it shines | Public APIs, simple CRUD, cache-heavy reads | Internal microservice-to-microservice calls | Aggregating many resources for varied clients |
The instinct to pick "the best one" misreads the trade-off. gRPC's protobuf schema and HTTP/2 multiplexing minimize latency and serialization cost for calls between services you control on both ends — exactly the setting where you can regenerate typed clients on every schema change and don't need browser support. GraphQL earns its complexity at a public or mobile edge where many different client shapes would otherwise mean either chronic over-fetching (send the whole order object when the UI needs three fields) or an explosion of purpose-built REST endpoints. REST wins when the API's read traffic dominates and HTTP-native caching, human-readability, and broad tooling compatibility matter more than either of the other two's advantages.
A system with meaningful scale commonly runs all three at once: gRPC between internal services for the low-latency call graph, REST for simple public integrations and webhooks, and a GraphQL layer at the edge specifically to let mobile and web clients aggregate order, user, and inventory data from several backing services in one round trip without those services being coupled to each other.
Webhook design: delivery, signatures, retries, replay protection
A webhook flips the usual request direction: your API becomes the client, and pushes an event to a URL the receiver registered. That inversion is what makes webhooks powerful — near-real-time notification with no polling waste — and also what makes them operationally harder than a normal endpoint: you can no longer assume the receiver is up, honest, or fast.
// Receiver side: verify signature over the raw body, then dedupe by event id
String rawBody = readRawBody(request); // must be the exact bytes, before any parsing
String expected = hmacSha256Hex(webhookSecret, rawBody);
if (!MessageDigest.isEqual(expected.getBytes(), request.getHeader("X-Signature").getBytes())) {
return status(401); // reject -- not from a holder of the shared secret
}
String eventId = request.getHeader("Event-Id");
if (processedEventStore.alreadySeen(eventId)) {
return status(200); // already handled -- ack without reprocessing
}
processedEventStore.markSeen(eventId, Duration.ofDays(7));
handle(parse(rawBody));
return status(200);
Key design decisions and interview talking points
Is an API that only uses POST and a single /api endpoint ever really "REST"?
No — that's Richardson Maturity Level 0, RPC tunneled over HTTP. It's a perfectly reasonable design, but calling it REST is where the term loses its meaning. The label only starts to matter in an interview when you can say which level you actually mean: resource URIs (Level 1), verbs plus status codes (Level 2), or hypermedia-driven (Level 3).
Why do almost no production APIs reach Richardson Level 3 (HATEOAS)?
Hypermedia links make the client generic — it follows _links instead of hardcoding URLs — but almost every real client (a mobile app, a frontend SPA) is written against a specific API version anyway, so the flexibility HATEOAS buys is rarely worth the extra payload size and server-side complexity of computing valid transitions per resource state. It survives mostly in payment and travel-booking-style domains where clients genuinely are decoupled from a single backend team's release cycle.
Why is offset pagination (LIMIT/OFFSET) dangerous under concurrent writes?
OFFSET counts rows from the start of the result set on every request; if a row is inserted or deleted between page 1 and page 2, every row after it shifts position, so the client either sees the same row twice or skips one entirely. It's not a rare edge case in a feed, an order list, or any actively-written table — it happens on essentially every page 2 request against a live table.
If cursor pagination is strictly more correct, why does offset pagination still show up everywhere?
Offset pagination gives you things cursors can't: jump straight to page 47, show a total page count, let a user click backward and forward freely. Cursor pagination only supports next/previous relative to where you are. For an admin table with filters and page numbers, offset is the right trade-off; for an infinite-scroll feed or any API returning live, frequently-mutated data, cursor pagination is worth the lost page-jumping.
How do you keep a cursor from leaking internal implementation details like primary key structure?
Treat the cursor as an opaque token from the client's point of view: base64-encode a small JSON or delimited payload of whatever the query actually needs (typically the sort key's value and the id as a tiebreaker), and never document its internal shape as part of the contract. That way you're free to change the underlying pagination implementation later without it being a breaking API change.
URI versioning, header versioning, or content-negotiation versioning — how do you actually choose?
URI versioning (/v2/orders) wins on visibility and cacheability — you can see the version in a log line, a browser tab, or a CDN cache key with zero extra tooling — which is why almost every public API uses it despite REST purists objecting that the URI is supposed to name a resource, not a protocol version. Header versioning keeps the URI clean and is common for internal APIs where clients are trusted services, not curl and browsers. Content-negotiation versioning is the most "correct" HTTP mechanism but is the hardest to debug and cache, which is exactly why it's the least used of the three in practice.
What does it mean to evolve an API "without versioning" at all?
It means committing to additive-only changes: new fields are always optional with sane defaults, new endpoints are added rather than existing ones repurposed, and enums are only ever extended, never removed from. Clients that ignore fields they don't recognize keep working forever. The moment you need to remove, rename, or change the meaning of something existing clients depend on, additive-only breaks down and you need an actual version boundary.
Is adding a new required field to a request body a safe change?
No — it's breaking, even though it only touches the request side. Every existing client that was built before the field existed will send a request without it and get rejected by validation. A new field is only additive if it's optional with a server-side default; making it required later is the same breaking change deferred, not avoided.
Is removing a field from a JSON response ever actually safe?
Only if you can prove no client reads it, which in practice you almost never can for a public API — clients parse fields you didn't expect them to, log them, or pass them through to another system. The safer sequence is: stop populating meaningful data in the field, watch usage metrics or wait out a documented deprecation window, and only physically remove it in a new major version once you're confident, or contractually allowed, to break old clients.
Why is changing a field's type from integer to string a breaking change even if every value still fits?
Statically typed clients (a Java DTO, a generated protobuf/OpenAPI client) deserialize by type, not by value — a field typed as long will fail to parse a JSON string even if that string is numeric. The values being representable doesn't matter; the wire format's declared type is part of the contract, and changing it breaks every client that generated code against the old type.
How exactly does an Idempotency-Key header prevent a duplicate POST from double-charging a customer?
The same pattern used in this site's payment/wallet mini project: before doing any work, the server checks whether a request with that key has already been processed, under a unique database constraint so a race between two near-simultaneous retries can't both slip through. If it's seen before, it returns the stored result immediately instead of re-executing the operation — so a client that times out and retries with the same key gets the same order or payment back, never a second one.
Which HTTP methods are naturally idempotent, and why does that matter beyond retries?
GET, PUT, DELETE, HEAD, OPTIONS are defined as idempotent — repeating the same request produces the same server state as doing it once. POST and (usually) PATCH are not. This matters beyond client retries: a load balancer, proxy, or HTTP client library is allowed to automatically retry an idempotent request on a network failure without asking the application, but is never allowed to auto-retry a POST for exactly this reason — which is precisely why POST is the method that needs an explicit idempotency key when retry-safety matters.
Is DELETE actually idempotent if the second call returns 404 instead of 204?
Yes — idempotency is about the resulting server state, not the response code. After the first DELETE, the resource is gone; after the second, it's still gone. The state converges to the same place either way, which is what the idempotency guarantee actually promises. Some APIs choose to keep returning 204 on a repeat DELETE specifically so retry logic doesn't have to special-case a 404 as a possible success, which is a defensible client-ergonomics choice layered on top of, not instead of, the underlying guarantee.
When would a real system use gRPC, REST, and GraphQL all at once?
A common shape: gRPC between internal microservices, where both ends are code you control and protobuf's strict schema plus HTTP/2 multiplexing and streaming minimize latency and serialization cost; then either REST or GraphQL at the public edge, where human-readable JSON, browser and curl compatibility, and HTTP caching matter more than shaving milliseconds off an internal call. GraphQL specifically earns its place at the edge when the public API genuinely serves many different client shapes (web, mobile, partner integrations) that would otherwise need several REST endpoints or chronic over-fetching.
What's the concrete cost of GraphQL's client-driven flexibility?
Two operational problems that REST mostly doesn't have: the N+1 problem, where a naive resolver fetches a parent list and then issues one query per child field instead of batching, and query-complexity attacks, where a client can nest fields deeply enough to make the server do enormous or even exponential work behind what looks like one small request. Both are solvable — DataLoader-style batching for the first, query depth/cost limits for the second — but they're solved per GraphQL server, not handed to you by the protocol the way REST's per-endpoint scoping mostly avoids them by construction.
Why is protobuf over HTTP/2 a better fit than JSON over HTTP/1.1 for internal service-to-service calls?
Protobuf is a binary, schema-defined format, so messages are smaller and parsing is faster than JSON's text parsing, and the .proto contract is generated into strongly-typed client and server code on both ends, catching mismatches at compile time instead of at runtime. HTTP/2 adds multiplexed streams over one connection and native bidirectional streaming, which matters for high-fan-out internal call graphs in a way it rarely does for a public API talking to a browser.
What's the trade-off REST gives up by being so cacheable, and why does that matter for a public API specifically?
REST's cacheability comes directly from HTTP semantics — a GET to a stable URI can sit behind a CDN or shared cache using standard Cache-Control/ETag headers with zero custom infrastructure. GraphQL's single endpoint and per-query-shaped POST requests defeat that out of the box; gRPC isn't meant to be browser- or CDN-cached at all. For a public API where read traffic dwarfs write traffic, that built-in cacheability is often worth more than either protocol's flexibility or performance advantage.
What's the difference between a webhook and just having clients poll an endpoint?
Polling means the client repeatedly asks "anything new?", trading latency for simplicity and wasting most requests on "no"; a webhook means the server pushes an event the moment it happens, trading that wasted polling traffic for the operational burden of reliable delivery — retries, signatures, and idempotent receivers — that a simple polling GET never has to solve. Webhooks earn their complexity when near-real-time notification actually matters to the receiver; for anything that just needs eventual freshness, polling is simpler and has no delivery-guarantee problem to solve at all.
How do you verify a webhook actually came from the sender and wasn't forged?
HMAC-sign the raw request body with a secret both sides share out of band, send the signature in a header (e.g. X-Signature), and have the receiver recompute the HMAC over the exact bytes it received and compare in constant time. Because the signature covers the payload itself, an attacker without the shared secret can't forge a valid event even if they know or guess the receiving URL — and because it's computed over the raw body, the receiver must verify before any JSON parsing or body transformation touches the bytes.
How do you protect a webhook receiver from processing the same event twice?
Give every event a unique, stable event id and have the receiver store processed ids (with a TTL) so a delivery it's already handled is a no-op on a repeat — the same idempotent-consumer discipline used for Kafka event handlers, applied to inbound HTTP instead of a broker. This covers both legitimate retries after an ambiguous timeout and a replayed capture of an old, valid payload, which is exactly the same threat model an Idempotency-Key solves on the inbound side of a POST.
If a webhook receiver is down for an hour, what happens to the events sent during that window?
A well-designed sender retries on any non-2xx response or timeout using exponential backoff over a bounded window (minutes, then tens of minutes, up to hours), not a fixed short interval that just hammers a struggling endpoint. After the retry budget is exhausted the event moves to a dead-letter queue or is flagged for manual/API-driven redelivery rather than being silently dropped — the sender's job is at-least-once delivery, and the receiver's job is making repeat delivery safe via idempotency.
If an interviewer asks for the single biggest mistake in real-world API design, what do you say?
Treating the response shape as free to change because "it's just JSON." A schema-less wire format makes it easy to forget that every field you ship becomes part of an implicit contract the moment a real client parses it — which is the same underlying issue behind broken pagination, breaking version bumps, and un-replay-safe webhooks throughout this whole topic. The fix isn't a specific protocol choice, it's discipline: decide up front what's additive-safe, document it, and enforce it in code review the same way you'd enforce a database migration policy.
Related guides
Update these hrefs to your published Blogger post URLs once each page is live.
Post a Comment
Add