What FDE coding rounds test sorted the rounds; this lesson builds the code a practical round, a take-home and a customer integration all lean on: a call to someone else’s API that fails, slows down or limits you. You will retry with backoff and , make the repeat safe, and pace it with a limiter so retries cannot flatten the dependency. On the job, that is the difference between a blip and a customer billed twice.
Vercel’s public fde-challenge-backend spec, checked by a black-box conformance suite, describes a service that forwards valid inventory events to a partner catalog API with the event ID as the , treats 429 and 503 as transient with at most three attempts in total, and keeps at most four catalog calls in flight; the file does not itself say it is an interview. Source 1HTTP contractPublisherVercel (vercel-solutions on GitHub)Source typecompany website Read “three attempts in total” as one try and two retries. That contract is this lesson in one paragraph: a key, a retry rule and a cap on calls in flight. Integrating with systems of record designed around it; here you write the code.
The running example is fictional: a water utility posts each meter’s monthly usage charge to its billing vendor with POST /usage-charges, not idempotent without a key. Before any code, get the contract. The first thing to say aloud:
Candidate: “Four questions for the vendor first. Does it take a key, and for how long? Does a 503 mean it did nothing, or might it have charged? Is the limit ours alone or shared with the utility’s other jobs? And who resolves an unknown charge?”
The answers: it takes a key and keeps it a day. The limit is ours alone: 20 requests a second with bursts of 5, and four attempts per charge, not Vercel’s three: read the number off each contract. Its slowest normal calls take 1.8 s. It answers HTTP 429 with Retry-After before doing any work, but its proxy sometimes sends HTTP 503 after it has charged. Operations resolves unknown charges.
When to retry and when not to
First, the decision. Ask two questions: did the server act, and is a repeat safe?
| What you saw | Did it act? | Retry? |
|---|---|---|
| Connect error, 408, 429 | No | Yes, after the wait |
| Read timeout, reset, 500, 502, 503, 504 | Maybe | Only if safe to repeat |
| 400, 401, 403, 404, 422 | No, it rejected it | Never; fix the request |
In short: turned away before the work, retry after the wait; may have run, retry only if a repeat is safe; rejected as wrong, never retry.
One exception: a 401 from an expired access token. Refresh once and resend; a second 401 is final.
HTTP’s own spec defines PUT, DELETE and safe methods such as GET as idempotent (RFC 9110); POST is not, nor is PATCH (RFC 5789). So repeat a POST only under an idempotency key, or after checking the first did not land. The row that catches people is 503: the spec says the server is currently unable to handle the request, and nothing about whether any work ran. No spec promises that a 429 ran nothing either; that row rests on the vendor’s contract. The retry jitter question builds the full classifier.
Candidate: “This POST creates a charge, so a read timeout is an unknown, not a failure. I’ll retry it only because it carries a key.”