Bound the buffer between producer and consumer; when it fills, the producer blocks, or the system drops, samples or spills to disk by a policy someone chose. In Python, await queue.put(item) on an asyncio.Queue(maxsize=1000) waits while the queue is full (Python docs, asyncio queues); in Node.js, writable.write() returns false once the buffer passes its highWaterMark, and the producer must stop until the stream emits drain, which stream.pipeline() does for you (Node.js, Backpressuring in streams); and in Reactive Streams, the subscriber requests n items at a time.

The fde-challenge-backend spec in Vercel’s vercel-solutions GitHub organization (repository created in July 2026; the file does not say it is an interview) allows at most four catalog calls in flight at once. Source 1HTTP contractPublisherVercel (vercel-solutions on GitHub)Source typecompany website The bug to spot is await asyncio.gather(*(embed(d) for d in docs)) over a whole corpus: it starts every call at once against a rate-limited API. Use a fixed pool of workers reading a bounded queue, or a semaphore sized to the provider’s limit.

A pull-based consumer such as Kafka’s gets backpressure for free by polling only when it is ready, and consumer lag is the metric to alert on. A consumer that goes longer than max.poll.interval.ms between polls is removed from its group (Apache Kafka, consumer configs), so to slow down, call pause() on its partitions and keep calling poll() rather than stop polling. Decide what happens when the buffer is full before writing code: pick the policy by what the business can afford to lose (block for payments, sample for telemetry, spill for audit logs), and pass the signal upstream as HTTP 429 or HTTP 503 with Retry-After rather than letting an unbounded queue grow until the process runs out of memory. Show where it applies in the bounded-buffer consumer question.