A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.

How to answer

The interviewer is testing three things: that memory stays bounded however fast the stream arrives, that you chose what happens when the buffer is full, and that the consumer stops cleanly. Say all three before writing code.

  1. Ask the question that decides the policy. Can the source wait? A pull-based source such as a log or a paginated API can: if you stop reading, nothing is lost, so the full-buffer policy is to block, and that is . A push source such as a webhook burst or a sensor feed can’t, so you shed load, drop and count it, or spill to disk. Also ask about ordering, and whether an item may be processed twice.
  2. State the memory bound out loud. A bounded queue, a reader that waits when the queue is full, and a fixed number of workers. The most items held at once is capacity + workers + 1: the buffer, one per worker and the one the reader is holding.
  3. Isolate failures per item. A handler error goes to a dead-letter path and is counted; the worker keeps going.
  4. Design shutdown as its own step. Stop reading, drain what is buffered, give up after a deadline, and report what was abandoned. Commit only the highest offset below which every item has finished, not each item’s offset as it finishes, because with several workers a later item can finish before an earlier one. Then anything abandoned is redelivered rather than lost.
  5. Test the bound, not the speed. Block the handler, then assert how many items were pulled from the source.

The trap is the unbounded buffer in disguise: a queue with no maximum, or a task spawned per message. Both pass a demo and hold the whole backlog in memory the first time the downstream slows. Say which one you are avoiding before the interviewer asks.

GlossaryBackpressureA signal from a slow consumer that makes a producer slow down instead of overflowing memory or queues.More on Backpressure

Follow-ups

What the interviewer may ask next, once your first answer is on the table.

  • Order must be kept per customer ID, but you have four workers. What changes?
  • The source is a Kafka topic and your handler stalls for ten minutes. What happens to the consumer group?
  • When do you commit offsets, so that a crash neither loses items nor repeats more than you can tolerate?
  • How would you test that memory stays bounded without relying on timing?

Where answers go wrong

  • An asyncio.Queue() with no maxsize, or a new task per message, which is an unbounded buffer under another name.
  • A shutdown that either drops what is buffered or waits forever on a handler that hangs.
  • Letting one bad record raise out of a worker, so the pool shrinks by one with every failure until nothing is consuming.

Answer this in two minutes

Write the answer you would say out loud. The clock starts with your first word.

Two minutes

Model answer

“Before I write it: is the source pull-based, can it wait, does order matter, and can an item be processed twice? You’ve said it’s a pull-based topic, at-least-once is fine and order doesn’t matter.