A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.
How to answer
Say the rule in one sentence, then spend the time on edges.
- Pin the inputs.
{role, content}messages, a budget, and acount_tokensfunction. The budget is the context window minus room for the reply and tool definitions. Real counts come from the model’s tokenizer or the provider’s counting endpoint; a labeled estimate is fine here. - State the rule. The system message always stays. Walk back from the newest, adding whole units while they fit; stop at the first that doesn’t. Skipping a large unit to fit older ones leaves a hole the model can’t see.
- Drop by unit, not by message. A unit is a single message, or an assistant tool call with every result that answers it; APIs reject a result whose call is gone. Trim whole units, then advance the start to a user message, which some APIs require.
- Fail loudly on the two cases trimming cannot fix. If the system message alone is over budget, or the newest user message doesn’t fit beside it, raise an error naming the numbers. The caller decides to summarize, chunk or refuse; your function never cuts text.
- Trim in steps, not one message per turn. Dropping one message per turn changes the prefix each turn, so prompt caching misses on everything after the system prompt. Over budget, cut to a low-water mark such as
0.7 * budget, so the prefix holds for several turns. Cache each message’s count; offer summaries or pinned facts next. - Test the boundaries. An exact fit, one token over, a tool pair at the cut line, and output order.
The trap is trimming a tool-using history by message. The cut splits a tool call from its result and the API rejects it, but only on conversations long enough to trim, so short tests never catch it.
Follow-ups
What the interviewer may ask next, once your first answer is on the table.
- The newest user message alone is over the budget. What does your function do?
- Trimming dropped the turn where the user gave their account number. How would you keep facts like that?
- After you shipped trimming, the prompt-cache hit rate on long conversations fell to zero. Why, and what do you change?
- This runs on every turn of a long conversation. What do you cache?
Where answers go wrong
- Truncating the text of the oldest message to fill the last few tokens, which is exactly the half-message the prompt forbids.
- Skipping a large message to fit smaller, older ones behind it, which leaves a gap the model cannot see.
- Cutting between a tool call and its result, so the request is rejected only once conversations get long.
Answer this in two minutes
Write the answer you would say out loud. The clock starts with your first word.
Model answer
“I’ll take messages in the chat-completions shape: role and content, an assistant message may carry tool_calls with ids, and each tool result names its tool_call_id. The budget is the context window minus room for the reply and the tool definitions.