A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.
How to answer
Start from what can go wrong for the customer, not from a list of metrics. A support fails in a few specific ways: it gives a wrong answer, takes a wrong action (a refund it should not issue), misses a handoff to a human, or leaks another customer’s data. The harness exists to catch those before a change ships, so build it in this order and say each step out loud.
- Clarify what a change is. A prompt edit, a model swap, a new tool, a retrieval change. All of them go through the same gate.
- The . Real conversations, anonymized, each with an expected outcome written by the support team. Say who labels, and how disagreements are settled.
- Simulated users. Multi-turn cases need a scripted customer with a goal and hidden facts, and tools that run against fixtures, not production.
- Graders, cheapest first. Deterministic checks on tool calls and policy, then a model-graded rubric that you have checked against human labels.
- Gates. Must-pass cases block on any failure; the aggregate score must not drop beyond run-to-run noise.
- Cost and time. A fast tier on every change, the full suite before release, and caching.
Close by naming the owner: who can override a failed gate, and where that decision is written down.
The trap is answering with metrics (helpfulness, faithfulness, latency) and a dashboard. Without a labeled set, a blocking gate and a budget, nothing stops a bad change, and you have no answer for what happens when the score goes up and a refund case breaks.
Golden sets, graders and release gates are taught end to end in Evaluation: proving it works, and multi-turn agents in The agent design round; both modules are in Pro.
Follow-ups
What the interviewer may ask next, once your first answer is on the table.
- What is in the golden set, and who labels it?
- The suite takes two hours. How do you gate every change?
- A change improves the score but breaks one high-risk case. Ship it?
Where answers go wrong
- Lists metrics without a golden set, gates or cost control.
- Gates on the average score, so a change that breaks a must-pass case ships because the number went up.
- Trusts a model grader that was never checked against human labels.
Answer this in two minutes
Write the answer you would say out loud. The clock starts with your first word.
Compare with the model answer
Model answer
“Before any metrics, I want to list what this can get wrong for a customer: a wrong answer, a wrong action, a missed handoff, or leaked data. The harness exists to stop those, so I’ll build it around them.”
Clarify. I’d ask what the agent can do (answer from the help center, look up orders, issue refunds up to a limit, hand off to a human) and what has gone wrong before. The actions set the risk tiers, so I’d design around them.
. Before I pull a single transcript, I get the customer’s data owner to approve using production conversations for evaluation, and the golden set lives in their tenancy with the same retention rules as the transcripts. Then I pull real conversations, replace with consistent fake values that match the fixture records (order 88213 in the transcript is order 88213 in the sandbox), start with about 200 cases, and have two senior support agents label each with the expected outcome: resolved, handed off, or refused, plus the required tool calls. Where they disagree, the support lead decides, and the rule she used goes into the labeling guide. The set is stratified: common intents, policy edges (a refund just over the limit), handoff triggers, in a customer message, and cross-account requests. High-risk cases are tagged must_pass. Every production incident adds a case.
Simulated users. A case is a script, not a single prompt:
id: refund-over-limit-017
persona: >-
Polite, repeats the request,
mentions a chargeback on turn 3
goal: Get a full refund for order 88213
hidden_facts:
order_total: 412.00
delivered: true
fixtures: orders/88213.json
expect:
final_state: # sandbox after the run
refunds_issued: []
ticket:
status: escalated
tools_called: [lookup_order]
tools_forbidden: [issue_refund]
outcome: handoff
must_pass: true
A model plays the customer from the persona. Tools run against a sandbox loaded from fixtures, so a run is repeatable and never touches production.
Graders. Deterministic checks first. I diff the sandbox’s end state against the expected state, which is how tau-bench grades its tasks: it compares the database at the end of the conversation with the annotated goal state. An end-state check survives the agent paraphrasing or calling tools in a different order, where a check on the exact tool sequence would not. Then forbidden tool calls, the refund amount against the policy limit, and whether a handoff fired. Then a model grader with a written rubric for tone and correctness. Before I trust it, I measure its agreement with the human labels on a held-out slice and re-check whenever the grader’s model changes.
Gates.
from dataclasses import dataclass
from statistics import mean
@dataclass
class Case:
id: str
stratum: str # intent, language or risk tier
must_pass: bool
runs: list[bool] # n=3 runs of the same case
@property
def passed(self) -> bool:
return all(self.runs) # pass^k: every run must pass
@dataclass
class Decision:
allow: bool
reason: str = ""
def score(results: dict[str, Case], ids: list[str]) -> float:
return mean(results[i].passed for i in ids)
def gate(baseline, candidate, noise_band) -> Decision:
broken = [c.id for c in candidate.values() if c.must_pass and not c.passed]
if broken:
return Decision(False, f"must_pass failed: {broken}")
for stratum in sorted({c.stratum for c in candidate.values()}):
# the candidate ran a sample; score the baseline on the same case IDs
ids = [i for i, c in candidate.items() if c.stratum == stratum]
delta = score(candidate, ids) - score(baseline, ids)
if delta < -noise_band(stratum, len(ids)):
return Decision(False, f"{stratum} dropped {delta:.2f}")
return Decision(True)
ok, flaky = [True] * 3, [True, False, True]
baseline = {f"order-{i}": Case(f"order-{i}", "orders", False, ok) for i in range(10)}
baseline["refund-017"] = Case("refund-017", "refunds", True, ok)
candidate = {**baseline}
for i in range(3):
candidate[f"order-{i}"] = Case(f"order-{i}", "orders", False, flaky)
band = lambda stratum, n: 0.2 # 2 sd of repeated baseline runs at this n
print(gate(baseline, candidate, band))
# Decision(allow=False, reason='orders dropped -0.30')
An average that holds steady can hide one intent falling apart, so the gate is per stratum, on the same cases on both sides. On a pull request the candidate runs a sample, so the baseline score is taken over those same case IDs, not over the full suite.
The agent is nondeterministic, so each case runs n=3 times and a must-pass case must pass all of them; that is the pass^k measure from the same tau-bench paper, and it is much stricter than passing once in three runs. The noise band for each stratum comes from running the unchanged baseline repeatedly, and it has to be measured at the pull request’s sample size, not the full suite’s: a stratum scored on 10 cases swings far more than one scored on 100, so a band taken from the full suite flags sampled runs as regressions when nothing has changed. With a dozen strata, some stratum is likely to cross a tight band by chance, so I set the band at about two standard deviations and rerun a blocked stratum once before the block stands.
A must-pass case that flakes, passing on one run and failing on the next with no change, goes to quarantine with a named owner and a fix-by date. It is never silently skipped, and while it sits in quarantine the release needs that owner’s sign-off.
Cost and time. If the full suite takes hours, every pull request runs all must_pass cases plus a stratified sample, in parallel, with results cached by a hash of the agent prompt, agent model, tool versions, retrieval index version, simulator model and persona prompt, grader model and rubric, and the case. Leave any of those out and a grader change passes against stale results. Once incidents have grown the set, the full suite is about 400 cases × 3 runs × ~8 turns, and the simulator and grader each add a call per turn: roughly 29k model calls, which is why it runs nightly and blocks the release. A run that hits its budget fails loudly instead of truncating. The cheaper model grades everything; cases it scores near the threshold, plus a random 10%, go to the stronger grader, and I track how often the two disagree.
First slice. In week one: 30 must-pass cases taken from last quarter’s escalations and refunds, deterministic checks only (end state, forbidden tools, refund limit, handoff fired), run in CI on every pull request and blocking merge. The simulated users, the model grader and the nightly full suite come after, in that order, each added once the one before it has caught something.
The follow-up. If a change raises the score but breaks one high-risk case, I don’t ship it. Must-pass cases are gates, not part of the average. Either the change is fixed, or the case owner signs off that the expectation was wrong and updates the label, on the record.