A regression suite is a fixed set of test cases, many of them past failures, rerun on every change to a prompt, model version, retrieval setting or code, to catch behavior that used to work and now does not. For model outputs the checks are graded (exact match, schema validity, a judge score) and compared with the last accepted baseline, not diffed byte for byte. It catches changes you make; drift, a shift in production inputs or outputs that nobody deployed, needs a monitor on live traffic as well.

In FDE interviews

Decagon’s Deployment Engineer posting expects the hire to design evaluation and regression-testing frameworks that validate agent behavior against ground truth and guard against performance drift. Source 1Agent Deployment Engineer @ Decagon (San Francisco)PublisherDecagon (Ashby job board)Source typecompany job posting Sierra’s Agent Development Life Cycle post (June 2024) says any annotated conversation can become a conversation test, simulated against mock APIs, and that these tests can be run during development and before every agent release as a regression suite. Source 2The Agent Development Life Cycle (Bret Taylor, Clay Bavor)PublisherSierraSource typecompany blog Mock APIs matter in customer work: run against them, a regression suite never touches the customer’s live systems.

Asked whether to ship a new model version, a strong candidate says every fixed bug becomes a case, a fast subset runs on each change and the full suite before release, and the gate blocks a release when any segment drops by more than its noise band, not only when the average does. Measure that band by rerunning the accepted baseline several times (k=5, say, at the production temperature) and taking the spread per segment, and pin the model snapshot so a provider update shows up as a change you test, not as drift.

Related: golden set, LLM-as-judge, canary release.

GlossaryAgentA system in which a model chooses steps and tool calls to complete a task, within limits the design sets.More on Agent