A golden set is a fixed, versioned collection of inputs with expert-agreed expected outputs or labels, sampled to cover the segments and failure cases that matter, and used to score every version of a system the same way. It stays out of prompt examples and training data, and you keep a dev split to iterate on and a held-out split you only score; tune against the whole set and a better score may mean you fitted the test, not that the system improved. New production failures go in as a new version of the set, so scores are only ever compared on the same version.
In FDE interviews
Snowflake’s Senior/Staff Applied AI posting makes the hire responsible for turning customer goals into quality metrics, evaluation frameworks and golden datasets. Source 1Senior/Staff Forward Deployed Engineer, Applied AI @ SnowflakePublisherSnowflake (Ashby job board)Source typecompany job posting In a case, it comes up when the customer wants proof the system works and has no labels. A strong candidate says how the labels will get made: write labeling guidelines with the customer’s domain experts first, check that they agree with each other (inter-rater agreement), stratify the sample by segment and risk so rare but costly cases are represented, and size the set for the decision it has to support. At a pass rate near 0.90, n=100 gives a 95% confidence interval of about ±0.06 and n=400 about ±0.03, so a set of 100 cannot reliably tell a 0.90 system from a 0.93 one. Count per segment too: a segment with 20 cases tells you almost nothing on its own.
Related: regression suite, LLM-as-judge, launch criteria.