Launch criteria are the conditions, agreed in writing with the customer before anyone sees evaluation results, that decide whether a system goes live: the metric for each task, the threshold each must meet, the evaluation set they are measured on, and who signs off. They include the limits the system must not break, such as an error rate on the riskiest segment, as well as the targets it must hit. Change them only in writing, and before the next run.
In FDE interviews
Among the job’s responsibilities, OpenAI’s (Healthcare) posting lists defining and operationalizing evaluations and launch criteria against customer-specific acceptance thresholds. Source 1Forward Deployed Engineer (FDE), Healthcare - SFPublisherOpenAI (Ashby)Source typecompany job posting In a decomposition or stakeholder scenario, a strong candidate proposes the criteria before the architecture:
“Before we build, let’s agree with your head of claims what good enough means: for example recall >= 0.95 on the fraud segment at a flag rate your reviewers can clear, with no increase in handling time, measured on a golden set we both trust. Let’s score your current process on that set first, so we know what we are beating, and agree now what happens if we miss: a narrower launch with review on every flag, or another iteration.”
That turns a later argument about whether the pilot worked into a check against numbers both sides signed.
Related: guardrail metric, regression suite, canary release.