In this post12 sections
  1. The five parts of a complete answer
  2. A weak answer and a strong one, word for word
  3. When the output is text, not a label
  4. Do FDE interviews test LLM evaluation?
  5. Building an eval set when the customer has no labels
  6. A short scoring script you can explain
  7. What you monitor after launch
  8. Follow-ups: LLM judges, drift and the cost of a wrong answer
  9. Practice it out loud
  10. Questions people ask
  11. Keep reading
  12. More from the blog

You built the demo. It answers the sample questions cleanly, the interviewer nods, and then asks: “How do you know it works?” The easy reply is to describe the demo again. This post shows you how to answer with evidence instead. It is one round of many; the FDE interview guide covers the rest.

The short version: a complete answer names five things. The metric that matches what the customer cares about, the data you measured it on, the baseline you compared against, the result, and what you monitor after launch. Then you say, without being asked, where the system still fails. “I tried a bunch of examples and they looked good” is not an answer; it is a description of a demo.

The five parts of a complete answer

This checklist is ours, not a rubric any employer publishes. It works because each part answers a doubt a careful listener has.

PartThe doubt it answers
Metric“Are you measuring what matters to the customer?”
Data“Is the test set real, and did you tune on it?”
Baseline“Better than what?”
Result“How good, and where does it fail?”
Monitoring“How would you know it broke next month?”

Metric. Start from the decision the system supports and the mistake that costs the most. For a support router, a ticket in the wrong queue might sit for a day. For a medical summary, a missed allergy costs far more than a clumsy sentence. Pick the metric that counts the expensive mistake, and say who agreed to it. The lesson on stakeholders and success metrics covers how to find the number the sponsor already trusts.

Data. Say where the test examples came from, how many there are, who wrote the correct answers, and that you did not tune the prompt on them. Real inputs beat invented ones.

Baseline. A number alone means nothing. Compare it to what the customer does today (a person, a keyword rule, the old system) or to the simplest version of your own system.

Result. Give the count as well as the rate, and break it down by category. “Missed 3 of 38 cancel emails” says more than a percentage, because the listener can see how small the category is.

Monitoring. Name what you watch in production, what would trigger an alarm, and what happens when it fires.

A weak answer and a strong one, word for word

The setup: in a live build, you made an LLM router that reads a subscription company’s support emails and sends each one to a queue: refund, cancel, billing or other. The interviewer asks how you know it works.

The weak answer.

“I tested it on a bunch of emails and it got almost all of them right. I also used GPT to grade the responses, and the prompt is pretty robust. In production I’d add logging and monitoring.”

What is wrong with it:

  • “A bunch” and “almost all” are not numbers. The listener cannot tell five emails from five hundred.
  • No baseline, so “almost all” could be worse than the keyword rules the customer already has.
  • The model grader is mentioned but not checked. Who checked the grader?
  • “Logging and monitoring” names a tool, not a signal or an action.
  • It says nothing about what fails. That reads as either not looking or hiding.

The strong answer. (The numbers are invented for the example.)

“The decision is where each email goes, and the costly mistake is a cancel request landing in billing, because the customer gets charged again. So I measured accuracy per queue, and separately recall on the cancel queue: how many real cancel emails ended up there.

“The data: 200 real emails from last month’s inbox, none of which I looked at while writing the prompt. The support lead labeled them, and she and I agreed on the rules for the edge cases before I scored anything.

“The baseline is their current keyword rules. On the same 200 emails, the rules routed 131 correctly and the LLM router 172. On cancel requests, the rules missed 14 of 38; the router missed 3.

“Where it fails: of the 28 it got wrong, 11 asked for two things at once, like a refund and a cancellation. I route those to a person for now.

“After launch, I’d log every routing decision with the email’s ID, have the support team re-queue anything misrouted, and treat each re-queue as a new labeled example. If the re-queue rate for cancel emails rises above what we agreed, the router falls back to the keyword rules while I find out why.”

Notice the order: decision, metric, data, baseline, result, failure, monitoring. It is short, it is specific, and each sentence closes one doubt. If the interviewer wants the vocabulary, it is all there: the cancel number is recall on one class, and the “which queue did it go to instead” pairs are the off-diagonal cells of a confusion matrix. Practice saying your own version against the question page for this exact prompt, which works through a customer asking it.

When the output is text, not a label

A router has one right answer per email. A summarizer, a assistant or a reply drafter does not, so “accuracy” has nothing to count. The fix is to turn “good” into checks that each pass or fail.

  1. Write the rubric as pass/fail checks. For a clinical note summary: “mentions every allergy in the note”, “cites a passage that exists”, “gives no advice outside policy”. Each check is a yes or no, not a score out of ten.
  2. Have people score a slice first. A domain expert grades a few dozen outputs against each check. That is your ground truth for the checks.
  3. Add an LLM judge per check, only after it agrees with them. One narrow question per judge call. If it disagrees with the expert too often on a check, a person keeps grading that check.
  4. Report failures per check. “Missed an allergy in 4 of 120 summaries” tells the customer what to worry about. An average score does not.

A strong spoken version, for a summarizer:

“I turned ‘a good summary’ into four pass/fail checks with the charge nurse, starting with ‘every allergy in the note appears in the summary’. She graded 60 summaries, and the LLM judge for each check matched her on all but a few before I let it grade the rest. Across 400 summaries, the allergy check failed 6 times, all on notes where the allergy was in a scanned attachment, so those notes now go to a person.”

Do FDE interviews test LLM evaluation?

The job descriptions ask for it, which is the best reason to prepare.

  • OpenAI’s (Healthcare) posting lists defining evaluations and that measure quality against customer-specific acceptance thresholds. Source 1Forward Deployed Engineer (FDE), Healthcare - SFPublisherOpenAI (Ashby)Source typecompany job posting
  • Snowflake’s Senior/Staff Applied AI FDE posting asks the hire to turn customer goals into quality metrics, evaluation frameworks and golden datasets, then run eval loops on quality. Source 2Senior/Staff Forward Deployed Engineer, Applied AI @ SnowflakePublisherSnowflake (Ashby job board)Source typecompany job posting
  • Scale AI’s Frontier Agents FDE postings list evaluation harnesses built from offline benchmarks, online experiments, golden datasets, regression suites and LLM-as-a-Judge. Source 3Frontier Agents Engineer (Forward Deployed Engineering)PublisherScale AI (Greenhouse)Source typecompany job posting

Postings describe the job, not the interview. The interview evidence is thinner, and it comes from candidates. One candidate reported on Blind, in August 2026, that their interviews for an L5 FDE role in Google’s Cloud AI division included a system design interview specifically for agents, covering evals and guardrails. Source 4Google FDE vs Remain Amazon SDEPublisherBlind (teamblind.com)Source typecandidate report on BlindSource 5Google L5 FDE vs Remain Amazon L5 SDEPublisherBlind (teamblind.com)Source typecandidate report on Blind One candidate reported, in a public GitHub repo in August 2026, what it describes as Hippocratic AI’s take-home for an AI Agent Deployment Engineer role: turn a simple bedtime-story request into a story for ages 5 to 10, and add an LLM judge. Source 6Hippocratic AI Coding AssignmentPublisherakshay-menta (GitHub)Source typecandidate’s take-home repository

One prep site calls it “the differentiator question” for OpenAI and cites no source. Source 7Forward Deployed Engineer Interview: The Definitive 2026 GuidePublisherExponent (Aced)Source typeinterview prep site We found no employer that publishes how this answer is scored. Prepare it because the job asks for it, not because of a rumor about its weight.

Building an eval set when the customer has no labels

The follow-up to a strong answer is often “the customer had no labeled data, so where did your correct answers come from?” Here is a sequence you can describe in under a minute.

An eval set from nothing

  • Pull real inputs from logs, tickets or documents. Invented examples miss the messy cases.
  • Sample to match the real mix, then add the hard cases on purpose: ambiguous requests, missing data, questions the system should refuse.
  • Have a domain expert on the customer’s side write the correct answers, not you.
  • Have two people label an overlapping slice. Where they disagree, a named person decides, and the reasons become the grading guide.
  • Freeze a test split you never tune on. Keep a separate slice for trying prompt changes.
  • Grow the set from production: every flagged or corrected output goes in.

The disagreement step matters more than it looks. If two experts agree on only some answers, treat their agreement rate as a rough ceiling. A score far above it usually means the system has learned one labeler’s habits, not that it is right more often. Say that out loud; it shows you know what the number can and cannot tell you. The post on building a golden set for LLM evals goes deeper on sampling and labeling.

A short scoring script you can explain

You may be asked to write the scoring code itself. One candidate reported, in April 2026, that Snorkel AI’s recruiter described an FDE technical screen built on a client brief, a ground-truth CSV and model outputs, in which they would write evaluation code and find error patterns; the candidate later said the interview was canceled. Source 8Got a Snorkel AI Forward Deployed Engineer interview next week, anyone been through this? Format is unusual (post by u/itskabeer)PublisherReddit r/csMajorsSource typecandidate report on RedditSource 9Got a Snorkel AI Forward Deployed Engineer interview next week, anyone been through this? Format is unusual (comment by u/itskabeer)PublisherReddit r/csMajorsSource typecandidate report on Reddit

You do not need a framework for this. Here is the whole idea, with the baseline beside the system:

from collections import Counter

rows = [  # (gold, baseline, system)
    ("refund", "refund", "refund"),
    ("cancel", "refund", "cancel"),
    ("billing", "billing", "billing"),
    ("cancel", "other", "cancel"),
    ("other", "refund", "billing"),
    # None: output failed to parse
    ("refund", "refund", None),
]

def score(col):
    right = sum(r[0] == r[col] for r in rows)
    errors = Counter(
        (r[0], r[col])
        for r in rows
        if r[0] != r[col]
    )
    return right / len(rows), errors

for name, col in [("baseline", 1), ("system", 2)]:
    acc, errors = score(col)
    print(f"{name}: {acc:.2f}")
    for (gold, got), n in errors.items():
        print(f"  {gold} -> {got}: {n}")

Run it and it prints:

baseline: 0.50
  cancel -> refund: 1
  cancel -> other: 1
  other -> refund: 1
system: 0.67
  other -> billing: 1
  refund -> None: 1

What to say while you write it:

  • The denominator is every row. The last row failed to parse, and it still counts, as a wrong answer. Dropping it would lift the system from 0.67 to 0.80, which is how a model that gives up on hard inputs ends up looking better.
  • The error pairs are the useful part. They are the off-diagonal cells of a confusion matrix. cancel -> refund tells you what to fix. A single accuracy number does not.
  • Six rows prove nothing. Say so. On a real set you would also report how many examples each queue has, because a rate on a handful of rows swings with one example.

To practice the full version, with error buckets and a ground-truth file, try the scoring script question.

What you monitor after launch

Offline scores describe the day you measured. Production drifts: users ask new things, documents change, and the model provider ships a new version. Name signals and actions, not tools.

  • Input mix. The share of each category or topic. A sudden new topic means your test set no longer matches reality.
  • Outcome signals. Re-queues, edits to drafted replies, user flags, hand-offs to a person. These are free labels.
  • Sampled review. A fixed number of outputs a week, read by someone on the customer’s side, scored with the same guide as the test set.
  • Regression runs. The frozen test set runs before every prompt change, model change or retrieval change, and a drop blocks the change.
  • Cost and latency. A slower or pricier answer can break the business case even when quality holds.

Each signal needs an owner and an action: “if cancel misroutes rise above the agreed level, fall back to the rules and page me.” The question on designing an eval harness for a support agent is the design-round version of this list.

Follow-ups: LLM judges, drift and the cost of a wrong answer

A strong first answer earns harder follow-ups. Have a sentence ready for each.

“You used an LLM to grade. How do you trust it?” Score a held-out slice by hand, then compare the judge’s verdicts with the human labels. If the judge disagrees with people more often than the system makes the mistake you care about, it cannot measure that mistake, and a person grades that category instead. Watch for judges that prefer longer answers or the first option shown (the MT-Bench judge paper by Zheng and colleagues, section 3.3 names both). The post on LLM-as-a-judge failure modes lists the rest.

“What happens when the provider updates the model?” Pin a fixed model version, never an alias that moves (Anthropic’s page on model IDs shows which of its IDs are fixed snapshots and which are aliases). Rerun the frozen set before switching, compare error pairs as well as the headline number, and track the provider’s retirement date so a forced switch never arrives untested.

“What if a wrong answer is expensive?” Then accuracy is the wrong single number. Split errors by cost, set a stricter bar for the expensive kind, and design the system to hand off when it is unsure. Explaining that trade to a non-technical buyer is its own skill; the lesson on explaining AI limits to non-technical leaders covers the words, and so does the question on why an assistant cannot be right every time.

“It’s a RAG system. What do you measure?” Measure retrieval and generation separately: did the right passage come back, and did the answer stay faithful to it? Otherwise you cannot tell which half to fix. See the RAG system design interview walkthrough and the question on checking that citations are real.

Practice it out loud

Reading the five parts is easy. Saying them in order, with numbers, under a follow-up, is not. Pick a project you have built, write the five parts on one card, and say the answer in two minutes. When you want a customer who pushes back, run the free AI practice case and see whether you reach for evidence under pressure.

Start with the question page for this exact prompt: write your answer, time it at two minutes, then compare it with the model answer. The rest of the AI-round questions below are built the same way.

GlossaryRetrieval-augmented generationAnswering with a model that is given passages retrieved from a document collection as context.More on Retrieval-augmented generationGlossaryForward deployed engineerA software engineer who builds and ships production systems inside a customer’s problem and environment, accountable to that customer’s outcome.More on Forward deployed engineerGlossaryLaunch criteriaThresholds agreed with a customer before building that decide whether a system goes live.More on Launch criteriaGlossaryAgentA system in which a model chooses steps and tool calls to complete a task, within limits the design sets.More on Agent

Questions people ask

What should I say when asked how I know my AI system works?

Name five things. The metric that matches the customer’s goal, the data you measured it on, the baseline you compared against, the result, and what you monitor after launch. Then say what the system still gets wrong and what you would do about it.

Do FDE jobs involve evaluation work?

Several postings name it. Snowflake’s Applied AI FDE posting asks the hire to turn customer goals into quality metrics, evaluation frameworks and golden datasets, and Scale AI’s Frontier Agents FDE postings list evaluation harnesses with golden datasets, regression suites and LLM-as-a-judge.Source 2Senior/Staff Forward Deployed Engineer, Applied AI @ SnowflakePublisherSnowflake (Ashby job board)Source typecompany job postingSource 3Frontier Agents Engineer (Forward Deployed Engineering)PublisherScale AI (Greenhouse)Source typecompany job posting

Is ‘how do you know it works?’ the key question in OpenAI’s FDE interview, as one prep site says?

One prep site calls it the differentiator question for OpenAI and cites no source. OpenAI’s FDE posting does say FDEs measure success partly through eval-driven feedback.Source 7Forward Deployed Engineer Interview: The Definitive 2026 GuidePublisherExponent (Aced)Source typeinterview prep siteSource 10Forward Deployed Engineer (FDE) - SF | OpenAIPublisherOpenAISource typecompany job posting

Keep reading