In this post13 sections
- What ‘we want an AI agent’ can hide
- What AI deployment teams say they start with
- Pick one workflow
- Measure the baseline before you build
- Build the eval set from real cases
- Where a person stays in the loop
- Worked example: an insurer’s claims request
- The launch criterion, and what to say in the interview
- Common mistakes and the fix
- Practice the cut out loud
- Questions people ask
- Keep reading
- More from the blog
The customer’s COO opens the kickoff with a slide titled “AI agents across every claims process” and asks when the first one goes live. You have a small team, a fixed budget and a list of wishes with no edges. This post shows how to turn that request into a first version that ships, and it belongs to the decomposition interview guide, which covers the round where you are asked to do the same thing on a whiteboard.
The short answer: to scope an AI project for a customer, pick one workflow where the customer can measure today’s cost, measure that baseline, build an eval set from real past cases, decide where a person reviews the output, and agree a launch criterion before anyone writes a prompt. Everything else on the slide goes on a list called “later”, with a reason beside each item.
What ‘we want an AI agent’ can hide
“” names a technology, not a problem. The same sentence can sit on top of very different needs:
- A cost problem. A backlog is growing faster than the team can hire.
- A quality problem. Errors are reaching customers or regulators.
- A mandate. The board asked for an AI plan, and someone has to show progress.
- A reaction. A competitor announced something.
- A workflow someone already hates. One manager has a painful process in mind and has not said which.
Each of these leads to a different first version. The mandate needs something visible and safe. The cost problem needs throughput on a high-volume task. Build before you know which, and you build the demo that wins the meeting and loses the renewal.
Four questions sort it out in the first conversation. Ask them in these words:
- “If this worked, which number on which report would change?”
- “Who does this work today, and how long does one case take them?”
- “When it goes wrong today, what happens, and who finds out?”
- “Who has to say yes before anything the system produces reaches a customer?”
The answers give you the metric, the baseline, the cost of an error and the approver. The lesson on going from a discovery call to a technical proposal in one sitting turns them into a written plan.
What AI deployment teams say they start with
OpenAI says a typical OpenAI Deployment Company engagement begins with a focused diagnostic of where AI can create the most value, followed by a small number of priority workflows selected with the customer’s leadership and operating teams. Source 1OpenAI launches the OpenAI Deployment Company to help businesses build around intelligence | OpenAIPublisherOpenAISource typecompany blog The post on what the OpenAI Deployment Company is covers the rest of that announcement.
The Pragmatic Engineer reported, in August 2025, that Colin Jarvis, then OpenAI’s head of forward deployed engineering, said what customers describe in scoping often does not match the data and system reality, so the team biases toward moving fast, proving out any brick walls and then adjusting the scope. Source 2What are Forward Deployed Engineers, and why are they so in demand? (Gergely Orosz)PublisherThe Pragmatic EngineerSource typenews report
Job postings name the same work. In September 2026, Cohere’s Agentic Platform posting asks for ambiguous business problems turned into agentic workflows with clear success criteria and evaluation methodologies, Source 3Forward Deployed Engineer, Agentic PlatformPublisherCohere (Ashby job board)Source typecompany job posting and Snowflake’s Senior/Staff Applied AI FDE posting asks for ambiguous customer goals turned into quality metrics, evaluation frameworks and golden datasets. Source 4Senior/Staff Forward Deployed Engineer, Applied AI @ SnowflakePublisherSnowflake (Ashby job board)Source typecompany job posting
The steps below are our method, built to match those descriptions. No employer publishes it as a rubric.
Pick one workflow
List every workflow on the customer’s slide, then check each against five tests:
- Volume. People do it often and in roughly the same way, so there is history to learn from and to test against.
- A cost you can see today. Time per case, backlog, error rate or money lost. If nobody can say what it costs now, nobody will agree it got better.
- Reachable data. You can get the inputs and the past outcomes without a new integration project or a fresh security review.
- Recoverable mistakes. A person can catch a wrong output before it harms a customer.
- An owner who wants it. A named operations lead who will use the first version and lend experts to label cases.
Pick the workflow that passes the most tests, not the one with the biggest headline. Then write the “later” list, with a sentence for each item on why it waits. That list is how you say no without saying no. The lesson on sequencing work by risk and value shows how to order it.
This step does not pick the technique: prompt, retrieval, fine-tuning or a multi-step agent comes after you know the task and the data. If an interviewer pushes, the question on fine-tuning versus RAG versus prompting walks through that choice.
Measure the baseline before you build
A baseline is today’s number for the workflow you picked, measured the same way you will measure the new system. Without it, “the AI is better” is an opinion.
Measure four things:
- Volume: cases per week.
- Time: from arrival to done, and the hands-on minutes inside that.
- Quality: how often today’s process gets it wrong, from a quality-assurance sample or from rework records.
- Cost of an error: what one mistake costs downstream, in time, money or risk.
Most of this sits in the customer’s system of record as timestamps and status changes. The rest you get by sitting with the team for a morning. Score the current human process on the same eval set you build next, so the launch bar is “at least as good as today”, not “perfect”.
Build the eval set from real cases
An eval set is a fixed sample of real past cases with the answers the customer’s experts agree are correct. Every version of the system is scored against it before it goes near a user. When it is versioned and agreed, it is a golden set.
Four rules keep it honest:
- Sample from real traffic, by segment. Random sampling buries the rare, expensive cases.
- Let the customer’s experts own the labels. You can draft them; they confirm.
- Add the hard cases on purpose. Missing fields, two requests in one message, a second language, the case that should go straight to a person.
- Hold part of it out. Tune on one split and score on the other.
Size it so a change is visible. ElevenLabs’ post on lessons from its forward deployed engineering says batches of fewer than 100 calls per branch produce too much variance to evaluate outcomes reliably. Source 5Building voice agents that last: some lessons learned from forward deployed engineeringPublisherElevenLabsSource typecompany blog That was written about voice-agent calls, not offline eval sets, but the arithmetic carries over: at 100 cases, one case moves accuracy by one point, and a result of 93% is consistent with anything from about 86% to 97%. If the bar sits close to today’s rate, you need far more cases, or real traffic, to tell the two apart. The question on designing an eval harness for a support agent shows how to run the scoring.
This work does not stop at launch. Cursor’s FDE postings, as of September 2026, make the FDE responsible for production quality, including tracing, evals, metrics, debugging model behavior and failure modes. Source 6Forward Deployed EngineerPublisherCursor (Ashby job board)Source typecompany job board The post on building a golden set when the customer has no labeled data covers sampling, labeling and agreement in detail.
Where a person stays in the loop
Decide where a person reviews the output by the cost of an error, not by how confident the model sounds. A first version runs in shadow, then starts at assist:
| Setting | What the person does |
|---|---|
| Shadow | Never sees it; you compare |
| Assist | Approves every output |
| Review on risk | Sees only flagged cases |
| Audit | Samples outputs after the fact |
Write the rules for sending a case to a person before you build:
- By segment: anything in a category the customer names as risky.
- By signal: missing required fields, a low extraction score, or conflicting values.
- By consequence: anything that reaches a customer, moves money or cannot be undone.
Every time a person overrides the system, record it. The override rate is your best live quality signal, and each override is a new eval case. Moving from assist to review on risk is a decision the customer makes on evidence, at a date you agree, not something that happens quietly. For agents that act rather than draft, the post on the agentic system design interview covers permissions and the hand-off to a person.
Worked example: an insurer’s claims request
The insurer below is fictional, and so are its numbers; they belong to the exercise.
A regional property and auto insurer asks for “AI agents across claims”. The COO’s slide lists five: intake of new claims, an adjuster assistant, customer status updates, fraud detection and subrogation recovery.
Discovery. The four questions turn up the following. An intake team of 12 handlers reads about 1,800 emailed claim notices a week. For each, a handler types the policy number, date of loss, loss type and location into the claims system and routes the claim to a queue. The median time from email to claim record is 26 hours. A quality sample shows 7% of claims routed to the wrong queue, and each misroute adds about 2 days before an adjuster sees it. The email archive and the final claim records go back years, so the right answers already exist. The head of claims operations owns the intake backlog and offers 2 senior handlers to label cases.
The cut. Intake passes all five tests. Here is where the rest lands, and the test each one fails:
| Request | First version? |
|---|---|
| Claim intake | Yes: draft and route |
| Adjuster assistant | Later: fails reachable data (intake first) |
| Status updates | Later: fails recoverable (customer-facing) |
| Fraud detection | Later: fails reachable data (legal review) |
| Subrogation | Later: fails volume |
Baseline. 26 hours to a claim record, about 7 hands-on minutes per email, and 93% routing accuracy. From the same QA sample, handlers type the policy number correctly 98% of the time and the date of loss 95%.
Eval set. 400 past emails, sampled by loss type and by format: plain email, attachment only, forwarded by a broker, and not in English. The labels come from the final claim records, with the 2 senior handlers settling the routing on the cases where the record was later corrected. 100 are held out as a final check the build never saw. The exact rates, routing above all, are scored on the shadow run’s real traffic, about 3,600 emails checked against the final claim records, because a 100-case holdout cannot tell 93% from 90%. Hard cases are tagged on purpose: 2 claims in one email, no policy number, a mention of injury, a letter from an attorney.
The first version. A walking skeleton runs end to end in shadow mode: emails in, a draft record on a side screen, handlers working as usual, and every draft compared with what they typed.
Then assist. Two handlers switch to assist mode: the system pre-fills the claim form and suggests a queue, and the handler confirms or corrects. Any email that mentions an injury or an attorney goes to the senior queue, whatever the system suggests.
The launch criterion, agreed with the head of claims operations before the build:
eval: intake-v1 holdout,
then shadow run
# field bars are today's rates
policy_number_accuracy: ">= 0.98"
date_of_loss_accuracy: ">= 0.95"
injury_or_attorney: all flagged
routing_accuracy: not below
baseline (shadow run, all emails)
shadow_run: 10 working days
shadow_goal: drafts match handlers
on fields and queue
assist_pilot: 2 handlers, 10 days
assist_goal: fewer hands-on minutes,
no rise in misroutes
signed_by: head of claims operations
The “agents everywhere” request is not refused. It becomes an ordered list in which the first item produces the clean claim data the adjuster assistant and the status updates will need.
The launch criterion, and what to say in the interview
A launch criterion is the condition the customer agrees, before launch, that the system must meet on the held-out eval set. It turns “is it good enough?” from an argument into a check. It names the metric per segment, the threshold, the baseline it is compared with, who signs, and what happens on a miss: another iteration, or a narrower launch with a person reviewing the risky segment. ElevenLabs’ April 2026 post on lessons from forward deployed engineering recommends test-driven development to keep agents aligned to core metrics throughout the build, which only works if the metrics exist before the build does. Source 5Building voice agents that last: some lessons learned from forward deployed engineeringPublisherElevenLabsSource typecompany blog
When an interviewer says “the customer wants AI agents, where do you start?”, here is an answer you can say in about a minute:
“I’d start by finding out what ‘agents’ is standing in for: which number they want to move, who does the work today, and what a mistake costs. Then I’d pick one workflow with volume, a cost we can measure, data we can reach, mistakes a person can catch, and an owner who wants it. Everything else goes on a later list with a reason. Before building, I’d measure today’s baseline and build an eval set from real past cases, labeled by their experts, with a held-out split. The first version runs in shadow next to the team, then with a person approving every output, and we agree the launch criterion up front: accuracy against the baseline on the holdout, with the risky cases always sent to a person. Once that ships, the next workflow is easier, because we know what the data really looks like.”
Expect three follow-ups:
- “The COO wants all five.” Show the ordered list and what each later item needs from the first. The lesson on explaining AI limits to non-technical leaders helps with the tone.
- “How do you know it works?” Point to the eval set and the launch criterion, then practice the full answer with the question a customer asks how you know your AI system works.
- “Why not build a platform first?” Say: “After intake ships, we’ll know which parts repeat: the email parsing, the policy lookup and the review queue. Building them generically now means guessing which ones the adjuster assistant needs.” In a November 2025 Altimeter interview, Colin Jarvis said the biggest mistake his OpenAI team made that year was generalizing too early. Source 7Colin Jarvis | Head of Forward Deployed Engineering at OpenAI: Trust. Product. Impact.PublisherAltimeter Capital (YouTube)Source typerecorded talk or interview He meant how OpenAI built its FDE team, not one customer project, but the lesson carries by analogy.
Common mistakes and the fix
| Mistake | Fix |
|---|---|
| Picking the flashiest idea | Pick the one you can measure |
| Choosing a model first | Choose the workflow first |
| No baseline | Measure today first |
| Demo examples as the test | Real, sampled, held out |
| Full automation on day one | Shadow, then assist |
| Criteria set after results | Sign them before the build |
Practice the cut out loud
The free practice case gives you the same problem with a different customer: a city whose mayor has already promised an AI review tool, a chief building official who will not let software approve anything, and a clock. It needs a sign-in. Scope the first version out loud, and put the later list on the board before the customer asks for it.
Questions people ask
How do you scope an AI project for a customer?
Pick one workflow where the customer can measure today’s cost, measure that baseline, build an eval set from real cases, decide where a person reviews the output, and agree a launch criterion before you build. OpenAI says a typical Deployment Company engagement begins with a focused diagnostic of where AI can create the most value, followed by a small number of priority workflows.Source 1OpenAI launches the OpenAI Deployment Company to help businesses build around intelligence | OpenAIPublisherOpenAISource typecompany blog
What is a launch criterion for an AI system?
A condition the customer agrees before launch that the system must meet on an eval set of real cases, such as a minimum accuracy on a held-out set and a ceiling on errors that reach a customer without review.
Why not build a general agent platform first?
Because a first workflow that ships shows what a platform would need; building the platform first is a guess. Colin Jarvis, then head of forward deployed engineering at OpenAI, said in a November 2025 interview that generalizing too early was the biggest mistake his team made that year; he was describing how OpenAI built its FDE team, but the lesson carries.Source 7Colin Jarvis | Head of Forward Deployed Engineering at OpenAI: Trust. Product. Impact.PublisherAltimeter Capital (YouTube)Source typerecorded talk or interview
Keep reading
Lessons
Questions
- A customer asks how you know your AI system works. Answer them.
- Design an evaluation harness that runs before every change to a customer-facing support agent.
- When would you fine-tune, when would you use retrieval, and when is prompting enough? Use a real example.
- A non-technical VP asks why the assistant cannot be right every time. Explain in two minutes.
More from the blog
Interview rounds
Choosing a success metric in a case interview: the number, its guardrail and who owns it
‘How would you measure success?’ exposes vague answers. Name the outcome metric, the guardrail and the person who will read it, with three worked cases.
Interview rounds
Decomposition interviews that come with a dataset: how to start from the columns
When the decomposition round hands you a table, read its grain, keys and timestamps first. A method, a worked fitness-app example and the first query.
Interview rounds
Palantir decomposition interview example: a full walkthrough of a reported prompt
A full walkthrough of a Palantir decomposition prompt one candidate reported, built on London taxi data, with what to say out loud at each step.