In this post14 sections
- Why ‘how would you measure success?’ exposes vague answers
- The outcome metric the customer already trusts
- The guardrail that catches gaming
- The owner who will read it
- Worked case: city bus reliability
- Worked case: hospital readmissions
- Worked case: an AI support agent
- What you can measure in the first weeks
- Follow-ups to expect, and what to say
- Why this matches the job
- Common mistakes and the fix
- Questions people ask
- Keep reading
- More from the blog
You have clarified the problem, named the users and sketched a first version, and it is going well. Then the interviewer asks “How would you measure success?”, and you hear yourself say “user satisfaction, adoption, and accuracy”. To choose a success metric in a case interview, name one outcome metric the customer already counts, one guardrail that catches the ways that number can improve while things get worse, and the person who will read both and act on them. Three parts, one breath: that is the difference between a list of words and a plan. This post is one narrow piece of the decomposition interview guide, which covers the whole round, and the free lesson on the decomposition method shows where the metric sits among the other steps.
Why ‘how would you measure success?’ exposes vague answers
Much of a case answer can stay abstract and still sound good. A metric cannot: either you can say what gets counted, from which table, compared with what, or you cannot.
Here is what a vague answer and a specific one sound like on the same case, a transit agency whose riders say the buses are unreliable:
- Vague: “We’d track on-time performance and rider satisfaction, and see if they improve.”
- Specific: “On the frequent routes, excess wait: the wait riders get minus the wait the schedule promises, per route, per week, against this week’s value. The guardrail is passenger-minutes on board, because holding buses to even out gaps adds minutes for riders already on the bus. The head of operations reads both every Monday and decides whether to extend the pilot to a second corridor.”
The second answer does four things the first does not. It picks a number that matches the problem. It has a baseline. It names a way the fix could backfire. And it says who decides what happens next.
The three-part pattern below (outcome, guardrail, owner) is our method, not a rubric any employer publishes.
The outcome metric the customer already trusts
Your first sentence names one number. Pick it with four tests:
- The customer already counts it, or would recognize it at once. The head of operations reports service to a board; the chief quality officer reports readmissions.
- It moves when the problem is solved, and only then. If riders leave because of bunching, on-time performance can look fine while they wait longer. Pick the number the problem lives in.
- You can write its definition down. What counts, what is excluded, over what window, from which system. When the case hands you a table, that definition is usually hiding in the columns; here is how to find the metric in the columns.
- It has a baseline. “Down from where?” is the next question. Say where the starting value comes from: last month, last year, the same routes before the pilot.
Then say whether it is leading or lagging. A lagging metric is the slow outcome the sponsor cares about; a leading metric changes within weeks and should move the lagging one. The sentence to use: “The outcome we’re judged on is X. Since that moves slowly, I’d watch Y weekly, because it should move first.”
Output is not outcome
“Number of dashboards shipped”, “model accuracy” and “tickets processed by the ” count your work, not the change at the customer. Accuracy can go up while nobody uses the tool. Name the output as a leading indicator if you like, but never as the success metric.
The guardrail that catches gaming
A guardrail metric is a number you watch beside the main one to catch harm the main one hides. Find it with one question: how could my number get better while the customer gets worse?
Look for three kinds of answer:
- Gaming. Someone hits the number without solving the problem. On-time departures improve because drivers skip stops, so the skipped-stop rate is the guardrail.
- Displacement. The problem moves somewhere you are not looking. Readmissions fall because returning patients are billed as observation stays instead, so observation returns are the guardrail.
- Quality loss. Speed rises because care drops. Loan decisions get faster because underwriters check less, so later defaults or reversed decisions are the guardrail.
A guardrail without a limit is just a second chart. Say what happens if it moves, and make sure both sides of the comparison are in the same units. For the buses, weight each minute by how many riders it affects: “We agree up front that if holding adds more passenger-minutes on board than it saves in passenger-minutes of waiting, we stop holding buses and try something else.” Agree the limit before launch, because afterwards everyone has a reason to explain a bad number away. One or two guardrails is plenty.
The owner who will read it
The third part is the easiest to forget: a named person who reads the number on a schedule and makes a decision with it. A metric nobody reads is never acted on, and an who builds one has built a chart.
Name three things:
- Who reads it. A role, not “the team”. The sponsor usually owns the outcome; the people doing the work often own the leading metric.
- How often. Daily for a work list, weekly for a pilot, monthly for a board number.
- What they decide. Extend the pilot, stop it, change the rule, add a second site.
Separate the sponsor from the user: the general manager reads the outcome; the dispatcher on shift reads whether the tool is helping this hour. Stakeholders and success metrics, a lesson in Pro, goes further on finding the number each person is already judged on and who would dispute yours.
Worked case: city bus reliability
The prompt, from the city bus reliability question: a transit agency says its buses are unreliable and riders are leaving, and you have a week with their team. Every figure below belongs to this fictional agency.
Outcome. On frequent routes, riders turn up without checking the timetable, so what they feel is how long they wait. Use excess wait: the average wait riders experience minus the wait the schedule promises. On infrequent routes, use timetable adherence, with early departures counted separately because an early bus strands people. Ridership by route is the lagging outcome the general manager cares about; excess wait is the leading metric you can move this month.
Guardrail. The first version flags buses running too close together and suggests holding the second one. Holding adds minutes for riders already on board, so passenger-minutes on board is the guardrail. Add skipped stops if operators are under pressure to recover time.
Owner. The head of operations reads excess wait and the guardrail for the pilot corridor every Monday and decides whether to extend it. Dispatchers read the live flags on shift. Say this out loud, and add a warning: excess wait will likely show service as worse than the on-time figure the head of operations reports to the board, and holding buses can itself lower on-time performance. So you agree the new metric with them on day one, not on Friday.
What to say, in one breath:
“Success is lower excess wait on the pilot corridor against this week’s value, with ridership on the routes that lost riders as the outcome behind it. The guardrail is time on board: if holding adds more passenger-minutes on board than it saves in passenger-minutes of waiting, we stop. The head of operations reads both every Monday and decides whether we extend.”
The model answer on the question page shows the query that computes excess wait from stop events, and one of its follow-ups: ridership fell but on-time performance did not. Your metric choice is the answer to that follow-up.
Worked case: hospital readmissions
The prompt, from the hospital readmissions question: a hospital group wants fewer patients readmitted soon after discharge. Again, every figure is the fictional customer’s.
Outcome. First, the definition. Ask whether leadership means the number Medicare penalizes them on or an internal one, then offer to write down the window, which stays count and the exclusions. For the pilot, use the unplanned readmission rate for heart failure discharges at one hospital, within the 30-day window you and the chief quality officer write down, counting only discharges whose window has closed. That is the lagging number. The leading one is the share of flagged patients the transitional-care nurses reach within two days of discharge, the target the nurse manager sets, which you can see move in the first week.
Guardrails. Readmissions can fall for bad reasons, and each one gets its own line beside the rate:
- longer length of stay, which keeps patients in hospital instead of preventing a return;
- returns billed as observation stays, which do not count as readmissions;
- emergency visits that end in discharge;
- deaths within the window, the worst way for a patient not to come back.
Owner. The chief quality officer owns the rate and signs off on extending the pilot, monthly. The nurse manager reads the daily list and the reach rate, and decides whether the list is the right size for the team.
What to say:
“Success is a lower unplanned readmission rate for heart failure at this hospital, by our written definition, against last year’s rate. Week to week, I’d watch how many flagged patients the nurses reach. Beside the rate go length of stay, observation returns, emergency visits and deaths. The chief quality officer reads it monthly; the nurse manager reads the reach rate daily.”
The worked healthcare and finance cases in Pro walk through the definition in full.
Worked case: an AI support agent
This is the prompt an AI-lab or AI-product candidate is most likely to meet: a customer has put an LLM agent in front of its support queue and wants to know whether it is working. Every figure is the fictional customer’s. If the prompt is only “we want AI”, scope it first, then come back to the metric.
Outcome. Containment rate: the share of conversations the agent resolves without a human. The customer already counts handoffs to agents, so they will recognize it.
Guardrails. Containment is easy to game: an agent that never hands off “contains” everything. So watch tickets reopened or escalated soon after the bot closed them. Add the error types the customer signed off on in the eval set, such as a wrong refund amount or an invented policy. That is where how you would show an AI system works comes in.
Owner. The head of support reads containment and reopens weekly and decides whether to widen the agent to a second queue. Team leads read escalations daily.
What to say:
“Success is containment on the billing queue against this week’s value. The guardrails are reopens after the bot closed a ticket and the eval error types you signed off on. The head of support reads them weekly and decides whether we widen to a second queue.”
What you can measure in the first weeks
The outcome often takes months to move. Readmission rates need the window to close; ridership lags service. So the interviewer may push: “You only have a few weeks. How do you know it’s working?”
Have early signals ready. These are ours:
- Adoption. Are the people meant to use it opening it? If the nurses stop opening the list, nothing else matters.
- Coverage. Of the cases the tool flagged, how many did someone act on?
- The leading metric. Excess wait on one corridor, or the reach rate, compared with the baseline.
- Data health. Is the feed fresh, and how many records failed to match?
Follow-ups to expect, and what to say
“How would you know it was your tool and not something else?” Say: “We compare shifts or corridors that used it with ones that didn’t, and we agree that comparison before launch.” In the readmissions case, patients the nurses happened to reach differ from those they did not, so the comparison cannot be chosen afterwards.
“The customer doesn’t collect that data.” Say: “Then the first week’s job is to start logging it. Until then I’d use a proxy, and I’ll say how it could mislead.” For the buses, scheduled headways stand in for real waits until stop events are logged, and they flatter the service.
“What if the number doesn’t move?” Say: “Then the guardrail and the leading metric tell us whether the tool isn’t being used or doesn’t work, and those have different fixes.” Low adoption is a workflow problem. High adoption with no change is a design problem.
Why this matches the job
Employers describe how forward deployed work is measured on the job, not how interviews are scored. Still, they point the same way. Palantir says Deltas, working on a team that directly supports one customer, measure success in terms of impact on the customer’s goal. Source 1Dev versus Delta: Demystifying engineering roles at PalantirPublisherPalantir BlogSource typecompany blog As of September 2026, OpenAI’s San Francisco FDE posting says FDEs measure success by production adoption, measurable workflow impact and eval-driven feedback that changes product and model roadmaps. Source 2Forward Deployed Engineer (FDE) - SFPublisherOpenAI CareersSource typecompany job posting ElevenLabs says its FDEs tie each deployment to outcomes such as resolution speed or containment rate, and stay engaged after go-live. Source 3Meet our Forward Deployed EngineersPublisherElevenLabsSource typecompany blog
Candidate reports are thin. One report on Aced, for a mid-level Palantir role (a separate role from the ) in June 2025, says the case discussions asked about metrics frameworks. Source 4Palantir Deployment Strategist Interview ExperiencePublisherAced (formerly Exponent)Source typecandidate’s personal write-up In a Blind thread about a Palantir FDSE open-ended round, one commenter, who did not name their role, wrote in January 2025 that the question moved from which features were P0 (the must-haves) to data models, which KPIs to track and what to present to executives. Source 5Palantir FDSE InterviewPublisherBlind (teamblind.com)Source typecandidate report on BlindSource 6Palantir FDSE Decomp/System Design InterviewPublisherBlind (teamblind.com)Source typecandidate report on Blind That is one report each, not a pattern, and neither tells you how the answer was judged.
Common mistakes and the fix
- Averaging across unlike segments. Excess wait on frequent routes and adherence on infrequent ones do not blend into one number. Report each where it applies.
- A ratio whose denominator can be gamed. Flagging fewer patients raises the reach rate. Pair the rate with the count flagged.
- A metric you cannot compute in time. If the data needs a quarter to arrive, it cannot judge a pilot this month. Name the leading metric instead.
Before your next case
- One outcome metric the customer already counts, with a written definition
- A baseline and where it comes from
- A leading metric you can see move within weeks
- One or two guardrails, each with a limit agreed up front
- A named owner, how often they read it, and what they decide
Practice the whole sentence on a different prompt, such as the faster small business loans question, where speed is the obvious number and quality is the obvious guardrail. Then try it live. In the free practice case, which needs a sign-in, the customer is a city whose building permits take months, and the mayor has already promised an AI review tool. Permit time is the obvious number. Before the customer asks, say its guardrail and who reads it.
Questions people ask
What is a guardrail metric?
A number you watch beside the main metric to catch harm the main metric hides. If on-time bus departures improve because drivers skip stops, the rate of skipped stops is the guardrail.
How do FDE employers say they measure success?
Palantir says Deltas measure success in terms of impact on the customer’s goal. As of September 2026, OpenAI’s San Francisco FDE posting says FDEs measure success through production adoption, measurable workflow impact and eval-driven feedback that changes product and model roadmaps.Source 1Dev versus Delta: Demystifying engineering roles at PalantirPublisherPalantir BlogSource typecompany blogSource 2Forward Deployed Engineer (FDE) - SFPublisherOpenAI CareersSource typecompany job posting
Should I give one success metric or several in a case interview?
Lead with one outcome metric so the customer can tell whether the work succeeded, then add a guardrail or two and name who reads them. A long list with no owner sounds as if you have not decided.
Keep reading
Lessons
Questions
- A city transit agency says its buses are unreliable and riders are leaving. You have a week with their team. How do you approach it?
- A hospital group wants fewer patients readmitted soon after discharge. Where do you start?
- A bank wants to approve small-business loans faster without taking more risk. How would you break this down?
More from the blog
Interview rounds
‘We want an AI agent’: decomposing a vague AI request into a first version
A customer asks for AI agents everywhere. Cut it to one workflow, an eval set, human review and a launch criterion, worked on an insurer’s claims request.
Interview rounds
Decomposition interviews that come with a dataset: how to start from the columns
When the decomposition round hands you a table, read its grain, keys and timestamps first. A method, a worked fitness-app example and the first query.
Interview rounds
Palantir decomposition interview example: a full walkthrough of a reported prompt
A full walkthrough of a Palantir decomposition prompt one candidate reported, built on London taxi data, with what to say out loud at each step.