In this post11 sections
  1. The demo went wrong: what the executive is really asking
  2. Why this answer matters in an FDE interview
  3. The explanation: why errors happen, what you measured, how you catch them, what a person still does
  4. Bad wording and good wording, side by side
  5. The version that fits on a slide
  6. Follow-up questions to expect
  7. Mistakes that lose the room
  8. Practice it out loud
  9. Questions people ask
  10. Keep reading
  11. More from the blog

The demo was going well until the assistant gave the VP of claims the wrong filing deadline, one she knows by heart. Now she has folded her arms and wants to know why the AI “lies”, and you can feel the project’s budget listening. To explain an AI mistake to a non-technical executive, answer in four plain parts: why this kind of error happens, what you measured on her own data, how the system catches errors, and what a person still decides. Skip the lecture on how language models work, and never promise it will be right every time. This post gives you the words; for where this moment sits in the whole loop, read the FDE interview guide.

The demo went wrong: what the executive is really asking

“Why does it lie?” sounds like a technical question. It isn’t. Under it sit three questions the VP will not say out loud:

  • Can I trust anything it says? One wrong answer makes every right answer suspect.
  • Will this embarrass me? She pictures the same mistake in front of her boss, a policyholder or a regulator.
  • Should I keep funding this? She is deciding whether the project survives the quarter.

A long answer about tokens and probabilities answers none of them. A defensive answer (“it’s usually right”) answers the third one badly. What works is showing that you understand the mistake better than she does, that you have a number for how often it happens, and that the design already assumes it will.

Your first sentence matters most. Agree with what she saw, specifically:

“You’re right, that deadline is wrong, and it’s the kind of mistake that matters in claims. Let me show you where it came from.”

Then look. Open the source the assistant used. In this fictional demo, it pulled last year’s version of the policy, which the document store still held beside the current one. That one click turns “the AI lies” into “the AI read an outdated document”, which is a cause you can fix and a fact she can check.

Don't correct her vocabulary first

“Lie” is the wrong word, since lying needs intent. But opening with “well, technically it can’t lie” sounds like you are defending the machine against the customer. Fix the word later, in passing, once she trusts the diagnosis.

Why this answer matters in an FDE interview

leaders put plain-language communication with executives at the center of the job. Sarah Khalid, an FDE director at Salesforce, said on Salesforce’s blog that FDEs engage business and executive stakeholders who “definitely don’t understand technical jargon”. Source 1Today's Hottest Role: Forward Deployed EngineerPublisherSalesforce 360 BlogSource typecompany blog In Salesforce’s newsroom, she recalled an insurance engagement where the work was more judgment, trust and communication skills than coding. Source 2How Forward Deployed Engineers Are Proving AI Makes Tech Jobs More HumanPublisherSalesforce NewsroomSource typecompany blog

OpenAI wants FDEs who are strong software engineers and also good communicators, able to explain the solution in a way that resonates with everyone. That’s how Gregor Ojstersek’s Engineering Leadership newsletter reported a conversation in August 2026 with Colin Jarvis, whom it described as OpenAI’s global head of forward deployed engineering. Source 3Inside OpenAI's Forward Deployed Engineer RolePublisherEngineering Leadership (Gregor Ojstersek)Source typenews report

That is what the job asks for, not a description of how any interviewer scores. In a client role-play, the prompt can be as short as ours: a non-technical VP asks why the assistant cannot be right every time, and you have two minutes. A strong answer shows three things at once: you can diagnose an AI failure, you can put a number on it, and you can say both in words a VP repeats to her boss.

The explanation: why errors happen, what you measured, how you catch them, what a person still does

This four-part structure is our method. Each part answers one of her unspoken questions, and together they take about two minutes to say.

Why errors happen

One or two sentences, no more. The goal is a mental model she can hold, not a course.

“The assistant doesn’t look facts up the way a database does. It writes the most likely answer from the documents we give it. So when it’s handed the wrong document, or none, it can sound completely sure and still be wrong. Today it was handed last year’s policy.”

Then name the kind of error, because “AI mistakes” is not one thing. For an assistant built on retrieved documents, the useful split is:

  • It found the wrong source. An outdated version, the wrong product line, a similar clause from another policy.
  • It found the right source and said something the source doesn’t support. This is what people call hallucination.
  • It found nothing and answered anyway. No clause covers the question, so it filled the gap with a plausible answer. The fix is to make it decline.
  • The question was ambiguous. “The deadline” could mean the filing deadline or the payment deadline.

Say which one it was, because each has its own cause and its own fix.

If the system is an agent

Split the same way, with an ’s categories. ElevenLabs, writing about lessons from its forward deployed engineering on voice agents, names prompt ambiguity, tool misuse and escalation drift as recurring failure patterns in agent deployments. Source 4Building voice agents that last: some lessons learned from forward deployed engineeringPublisherElevenLabsSource typecompany blog So tell the executive which one it was and which instruction or tool you’re changing.

What you measured

This is the part that keeps the project alive, because it replaces one vivid mistake with a rate. (The numbers here are invented for the example.)

“Before the demo, we ran the assistant on 312 real questions your adjusters asked last quarter, and two of your senior adjusters checked every answer. It was wrong on 9. After what we just saw, I traced each of those nine, and six had the same cause: an old policy version still in the document store.”

Three habits make this credible:

  • Use her data, not a benchmark. A score on a public test set, or another customer’s result, says nothing about her policies.
  • Give the range, not just the rate. 9 of 312 is about 2.9%, and with a sample that size a 95% Wilson interval puts the true rate between about 1.5% and 5.4%. Don’t say “interval” to a VP. Say: “On questions like these, expect it to be wrong on somewhere between 1 and 6 in every 100 until we fix this cause.” Round outward and include the upper end, because that is the one she will be held to.
  • Admit the missing baseline. If you don’t know how often her team gets the same questions wrong today, say so, and offer to measure it. That comparison is the one she cares about.

If you haven’t measured anything yet, don’t invent a number to sound concrete. Say what you will measure, on how many of her items, checked by whom, and when she gets the result:

“I don’t have your error rate yet, and I won’t guess. By Friday we’ll run it on 300 of your adjusters’ real questions, two of your senior adjusters will check every answer, and you’ll get the count, the range and the causes.”

This is the version you’ll need most in an interview role-play, where there are no real numbers. Practice it on the two-minute VP question. How to answer “how do you know it works?” covers the measurement in more depth, and building a golden set for LLM evals covers where the test questions come from.

How you catch them

Now show her the design assumes errors will happen. Name each check in terms of what she would see:

  • Every answer shows its source. The adjuster clicks through to the clause. A wrong answer with a visible source is caught in seconds. The verifiable citations question works through how to build and check this.
  • It declines when it can’t find a source. “I couldn’t find this in the current policy” is a safe answer.
  • Superseded documents come out. The fix for today’s error is to load only current policy versions, then rerun all the test questions before anyone sees it again.
  • The test set grows. Today’s question goes into it, so this exact mistake is checked before every future change.

What a person still does

End with the decisions that stay human. This is what lets her tell her boss the risk is contained.

“The assistant drafts; your adjusters decide. Anything that changes a payout or a deadline on a live claim is confirmed by an adjuster against the linked clause before it goes to a policyholder.”

Be specific about which decisions. “There’s a human in the loop” is a slogan; “an adjuster confirms every deadline before it’s sent” is a control she can audit.

The whole answer, spoken

Put together, it runs well under two minutes:

“You’re right, that deadline is wrong. It came from last year’s version of the policy, which was still in the documents we gave the assistant. It doesn’t look facts up; it writes the most likely answer from what it’s given, so a stale document gives a confident wrong answer.

“We tested it on 312 of your adjusters’ real questions, checked by two of your senior adjusters. It was wrong on 9, and when I traced them, six had this same cause. On questions like these, expect it to be wrong on somewhere between 1 and 6 in every 100 until this cause is fixed. We’re removing the old versions and rerunning all 312 before your team sees it again.

“Every answer links to the clause it used, it says so when it can’t find one, and an adjuster confirms any deadline before it reaches a policyholder. I’d also like to measure how often the current process gets these same questions wrong, so you can compare like with like.”

Bad wording and good wording, side by side

The same facts can sink or save the project. Swap the phrase on the left for the one on the right.

Instead ofSay
“It hallucinated”“It said something its source doesn’t support”
“That’s just how LLMs work”“Here’s why this one went wrong, and the fix”
“It’s highly accurate”“It was wrong on 9 of 312 of your questions”
“It’s still learning”“We found the cause and we’re retesting”
“That’s an edge case”“We hadn’t tested that; it’s in the test set now”
“We’ll get it to perfect”“Nothing here is right every time, so every answer is checkable”
“Trust me”“Click the source on any answer and check it”

“It’s still learning” deserves a note, because it misleads. A deployed assistant generally doesn’t learn from its mistakes on its own. It changes when someone changes its documents, its prompt or its model, and each change should be retested. Promising it will learn is promising something no one is doing.

For more on plain wording in writing, not just speech, see the lesson on writing for customers.

The version that fits on a slide

Executives forward slides. If the VP has to brief her boss tomorrow, give her one she can reuse, in her words, not yours:

Claims assistant:
wrong deadline in the demo

Why:    used last year's policy
Tested: 312 real adjuster
        questions, wrong on 9
        (6 had this same cause)
Range:  likely 1 to 6 wrong
        per 100
Catch:  every answer links its
        clause; declines when
        no clause is found
People: adjusters confirm every
        deadline before it
        reaches a policyholder
Next:   remove old versions,
        rerun all 312, report
        results to you

Five lines of substance and one line of next step. No model names, no architecture, no percentages without the count behind them. If a line needs a footnote to be understood, rewrite the line.

Follow-up questions to expect

A good explanation earns harder questions. Have an answer ready for each.

“So how often will it be wrong?” Give the rate with its range, on her data. If you don’t have it yet: “I don’t have your number yet. I’ll have it after we run it on a few hundred of your real questions.”

“Can you guarantee it for the auditor?” Ask what the guarantee is for. What an auditor can check is evidence of control: the test results, the source on every answer, and who confirms what. Offer that.

“Why didn’t you catch this before the demo?” Say it plainly: “We counted the wrong answers but hadn’t traced each one to its cause, so we missed that several shared one. That was a gap in our testing, and tracing causes is now part of every run.” Owning a gap costs less than explaining it away.

“Would a newer model fix it?” Not this error. A better model given last year’s policy still gives last year’s deadline. Say which fix matches which cause.

“Should we stop here?” Offer a narrower start instead: questions whose answer is one clause the adjuster can check, with deadlines and payouts confirmed by a person, and widen once the rerun of all 312 comes back clean.

The broader version of these questions, about the risks of the technology itself, is in the biggest risks question.

Mistakes that lose the room

Tripwire: The lecture

Five minutes on tokens, embeddings and training data. She hears that you find the machine more interesting than her problem. Fix: one sentence of why, then the number.

Tripwire: Defending the model

“It’s actually right most of the time.” You have just argued with the person paying for the project about something she saw with her own eyes. Fix: agree first, then diagnose.

Tripwire: Blaming her data

“Your documents are a mess.” They may be, but say it as a fix you own: “We loaded two versions of the policy; we’ll load only the current one.”

Tripwire: The borrowed number

Quoting a benchmark score or another customer’s accuracy. It isn’t her data, and she can tell. Fix: her questions, her experts, her result.

Tripwire: Killing the project yourself

Over-apologizing until you say “maybe AI isn’t ready for this”. You have made her decision for her, and made it badly. Fix: size the problem, name the control, propose the next step.

A similar structure works for other bad news. How to tell a customer the project is late applies it to a missed date.

Practice it out loud

Reading the four parts is not the same as saying them to someone frowning at you. Start with the two-minute VP question: answer it out loud, time yourself, and check that you gave a cause, a number, a check and a human decision. The lesson on explaining AI limits to non-technical leaders, which is in Pro, goes further, including a worked exchange with a VP who wants a guarantee. Pro starts with a free trial.

Then do it live. The free practice case puts you with a city customer whose mayor has already promised an AI review tool. The customer holds facts back until you ask the right questions, and your score quotes what you said. It’s free with a sign-in.

GlossaryForward deployed engineerA software engineer who builds and ships production systems inside a customer’s problem and environment, accountable to that customer’s outcome.More on Forward deployed engineerGlossaryAgentA system in which a model chooses steps and tool calls to complete a task, within limits the design sets.More on Agent

Questions people ask

How do you explain AI hallucinations to a non-technical executive?

Say that the model predicts likely text rather than looking facts up, so it can sound sure and be wrong. Then say how often it was wrong on the customer’s own test cases, how the system catches errors, and which decisions a person still checks.

Should you promise the AI will be accurate every time?

No. Promise what you measured and what you control: the results on the customer’s test cases, the checks that catch errors, and the process for reviewing and fixing them.

Why does explaining AI to executives matter in FDE roles?

Because the job needs it. A Salesforce FDE director said FDEs engage business and executive stakeholders who do not understand technical jargon, and the same director recalled an engagement where the work was more judgment, trust and communication than coding.Source 1Today's Hottest Role: Forward Deployed EngineerPublisherSalesforce 360 BlogSource typecompany blogSource 2How Forward Deployed Engineers Are Proving AI Makes Tech Jobs More HumanPublisherSalesforce NewsroomSource typecompany blog

Keep reading