A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.
How to answer
The VP is not asking for a lesson on language models. He is asking whether he can trust the thing, and what happens when it is wrong. Answer those two questions in about two minutes:
- Say yes to the premise, plainly. “It will be wrong sometimes. The comparison that matters is how often your team gets the same questions wrong today, and we can score both against the same answer key, with your people answering blind.”
- Give one analogy and stop. A sharp new analyst with the policy binder: fast, reads the pages it finds, and sometimes finds the wrong page or misreads the right one. One analogy, used once. Don’t add a second.
- Name the ways it goes wrong, in his terms. It pulls the wrong document, misreads the right one, finds one that is out of date, or picks between two that disagree. Each points to a different fix (better search, a better model, an owner for each document), which is why you name them.
- Say how you measure it. A test set of real questions with answers his team has signed off, scored before launch and rechecked after every change. You give the number from that test, never a promise.
- Say what catches errors. It shows its sources, it says “I don’t know” when it finds nothing, it asks which one he meant when a question is ambiguous, a person approves anything costly, and you review a sample every week.
- Hand him the decision. Which mistakes are expensive, and how many of them he can accept before he wants a person in the loop. Show him the dial: fewer wrong answers costs more “I don’t know”.
Leave out tokens, probabilities, embeddings and “hallucination”. A clear, correct explanation of how the model works that never reaches the error rate or the controls has not answered him. Ask what the VP needs to decide, and answer that.
In Pro, Explaining AI limits to non-technical leaders turns this into a method you can reuse, with the dial shown on a worked contract-review case.
Follow-ups
What the interviewer may ask next, once your first answer is on the table.
- He asks: ‘So how often will it be wrong?’ Give a number or say how you will get one.
- He wants a guarantee for the regulator. What do you offer instead?
- Which control would you add if errors cost money?
Where answers go wrong
- Explains how language models work instead of the error rate and the controls.
Answer this in two minutes
Write the answer you would say out loud. The clock starts with your first word.
Compare with the model answer
Model answer
“Short answer: it will be wrong sometimes. The useful comparison isn’t perfection, it’s how your team answers the same questions today, and we can measure both side by side. What we control is how often, on which questions, and what catches the mistakes before they cost you anything.
“Think of it as a sharp new analyst with your policy binder. You ask a question, it finds the pages that look relevant, reads them and writes an answer. It’s fast. When it goes wrong, it goes wrong in four ways. It pulls the wrong page. It reads the right page and gets it wrong. The page is out of date. Or two pages disagree and it picks one. Better search and a better model help with the first two. The other two are about your documents, and fixing them helps your people too.
“So we don’t ask you to trust it. We measure it. Your team gave us real questions with the answers they signed off. We score the assistant against that key before every release, and two of your analysts answer a sample of the same questions blind, scored the same way. That gives you the side-by-side, broken down by the kind of question, so you can see where it is strong and where it isn’t.
“Then there are the controls. Every answer shows the document and section it came from, so your people can check it in a click. When it can’t find a source, it says so instead of guessing. When a question is ambiguous, it asks which one you meant. Anything that changes a customer’s account goes to a person to approve. And we review a sample of real answers every week, so a pattern of mistakes shows up in days, not when a customer complains.
“The decision I need from you is which mistakes are expensive. That tells us where to keep a person in the loop.”
If he asks, “So how often will it be wrong?”: “I won’t guess. This week I’ll give you three counts from the test set: answered correctly, answered wrongly, and declined because it found no source. I’ll say how many questions they rest on, because a small test set moves a lot. The useful part is that we can trade one for another: make it decline more, and it’s wrong less often but helps less often. To show the shape, with made-up numbers: say that on 300 questions it gets 255 right, 15 wrong and declines 30. Made stricter, it might get 240 right, 6 wrong and decline 54. Going stricter turns 9 wrong answers and 15 right ones into 24 more questions your team answers by hand. Where to set that is your call, question type by question type. After launch, you get the same counts from the weekly review, so you are watching a measured rate, not a promise.”
If he wants a guarantee for the regulator: “No one can guarantee an answer is always right, and a guarantee I can’t back up is worth less to you than evidence you can show. What I can give you is the test set and its results, a log of every question, answer and source, a record of every change to the prompt or the model with the test result before and after, and the rule that a person signs off on anything that affects a customer’s money or rights.”
If errors cost money: “I’d add an approval step on every answer that triggers a payment or a refund, with the source shown beside it. The assistant drafts, a person approves, and we log who approved what. That turns an expensive error into a slower correct one.”