In this post14 sections
  1. The one-minute answer to the direct question
  2. ‘Just fine-tune it’: what the interviewer is really testing
  3. What each option changes: knowledge, behavior or cost
  4. The decision, question by question
  5. Scenario: a policy assistant whose documents keep changing
  6. Scenario: output that must match a strict format
  7. Scenario: a high-volume task where cost and latency matter
  8. Scenario: a customer with little data and no evaluation
  9. The words to say no, and what you offer instead
  10. Follow-ups to expect
  11. Practice it out loud
  12. Questions people ask
  13. Keep reading
  14. More from the blog

The customer’s head of operations leans back and says it: “The model doesn’t know our business. Can’t we just fine-tune it on our data?” In an AI round, or a client simulation built around an AI product, how you answer that one sentence says more about you than any architecture diagram. Here is how to answer it, in the room and to the customer. For every other round, see the FDE interview guide.

The short answer: start with prompting, add retrieval when the model lacks knowledge or the knowledge changes, and fine-tune only to change behavior, format or cost, once you have good examples and an evaluation that proves it helps. Then say that decision in the customer’s words. If fine-tuning won’t fix their problem, say no, and offer what will.

The one-minute answer to the direct question

An engineer on the panel may skip the role-play and ask: “When would you fine-tune, and when would you use ?” Here is an answer to say out loud:

“I separate two failures: the model doesn’t know something, or it behaves wrongly. Retrieval fixes the first, because I can update the source and cite it. Fine-tuning fixes the second: a strict format, a stable classification, or a narrow task I want a smaller, cheaper model to do. I start with a prompt and a held-out evaluation set, because the prompt is the baseline everything else has to beat. And I name the upkeep: a fine-tune is tied to one base model, so every upgrade means retraining and re-evaluating.”

It splits knowledge from behavior, puts measurement first, and prices the upkeep. The rest of this post applies it to a real customer.

‘Just fine-tune it’: what the interviewer is really testing

We have not seen an employer publish how this question is scored, so what follows is our reading: “just fine-tune it” tests three things at once.

  • Do you diagnose before you prescribe? The customer named a solution. Your job is to find the problem under it: missing facts, wrong format, too slow, too expensive, or no agreed idea of “good”.
  • Can you say no to a customer? An who agrees to every request builds the wrong thing on time. The skill is refusing the method while keeping the goal.
  • Do you let evidence decide? Anyone can recite “RAG for knowledge, fine-tuning for behavior”. The strong answer says what you would measure to choose.

The postings name this work directly: Databricks’ AI FDE posting lists RAG and fine-tuning, Source 1AI Engineer – Forward Deployed Engineering (AI FDE)PublisherDatabricks (careers site / Greenhouse)Source typecompany job posting an archived xAI Government Engineering posting, for a customer-facing applied engineer rather than an FDE title, lists system-prompt tuning or fine-tuning, Source 2Member of Technical Staff - Government Engineering (archived)PublisherxAI (Greenhouse, via Wayback Machine)Source typearchived job posting and Cohere asks for the ability to build evaluation frameworks that measure accuracy, safety and latency. Source 3Forward Deployed Engineer, Agentic PlatformPublisherCohere (Ashby job board)Source typecompany job posting

What each option changes: knowledge, behavior or cost

Get this sentence into your answer early, because everything else follows from it: prompting changes the instructions, retrieval changes what the model can see, and fine-tuning changes how the model behaves by default.

OptionWhat it changesWhat it costs you later
PromptingInstructions and examples sent with each callAlmost nothing; edit or roll back in minutes
RetrievalThe documents the model reads at answer timeAn index to build, secure and keep fresh
Fine-tuningThe model’s weights, so its default behaviorLabeled data, training runs, and re-evaluation with every base model change

Three consequences are worth saying out loud in the interview:

  1. Fine-tuning is a poor way to add facts. The model absorbs them unreliably, can’t point to where an answer came from, and keeps the old version after a document changes. Retrieval-augmented generation exists for exactly that job. In one comparison, retrieval consistently beat unsupervised fine-tuning at adding knowledge (Ovadia et al.).
  2. Retrieval does not fix behavior. Better documents won’t make the model follow a strict output format or use the customer’s labels consistently.
  3. Prompting is the baseline for both. Whatever you build next has to beat it on the same test set.

The decision, question by question

Start with the first question, because it gates the rest. Each of the others points to an option, and together they give you a recommendation you can defend.

AskPoints toWhy
Can you measure success on a held-out set?If no: build that firstOtherwise nothing can be compared
Does the knowledge change, or must answers cite a source?RetrievalUpdate the index, not the model
Is the problem format, tone or labels, not facts?Better prompt, then fine-tuningThis is behavior
Do you have enough good, labeled examples to train and test?Fine-tuning becomes possibleWithout them there is nothing to train on or check against
Do cost per call or latency break the business case?A smaller model, then a tuned oneCache the prompt, shrink the model, and tune only if accuracy drops

Evaluation. Without a golden set and agreed launch criteria, you can’t tell whether the fine-tune beat the prompt, and neither can the customer. The post on building a golden set for LLM evals walks through making one.

Freshness. If the facts change weekly, fine-tuning for knowledge is ruled out: you would retrain every week and still be behind. If the whole corpus is small and stable, you may not need an index at all: put it in the prompt and see whether that passes.

Behavior or format. Try the cheap fixes first: a clear instruction, a few worked examples, and, if the model API supports it, output constrained to a schema with a validator behind it. Fine-tuning earns its cost when those plateau.

Data. “We have lots of data” usually means documents, which feed retrieval. Supervised fine-tuning, the kind most hosted APIs offer, needs input and output pairs that show the behavior you want, checked by someone who knows the task.

Latency and cost. A long prompt stuffed with examples is paid for on every call. At high volume, moving those examples into a smaller fine-tuned model can cut cost per call and response time. OpenAI’s model optimization guide makes the same point: fine-tuning allows shorter prompts, which saves token costs at scale and can lower latency. Check it against the customer’s latency budget, not in the abstract.

Scenario: a policy assistant whose documents keep changing

The setup. An insurance broker wants an assistant that answers staff questions about internal underwriting guidelines. Compliance revises the guidelines every month. The customer says: “Fine-tune a model on the guidelines so it knows them.”

The diagnosis. This is knowledge, and it changes. Staff also need to see which clause an answer came from, because they may have to justify a decision to a client.

The recommendation. Retrieval over the current guidelines, chunked by clause, with the clause cited in every answer and a refresh job that runs when compliance publishes a revision. Evaluate it with a set of real staff questions whose answers compliance signed off.

The words.

“Fine-tuning wouldn’t reliably teach the model even this month’s guidelines, and next month it would be mixing old rules with new ones. It also couldn’t show your staff which clause it used. I’d rather have the model look up the current guidelines every time it answers and cite the clause, so when compliance changes a rule, the assistant changes the same day, with no retraining.”

For the design follow-up (how you chunk, how you rank, how you filter by permission), the RAG system design interview walkthrough has the full answer. If the customer then asks “why can’t it just be right every time?”, the question on explaining why an assistant can’t be right every time has the answer.

Scenario: output that must match a strict format

The setup. A freight company extracts fields from scanned shipping documents into its ERP. The model’s answers are mostly right, but it sometimes renames a field, writes dates in two formats, or adds commentary, and the import fails. The customer says: “It needs to learn our format. Fine-tune it.”

The diagnosis. This is behavior, so fine-tuning is a fair option. But the first fixes are cheaper.

The recommendation. First, constrain the output to the ERP’s schema if the model API supports structured output, validate every response, and retry once on failure. Add a few worked examples covering the awkward documents. Measure field-level accuracy on a held-out set of real documents. If the format errors are gone and the remaining errors are misread values, fine-tuning for format was never the fix. If specific field conventions still fail after that, and the customer has a few thousand corrected extractions, a fine-tune is worth testing against the prompted version on the same set.

The words.

“You’re right that this is about behavior, not knowledge, so fine-tuning is on the table. Before we train anything, I’ll lock the output to your schema and add validation, which should stop the failed imports this week. Then we’ll know exactly which errors are left, and whether a fine-tune is worth the upkeep.”

Scenario: a high-volume task where cost and latency matter

The setup. A marketplace routes every support ticket to one of several queues. A large model with a long prompt of examples does it well, but it runs on forty thousand tickets a day, the monthly bill worries finance, and the agents wait on it. The team already has two years of tickets with the queue a human finally chose.

The diagnosis. Quality is fine. Cost and latency are the problem, the task is narrow and stable, and labeled examples already exist. This is where fine-tuning, or no LLM at all, earns its place.

The recommendation. Build a held-out test set from recent tickets. Compare four options on it: today’s prompted model, a plain classifier trained on embeddings of the ticket text, a smaller model given the most similar past tickets as examples, and a smaller fine-tuned model. Report accuracy per queue, cost per ticket and p95 latency side by side, and ship the cheapest option that holds accuracy on the queues that matter. When asked “why an LLM at all?”, you have the classifier’s numbers. Say the upkeep too: when the base model is retired, you retrain and rerun the same evaluation.

The words.

“The current setup gets the answer right, so we’re not fixing quality, we’re fixing cost and speed. You already have the labeled tickets a fine-tune needs. I’ll test a smaller tuned model against what you run today on the same tickets, and you’ll see accuracy, cost per ticket and response time in one table before anything changes. If a simple classifier matches it, we ship that.”

Scenario: a customer with little data and no evaluation

The setup. A software startup wants its assistant to write sales emails “in our voice”. They have about thirty emails they like and no written definition of what good looks like. The founder says: “Fine-tune it so it sounds like us.”

The diagnosis. Thirty emails is enough to start a training job, but not enough to both train on and keep a held-out set that shows whether it worked. And “sounds like us” is not yet a measurable goal. Like any vague request, it needs scoping with the customer before anyone picks a technique.

The recommendation. Don’t fine-tune. Work with the founder to write a short style guide, turn the good emails into examples in the prompt, and build a small review set: new prompts, scored by the founder against the style guide. Collect every email the team approves or edits. If, a few months on, there are many approved emails and the prompt has stopped improving, revisit fine-tuning then, with an evaluation already in place.

The words.

“With these emails, a fine-tune would copy their surface and miss what makes them yours, and we’d have no way to check. Let’s write down what ‘our voice’ means, put your best emails in the prompt, and have you score a batch each week. Every email you approve becomes training data, so if we do fine-tune later, we’ll do it with proof it helps.”

Mistakes that sink the answer

  • Recommending fine-tuning to add knowledge. A listener who knows the field spots it in one sentence.
  • Picking an option before naming the evaluation. Without a test set, every choice is a guess.
  • Comparing accuracy only. Cost per call, latency and upkeep decide real deployments.
  • Saying “it depends” and stopping. Say what it depends on, then choose.
  • Talking to the customer in the model’s terms. Talk about their documents, their format and their bill, not weights and embeddings.

The words to say no, and what you offer instead

Saying no to fine-tuning is not saying no to the customer: the lesson on saying no, and holding scope covers the skill in depth. The pattern is the same every time: agree with the goal, name the real problem, offer the fix, and leave the door open with a condition.

“I want the same result you do: answers your team can trust. The problem isn’t that the model needs retraining; it’s that it can’t see your latest documents. So I’d start by connecting it to them, which we can test this week. If we find behavior a prompt can’t fix, fine-tuning is the next step, and we’ll have the test set to prove it’s worth it.”

Four moves are in that paragraph:

  1. Agree with the goal, not the method.
  2. Name the problem in their terms: stale documents, failed imports, the monthly bill.
  3. Offer something testable soon, so “no” comes with progress.
  4. State the condition for yes, so fine-tuning stays on the table.

Follow-ups to expect

In an AI round, be ready for these, each with a one-line answer:

  • “Full fine-tune or LoRA?” LoRA trains small adapter weights and leaves the base model frozen, so it is cheaper to train and to store (Hu et al.).
  • “How many examples do you need?” OpenAI’s supervised fine-tuning guide sets a minimum of 10 examples, reports improvements from 50 to 100, and says the right number varies by use case. You also need a separate held-out set to prove it worked.
  • “Why not just use a long context window?” Fine for a small, stable corpus. Once the corpus changes, grows or needs per-user permissions, you want an index.
  • “Can the customer even tune their model?” Check that the approved model and cloud allow fine-tuning before you propose it.

Practice it out loud

Answer the fine-tuning vs RAG vs prompting question aloud in two minutes, using a project you have built, then take the follow-up it asks: “The facts change weekly. What does that rule out?” Then practice the evidence half with how do you know your AI system works?

To test the customer half, run the free practice case: a city whose mayor has already promised an AI review tool. It is “just fine-tune it” at city scale, and the test is whether you find the real bottleneck before you accept the solution. Then work through the AI round’s questions below. Each has a model answer and the follow-ups an interviewer is likely to ask next.

GlossaryRetrieval-augmented generationAnswering with a model that is given passages retrieved from a document collection as context.More on Retrieval-augmented generationGlossaryForward deployed engineerA software engineer who builds and ships production systems inside a customer’s problem and environment, accountable to that customer’s outcome.More on Forward deployed engineerGlossaryAgentA system in which a model chooses steps and tool calls to complete a task, within limits the design sets.More on Agent

Questions people ask

When should you fine-tune instead of using RAG?

Fine-tune to change how a model behaves, such as its format, tone or a narrow task done more cheaply, when you have enough good examples and an evaluation to prove it helps. Use retrieval when the model needs facts it does not have or that change, and when answers must cite their source.

Can you combine fine-tuning and RAG?

Yes. A model can be fine-tuned to follow a format or a domain’s vocabulary and still retrieve current documents when it answers. Add each piece only when your evaluation shows the simpler setup falls short.

Why start with prompting?

It is the cheapest change to make and to undo, and it gives you a baseline. If a better prompt with a few examples meets the acceptance bar, you avoid building a training pipeline and keeping it current.

Keep reading