In this post12 sections
- What OpenAI says about take-homes and AI tools
- What candidates report: one FDE account and two looser ones
- One routine for any format: scope, build, evaluate, defend
- Scope: write the brief back before you build
- Build: a thin working version first
- Evaluate: show how you know it works
- Defend: the live discussion of your submission
- The submission checklist
- Practice the defend step before it counts
- Questions people ask
- Keep reading
- More from the blog
You have an OpenAI process starting, and the first question in your head is what the take-home will look like. Search for it and you get a jumble: one FDE candidate’s week-long build, a five-hour build with a video that someone expected for a sister role, and a design doc from a commenter who never says which company sent it. OpenAI does not publish an FDE take-home format, and its general interview guide says formats vary by team, so the useful move is to prepare one routine that works for all three: scope, build, evaluate, defend. This post walks through that routine, with the words to use and a checklist for the night you submit. For the whole loop, round by round, start with the FDE interview guide.
What OpenAI says about take-homes and AI tools
OpenAI’s interview guide is a general one, for all its roles, and it presents the process as an example of what to expect. It says skills-based assessment formats vary by team and may include pair coding interviews, take-home projects or technical tests, and that you may be asked to complete more than one assessment depending on the role. Source 1Interview Guide | OpenAIPublisherOpenAISource typecompany hiring page
Read that carefully. It tells you a take-home is possible. It does not tell you the length, the deliverable, or whether a take-home is part of any given FDE loop.
On AI tools, the same guide says expectations vary by interview: some formats intentionally allow them, and others are designed to assess your problem-solving without AI tools. It says your interview preparation materials explain what is allowed, and to ask your recruiter before the interview if you are unsure. Source 1Interview Guide | OpenAIPublisherOpenAISource typecompany hiring pageSource 2OpenAI interview guide (Wayback Machine capture)PublisherOpenAI (archived by Internet Archive)Source typearchived company page
That sentence is recent. It is absent from the archived copies of the guide we checked from February to July 2026, and present in the August 2026 copy. Source 2OpenAI interview guide (Wayback Machine capture)PublisherOpenAI (archived by Internet Archive)Source typearchived company pageSource 3OpenAI interview guide (Wayback Machine capture)PublisherOpenAI (archived by Internet Archive)Source typearchived company page So a write-up from earlier in the year may describe AI rules that predate it. Go by your own prep materials.
The guide also says what OpenAI generally looks for in engineering interviews: well-designed solutions, high-quality code, optimal performance and good test coverage, and that it evaluates communication and collaboration. Source 1Interview Guide | OpenAIPublisherOpenAISource typecompany hiring pageSource 2OpenAI interview guide (Wayback Machine capture)PublisherOpenAI (archived by Internet Archive)Source typearchived company page The guide says this of engineering interviews in general and doesn’t say how take-homes are graded. Our reading: they are the closest thing to a published standard, so give a reviewer evidence for each one.
The job posting adds a line worth reading. As of September 2026, OpenAI’s Forward Deployed Engineer posting for San Francisco says FDEs measure success through production adoption, measurable workflow impact, and eval-driven feedback that changes product and model roadmaps. Source 4Forward Deployed Engineer (FDE) - SF | OpenAIPublisherOpenAISource typecompany job posting Our reading: a take-home for this job should show how you know the thing works, not only that it runs.
What candidates report: one FDE account and two looser ones
Here is who reported what, and when.
A week-long build (an OpenAI FDE role, onsite reached). One candidate’s report on Aced, for an OpenAI Forward Deployed Engineer role with the interview month listed as May 2026, says the main stages were a one-week take-home to build semantic search over Amazon products for ChatGPT, a live team discussion of that case study, and a separate AI-enabled coding screen. Source 5OpenAI Forward Deployed Engineer Interview ExperiencePublisherAced (formerly Exponent)Source typecandidate’s personal write-up The candidate says they reached the onsite stage.
Five hours with a video (a sister role, before any round). In March 2026, a Blind commenter with an HR call scheduled for OpenAI’s AI Deployment Engineer role wrote that the process had a take-home project of about 5 hours including a video walkthrough of the solution, then a virtual onsite. Source 6Anthropic/OpenAI FDE Interview loop -- same as SWE?PublisherBlind (teamblind.com)Source typecandidate report on Blind They do not say they had completed any round, so this is an expectation, and the role is not FDE.
A design doc (company and role unnamed). On a Blind thread whose original poster asked how to prepare for OpenAI’s FDE role, a commenter wrote, in June 2026, that “they” sent a take-home that was not coding but asked for a design doc. Source 7What to prepare for a Forward Deployed Engineer role at OpenAI?PublisherBlindSource typecandidate report on Blind The commenter does not name the company or the role.
An anonymous write-up on gaijineer.co, a blog that sells interview prep, describes an OpenAI FDE take-home of about five hours of work building something with OpenAI’s APIs. Source 8OpenAI Forward Deployed Engineer Interview ProcessPublishergaijineer.co (a blog that promotes the furustack interview-prep product)Source typeprep-product blog (anonymous, may be compiled) Treat it as the weakest of the four.
What you can safely conclude is small: a handful of people describe different shapes, at different times, for roles that are not always the same. None of it tells you what your recruiter will send, and none of it says how often any format is used. The evidence map lesson shows how to sort any company’s facts this way: what the employer publishes, what candidates report (who, when, which role), and what nobody knows.
What the reports do share is more useful than their differences. Each format asks you to take a loose problem, decide what to build or propose, show that it works, and then explain your choices to people who will push on them. That is the routine.
One routine for any format: scope, build, evaluate, defend
This is our method, not a rubric OpenAI publishes. It is built from the guide’s own words (solutions, quality, performance, tests, communication) and from the posting’s emphasis on evals.
| Step | Code | Design doc |
|---|---|---|
| Scope | Brief in README | Problem + non-goals |
| Build | Thin end-to-end build | Simplest architecture |
| Evaluate | Labeled set + score | Eval plan + ship bar |
| Defend | Video, then live Q&A | Trade-offs, then live Q&A |
The steps keep their order in every format; only the time each gets changes. With a week, the build step grows, but a week is also enough rope to over-build, so scope matters more, not less.
How to split the time
Our suggested split, not anything OpenAI publishes.
Five hours:
- Thirty minutes: scope, and send your questions
- Two hours: the thin end-to-end skeleton
- One hour: build and score the eval set
- Forty-five minutes: fix the worst miss
- Forty-five minutes: README and video
A week:
- Day one: scope, questions and the skeleton
- Days two to four: improve against the eval
- Day five: read the misses and write them up
- Day six: README and video
- Day seven: buffer, and a clean-clone test
Scope: write the brief back before you build
Before any code, write the problem back in your own words. Put it at the top of your README or doc.
Take a prompt like the one a candidate described for an interview in May 2026: semantic search over a product catalog for a chat assistant. Source 5OpenAI Forward Deployed Engineer Interview ExperiencePublisherAced (formerly Exponent)Source typecandidate’s personal write-up Here is a filled-in brief for it:
- Goal: a shopper types a plain-language request and gets up to five real products with title, price and link.
- User: a shopper in a chat assistant, often with a constraint (price, size).
- In scope: a sample of the catalog, one model, top-five retrieval, a twenty-query labeled set.
- Out of scope: personalization, margin re-ranking, a web UI (a CLI shows the same thing).
- Assumptions: the catalog is static; with daily updates I’d re-embed changed items only.
- Success: hit rate at five and MRR on the labeled set, plus one line per miss.
A weak brief says “build semantic search”. This one tells a reviewer what you built, what you skipped and how to judge it.
The assumptions line is where your judgment shows most. A reviewer can disagree with an assumption and still credit you for making it visible.
Send your questions early. If the prep materials leave something open, write to your recruiter the day the prompt arrives:
“Thanks for the take-home. Two quick questions so I scope it right: are AI coding assistants allowed for this assessment, and is a video walkthrough expected with the submission? If there’s no preference, I’ll disclose any AI use in the README.”
That message follows OpenAI’s own advice to ask when unsure. Source 1Interview Guide | OpenAIPublisherOpenAISource typecompany hiring pageSource 2OpenAI interview guide (Wayback Machine capture)PublisherOpenAI (archived by Internet Archive)Source typearchived company page For how other employers word their AI rules, and how to disclose your use, see the post on using AI on a take-home.
Build: a thin working version first
Build the whole path end to end before you make any part of it good. For the search example that means: load a slice of the catalog, embed it, answer one query, and return results through whatever interface the prompt asks for. Ugly is fine. Working is the point.
The walking skeleton lesson covers this pattern in depth. In a take-home it protects you from an easy failure: a beautiful ingestion pipeline and no way to run a query when time runs out.
Once the skeleton runs, spend the rest of the build on whatever your eval shows is weakest. If you want a warm-up before the real thing, build TF-IDF retrieval with no libraries: it is the same retrieve-and-rank loop at a size you can finish in one sitting.
Mistakes to avoid:
- Swapping vector stores in hour two. Use an in-memory array of embeddings; a few thousand products fit in memory.
- Embedding the whole catalog first. Start with a sample, get one query to answer end to end, then scale.
- A React front end. A CLI that prints the top five with scores shows everything the reviewer needs.
- Hiding the rough edges. List known gaps in the README. Finding them yourself reads as judgment; a reviewer finding them reads as a miss.
For a design-doc prompt, the build step is the architecture: one diagram, the data flow from input to answer, and the few components you would build first. Resist listing every service you know.
Evaluate: show how you know it works
This is the step that is easiest to drop when time runs short, and the one the posting’s language points at. Source 4Forward Deployed Engineer (FDE) - SF | OpenAIPublisherOpenAISource typecompany job posting A working demo shows it runs on the queries you chose. An evaluation shows how it behaves on queries you did not tune for.
Keep it small and honest in form: a short list of realistic queries, the products a person would call correct for each, and a score. Here is the whole harness for the search example:
def hit_rate(runs, gold, k):
pairs = zip(runs, gold)
s = [bool(set(r[:k]) & g)
for r, g in pairs]
return sum(s) / len(s)
def rr(r, g):
for n, doc in enumerate(r):
if doc in g:
return 1 / (n + 1)
return 0.0
def mrr(runs, gold):
pairs = zip(runs, gold)
s = [rr(r, g)
for r, g in pairs]
return sum(s) / len(s)
# Top 3 for three queries, and
# the products marked relevant.
runs = [["p7", "p2", "p9"],
["p4", "p1", "p3"],
["p8", "p6", "p5"]]
gold = [{"p2"}, {"p4"}, {"p0"}]
hr = hit_rate(runs, gold, 3)
print(round(hr, 2))
print(round(mrr(runs, gold), 2))
# 0.67, then 0.5
Hit rate says how often a correct product shows up near the top at all. Mean reciprocal rank (MRR) rewards putting it first. Report both, then do the part that matters more than either number: read the misses. The third query above found nothing relevant. Why? A typo, a price cap that similarity search can’t enforce, a product missing from the index? Write one line per miss.
That short failure table is what turns a demo into an engineering submission:
| Query type | What went wrong | Next fix |
|---|---|---|
| Price limit (“under a budget”) | Similarity can’t enforce a price cap | Parse the limit, filter first |
| Brand misspelled | No close match | Add fuzzy match on brand |
For a design-doc prompt, write the evaluation plan instead: the test set you would build, the metric, the threshold at which you would ship, and what you would monitor after launch. How do you know your AI system works? is the same question asked out loud, with a framework for the answer.
Defend: the live discussion of your submission
One candidate reported, for May 2026, a live team discussion of the submission, and a Blind commenter wrote, in March 2026, of a video walkthrough, so plan for both. Source 5OpenAI Forward Deployed Engineer Interview ExperiencePublisherAced (formerly Exponent)Source typecandidate’s personal write-upSource 6Anthropic/OpenAI FDE Interview loop -- same as SWE?PublisherBlind (teamblind.com)Source typecandidate report on Blind
For the video, lead with the result, not the code: what it does, one live query, the eval score, the biggest gap. The take-home walkthrough video post has a script for it.
For the live discussion, expect questions like these, and have a sentence ready for each:
- “Why this approach and not the obvious alternative?” Say what you chose, what it cost, and when you would switch: “I used embeddings alone because the catalog is small and static. The cost is that price constraints fail, as the eval shows. With more time I’d add a structured filter before ranking.” Practice it with defend your take-home design.
- “What would you change with another day?” Name one fix, tied to a failure you measured, and one thing you left out on purpose. What would you change in your take-home? has a model answer.
- “How would this hold up with a real customer?” Talk about what breaks first: latency at larger catalogs, stale data, queries in other languages, and how you would find out.
- “How did you use AI tools?” If they were allowed, say exactly where: “I used an assistant for the boilerplate and the test scaffolding. The retrieval design, the eval set and the failure analysis are mine.”
The trap in this conversation is defending everything. When a reviewer finds a real weakness, agree, say how you would fix it, and move on. That shows the communication and collaboration the guide says it evaluates in engineering interviews. Source 1Interview Guide | OpenAIPublisherOpenAISource typecompany hiring pageSource 2OpenAI interview guide (Wayback Machine capture)PublisherOpenAI (archived by Internet Archive)Source typearchived company page
The submission checklist
Run this the night before you submit, whatever the format.
Before you submit
- The brief back is at the top: goal, user, scope, out of scope, assumptions, success
- It runs from a clean clone with the commands in the README, and you have tried them
- Secrets and API keys are out of the repo, and the README says which variables to set
- There is a test set, a score, and a short list of the misses with a reason for each
- Tests cover the core path (the guide names test coverage for engineering interviews)
- Known gaps and the next fixes are written down before a reviewer finds them
- AI use is disclosed in the way your prep materials or recruiter asked
- The video, if asked for, leads with the result and fits the time you were given
- You can explain every file, including any code an assistant wrote
- You have said out loud why you chose your approach over the obvious alternative
For a design doc, swap the first four items for: problem and non-goals up top, one architecture diagram, the evaluation plan with a ship threshold, and a trade-offs section that names what you rejected. The broader post on FDE take-homes covers what other companies’ prompts ask for, if your process includes more than one.
Practice the defend step before it counts
The build is the part you can do alone. The defend step is the part that goes wrong under pressure, because a person is asking follow-ups you did not plan for. The free practice case puts you in front of an AI customer with a vague problem on a 10-minute clock: you scope it, propose a first version and say how you’d know it works, then a scorecard shows what you asked and what you missed. It is free with a sign-in. Do one run before the prompt lands, so the first brief back you write isn’t the one that counts.
Questions people ask
How long is the OpenAI FDE take-home, and what have candidates reported?
OpenAI does not publish an FDE take-home length. Its general interview guide, written for all roles, says formats vary by team and may include take-home projects. Reports differ: a Blind commenter with an HR call scheduled for OpenAI’s AI Deployment Engineer role, who did not say they had taken any round, wrote in March 2026 that the process had a take-home project of about 5 hours including a video walkthrough, and one candidate’s report for an OpenAI FDE role, interview month May 2026, described a one-week take-home.Source 1Interview Guide | OpenAIPublisherOpenAISource typecompany hiring pageSource 5OpenAI Forward Deployed Engineer Interview ExperiencePublisherAced (formerly Exponent)Source typecandidate’s personal write-upSource 6Anthropic/OpenAI FDE Interview loop -- same as SWE?PublisherBlind (teamblind.com)Source typecandidate report on Blind
Can I use ChatGPT or other AI tools on the OpenAI take-home?
It depends on the assessment. OpenAI’s interview guide says expectations for AI tools vary by interview: some formats allow them and others are designed to assess problem-solving without them. It says the preparation materials explain what is allowed, and to ask your recruiter if you are unsure.Source 1Interview Guide | OpenAIPublisherOpenAISource typecompany hiring pageSource 2OpenAI interview guide (Wayback Machine capture)PublisherOpenAI (archived by Internet Archive)Source typearchived company page
What does OpenAI look for in engineering interviews?
OpenAI’s interview guide says that for engineering interviews it generally looks for well-designed solutions, high-quality code, optimal performance and good test coverage, and that it evaluates communication and collaboration.Source 1Interview Guide | OpenAIPublisherOpenAISource typecompany hiring pageSource 2OpenAI interview guide (Wayback Machine capture)PublisherOpenAI (archived by Internet Archive)Source typearchived company page
Keep reading
Company guides
Lessons
Questions
- Looking at your take-home submission, what would you change with another day, and what did you leave out on purpose?
- A customer asks how you know your AI system works. Answer them.
- In your take-home you chose one approach over the obvious alternative. Defend that choice as if I were the customer’s architect.
- Build retrieval over a folder of text files using TF-IDF with no external libraries. Return the top passages with scores.
More from the blog
Company loops
Cognition’s deployed engineer take-home: what candidates built with the Devin API
What candidates built for the Cognition take-home, the two Devin API calls you need, and how to scope, guard and present the integration.
Company loops
C3 AI’s AI screening interview and the coding rounds candidates describe
Two candidates report C3 AI’s FDE process opening with an AI interview, then coding and system design. What it asks and how to prepare for each step.
Company loops
Databricks AI FDE coding round: preparing for applied data science and machine learning
What Databricks publishes, what one candidate was told about the AI FDE coding round, and a prep plan: a baseline model, leakage checks and trade-offs.