Reranking re-scores a short list of retrieved candidates with a slower, more accurate model before the best few go into the prompt. The usual reranker is a cross-encoder, which reads the query and a passage together and outputs a relevance score, as in Nogueira and Cho, Passage Re-ranking with BERT; unlike embeddings, its scores cannot be computed ahead of time, so it runs once per candidate per query.
In an FDE interview
One candidate reported, in a public repo created in August 2026 that they describe as a ‘NewPage Solutions AI take-home assignment’, building a chat-with-your-docs assistant that retrieves with embeddings, a vector index and a reranker. Source 1newpage-fde-takehome (README)PublisherGitHub user cagkanacarbaySource typecandidate’s take-home repository
A strong candidate places the reranker after a cheap first stage (top_k = 50 in, top_n = 5 out, say), notes that it cannot recover a passage the first stage never returned, and measures retrieval at the cutoff actually sent to the prompt (recall@5, nDCG@5) with and without it. They state its cost in added latency per query and keep it only if it moves the metric.
A reranker is also the practical way to say “nothing relevant found”: set a score threshold on labeled queries, because the raw scores are not calibrated probabilities, and below it the system says it has no answer. A cross-encoder reads a limited number of tokens (512 for a BERT-based model), so rerank the chunk, not the whole document.