Course Content
Advanced RAG
3 sections · 38 lessons
How does FLARE decide when to fetch external information?
What you need to know
Why retrieve during generation
Standard RAG retrieves once, using the question. That works when the question names everything the answer needs. It fails for long outputs. Ask for "a report on stability-testing rules in three regions", and paragraph four is about a specific guideline the question never mentioned. The first retrieval could not have known to fetch it.
The loop, sentence by sentence
- Draft — generate the next sentence without retrieving.
- Check — look at each token's probability. If all are above the threshold (say 0.4), accept the sentence.
- Build a query — if any token is below it, mask the unsure tokens, or ask the model to write a question about them, and use that as the search query.
- Retrieve and regenerate — fetch documents and rewrite that sentence with them in context.
- Repeat until the answer is done.
Two ideas make it work. First, the query is forward-looking: it is built from the upcoming sentence, which is far more specific than the original question. Second, retrieval is conditional: confident sentences cost nothing extra.
The paper had two variants. FLARE-direct uses token probabilities as described. FLARE-instruct prompts the model to write explicit search queries inside its output when it needs facts.
Costs and limits
- You need log-probabilities. Not every API exposes them, and reasoning models often don't. Without them, only the instruct-style variant is possible.
- Latency multiplies. Worst case is one retrieval plus one regeneration per sentence. Streaming also becomes harder, because a sentence may be rewritten.
- Confidence is not correctness. A model can be confidently wrong, so FLARE catches uncertainty, not every error.
How the idea shows up today
The spirit of FLARE — retrieve when you discover you need something — now usually appears as agentic search with tool calling. You give the model a search tool, and it decides mid-answer to call it, often several times. This works with any modern API and needs no log-probabilities. Knowing FLARE helps you explain why that design works and what it costs.
A real-life example
A pharma company's regulatory team asks its assistant: "Draft a two-page summary of stability-testing requirements for a new film-coated tablet in the EU, US and India."
Up-front retrieval uses the question and fetches the general ICH stability guideline — good for the first paragraph. When the draft reaches India, the model writes: "For zone IVb markets, long-term testing must be conducted at ___ °C and ___% relative humidity." The tokens for the numbers come out with low probability.
FLARE masks the unsure numbers and searches with "zone IVb long-term stability testing conditions India". It retrieves the internal SOP that lists the climatic-zone conditions and rewrites the sentence with the correct values and a citation. The confident sentences about the EU section trigger no retrieval at all. Over a 20-report test, the team counts how many retrievals were triggered per report and how many cited values matched the source, and decides the extra latency — about a minute per report — is acceptable for a document that is reviewed for hours anyway.
Follow-up questions to expect
- "How do you pick the confidence threshold?" — Sweep it on a small set: lower thresholds retrieve less and risk more guesses; higher thresholds retrieve on nearly every sentence. Choose by accuracy and latency together.
- "Why not just retrieve more up front?" — For long outputs, you cannot know the later topics in advance, and fetching everything floods the context with distractors.
- "Would you use FLARE for a chatbot?" — Rarely. For two-sentence answers, one retrieval before generation is simpler and faster.