Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 4: Missing Context Handoff
What you need to know
The scenario: the final agent produces an answer missing details that an earlier agent definitely found — for example, the contract end date found by the researcher never reaches the writer.
How context flows in a crew
By default, in a sequential process, each task receives the output of the task just before it. If task C needs something from task A, it only gets it if B happened to repeat it. And each output is usually a summary, so specifics like dates, IDs and amounts are the first things lost.
Three fixes
- Type the handoff —
output_pydanticon every task whose output another task uses. - Wire context explicitly —
context=[task_a, task_c]lists everything the consuming task needs, not just the previous task. - Validate at the boundary — a guardrail checks that required fields are present and non-empty, and retries if not.
1class ContractFacts(BaseModel):2 counterparty: str3 end_date: date4 renewal: Literal["auto", "manual", "none"]5 notice_days: int | None # None means "not stated", not "forgotten"67def facts_complete(result: TaskOutput):8 f = result.pydantic9 if f.renewal == "auto" and f.notice_days is None:10 return (False, "Auto-renewal found but notice period missing; check the termination clause.")11 return (True, result)1213extract = Task(description="Extract key facts from {contract_id}.", expected_output="ContractFacts",14 output_pydantic=ContractFacts, guardrail=facts_complete, agent=extractor)15advise = Task(description="Advise whether to renew.", expected_output="A recommendation",16 agent=advisor, context=[extract, pricing]) # both, explicitlyLarge artifacts: pass references
A scraped website or a 60-page document should not travel through the prompt chain. Write it to a shared store and pass an ID; the next agent gets a tool to read the parts it needs. Passing 20,000 tokens between agents is both lossy and expensive.
Where crew memory fits
memory=True gives agents retrieval over earlier results, which helps recall. It is not a reliable data channel: it has its own miss rate. Use explicit, typed context for anything correctness depends on.
The metric
Schema validation failures per handoff. That one number shows which boundary is broken, which is usually the hardest part of debugging.
A real-life example
Scenario, numbers made up. A procurement crew has an extractor, a pricing analyst and an advisor. In 14% of runs, the advisor recommends renewing contracts that auto-renew with a 90-day notice period already passed — the notice period was found by the extractor but summarised away by the pricing analyst.
The team adds the ContractFacts schema, passes the extractor's task directly to the advisor through context, and adds a completeness guardrail. Wrong renewal advice on a 200-contract test set falls from 14% to 1%, and the guardrail fires on 6% of runs, each time prompting the extractor to look again at the termination clause.
Follow-up questions to expect
- "Why not pass everything to every task?" — Longer prompts cost more and distract the model; pass exactly what each task needs.
- "What if the upstream fact truly isn't in the document?" — Make that representable (
Nonewith a reason) so "not stated" is different from "forgot to extract". - "Is a hierarchical process better for context?" — It changes who assigns work, not how reliably facts are passed; typed handoffs are still needed.