Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 1: Inconsistent Multi-Agent Output
What you need to know
The scenario: the same crew, given the same input, produces reports with different structure, missing sections or different conclusions from run to run.
Where the variance comes from
A CrewAI task has a description and an expected_output. If expected_output says "a detailed research summary", the model decides each time what "detailed" means, which fields to include and in what order. The next agent then receives a different shape each run, and the variation compounds down the crew.
Fix the contract
1from pydantic import BaseModel, Field2from crewai import Task, TaskOutput34class Findings(BaseModel):5 claims: list[str] = Field(min_length=3)6 sources: list[str]7 confidence: float = Field(ge=0, le=1)89def every_claim_sourced(result: TaskOutput):10 f = result.pydantic11 if len(f.sources) < len(f.claims):12 return (False, "Each claim needs a source URL or document ID.")13 return (True, result)1415research = Task(16 description="Research {company}'s EV battery suppliers.",17 expected_output="Findings with at least 3 claims, each with a source.",18 output_pydantic=Findings,19 guardrail=every_claim_sourced,20 agent=researcher,21)A malformed result now fails at its own task, not three tasks later.
Three more variance cuts
| Change | Why it helps |
|---|---|
| Temperature 0 for agents feeding code or other agents | Less sampling randomness; keep higher temperature only for creative tasks |
| A dedicated assembler task | Deciding content and formatting in one task is a big source of variation |
| Pin the model version | Provider aliases move, and the crew inherits the change |
Measure it
- Pick 10 real inputs.
- Run each 10 times.
- Track — schema validity rate and agreement on key fields (for example, the recommendation).
- Gate — run this in CI, so you can say "this change cut variance" with numbers.
A real-life example
Scenario, numbers made up. A consulting firm's crew of three agents writes supplier-risk reports. Run ten times on the same company, it produces four different section layouts, and the final risk rating differs in three runs.
The team adds Pydantic outputs to both upstream tasks, a guardrail requiring a source per claim, temperature 0 for the researcher and analyst, and an assembler task that only formats. Over ten runs on ten companies, schema validity rises from 71% to 100% and the risk rating agrees in 96 of 100 runs. The four disagreements are borderline cases the analyst flags with low confidence, which is useful rather than noisy.
Follow-up questions to expect
- "Why not just write a more detailed
expected_output?" — It helps, but prose is still interpreted; a schema is enforced. - "Won't temperature 0 make writing dull?" — Only use it for agents that produce data. The final writer can keep some temperature, working from validated inputs.
- "How many retries should a guardrail allow?" — Two or three; if it still fails, surface the error rather than looping.