Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 1: Inconsistent Multi-Agent Output


What you need to know

The scenario: the same crew, given the same input, produces reports with different structure, missing sections or different conclusions from run to run.

Where the variance comes from

A CrewAI task has a description and an expected_output. If expected_output says "a detailed research summary", the model decides each time what "detailed" means, which fields to include and in what order. The next agent then receives a different shape each run, and the variation compounds down the crew.

Fix the contract

Python
from pydantic import BaseModel, Fieldfrom crewai import Task, TaskOutputclass Findings(BaseModel):    claims: list[str] = Field(min_length=3)    sources: list[str]    confidence: float = Field(ge=0, le=1)def every_claim_sourced(result: TaskOutput):    f = result.pydantic    if len(f.sources) < len(f.claims):        return (False, "Each claim needs a source URL or document ID.")    return (True, result)research = Task(    description="Research {company}'s EV battery suppliers.",    expected_output="Findings with at least 3 claims, each with a source.",    output_pydantic=Findings,    guardrail=every_claim_sourced,    agent=researcher,)

A malformed result now fails at its own task, not three tasks later.

Three more variance cuts

ChangeWhy it helps
Temperature 0 for agents feeding code or other agentsLess sampling randomness; keep higher temperature only for creative tasks
A dedicated assembler taskDeciding content and formatting in one task is a big source of variation
Pin the model versionProvider aliases move, and the crew inherits the change

Measure it

  1. Pick 10 real inputs.
  2. Run each 10 times.
  3. Track — schema validity rate and agreement on key fields (for example, the recommendation).
  4. Gate — run this in CI, so you can say "this change cut variance" with numbers.

A real-life example

Scenario, numbers made up. A consulting firm's crew of three agents writes supplier-risk reports. Run ten times on the same company, it produces four different section layouts, and the final risk rating differs in three runs.

The team adds Pydantic outputs to both upstream tasks, a guardrail requiring a source per claim, temperature 0 for the researcher and analyst, and an assembler task that only formats. Over ten runs on ten companies, schema validity rises from 71% to 100% and the risk rating agrees in 96 of 100 runs. The four disagreements are borderline cases the analyst flags with low confidence, which is useful rather than noisy.

Follow-up questions to expect

  • "Why not just write a more detailed expected_output?" — It helps, but prose is still interpreted; a schema is enforced.
  • "Won't temperature 0 make writing dull?" — Only use it for agents that produce data. The final writer can keep some temperature, working from validated inputs.
  • "How many retries should a guardrail allow?" — Two or three; if it still fails, surface the error rather than looping.