CrewAI Multi-Agents

Course Content

CrewAI Multi-Agents

9 sections · 53 lessons

How do you control an agent's reasoning style and verbosity?


What you need to know

Knob 1: trace verbosity (for humans watching)

  • verbose=True on an Agent or Crew prints the agent's steps, tool calls and results.
  • output_log_file=True (or a path) on the Crew writes task logs to a file, and a path ending in .json gives JSON.
  • For production, event listeners and tracing integrations capture LLM calls and tool usage without printing.

Knob 2: reasoning behaviour

Python
from crewai import Agent, LLMfrom crewai.agent.planning_config import PlanningConfigscorer = Agent(    role="Lead Fit Scorer",    goal="Score each lead 0-100 against the ideal customer profile",    backstory="You explain every score with evidence.",    llm=LLM(model="openai/gpt-4o-mini", temperature=0.1),    planning_config=PlanningConfig(reasoning_effort="low", max_steps=5),    max_iter=6,    max_execution_time=90,   # seconds    max_rpm=30,)
  • Planning. planning=True gives a short, bounded plan (low effort, one attempt). planning_config=PlanningConfig(...) lets you set reasoning_effort ("low", "medium", "high"), max_steps and a separate planning llm. Higher effort adds LLM calls after steps to check progress and re-plan. The old reasoning=True and max_reasoning_attempts still work but print a deprecation warning.
  • Model settings. temperature on the LLM controls randomness; low values (0 to 0.2) make scoring and extraction more repeatable. For reasoning models, LLM(..., reasoning_effort="low") sets how much hidden thinking the model does.
  • Loop limits. max_iter (default 25) caps steps, max_execution_time caps seconds, max_rpm caps request rate.

Crew-level planning=True is different: before the run, a planner writes a step-by-step plan for every task and adds it to the task descriptions.

Knob 3: output length and shape

"Be concise" in a backstory is weak. Instead:

  • expected_output="Exactly 3 bullets, each under 20 words".
  • output_pydantic=LeadScore so the answer must fit a schema.
  • A guardrail function that rejects outputs over a word limit and sends the reason back.

A real-life example

A B2B SaaS lead-qualification crew scored leads inconsistently: the same company got 62 one day and 78 the next, and each explanation was 400 words long.

Three changes, one per knob:

  1. verbose=True during testing showed the scorer re-searching the web although research was already in its context. The team removed its search tool.
  2. Temperature dropped from the default to 0.1, and planning=True made it list the five scoring criteria before scoring. Repeat runs on 50 leads now differed by at most 5 points.
  3. expected_output became "score, and one reason per criterion under 20 words", with output_pydantic=LeadScore. Explanations shrank to about 80 words, and the CRM could read the fields directly.

In production they set verbose=False and kept output_log_file="logs/scoring.json" for audits.

Follow-up questions to expect

  • "Does verbose=True change the answer?" — No. It only changes what is printed; the prompts and the model's behaviour are the same.
  • "When would you avoid planning?" — On simple, one-tool tasks, where the plan is an extra LLM call that adds cost and latency without improving results.
  • "How do you make an agent more deterministic?" — Low temperature, a tight expected_output, a schema, fewer tools and a low max_iter. Exact repeatability is still not guaranteed with LLMs.