Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
How do you control an agent's reasoning style and verbosity?
What you need to know
Knob 1: trace verbosity (for humans watching)
verbose=Trueon anAgentorCrewprints the agent's steps, tool calls and results.output_log_file=True(or a path) on theCrewwrites task logs to a file, and a path ending in.jsongives JSON.- For production, event listeners and tracing integrations capture LLM calls and tool usage without printing.
Knob 2: reasoning behaviour
Python
1from crewai import Agent, LLM2from crewai.agent.planning_config import PlanningConfig34scorer = Agent(5 role="Lead Fit Scorer",6 goal="Score each lead 0-100 against the ideal customer profile",7 backstory="You explain every score with evidence.",8 llm=LLM(model="openai/gpt-4o-mini", temperature=0.1),9 planning_config=PlanningConfig(reasoning_effort="low", max_steps=5),10 max_iter=6,11 max_execution_time=90, # seconds12 max_rpm=30,13)- Planning.
planning=Truegives a short, bounded plan (low effort, one attempt).planning_config=PlanningConfig(...)lets you setreasoning_effort("low","medium","high"),max_stepsand a separate planningllm. Higher effort adds LLM calls after steps to check progress and re-plan. The oldreasoning=Trueandmax_reasoning_attemptsstill work but print a deprecation warning. - Model settings.
temperatureon theLLMcontrols randomness; low values (0 to 0.2) make scoring and extraction more repeatable. For reasoning models,LLM(..., reasoning_effort="low")sets how much hidden thinking the model does. - Loop limits.
max_iter(default 25) caps steps,max_execution_timecaps seconds,max_rpmcaps request rate.
Crew-level planning=True is different: before the run, a planner writes a step-by-step plan for every task and adds it to the task descriptions.
Knob 3: output length and shape
"Be concise" in a backstory is weak. Instead:
expected_output="Exactly 3 bullets, each under 20 words".output_pydantic=LeadScoreso the answer must fit a schema.- A guardrail function that rejects outputs over a word limit and sends the reason back.
A real-life example
A B2B SaaS lead-qualification crew scored leads inconsistently: the same company got 62 one day and 78 the next, and each explanation was 400 words long.
Three changes, one per knob:
verbose=Trueduring testing showed the scorer re-searching the web although research was already in its context. The team removed its search tool.- Temperature dropped from the default to 0.1, and
planning=Truemade it list the five scoring criteria before scoring. Repeat runs on 50 leads now differed by at most 5 points. expected_outputbecame "score, and one reason per criterion under 20 words", withoutput_pydantic=LeadScore. Explanations shrank to about 80 words, and the CRM could read the fields directly.
In production they set verbose=False and kept output_log_file="logs/scoring.json" for audits.
Follow-up questions to expect
- "Does
verbose=Truechange the answer?" — No. It only changes what is printed; the prompts and the model's behaviour are the same. - "When would you avoid planning?" — On simple, one-tool tasks, where the plan is an extra LLM call that adds cost and latency without improving results.
- "How do you make an agent more deterministic?" — Low temperature, a tight
expected_output, a schema, fewer tools and a lowmax_iter. Exact repeatability is still not guaranteed with LLMs.