CrewAI Multi-Agents

Course Content

CrewAI Multi-Agents

9 sections · 53 lessons

How do you reduce token usage in CrewAI workflows?


What every step of one agent re-sendsRole, goal,backstoryTaskdescription and contextAll tool schemasScratchpad:every tool resultRetrieved memorytopbottomAn 8-step task pays for the persona and tool schemas 8 times; tool results stay for every later step.
The biggest savings come from what repeats — long tool outputs and personas — not from the task description.

What you need to know

To reduce tokens, you need to know where they go. Every LLM call an agent makes contains roughly:

  • the system prompt built from role, goal and backstory;
  • the task description, expected_output, and any context from earlier tasks;
  • the schemas of all its tools;
  • the growing scratchpad: every earlier thought, tool call and tool result in this task;
  • any memory or knowledge retrieved for this step.

An agent that takes 8 steps sends the first four items 8 times, and the scratchpad gets longer each step. That is why small repeated items matter.

The main levers

LeverWhy it helps
Short tool outputsA web page can be 20,000 tokens; the top 5 snippets are 800. Tool output stays in the scratchpad for every later step.
Prune contextPass only the tasks this task needs, ideally as a small schema, not the raw upstream text.
Short personaA 300-word backstory in an 8-step task costs about 3,200 tokens for nothing.
Cap max_iterThe default is 25 steps. Most well-specified tasks finish in 3–6, so set it explicitly.
Fewer toolsEvery tool's schema is in every request for that agent.
output_pydanticA clear target means the agent stops once the fields are filled.
Memory off by defaultMemory retrieval adds text to prompts, and the unified memory also uses an LLM when saving.
Tool cacheOff by default. Crew(cache=True) reuses a tool's result for identical calls in a run. Avoid it for live-data or state-changing tools unless you gate writes with a cache_function.

Measuring

Python
result = crew.kickoff(inputs={"topic": "EV two-wheelers in India"})print(result.token_usage)print(crew.usage_metrics)   # total, prompt, completion, cached tokens, requests

Record these before and after each change. For a per-agent breakdown, use a tracing integration or an event listener; the totals alone do not tell you which agent is expensive.

Prompt caching helps too

Most providers discount a repeated prompt prefix. If role, goal, backstory and tool schemas stay identical between calls, many of those tokens are billed at a lower "cached" rate. usage_metrics reports cached prompt tokens so you can see it working. Do not put changing values (like a timestamp) at the start of a backstory, or the cache misses.

A real-life example

A market-research crew used about 180,000 tokens per report, roughly ₹45 with its models. A trace showed where:

SourceBeforeAfter
Web-page tool results110,00018,000 (top 5 snippets, 300 words each)
Personas repeated each step22,0006,000 (backstories cut to 2 sentences)
Writer's context (raw research)30,0008,000 (analyst's structured summary only)
Everything else18,00016,000
Total180,00048,000

Quality on their 20-case golden set did not change, and the cost fell to about ₹12 per report. The single biggest win was the search tool, which nobody had suspected.

Follow-up questions to expect

  • "Does respect_context_window save tokens?" — It stops a crash when the context is too long by summarising history. It is a safety net, not a saving strategy; shorter inputs are better.
  • "Should you merge agents to save tokens?" — Often yes. Two agents that always run together pay for two personas and a handoff. Merge them unless they need different tools or models.
  • "Is YAML config cheaper than Python?" — No, the prompt is the same. YAML just makes the text easier to review.