Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
How do you reduce token usage in CrewAI workflows?
What you need to know
To reduce tokens, you need to know where they go. Every LLM call an agent makes contains roughly:
- the system prompt built from role, goal and backstory;
- the task description,
expected_output, and any context from earlier tasks; - the schemas of all its tools;
- the growing scratchpad: every earlier thought, tool call and tool result in this task;
- any memory or knowledge retrieved for this step.
An agent that takes 8 steps sends the first four items 8 times, and the scratchpad gets longer each step. That is why small repeated items matter.
The main levers
| Lever | Why it helps |
|---|---|
| Short tool outputs | A web page can be 20,000 tokens; the top 5 snippets are 800. Tool output stays in the scratchpad for every later step. |
Prune context | Pass only the tasks this task needs, ideally as a small schema, not the raw upstream text. |
| Short persona | A 300-word backstory in an 8-step task costs about 3,200 tokens for nothing. |
Cap max_iter | The default is 25 steps. Most well-specified tasks finish in 3–6, so set it explicitly. |
| Fewer tools | Every tool's schema is in every request for that agent. |
output_pydantic | A clear target means the agent stops once the fields are filled. |
| Memory off by default | Memory retrieval adds text to prompts, and the unified memory also uses an LLM when saving. |
| Tool cache | Off by default. Crew(cache=True) reuses a tool's result for identical calls in a run. Avoid it for live-data or state-changing tools unless you gate writes with a cache_function. |
Measuring
result = crew.kickoff(inputs={"topic": "EV two-wheelers in India"})print(result.token_usage)print(crew.usage_metrics) # total, prompt, completion, cached tokens, requestsRecord these before and after each change. For a per-agent breakdown, use a tracing integration or an event listener; the totals alone do not tell you which agent is expensive.
Prompt caching helps too
Most providers discount a repeated prompt prefix. If role, goal, backstory and tool schemas stay identical between calls, many of those tokens are billed at a lower "cached" rate. usage_metrics reports cached prompt tokens so you can see it working. Do not put changing values (like a timestamp) at the start of a backstory, or the cache misses.
A real-life example
A market-research crew used about 180,000 tokens per report, roughly ₹45 with its models. A trace showed where:
| Source | Before | After |
|---|---|---|
| Web-page tool results | 110,000 | 18,000 (top 5 snippets, 300 words each) |
| Personas repeated each step | 22,000 | 6,000 (backstories cut to 2 sentences) |
| Writer's context (raw research) | 30,000 | 8,000 (analyst's structured summary only) |
| Everything else | 18,000 | 16,000 |
| Total | 180,000 | 48,000 |
Quality on their 20-case golden set did not change, and the cost fell to about ₹12 per report. The single biggest win was the search tool, which nobody had suspected.
Follow-up questions to expect
- "Does
respect_context_windowsave tokens?" — It stops a crash when the context is too long by summarising history. It is a safety net, not a saving strategy; shorter inputs are better. - "Should you merge agents to save tokens?" — Often yes. Two agents that always run together pay for two personas and a handoff. Merge them unless they need different tools or models.
- "Is YAML config cheaper than Python?" — No, the prompt is the same. YAML just makes the text easier to review.