CrewAI Multi-Agents

Course Content

CrewAI Multi-Agents

9 sections · 53 lessons

How do you monitor cost across a running crew?


Cost per claim, one night's sample, in rupees1415131601416012345fraud agentlooping to max_iterThe average moved from 14 to 16; the p95 moved from 30 to 160.
Loops live in the tail, so an alert on p95 tokens per run finds them weeks before the monthly bill does.

What you need to know

Cost has two parts: how many tokens each run uses, and how many runs there are. Traffic growth raises total spend, which is fine. Tokens per run rising is the warning sign, because it means something in the crew changed.

Where the numbers come from

Python
result = crew.kickoff(inputs=inputs)usage = crew.usage_metrics   # total_tokens, prompt_tokens, completion_tokens,                             # cached_prompt_tokens, successful_requestscost_inr = (usage.prompt_tokens * PRICE_IN + usage.completion_tokens * PRICE_OUT)metrics.record("crew_cost_inr", cost_inr,               tags={"crew": "claims", "run_id": RUN_ID, "model": MODEL})

If agents use different models, one total is not enough, because prices differ. Use per-call data from a tracing tool (Langfuse, Datadog, Opik, CrewAI's built-in tracing and others show tokens and cost per span) or an event listener on LLM calls, tagged with the agent's role.

What to watch

  • p95 and p99 tokens per run. One looping run can use 20 times the normal tokens. The average hides it; the tail shows it.
  • Cost per agent. One agent is often most of the bill, and it is often not the one you expected.
  • Steps per task. More steps means more calls; it is an early signal of a loop.
  • Cached token share. If it drops, a prompt change broke provider caching.

Limits, not just dashboards

  • max_iter and max_execution_time on agents.
  • max_rpm on the crew for rate limits.
  • A @before_llm_call hook can return False to block further calls, for example when a run's running token count crosses a budget you track in an event listener.

A real-life example

An insurance-claims review crew processes about 2,000 claims a night. Average cost is ₹14 per claim, steady for weeks. Then the p95 goes from ₹30 to ₹160 while the average moves only to ₹16.

Per-agent traces show the fraud agent hitting its 25-step limit on some claims. A new photo-analysis tool returned an error string the agent did not understand, so it kept retrying with small changes until max_iter. The team fixed the tool's error message, lowered the fraud agent's max_iter from the default 25 to 8, and added a budget hook that stops any claim run above 150,000 tokens and sends it to a person. The p95 went back to ₹28. Without the p95 alert, the problem would have shown up only in the month-end bill.

Follow-up questions to expect

  • "Why not alert on total daily spend?" — It moves with traffic. A 20% traffic rise and a new loop look the same there. Tokens per run separates them.
  • "How do you charge cost back to teams or customers?" — Tag each run with the tenant or team, and sum cost per tag.
  • "Do tools cost money too?" — Yes, search and scraping APIs often cost per call. Track tool calls per run next to tokens.