CrewAI Multi-Agents

Course Content

CrewAI Multi-Agents

9 sections · 53 lessons

How do you choose appropriate models for different agents?


Swapping one agent at a time to the small model97%98%samesame71%89%Small modelStrong modelClassifierWriterInvestigatorCost fell from about Rs 9 to Rs 3.50 per ticket.
The agent everyone expected to need the big model did not, and the one that only looks things up did.

What you need to know

Models differ a lot in price and speed. A small model can cost 10–20 times less per token than a frontier model and answer faster. Using the big model everywhere is simple but wasteful; using the small one everywhere often fails on the hard steps.

Where CrewAI lets you set a model

Python
from crewai import Agent, Crew, LLMBIG = LLM(model="openai/gpt-4.1", temperature=0.2)          # example idsSMALL = LLM(model="openai/gpt-4.1-mini", temperature=0)classifier = Agent(role="Ticket Classifier", goal="...", backstory="...",                   llm=SMALL)investigator = Agent(role="Escalation Investigator", goal="...", backstory="...",                     llm=BIG, tools=[crm_lookup, order_lookup],                     function_calling_llm=SMALL)   # cheaper model picks tool argswriter = Agent(role="Reply Writer", goal="...", backstory="...", llm=SMALL)crew = Crew(agents=[classifier, investigator, writer], tasks=tasks,            planning=False)
  • llm on an agent — the model for its reasoning.
  • function_calling_llm on an agent or crew — the model that formats tool calls, if different.
  • manager_llm — the manager in a hierarchical crew.
  • planning_llm — the AgentPlanner when planning=True.

Always use the provider prefix (openai/, anthropic/, gemini/) so CrewAI routes to the right provider.

A starting split

WorkModelWhy
Manager, plannerstrongits mistakes spread to every agent
Multi-step investigation, judgementstrongambiguity and tool choice
Extraction or classification with a schemasmallthe schema does the hard part
Formatting, summarising tool outputsmalllow difficulty
Final user-facing textdependsmeasure it; often small is fine

Then measure, one change at a time

Change one agent's model, run the golden set, record score and cost, and keep it only if the score holds. Changing several at once hides which change hurt.

A real-life example

An e-commerce company's customer-support escalation crew runs about 3,000 escalations a day. It started with the strong model on all three agents, at about ₹9 per ticket.

The team tried the small model on each agent in turn:

  • Classifier: accuracy 97% vs 98%. Switched.
  • Writer: tone and accuracy scores the same. Switched.
  • Investigator: correct root cause 71% vs 89%. Kept the strong model, but set a small function_calling_llm.

Cost fell to about ₹3.50 per ticket, roughly ₹16,000 saved per day, with no measurable quality loss. Surprise: the writer, which they expected to need the big model, did not; the investigator, which "only looks things up", did.

Follow-up questions to expect

  • "Would you use a reasoning model?" — For the investigator or planner, possibly, if the golden set shows a gain. They are slower and cost more, so not for simple steps.
  • "Can you mix providers in one crew?" — Yes, each agent can use a different provider. Watch for different tool-calling behaviour and keep one tracing tool across them.
  • "How do you avoid surprises when a provider updates a model?" — Pin dated model versions where the provider offers them and rerun the golden set before moving.