Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
How do you choose appropriate models for different agents?
What you need to know
Models differ a lot in price and speed. A small model can cost 10–20 times less per token than a frontier model and answer faster. Using the big model everywhere is simple but wasteful; using the small one everywhere often fails on the hard steps.
Where CrewAI lets you set a model
1from crewai import Agent, Crew, LLM23BIG = LLM(model="openai/gpt-4.1", temperature=0.2) # example ids4SMALL = LLM(model="openai/gpt-4.1-mini", temperature=0)56classifier = Agent(role="Ticket Classifier", goal="...", backstory="...",7 llm=SMALL)8investigator = Agent(role="Escalation Investigator", goal="...", backstory="...",9 llm=BIG, tools=[crm_lookup, order_lookup],10 function_calling_llm=SMALL) # cheaper model picks tool args11writer = Agent(role="Reply Writer", goal="...", backstory="...", llm=SMALL)1213crew = Crew(agents=[classifier, investigator, writer], tasks=tasks,14 planning=False)llmon an agent — the model for its reasoning.function_calling_llmon an agent or crew — the model that formats tool calls, if different.manager_llm— the manager in a hierarchical crew.planning_llm— the AgentPlanner whenplanning=True.
Always use the provider prefix (openai/, anthropic/, gemini/) so CrewAI routes to the right provider.
A starting split
| Work | Model | Why |
|---|---|---|
| Manager, planner | strong | its mistakes spread to every agent |
| Multi-step investigation, judgement | strong | ambiguity and tool choice |
| Extraction or classification with a schema | small | the schema does the hard part |
| Formatting, summarising tool output | small | low difficulty |
| Final user-facing text | depends | measure it; often small is fine |
Then measure, one change at a time
Change one agent's model, run the golden set, record score and cost, and keep it only if the score holds. Changing several at once hides which change hurt.
A real-life example
An e-commerce company's customer-support escalation crew runs about 3,000 escalations a day. It started with the strong model on all three agents, at about ₹9 per ticket.
The team tried the small model on each agent in turn:
- Classifier: accuracy 97% vs 98%. Switched.
- Writer: tone and accuracy scores the same. Switched.
- Investigator: correct root cause 71% vs 89%. Kept the strong model, but set a small
function_calling_llm.
Cost fell to about ₹3.50 per ticket, roughly ₹16,000 saved per day, with no measurable quality loss. Surprise: the writer, which they expected to need the big model, did not; the investigator, which "only looks things up", did.
Follow-up questions to expect
- "Would you use a reasoning model?" — For the investigator or planner, possibly, if the golden set shows a gain. They are slower and cost more, so not for simple steps.
- "Can you mix providers in one crew?" — Yes, each agent can use a different provider. Watch for different tool-calling behaviour and keep one tracing tool across them.
- "How do you avoid surprises when a provider updates a model?" — Pin dated model versions where the provider offers them and rerun the golden set before moving.