CrewAI Multi-Agents

Course Content

CrewAI Multi-Agents

9 sections · 53 lessons

How do you provide domain-specific knowledge to agents?


How a policy PDF reaches the promptPDF placedin theknowledge folderChunked andembeddedinto ChromaDBTask rewritteninto a search queryTop chunksabove thescore thresholdAdded to thattask's promptAbout 1,500 policy tokens per task instead of 15,000 in the backstory.
Retrieval sends only the rules a task needs, which is both cheaper and harder for the model to overlook.

What you need to know

Four options

KnowledgeMechanismWhy
A few rules, rarely changesbackstory or task textalways present, no retrieval risk
Tens of documents, changes monthlyknowledge_sourcesbuilt in, little code
Thousands of documents, changes dailyyour own retrieval toolfull control and evaluation
A style or formatfine-tuning (rarely)teaches form, not reliable facts

Knowledge sources in code

Python
from crewai import Agentfrom crewai.knowledge.source.pdf_knowledge_source import PDFKnowledgeSourcefrom crewai.knowledge.source.string_knowledge_source import StringKnowledgeSourcefrom crewai.knowledge.knowledge_config import KnowledgeConfigpolicy = PDFKnowledgeSource(file_paths=["lending_policy_2026.pdf"])rules = StringKnowledgeSource(content="FOIR limit is 50% for salaried applicants.")policy_checker = Agent(    role="Policy Checker",    goal="Check each application against the lending policy",    backstory="You quote the policy section for every decision.",    knowledge_sources=[policy, rules],    knowledge_config=KnowledgeConfig(results_limit=4, score_threshold=0.5),)
  • File paths are relative to a knowledge/ folder at the project root.
  • Sources exist for PDF, text, CSV, Excel, JSON and plain strings.
  • Knowledge is stored in a local vector database (ChromaDB by default) and embedded with OpenAI unless you set embedder.
  • results_limit and score_threshold control how many chunks are added and how relevant they must be.

When to build your own retrieval tool

Built-in knowledge is easy but gives you little control. A custom BaseTool over your vector store lets you:

  • filter by metadata (product, region, date);
  • use hybrid keyword plus vector search and a reranker;
  • return document IDs and page numbers for citations;
  • measure retrieval quality (did the right chunk come back?) separately from the answer.

A real-life example

An NBFC's loan-document review crew needed its policy checker to apply a 40-page lending policy. Pasting the policy into the backstory cost about 15,000 tokens per step and the agent still missed rules in the middle.

They moved the policy PDF to knowledge_sources with results_limit=4. Each task now carries about 1,500 tokens of relevant policy text. Rules applied correctly on 46 of 50 test applications, up from 39.

Separately, the team's product catalogue — 3,000 loan schemes across 200 partner banks, updated daily — did not fit this pattern. They built a scheme_search tool over their own vector index with metadata filters for state and loan type. That kept daily updates in their normal data pipeline, and let them track retrieval accuracy on its own.

Follow-up questions to expect

  • "Knowledge or memory?" — Knowledge is reference material you provide; memory is what the crew learns while running. Policies are knowledge; "this customer already applied in May" is memory.
  • "Why not just use a long-context model?" — Cost and attention: long prompts are re-sent on every step and the model misses details in the middle. Retrieval sends only what the task needs.
  • "How do you know retrieval is working?" — Keep a test set of questions with the correct document sections and measure how often the right chunk comes back, before judging the final answer.