Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
How do you provide domain-specific knowledge to agents?
What you need to know
Four options
| Knowledge | Mechanism | Why |
|---|---|---|
| A few rules, rarely changes | backstory or task text | always present, no retrieval risk |
| Tens of documents, changes monthly | knowledge_sources | built in, little code |
| Thousands of documents, changes daily | your own retrieval tool | full control and evaluation |
| A style or format | fine-tuning (rarely) | teaches form, not reliable facts |
Knowledge sources in code
1from crewai import Agent2from crewai.knowledge.source.pdf_knowledge_source import PDFKnowledgeSource3from crewai.knowledge.source.string_knowledge_source import StringKnowledgeSource4from crewai.knowledge.knowledge_config import KnowledgeConfig56policy = PDFKnowledgeSource(file_paths=["lending_policy_2026.pdf"])7rules = StringKnowledgeSource(content="FOIR limit is 50% for salaried applicants.")89policy_checker = Agent(10 role="Policy Checker",11 goal="Check each application against the lending policy",12 backstory="You quote the policy section for every decision.",13 knowledge_sources=[policy, rules],14 knowledge_config=KnowledgeConfig(results_limit=4, score_threshold=0.5),15)- File paths are relative to a
knowledge/folder at the project root. - Sources exist for PDF, text, CSV, Excel, JSON and plain strings.
- Knowledge is stored in a local vector database (ChromaDB by default) and embedded with OpenAI unless you set
embedder. results_limitandscore_thresholdcontrol how many chunks are added and how relevant they must be.
When to build your own retrieval tool
Built-in knowledge is easy but gives you little control. A custom BaseTool over your vector store lets you:
- filter by metadata (product, region, date);
- use hybrid keyword plus vector search and a reranker;
- return document IDs and page numbers for citations;
- measure retrieval quality (did the right chunk come back?) separately from the answer.
A real-life example
An NBFC's loan-document review crew needed its policy checker to apply a 40-page lending policy. Pasting the policy into the backstory cost about 15,000 tokens per step and the agent still missed rules in the middle.
They moved the policy PDF to knowledge_sources with results_limit=4. Each task now carries about 1,500 tokens of relevant policy text. Rules applied correctly on 46 of 50 test applications, up from 39.
Separately, the team's product catalogue — 3,000 loan schemes across 200 partner banks, updated daily — did not fit this pattern. They built a scheme_search tool over their own vector index with metadata filters for state and loan type. That kept daily updates in their normal data pipeline, and let them track retrieval accuracy on its own.
Follow-up questions to expect
- "Knowledge or memory?" — Knowledge is reference material you provide; memory is what the crew learns while running. Policies are knowledge; "this customer already applied in May" is memory.
- "Why not just use a long-context model?" — Cost and attention: long prompts are re-sent on every step and the model misses details in the middle. Retrieval sends only what the task needs.
- "How do you know retrieval is working?" — Keep a test set of questions with the correct document sections and measure how often the right chunk comes back, before judging the final answer.