Course Content
AI Agent Frameworks
4 sections · 15 lessons
Why Frameworks Matter for Agent Development
A developer we will call Priya wrote her first agent on a Tuesday afternoon. It was 118 lines of Python: a loop that sent the conversation to a model, checked whether the reply contained a tool call, ran the tool, appended the result, and looped again. It answered "What was Tesla's revenue last quarter, and how does that compare to Ford?" beautifully. She demoed it. Everyone clapped.
On Thursday it went to a pilot group of forty people. By Friday morning the bill was 214 dollars and one support ticket read: "It just kept saying 'Let me search for that' over and over for two minutes and then stopped."
The logs showed what happened. A user asked about a company whose name the search tool returned nothing for. The model saw an empty result, decided it had phrased the query badly, searched again with a slightly different phrasing, got nothing again, and repeated — forty-one times before hitting the request timeout. Each iteration re-sent the entire growing conversation. The final request carried roughly 380,000 input tokens. At an example rate of 3 dollars per million input tokens (prices change often; check your provider's price list), that last request alone cost about 1.14 dollars, the forty calls before it cost roughly twenty times more between them, and the question never got an answer.
Priya's loop was not wrong. It was incomplete. Every piece it was missing — a step budget, loop detection, history trimming, a retry policy, a trace you can read afterwards — is a piece that an agent framework hands you on day one. That is the entire case for frameworks, and it is worth understanding precisely rather than as a slogan, because frameworks also cost you something and you need to know what.
The scale-up problem: what the 118 lines cannot do
An agent loop is deceptively easy to start. The core is genuinely about twenty lines. The difficulty is not in the loop — it is in everything that surrounds the loop once real users touch it.
Here is the honest comparison between the demo version and the version that survives a month in production.
| Concern | Demo agent | Production agent |
|---|---|---|
| Step limit | None — loop until the model stops | Hard cap (typically 8–15), plus a wall-clock deadline |
| Repeated tool calls | Allowed forever | Detected; identical call twice in a row is short-circuited |
| Conversation growth | Full history re-sent every step | Trimmed, summarised, or windowed |
| Tool failure | Exception kills the request | Caught, converted to a message the model can act on, retried with backoff |
| Malformed tool arguments | Crash on KeyError | Schema-validated; validation error fed back for a repair attempt |
| Observability | print() | Structured trace: every prompt, tool call, latency, token count |
| Cost control | None | Per-request token budget, enforced mid-run |
| Concurrency | One user at a time | Async, connection pooling, per-user rate limits |
Each row is a few hours of work. Together they are three to six weeks. And critically, none of them are interesting problems — they are the same problems every agent has, solved the same way every time. That is exactly the shape of thing a library should own.
Worked example: what history trimming is actually worth
Take a concrete case. An agent runs a multi-step research task. Each step adds roughly 800 tokens to the conversation — the model's reasoning plus the tool result. Because the whole conversation is re-sent on every API call, the input tokens for step n are about 800n.
Over 20 steps with no trimming, total input tokens are:
At 3 dollars per million input tokens, that is 0.504 dollars for one question.
Now add a sliding window that keeps only the last six steps. Steps 1–6 grow normally (16,800 tokens total, by the same sum formula with n=6: 800×21). Steps 7–20 each carry a fixed six-step window, so 800×6=4,800 tokens each, for 14×4,800=67,200. Total: 84,000 tokens, or 0.252 dollars.
Exactly half, from one feature. Multiply the 0.252-dollar saving by 40,000 requests a month and the sliding window is worth about 10,000 dollars a month, or roughly 120,000 dollars a year. This is the sort of thing frameworks give you as a configuration flag.
A framework's value is not that it lets you build an agent. It is that it stops you rebuilding, badly, the twelve boring things every agent needs.
The six components every agent framework provides
Frameworks differ enormously in style, but they all supply the same six capabilities. Knowing the list lets you evaluate any new framework in fifteen minutes.
1. Agent orchestration
The control loop itself: decide, act, observe, repeat. What varies between frameworks is how much control you get over the loop's shape. Some give you a single opaque run() call. Some let you define the loop as an explicit graph of steps with conditional edges. Some let you pause the loop, persist it, and resume it three days later after a human approves something.
The question to ask: can I insert a step in the middle of the loop without forking the library? If the answer is no, you will eventually fork the library.
2. Tool and skill registration
Models call tools by name with JSON arguments. Something has to turn your Python function into a JSON schema the model can read, validate the model's arguments against that schema, dispatch the call, and format the result back into the conversation.
1from pydantic import BaseModel, Field23class SearchArgs(BaseModel):4 query: str = Field(description="Search terms, 2-8 words, no punctuation")5 max_results: int = Field(default=5, ge=1, le=20)67# The framework turns this into:8# {"name": "search", "description": "...",9# "parameters": {"type": "object",10# "properties": {"query": {"type": "string", ...},11# "max_results": {"type": "integer",12# "minimum": 1, "maximum": 20}},13# "required": ["query"]}}The ge=1, le=20 constraint matters more than it looks. Without it, a model that asks for max_results: 500 gets a 500-result payload, blows the context window, and the run dies at step three. With it, the framework rejects the call before it reaches your code and hands the model a validation error it can correct.
3. Memory management
"Memory" covers three genuinely different things that people constantly confuse:
| Kind | Holds | Lives for | Retrieved by |
|---|---|---|---|
| Working memory | The current conversation and tool results | One request | Being in the prompt |
| Session memory | Earlier turns with this user today | A session | Recency window or summary |
| Long-term memory | Facts, documents, past outcomes | Indefinitely | Semantic search over embeddings |
A framework that only gives you a conversation buffer has given you working memory and nothing else. That is fine for a chatbot and useless for an agent that must remember, three weeks later, that this customer's refund was already processed.
4. Error handling and resilience
Tools fail. APIs rate-limit. Networks time out. The interesting design question is who the error is reported to. There are two answers and they produce very different agents.
| Strategy | Error goes to | Good for | Risk |
|---|---|---|---|
| Raise | Your code / the user | Auth failures, bad config, anything the model cannot fix | Fragile; one flaky API kills the run |
| Feed back | The model, as a tool result | Bad arguments, empty results, transient 503s | Model may loop forever "fixing" it |
Priya's forty-one searches were the second failure mode. The correct policy is: feed the error back, but count consecutive failures per tool and raise after two or three. Most frameworks expose this as a setting — a retry count, an error handler or a tool-call limit — and most people never change it from the default.
5. Observability
An agent that fails is not like a function that fails. There is no stack trace pointing at line 47. There is a chain of eleven model calls where the third one made a slightly odd decision that poisoned everything after it. Without a stored trace of every prompt, every tool call, every raw response and every latency, you cannot debug it — you can only re-run it and hope.
If you cannot answer "what exactly did the model see at step 4?" from your logs, you do not have an observable agent, and every bug report will be unreproducible.
6. Multi-agent coordination
Once one agent works, someone will want three: a researcher, a writer, a fact-checker. Coordination brings its own machinery — message passing between agents, shared state, a scheduler deciding who runs next, and a termination condition so the reviewer and writer do not revise each other forever.
The time-to-production paradox
Here is the thing nobody tells you: frameworks massively shorten time-to-demo and only modestly shorten time-to-production. Sometimes they lengthen it.
| Milestone | Hand-rolled | With a framework |
|---|---|---|
| First working agent | 1–2 days | 1–2 hours |
| Three tools, memory, error handling | 1–2 weeks | 1 day |
| Debugging a wrong answer in production | Hours — it is your code | Hours to days — it is inside three layers of abstraction |
| Changing behaviour the framework did not anticipate | Edit your loop | Subclass, monkey-patch, or fork |
| Total to a hardened v1 | 5–7 weeks | 3–4 weeks |
The paradox: the 100x speed-up on the first afternoon shrinks to roughly 1.5x by launch. That is still a good trade — but it means choosing a framework because the tutorial was fast is a mistake. You are optimising the one phase that was never the bottleneck.
The corollary is that the framework property that matters most is not "how quickly can I build a demo" but "how easily can I get out when I need to do something it did not anticipate?" Frameworks with clean escape hatches — where you can drop to raw model calls for one step and keep everything else — age far better than frameworks that own the whole loop.
Three systems, three different answers
A customer support bot
Roughly four tools (order lookup, refund status, knowledge-base search, escalate-to-human), a strict policy about what it may promise, and 30,000 conversations a month. Latency matters — users watch a typing indicator.
What this needs: fast simple tool calling, tight step limits (three or four), an audit log for compliance, and a hard rule that any refund above a threshold routes to a human. What it does not need is autonomous planning. A support bot that decides to "explore alternative approaches" is a liability.
A data analysis agent
One tool that matters — execute SQL or Python — plus schema introspection. Steps are few but each is expensive and dangerous. Here the framework requirements invert: you want sandboxed execution, result-size limits, and the ability to show the user the generated query before it runs against production.
A multi-agent research pipeline
A planner splits a question into sub-questions, three researchers work in parallel, a synthesiser merges. Now you need genuine coordination: parallel execution, shared state, and a merge step. Running this sequentially takes three times as long for no benefit; running it in parallel without shared state means three researchers independently searching the same thing.
How to actually choose
Split your criteria into four buckets and score honestly. Most teams only score the first two and are surprised later by the third and fourth.
| Bucket | Ask | How to test it in an afternoon |
|---|---|---|
| Performance | Overhead per step? Async support? Streaming? | Time 100 runs of a trivial one-tool agent; compare to raw API calls |
| Features | Memory types? Branching? Parallelism? Human-in-the-loop? | Try to build your second-hardest use case, not your easiest |
| Operational | Trace quality? Deterministic replay? Version stability? | Deliberately break a tool and see what the logs tell you |
| Deployment | Cold-start size? Serverless-friendly? Stateful requirements? | Check installed size and import time; some agent stacks exceed 200 MB |
On the last row: if you plan to deploy on a serverless platform with a 250 MB unzipped limit, a framework that pulls in a large dependency tree is disqualified before you write a line. People discover this in week five.
The shape of the landscape
| Style | Control model | Best fit | Weakness |
|---|---|---|---|
| Chain/agent libraries (LangChain-style) | Linear loop, model-driven | Tool-using assistants, RAG, prototypes | Hard to express branching or retries precisely |
| Graph runtimes (LangGraph-style) | Explicit state machine you define | Multi-step workflows, cycles, approval gates | More upfront design; overkill for one tool call |
| Autonomous loops (AutoGPT, BabyAGI style) | Self-generated goals and tasks | Open-ended exploration, research sweeps | Cost and drift; weak reliability guarantees |
| Role-based crews (CrewAI-style) | Multiple specialists with handoffs | Content pipelines, review workflows | Context loss at handoffs; coordination overhead |
| Enterprise SDKs (Semantic Kernel, now Microsoft Agent Framework) | Plugin registry + automatic function calling | Existing .NET/Java estates, governance needs | More ceremony; smaller community than Python-first tools |
| Lightweight agent SDKs (OpenAI Agents SDK-style) | One agent loop plus handoffs and guardrails | Single agents and simple handoffs with little code | Fewer workflow controls than a graph runtime |
Three things people get wrong about frameworks
Wrong idea 1: "The framework will make my agent smart." It will not. Agent quality is dominated by prompt design, tool descriptions and the model itself. A framework moves plumbing off your plate; it does not improve reasoning. If your agent picks the wrong tool, the fix is almost always a better tool description, not a different framework. A tool described as "Search" and one described as "Search the live web for current facts. Use for anything after the model's training cutoff, especially prices, news and people's current roles. Not for arithmetic." produce measurably different behaviour with identical code around them.
Wrong idea 2: "Pick one framework and standardise on it." Reasonable for a team of fifty. Actively harmful for a team of five with three different problems. Using a graph runtime for a stateful approval workflow and a plain tool-calling loop for a simple lookup service is not inconsistency — it is fit.
Wrong idea 3: "The framework is the system." It is one component among many. A production agent also needs a vector store, a model provider with fallback, a queue, a cache, a rate limiter, an eval harness and a deployment target. In a typical build the framework is perhaps 15 per cent of the code and 5 per cent of the operational surface. Choose it carefully, but do not expect it to carry the system.
What this means the day you start building
Write the raw loop first. Twenty lines, one tool, no library. Do it once, deliberately, before you install anything — because that hour teaches you what every framework is abstracting, and afterwards you can read framework documentation as "ah, this is how they solved the thing I just did by hand" instead of as magic.
Then pick a framework against your second-hardest requirement, not your easiest. The easy case works everywhere; the hard case is where frameworks differ. If you need an approval gate in the middle of a run, prototype exactly that on day one.
Set three limits before you set anything else: maximum steps, maximum tokens per request, and maximum consecutive tool errors. Priya's 214-dollar Friday was three configuration lines away from being a 12-dollar Friday. Whatever framework you pick, find those three settings in its documentation before you find anything else — they are the difference between an agent that fails cheaply and one that fails expensively.
Finally, instrument from the first commit. Log every prompt, every tool call with arguments, every raw response, every latency and token count, keyed by a request ID. You will not regret the disk space. You will absolutely regret the bug report you cannot reproduce.