AI Agent Frameworks

Why Frameworks Matter for Agent Development


A developer we will call Priya wrote her first agent on a Tuesday afternoon. It was 118 lines of Python: a loop that sent the conversation to a model, checked whether the reply contained a tool call, ran the tool, appended the result, and looped again. It answered "What was Tesla's revenue last quarter, and how does that compare to Ford?" beautifully. She demoed it. Everyone clapped.

On Thursday it went to a pilot group of forty people. By Friday morning the bill was 214 dollars and one support ticket read: "It just kept saying 'Let me search for that' over and over for two minutes and then stopped."

The logs showed what happened. A user asked about a company whose name the search tool returned nothing for. The model saw an empty result, decided it had phrased the query badly, searched again with a slightly different phrasing, got nothing again, and repeated — forty-one times before hitting the request timeout. Each iteration re-sent the entire growing conversation. The final request carried roughly 380,000 input tokens. At an example rate of 3 dollars per million input tokens (prices change often; check your provider's price list), that last request alone cost about 1.14 dollars, the forty calls before it cost roughly twenty times more between them, and the question never got an answer.

Priya's loop was not wrong. It was incomplete. Every piece it was missing — a step budget, loop detection, history trimming, a retry policy, a trace you can read afterwards — is a piece that an agent framework hands you on day one. That is the entire case for frameworks, and it is worth understanding precisely rather than as a slogan, because frameworks also cost you something and you need to know what.

What the 118-line loop is missingSix components aframework shipsOrchestration andthe step budgetTool and skill registrationMemory and history trimmingRetries,timeouts, resilienceObservability of every stepMulti-agent coordination
The loop is the easy part; the six things around it are what the second month of the project is spent writing.

The scale-up problem: what the 118 lines cannot do

An agent loop is deceptively easy to start. The core is genuinely about twenty lines. The difficulty is not in the loop — it is in everything that surrounds the loop once real users touch it.

Here is the honest comparison between the demo version and the version that survives a month in production.

ConcernDemo agentProduction agent
Step limitNone — loop until the model stopsHard cap (typically 8–15), plus a wall-clock deadline
Repeated tool callsAllowed foreverDetected; identical call twice in a row is short-circuited
Conversation growthFull history re-sent every stepTrimmed, summarised, or windowed
Tool failureException kills the requestCaught, converted to a message the model can act on, retried with backoff
Malformed tool argumentsCrash on KeyErrorSchema-validated; validation error fed back for a repair attempt
Observabilityprint()Structured trace: every prompt, tool call, latency, token count
Cost controlNonePer-request token budget, enforced mid-run
ConcurrencyOne user at a timeAsync, connection pooling, per-user rate limits

Each row is a few hours of work. Together they are three to six weeks. And critically, none of them are interesting problems — they are the same problems every agent has, solved the same way every time. That is exactly the shape of thing a library should own.

Worked example: what history trimming is actually worth

Take a concrete case. An agent runs a multi-step research task. Each step adds roughly 800 tokens to the conversation — the model's reasoning plus the tool result. Because the whole conversation is re-sent on every API call, the input tokens for step n are about 800n800n.

Over 20 steps with no trimming, total input tokens are:

800×(1+2+⋯+20)=800×20×212=800×210=168,000800 \times (1 + 2 + \cdots + 20) = 800 \times \frac{20 \times 21}{2} = 800 \times 210 = 168{,}000

At 3 dollars per million input tokens, that is 0.504 dollars for one question.

Now add a sliding window that keeps only the last six steps. Steps 1–6 grow normally (16,800 tokens total, by the same sum formula with n=6n=6: 800×21800 \times 21). Steps 7–20 each carry a fixed six-step window, so 800×6=4,800800 \times 6 = 4{,}800 tokens each, for 14×4,800=67,20014 \times 4{,}800 = 67{,}200. Total: 84,000 tokens, or 0.252 dollars.

Exactly half, from one feature. Multiply the 0.252-dollar saving by 40,000 requests a month and the sliding window is worth about 10,000 dollars a month, or roughly 120,000 dollars a year. This is the sort of thing frameworks give you as a configuration flag.

A framework's value is not that it lets you build an agent. It is that it stops you rebuilding, badly, the twelve boring things every agent needs.

The six components every agent framework provides

Frameworks differ enormously in style, but they all supply the same six capabilities. Knowing the list lets you evaluate any new framework in fifteen minutes.

1. Agent orchestration

The control loop itself: decide, act, observe, repeat. What varies between frameworks is how much control you get over the loop's shape. Some give you a single opaque run() call. Some let you define the loop as an explicit graph of steps with conditional edges. Some let you pause the loop, persist it, and resume it three days later after a human approves something.

The question to ask: can I insert a step in the middle of the loop without forking the library? If the answer is no, you will eventually fork the library.

2. Tool and skill registration

Models call tools by name with JSON arguments. Something has to turn your Python function into a JSON schema the model can read, validate the model's arguments against that schema, dispatch the call, and format the result back into the conversation.

Python
from pydantic import BaseModel, Fieldclass SearchArgs(BaseModel):    query: str = Field(description="Search terms, 2-8 words, no punctuation")    max_results: int = Field(default=5, ge=1, le=20)# The framework turns this into:# {"name": "search", "description": "...",#  "parameters": {"type": "object",#                 "properties": {"query": {"type": "string", ...},#                                "max_results": {"type": "integer",#                                                "minimum": 1, "maximum": 20}},#                 "required": ["query"]}}

The ge=1, le=20 constraint matters more than it looks. Without it, a model that asks for max_results: 500 gets a 500-result payload, blows the context window, and the run dies at step three. With it, the framework rejects the call before it reaches your code and hands the model a validation error it can correct.

3. Memory management

"Memory" covers three genuinely different things that people constantly confuse:

KindHoldsLives forRetrieved by
Working memoryThe current conversation and tool resultsOne requestBeing in the prompt
Session memoryEarlier turns with this user todayA sessionRecency window or summary
Long-term memoryFacts, documents, past outcomesIndefinitelySemantic search over embeddings

A framework that only gives you a conversation buffer has given you working memory and nothing else. That is fine for a chatbot and useless for an agent that must remember, three weeks later, that this customer's refund was already processed.

4. Error handling and resilience

Tools fail. APIs rate-limit. Networks time out. The interesting design question is who the error is reported to. There are two answers and they produce very different agents.

StrategyError goes toGood forRisk
RaiseYour code / the userAuth failures, bad config, anything the model cannot fixFragile; one flaky API kills the run
Feed backThe model, as a tool resultBad arguments, empty results, transient 503sModel may loop forever "fixing" it

Priya's forty-one searches were the second failure mode. The correct policy is: feed the error back, but count consecutive failures per tool and raise after two or three. Most frameworks expose this as a setting — a retry count, an error handler or a tool-call limit — and most people never change it from the default.

5. Observability

An agent that fails is not like a function that fails. There is no stack trace pointing at line 47. There is a chain of eleven model calls where the third one made a slightly odd decision that poisoned everything after it. Without a stored trace of every prompt, every tool call, every raw response and every latency, you cannot debug it — you can only re-run it and hope.

If you cannot answer "what exactly did the model see at step 4?" from your logs, you do not have an observable agent, and every bug report will be unreproducible.

6. Multi-agent coordination

Once one agent works, someone will want three: a researcher, a writer, a fact-checker. Coordination brings its own machinery — message passing between agents, shared state, a scheduler deciding who runs next, and a termination condition so the reviewer and writer do not revise each other forever.

The time-to-production paradox

Here is the thing nobody tells you: frameworks massively shorten time-to-demo and only modestly shorten time-to-production. Sometimes they lengthen it.

MilestoneHand-rolledWith a framework
First working agent1–2 days1–2 hours
Three tools, memory, error handling1–2 weeks1 day
Debugging a wrong answer in productionHours — it is your codeHours to days — it is inside three layers of abstraction
Changing behaviour the framework did not anticipateEdit your loopSubclass, monkey-patch, or fork
Total to a hardened v15–7 weeks3–4 weeks

The paradox: the 100x speed-up on the first afternoon shrinks to roughly 1.5x by launch. That is still a good trade — but it means choosing a framework because the tutorial was fast is a mistake. You are optimising the one phase that was never the bottleneck.

The corollary is that the framework property that matters most is not "how quickly can I build a demo" but "how easily can I get out when I need to do something it did not anticipate?" Frameworks with clean escape hatches — where you can drop to raw model calls for one step and keep everything else — age far better than frameworks that own the whole loop.

Three systems, three different answers

A customer support bot

Roughly four tools (order lookup, refund status, knowledge-base search, escalate-to-human), a strict policy about what it may promise, and 30,000 conversations a month. Latency matters — users watch a typing indicator.

What this needs: fast simple tool calling, tight step limits (three or four), an audit log for compliance, and a hard rule that any refund above a threshold routes to a human. What it does not need is autonomous planning. A support bot that decides to "explore alternative approaches" is a liability.

A data analysis agent

One tool that matters — execute SQL or Python — plus schema introspection. Steps are few but each is expensive and dangerous. Here the framework requirements invert: you want sandboxed execution, result-size limits, and the ability to show the user the generated query before it runs against production.

A multi-agent research pipeline

A planner splits a question into sub-questions, three researchers work in parallel, a synthesiser merges. Now you need genuine coordination: parallel execution, shared state, and a merge step. Running this sequentially takes three times as long for no benefit; running it in parallel without shared state means three researchers independently searching the same thing.

How to actually choose

Split your criteria into four buckets and score honestly. Most teams only score the first two and are surprised later by the third and fourth.

BucketAskHow to test it in an afternoon
PerformanceOverhead per step? Async support? Streaming?Time 100 runs of a trivial one-tool agent; compare to raw API calls
FeaturesMemory types? Branching? Parallelism? Human-in-the-loop?Try to build your second-hardest use case, not your easiest
OperationalTrace quality? Deterministic replay? Version stability?Deliberately break a tool and see what the logs tell you
DeploymentCold-start size? Serverless-friendly? Stateful requirements?Check installed size and import time; some agent stacks exceed 200 MB

On the last row: if you plan to deploy on a serverless platform with a 250 MB unzipped limit, a framework that pulls in a large dependency tree is disqualified before you write a line. People discover this in week five.

The shape of the landscape

StyleControl modelBest fitWeakness
Chain/agent libraries (LangChain-style)Linear loop, model-drivenTool-using assistants, RAG, prototypesHard to express branching or retries precisely
Graph runtimes (LangGraph-style)Explicit state machine you defineMulti-step workflows, cycles, approval gatesMore upfront design; overkill for one tool call
Autonomous loops (AutoGPT, BabyAGI style)Self-generated goals and tasksOpen-ended exploration, research sweepsCost and drift; weak reliability guarantees
Role-based crews (CrewAI-style)Multiple specialists with handoffsContent pipelines, review workflowsContext loss at handoffs; coordination overhead
Enterprise SDKs (Semantic Kernel, now Microsoft Agent Framework)Plugin registry + automatic function callingExisting .NET/Java estates, governance needsMore ceremony; smaller community than Python-first tools
Lightweight agent SDKs (OpenAI Agents SDK-style)One agent loop plus handoffs and guardrailsSingle agents and simple handoffs with little codeFewer workflow controls than a graph runtime

Three things people get wrong about frameworks

Wrong idea 1: "The framework will make my agent smart." It will not. Agent quality is dominated by prompt design, tool descriptions and the model itself. A framework moves plumbing off your plate; it does not improve reasoning. If your agent picks the wrong tool, the fix is almost always a better tool description, not a different framework. A tool described as "Search" and one described as "Search the live web for current facts. Use for anything after the model's training cutoff, especially prices, news and people's current roles. Not for arithmetic." produce measurably different behaviour with identical code around them.

Wrong idea 2: "Pick one framework and standardise on it." Reasonable for a team of fifty. Actively harmful for a team of five with three different problems. Using a graph runtime for a stateful approval workflow and a plain tool-calling loop for a simple lookup service is not inconsistency — it is fit.

Wrong idea 3: "The framework is the system." It is one component among many. A production agent also needs a vector store, a model provider with fallback, a queue, a cache, a rate limiter, an eval harness and a deployment target. In a typical build the framework is perhaps 15 per cent of the code and 5 per cent of the operational surface. Choose it carefully, but do not expect it to carry the system.

What this means the day you start building

Write the raw loop first. Twenty lines, one tool, no library. Do it once, deliberately, before you install anything — because that hour teaches you what every framework is abstracting, and afterwards you can read framework documentation as "ah, this is how they solved the thing I just did by hand" instead of as magic.

Then pick a framework against your second-hardest requirement, not your easiest. The easy case works everywhere; the hard case is where frameworks differ. If you need an approval gate in the middle of a run, prototype exactly that on day one.

Set three limits before you set anything else: maximum steps, maximum tokens per request, and maximum consecutive tool errors. Priya's 214-dollar Friday was three configuration lines away from being a 12-dollar Friday. Whatever framework you pick, find those three settings in its documentation before you find anything else — they are the difference between an agent that fails cheaply and one that fails expensively.

Finally, instrument from the first commit. Log every prompt, every tool call with arguments, every raw response, every latency and token count, keyed by a request ID. You will not regret the disk space. You will absolutely regret the bug report you cannot reproduce.