Course Content
AI Agent Frameworks
4 sections · 15 lessons
Framework Selection Guide — Choosing by Constraint
A team spent five weeks migrating their support agent from one framework to another. The reasoning was sound on paper: the new one had more stars, a bigger release cadence, and a comparison blog post said it was faster.
After the migration, they measured. Median latency: 3.1 seconds before, 2.9 seconds after. Answer accuracy on their eval set: 82 per cent before, 81 per cent after. The one thing that genuinely improved was the trace viewer.
Five weeks for a 0.2-second latency improvement and a better debugging experience. Meanwhile the actual problem — that 18 per cent of answers were wrong because four tool descriptions were ambiguous — went untouched. A colleague fixed those descriptions in an afternoon two months later and accuracy went to 91 per cent.
This is the pattern worth guarding against: framework choice is a real decision with real consequences, and it is almost never the decision that determines whether your agent works. Prompt quality, tool design and evaluation discipline dominate. The framework determines how much plumbing you write and how painful the awkward requirement is going to be.
Choose a framework the way you choose a database: by the constraints it makes cheap and the constraints it makes expensive, not by which one is winning.
"Which framework is best?" has no answer
The question has no answer for the same reason "which vehicle is best?" has none. A van and a motorbike are not competing; they are answers to different questions about load and traffic.
Reframe it as three questions that do have answers:
- Is my control flow known in advance? If you can write the steps down before seeing the input, you have a workflow. If not, you have an agent.
- Do I need guarantees, or flexibility? "This never runs without approval" is a guarantee. "It figures out what to do" is flexibility. You cannot have both in the same step.
- What already exists in my organisation? A hundred C# services, a compliance team, and a Kubernetes estate constrain the answer far more than any feature comparison.
The comparison that matters
| Framework style | Control model | Learning curve | Guarantees you get | Deployment weight | Best fit |
|---|---|---|---|---|---|
| Chain/agent library (LangChain) | Model-driven loop | Low | Step and time caps | Medium — large dependency tree | Tool-using assistants, RAG, prototypes |
| Graph runtime (LangGraph) | You define a state machine | Medium | Ordering, branching, retry counts, approval gates | Medium | Multi-step workflows with cycles |
| Goal-stack loop (AutoGPT style) | Self-decomposing tree | Medium | Very few — budgets only | Light if hand-rolled | Open-ended research where the tree is unknown |
| Task-queue loop (BabyAGI style) | Self-generating priority queue | Low | Very few — budgets only | Light | Broad sweeps with reprioritisation |
| Role-based crew (CrewAI) | Specialists with handoffs | Low | Role separation, independent verification | Medium | Content pipelines, review workflows |
| Enterprise SDK (Semantic Kernel, now Microsoft Agent Framework) | Function registry + auto calling | Medium | Registry scoping, filters, audit hooks | Medium; .NET available (Semantic Kernel also Java) | Existing service estates, governance |
| Lightweight agent SDK (OpenAI Agents SDK-style) | One agent loop plus handoffs and guardrails | Low | Step caps, guardrail checks | Light | Single agents or simple handoffs with little code |
| No framework | Your loop | Low to write, high to harden | Whatever you build | Minimal | One or two tools, tight latency or size limits |
Note the last row. "No framework" is a legitimate choice and it wins more often than people expect. A single-tool agent behind a 250 MB serverless size limit with a 2-second latency budget is better served by 80 lines of your own code than by any library.
A decision procedure that takes ten minutes
1. Can you write the steps down before seeing the input? YES -> go to 2 NO -> go to 52. Does the workflow have cycles, branches or approval gates? YES -> graph runtime NO -> go to 33. Are there 3 or fewer steps and 2 or fewer tools? YES -> no framework: write the loop NO -> go to 44. Do distinct stages need independent verification? YES -> role-based crew NO -> chain/agent library5. Do you have an existing registry of services to expose? YES -> enterprise SDK with a function registry NO -> go to 66. Is the task open-ended exploration with an unknown structure? YES -> goal-stack or task-queue loop, with hard budgets NO -> chain/agent libraryFour hard constraints override every branch above, and they are worth checking first because they disqualify options before you evaluate anything else:
| Constraint | Consequence |
|---|---|
| Deployment size limit under about 250 MB | Check installed size before choosing; several agent stacks exceed it once transitive dependencies land |
| Non-Python runtime required | The field narrows: LangChain and LangGraph have TypeScript versions, Microsoft Agent Framework covers .NET — check that the features you need exist in that port |
| Hard sub-second latency budget | Any framework loop adds overhead; consider a single model call with tools and no loop |
| Regulated audit requirement | You need per-invocation hooks, not just callbacks — check this explicitly |
Mapping real situations
Simple: an internal FAQ assistant
Three tools, one or two steps per request, 5,000 requests a day. Use a chain/agent library, or no framework at all. A graph here is 35 lines of wiring to express a sequence with no branches — pure overhead. The thing that will actually determine quality is retrieval, not orchestration.
Complex: a claims-processing workflow
Validate, look up policy, assess, and either auto-approve under 500 pounds or route to a human, with the human's decision resuming the flow days later. Every one of those is a graph feature: conditional edges, a checkpointer, an interrupt before the payout node. Building it as a model-driven loop means the approval threshold lives in a prompt, which means it is advisory, which means someone eventually gets an unapproved payout.
Autonomous: a competitive landscape sweep
"Find everything relevant about the UK meal-kit market." You cannot enumerate the steps because the steps depend on what turns up. This is where goal-stack or task-queue loops earn their place — and where budgets are not optional. Set maximum depth, maximum tasks, a token ceiling and a wall-clock deadline before you write the prompt.
Team: a research briefing pipeline
Research, write, fact-check. The reason to use a crew is not speed — sequential crews are slower than one agent. It is that a single agent cannot check its own work: the checker's evidence is the writer's output, so it approves everything. Separate agents with the checker receiving the sources, not just the draft, is the only structure that catches a wrong figure.
Enterprise: an assistant over 340 internal services
Governance dominates. You need one inspectable place listing what the assistant may call, per-invocation hooks for audit, the ability to scope capability per request without redeploying, and probably C# support. A function-registry SDK is built for exactly this shape; a Python-first agent library will have you writing the registry yourself.
Twelve questions before you commit
| Question | Why it decides something |
|---|---|
| Can I write the steps down in advance? | Workflow versus agent — the primary fork |
| Does anything need human approval mid-run? | Requires pause-and-resume, which most loops lack |
| Does anything need to retry a bounded number of times? | A prompt cannot guarantee a count |
| How many tools, and will that grow past 15? | Beyond ~15 you need routing regardless of framework |
| What is my latency budget? | Each loop step is a round trip; a 4-step agent will not fit 1 second |
| What is my per-request cost ceiling? | Determines step caps and whether autonomous loops are viable |
| What is my deployment size limit? | Can disqualify a framework outright |
| Do I need a language other than Python? | Narrows the field to a few ports; check their feature parity |
| Will an auditor ask what was called and when? | Needs invocation hooks, not print statements |
| Does state need to survive a restart? | Requires checkpointing |
| Who maintains this in a year? | Weight community size and API stability accordingly |
| What is my second-hardest requirement? | The easy case works everywhere; prototype the hard one |
That last question is the single most useful one. Frameworks are indistinguishable on the tutorial case. They differ entirely on the awkward requirement — the approval gate, the resumable run, the C# service, the 200 MB limit. Prototype that on day one, not in week six.
Five ways teams get this wrong
Choosing before understanding the problem
The commonest and most expensive. A framework chosen in week zero encodes assumptions about control flow that you discover are wrong in week four. The cheap alternative: spend two days building the crudest possible version with no framework — a loop, one tool, a print statement. You will learn which parts are genuinely hard, and that knowledge makes the choice almost automatic.
Choosing by popularity
Star counts measure interest, not fit. The team in the opening story migrated to a more popular option and got 0.2 seconds. Popularity is worth exactly one thing — the probability that your specific error message appears in a search result — and that is a real benefit, just a much smaller one than it feels like at 2am.
One framework for everything
Standardisation makes sense at organisational scale, where consistency across forty engineers beats local optimisation. At team scale it is usually a mistake. A resumable approval workflow and a two-tool lookup service have almost nothing in common; forcing them into one shape means one of them is badly served.
Ignoring the maintenance question
Agent frameworks move fast, and fast-moving APIs break. Before committing, check three things: how often the public API has changed in the last six months, whether there is a deprecation policy, and how long issues stay open. A framework that renamed its core abstraction twice this year will cost you a week a year in upgrade work, forever.
Not thinking about the scaled version
Everything works at ten requests a day. Ask what changes at ten thousand.
| Concern | At 10/day | At 10,000/day |
|---|---|---|
| In-memory conversation state | Fine | Broken — multiple replicas, no shared store |
| Synchronous execution | Fine | Needs async, or you buy machines to sit idle on I/O |
| Unbounded steps | Cheap | One runaway request is a real bill |
print() logging | Adequate | Unreproducible bug reports |
| No caching | Fine | You pay repeatedly for identical queries |
| Cold-start time | Irrelevant | Dominant if serverless |
Frameworks are indistinguishable on the tutorial case. Prototype your second-hardest requirement, because that is the only place they actually differ.
Mixing frameworks deliberately
The most effective production systems are usually hybrids, because different parts of one product have different shapes. A worked example for a customer-service platform:
| Component | Shape | Why |
|---|---|---|
| Intent classification | One model call, no framework | Latency-critical, no orchestration needed |
| FAQ answering | Retrieval plus one call | Two steps, fully known |
| Order investigation | Chain/agent library | Three tools, unpredictable order |
| Refund processing | Graph runtime | Approval gate, must resume after a human |
| Escalation summary | Role-based crew | Writer plus independent checker |
The rule that makes hybrids maintainable: each component exposes a plain function boundary — def handle_refund(request: RefundRequest) -> RefundResult — and nothing outside knows what is inside. The router calls functions. Swapping the refund component's internals from one framework to another then touches one file, not the system.
This is also what makes framework choice reversible, which is the property that matters most. If your entire application is written in one framework's idioms, changing it is the five-week migration in the opening story. If each component is a function with a typed signature, changing one is a day.
What to do on Monday
If you are starting something new, do not choose a framework yet. Spend two days building the crudest version — a loop, one or two tools, a system prompt, and a fixed set of twenty test questions with known correct answers. That prototype gives you three things no comparison article can: a real measurement of how long each step takes, a real sense of where the model goes wrong, and a list of your actual requirements rather than your imagined ones.
Then take your second-hardest requirement — the approval gate, the resumable run, the twenty-tool routing problem — and spend half a day prototyping just that in each of your two shortlisted options. You will usually find one of them makes it a configuration flag and the other makes it a fork. That half-day is worth more than every feature matrix, including the one in this lesson.
And whichever you pick, build the eval set first and keep it. Twenty questions with expected tool sequences and expected answers, run after every change. It is the only instrument that tells you whether a change helped, and it is the reason the team in the opening story could eventually prove that fixing four tool descriptions was worth nine percentage points of accuracy while a five-week migration was worth 0.2 seconds.