Course Content
AI Agent Fundamentals
5 sections · 13 lessons
Definition and Core Concepts (Agents vs LLMs)
Here is a real prompt someone typed into a chat model last week: "Find me the cheapest direct flight from Bengaluru to Singapore on 14 March, book it with my saved card, and put it in my calendar."
The model produced a beautiful, confident answer. It listed three flights with times and prices. It said the booking was confirmed and gave a reference number, SQ-8842-XR. It described the calendar entry it had created.
None of it happened. There was no flight lookup, because the model has no connection to any airline. The reference number was invented — it has the right shape because the model has seen thousands of booking confirmations during training, and SQ-8842-XR looks exactly like one. The calendar entry does not exist. The user found out at the airport.
The failure here is not that the model is stupid. It is that the model was asked to do something and all it can do is say something. Closing that gap — between saying and doing — is the entire subject of agents.
What a language model actually does
Strip away the interface and a large language model is a function. You hand it a sequence of tokens (roughly, word fragments) and it returns a probability distribution over what token comes next. Sample from that distribution, append the result, feed the whole thing back in, repeat. That is the machine.
This function has three properties that matter enormously for what follows:
- It is stateless. Each call starts from nothing but the text you passed in. The model does not "remember" your last conversation; the interface re-sends the history every time, which is a very different thing.
- It has no side effects. Calling it changes no file, sends no email, moves no money. It produces text.
- It runs exactly once per request. There is no loop inside it. It cannot decide "actually, I need to check something first" and go check.
Every one of those properties is a wall the flight-booking request ran into. The model could not check prices (no side effects, no external reach), could not verify its own answer (no loop), and could not carry the booking state forward (no state).
A language model is a text-in, text-out function with no memory, no reach into the world, and no second chance to correct itself. An agent is what you build around it to remove those three limits.
The definition
An AI agent is a system that pursues a goal by repeatedly observing its environment, deciding what to do next, and taking actions that change that environment — continuing until the goal is met or it gives up.
Read that again and notice what is doing the work. Not "intelligence". Not "the model". The load-bearing words are goal, repeatedly, and actions that change the environment.
A useful longer form: an agent is an LLM placed inside a control loop, given a set of tools it can invoke, a memory that persists across steps, and a termination condition. The LLM supplies judgement. The loop supplies persistence. The tools supply reach. The memory supplies continuity.
Note what this definition does not require. It does not require the model to be large, or the task to be hard, or the system to be "autonomous" in any dramatic sense. A twenty-line Python loop that calls a model, parses one function call, runs it, and feeds the result back is an agent. A thousand-line prompt with no loop is not.
Agents versus plain LLM calls, concretely
Take the same request and trace it through both systems.
Plain LLM call
USER: Cheapest direct BLR to SIN on 14 March?[one forward pass through the model]MODEL: The cheapest direct flight is Scoot TR509 at 6,412 rupees, departing 23:55.[done — total external calls made: 0]The number 6,412 came from the model's statistical sense of what a Bengaluru–Singapore fare looks like. It might be close. It might be off by a factor of three. There is no mechanism by which it could be right except coincidence.
Agent
USER: Cheapest direct BLR to SIN on 14 March?Step 1 THINK I need live fare data. I have search_flights(). ACT search_flights(from="BLR", to="SIN", date="2026-03-14", direct_only=true) OBSERVE 4 results. Cheapest: Scoot TR509, INR 11,240, dep 23:55. Next: IndiGo 6E1015, INR 12,890, dep 09:20.Step 2 THINK TR509 is cheapest but arrives 06:30 local. The user has not stated a time preference, so I should report both and flag the red-eye. ACT finish()ANSWER: Cheapest direct is Scoot TR509 at INR 11,240, but it is a red-eye landing 06:30. IndiGo 6E1015 at INR 12,890 lands 16:05.[done — total external calls made: 1, answer grounded in real data]The agent's answer is 11,240 because a real system returned 11,240. The difference between the two runs is not model quality. Both used the same model. The difference is that one of them was allowed to go and look.
| Dimension | Plain LLM call | Agent |
|---|---|---|
| Execution shape | One forward pass | Loop, N passes until done |
| Access to the world | None — training data only | Tools: APIs, databases, shell, files |
| Knowledge freshness | Frozen at training cutoff | As fresh as the tools it calls |
| Memory across steps | Only what fits in the prompt | Explicit scratchpad plus external store |
| Error recovery | None — a wrong answer stands | Sees the error, retries or changes plan |
| Cost per request | One call, predictable | N calls, variable and harder to bound |
| Latency | Seconds | Seconds to minutes |
| Failure style | Confident fabrication | Loops, tool errors, wrong tool chosen |
| Debuggability | Easy — one input, one output | Hard — must trace the whole trajectory |
Look at the last three rows especially. Agents do not simply dominate plain calls. They trade a cheap, fast, predictable system with one nasty failure mode for an expensive, slow, unpredictable system with a dozen different failure modes — and buy correctness with that trade. When the plain call would have been right, the agent is pure overhead.
The six components
Every agent, whatever the framework, decomposes into the same six parts. If you can name all six in a system you are looking at, you understand that system.
1. Goal
What the agent is trying to achieve, and — critically — how it knows it has achieved it. A goal without a termination test is not a goal; it is a mood. "Be helpful" cannot be checked. "Return a flight number and price, or state that no direct flight exists" can be.
2. Perception
How information enters the agent. For most LLM agents this is unglamorous: the user's message, tool return values, error strings, retrieved documents. It is still perception in the technical sense — it is the only channel through which the outside world reaches the decision-maker. Its quality caps everything downstream. An agent that receives Error: request failed is blind in a way that an agent receiving HTTP 429: rate limited, retry after 30s is not.
3. Reasoning
The part that turns perceptions into a decision about the next action. In an LLM agent this is the model call, and the prompt is its programming. This is where the goal, the tool descriptions, the memory, and the latest observation get assembled into one question: given all this, what should happen next?
4. Action
The mechanism that converts a decision into an effect. Concretely: parsing the model's output into a function name and arguments, validating those arguments, calling the function, capturing what comes back — including exceptions.
5. Memory
What persists. Split it into two kinds, because they behave differently:
- Working memory — the current run's trace of thoughts, actions and observations. Lives in the prompt. Bounded by the context window, so it eventually needs trimming or summarising.
- Long-term memory — facts, preferences and past outcomes stored outside the model, in a database or vector store, retrieved when relevant. This is what lets an agent know on Tuesday that you preferred aisle seats on Monday.
6. Environment interface
The boundary between agent and world: the tool registry, the permissions on each tool, the sandbox, the rate limits, the audit log. This is the component people skip when prototyping and then rebuild in a panic after an agent deletes something. It is where you decide what the agent is allowed to do, as distinct from what it is capable of doing.
Capability and permission are different axes. An agent that can call
delete_records()because you registered the tool, and is allowed to call it because you never restricted it, will eventually call it.
The loop
Those six components arrange themselves into one cycle:
┌──────────────────────────────────┐ │ │ ▼ │ PERCEIVE ──▶ REASON ──▶ ACT ──▶ OBSERVE (input, (model (run (capture tool call) tool) result, results) errors) │ │ │ goal met, or budget │ └──────── exhausted? ──────────────┘ │ ▼ FINISHIn code, the skeleton is genuinely this small:
1def run_agent(goal, tools, model, max_steps=10):2 memory = [f"Goal: {goal}"]34 for step in range(max_steps):5 prompt = build_prompt(goal, tools, memory)6 decision = model(prompt) # REASON78 if decision.is_final: # termination test9 return decision.answer1011 try: # ACT12 result = tools[decision.tool](**decision.args)13 except Exception as e: # OBSERVE (failure)14 result = f"Tool raised {type(e).__name__}: {e}"1516 memory.append(f"Action: {decision.tool}({decision.args})")17 memory.append(f"Observation: {result}") # PERCEIVE, next round1819 return "Step budget exhausted without reaching the goal."Three details in that skeleton matter more than they look.
max_steps is not optional. Without it, an agent that keeps choosing the same failing action runs until your API budget is gone. Ten steps at roughly three cents per call is thirty cents; an unbounded loop overnight is a support ticket.
The exception becomes an observation. Notice that a tool failure is not a crash — it is fed back as text the model reads on the next pass. This single line is what turns a brittle pipeline into something that can recover.
The loop returns a real value on exhaustion. An agent that hits its budget must say so. Agents that silently return their best guess after failing are the ones that produce SQ-8842-XR.
Four agents, and what each one actually needs
| Agent | Goal & stop test | Characteristic tools | Hardest part |
|---|---|---|---|
| Research | Answer a question with citations; stop when every claim has a source | web search, page fetch, PDF extract, summarise | Knowing when it has read enough — the natural instinct is to keep searching |
| Customer service | Resolve or correctly escalate; stop on resolution or handoff | order lookup, refund, ticket create, KB search | Recognising the boundary of its own authority before issuing a refund it should not |
| Code review | Flag real defects in a diff; stop when every changed hunk is examined | read file, run tests, static analysis, git blame | Precision — an agent that flags 40 style nits gets muted and misses the real bug |
| Trading | Execute a stated strategy within risk limits; stop at market close | price feed, position query, order placement, risk check | Irreversibility — a wrong action costs money instantly and cannot be undone |
The pattern across that table: the tools are the easy part. The goal definition and the stopping rule are where the design effort goes, and the "hardest part" column is in every case a judgement problem, not an engineering one.
When an agent is the wrong answer
This is the section most people skip, and it is the one that saves the most money.
| Situation | Use | Why |
|---|---|---|
| Summarise this document | Plain LLM | Everything needed is already in the prompt |
| Translate this paragraph | Plain LLM | No external fact is required |
| Rewrite in a formal tone | Plain LLM | Pure transformation of given text |
| Classify these 50,000 tickets | Plain LLM, batched | An agent loop multiplies cost by N for zero gain |
| Fetch data, then summarise it | Fixed pipeline, not an agent | The steps are known in advance; a loop adds only unpredictability |
| What is our Q3 churn? | Agent | Needs a live database query |
| Debug this failing test | Agent | Number of steps is unknown; each result changes the next move |
| Reconcile these two systems | Agent | Requires reading, comparing, acting, verifying |
The single sharpest test: can you write down the sequence of steps in advance? If yes, write the pipeline — it will be cheaper, faster, and you will be able to debug it. Agents earn their cost only when step n+1 genuinely depends on what step n returned.
Reach for an agent when the path is unknown, not when the task is hard. Hard-but-known is a pipeline's job.
Three ways people get this wrong
Mistaking a long prompt for an agent. A prompt saying "you are an autonomous research agent, think step by step, use your tools" wired to a single model call with no tool execution is theatre. The model will happily narrate calling tools it never called. If nothing in your code parses the output and executes something, there is no agent, only prose about one.
Building a loop for a task with no branch point. If your agent always calls tool A, then tool B, then answers, you have written an unreliable, expensive version of b(a(x)). Every model call is a chance to deviate. Removing the loop removes that risk for free.
Registering tools without registering limits. Prototypes get built with a database tool that has write access because it was easier than making a read-only role. Then the agent, reasoning perfectly reasonably from a badly worded goal, decides the way to fix a duplicate row problem is to delete rows. The reasoning was fine. The permission was the bug.
What this means when you build one
The practical consequence of all of the above is that agent design is mostly not model work. Assume the model is fixed and competent. Your leverage is in four places, roughly in order of impact:
- The stopping rule. Decide, before you write a line, what a finished run looks like and what the step budget is. Most bad agents are bad because nobody answered this.
- The observation text. Whatever your tools return is the agent's entire view of reality. Return structured, specific, actionable strings — including on failure.
Error: 404, no order with id 8842; check the id format (expected ORD-XXXXX)lets an agent recover.Errordoes not. - The tool surface. Fewer, sharper tools beat many overlapping ones. Every extra tool is another chance to pick the wrong one.
- The permission boundary. Give each tool the narrowest access that works. Read-only by default. Anything irreversible — payments, deletions, outbound messages — gets a confirmation gate or a hard cap.
And keep the flight booking in mind as your calibration example. The model that invented SQ-8842-XR was not malfunctioning. It was doing exactly what it does: producing the most plausible continuation of the text it was given. Plausible is not the same as true, and the only thing that closes that gap is letting the system go and check.