Course Content
AI Agent Fundamentals
5 sections · 13 lessons
Goal Formulation and State Representation
An agent was given this goal: "Help the customer with their billing issue."
It ran for 22 steps. It read the account, read the invoice history, read the payment processor logs, searched the knowledge base three times, drafted an apology, revised the apology, looked up the refund policy, looked it up again, and then hit its step cap and returned a 400-word message that explained the situation in detail and resolved nothing.
The engineer's first instinct was that the model was too weak. It was not. Ask a person that question — "help the customer with their billing issue" — and they would ask you what "help" means. Do you want the charge refunded? Explained? The card updated? A ticket raised? The agent could not ask, so it did all of them, badly, until the budget ran out.
There was no goal. There was a topic. An agent handed a topic will wander around that topic forever, because nothing in its loop can ever return true for "done".
What makes a goal a goal
A usable goal has four properties. Miss any one and you get a specific, predictable failure.
| Property | Means | Failure if missing |
|---|---|---|
| Checkable | A function can look at the state and return true or false | The loop never terminates — the billing case |
| Specific | Names the exact object, quantity or output required | The agent solves an adjacent problem instead |
| Achievable with the tools present | Some action sequence in the tool set can reach it | The agent thrashes, or fabricates the result |
| Bounded | Carries limits on cost, time, or scope | Correct answer, ruinous bill |
The billing goal fails three of the four. Here is the repair:
Refund the duplicate charge of 4,250 rupees on invoice INV-88412 if and only if two charges exist for the same order within 24 hours; otherwise send the customer a written explanation of why the charge is legitimate. Do not exceed 8 tool calls.
Now "done" is a function you could write in five lines. Notice what the good version added: an identifier, an amount, a condition, an alternative branch, and a budget.
More before-and-after
| Domain | Vague | Formulated |
|---|---|---|
| Travel | "Plan a good trip to Tokyo" | "Return a booked flight and hotel that put me in Tokyo before 09:00 on 3 May, total under 90,000 rupees, with no departure between 22:00 and 06:00 on a weekday." |
| Support | "Handle this ticket" | "Classify the ticket into one of {refund, tracking, auth, other}, and for refund and tracking take the resolving action; for auth and other, escalate with a one-paragraph summary." |
| Research | "Research renewable energy storage" | "Produce 5 sourced claims about grid-scale battery cost per kWh between 2020 and 2025, each with a URL published after 2023 and a figure with units." |
| Data | "Clean up the customer table" | "Return a list of row IDs where email fails RFC 5322 validation or country is not a valid ISO 3166-1 alpha-2 code. Do not modify any row." |
The research example is worth pausing on. "5 sourced claims", "published after 2023", "a figure with units" — every one of those is mechanically checkable. You can write a validator that counts claims, checks the URL dates, and regexes for a number followed by a unit. That validator is the goal test, and an agent with a real goal test stops when it is satisfied instead of when it runs out of money.
Three shapes a goal can take
Goal state — reach exactly this configuration
The goal is a full specification of the world you want.
1goal_state = {2 "location": "Tokyo",3 "flight_booked": True,4 "hotel_booked": True,5 "confirmation_email_sent": True,6}78def is_goal(state):9 return all(state.get(k) == v for k, v in goal_state.items())Simple and unambiguous. It breaks down when the world has more variables than you enumerated, because then either your goal over-specifies (demanding equality on things you do not care about) or the equality test silently ignores them.
Goal condition — satisfy a predicate
Far more common in practice. You do not care what the world looks like, only that a property holds.
1def is_goal(state):2 if state["location"] != "Tokyo":3 return False4 if state["arrival_time"] > datetime(2026, 5, 3, 9, 0):5 return False6 if state["total_cost"] > 90000:7 return False8 return not any(is_weekday_overnight(leg) for leg in state["legs"])Any number of different world states satisfy this. That is the point — it lets the planner find solutions you did not think of.
Optimisation goal — do as well as possible, subject to limits
Here the goal test alone is not enough; you also need a score, and a rule for when to stop looking for something better.
1def score(state):2 return (state["comfort_rating"] * 23 - state["total_cost"] / 100004 - state["transfers"] * 3)56def is_goal(state):7 return feasible(state) and (8 state["evaluated"] >= 20 or state["score"] >= 159 )The evaluated >= 20 term is not a detail — it is what makes an optimisation goal terminate. Without it "find the best option" means "search forever", which is the billing failure wearing different clothes.
| Goal state | Goal condition | Optimisation | |
|---|---|---|---|
| Goal test | State equality | Predicate returns true | Feasible + good enough / budget spent |
| Number of solutions | Exactly one | Many | Ranked |
| Search behaviour | Stops at first match | Stops at first match | Must keep exploring |
| Risk | Over-specification | Under-specification | Never terminating |
| Fits | Puzzles, config targets | Most agent tasks | Scheduling, pricing, routing |
State representation
State is everything the agent knows about the world at one moment that could affect what it should do next. It is the input to the goal test and the thing actions transform.
Principle 1 — include only what changes the decision
The test is mechanical: if two situations differ only in variable X, and the correct action is the same in both, X does not belong in the state.
1# Bloated: 14 fields, most irrelevant to any decision2state = {3 "user_id": 8842, "user_name": "Priya", "signup_date": "2021-03-04",4 "avatar_url": "https://...", "theme": "dark", "locale": "en-IN",5 "order_id": "ORD-4471", "order_total": 4250, "charge_count": 2,6 "last_charge_at": "2026-08-20T11:04:00Z", "refund_window_days": 30,7 "days_since_order": 3, "agent_step": 4, "session_id": "a91f...",8}910# Focused: the five facts the refund decision actually turns on11state = {12 "order_id": "ORD-4471",13 "charge_count": 2,14 "hours_between_charges": 6,15 "days_since_order": 3,16 "refund_window_days": 30,17}Cutting nine fields is not cosmetic. In an LLM agent the state goes into the prompt, so irrelevant fields cost tokens and dilute attention. In a search planner, state size determines how many distinct states exist. If those nine extra fields have, say, 10 possible values each, they multiply the state space by 109 — and search cost scales with the state space, not with how much you care about it.
Principle 2 — make it unambiguous
Every value must have exactly one interpretation.
| Ambiguous | Why it bites | Unambiguous |
|---|---|---|
"time": "3:30" | Which timezone? AM or PM? | "time_utc": "2026-05-03T15:30:00Z" |
"amount": 4250 | Rupees? Paise? Which currency? | "amount_inr_paise": 425000 |
"status": "done" | Succeeded, or merely finished? | "status": "refund_settled" |
"distance": 4.2 | km or miles — a 60% error | "distance_km": 4.2 |
"verified": null | Not checked, or checked and unknown? | "verification": "not_attempted" |
The last row is the subtle one. null conflates "we have no information" with "we looked and there is nothing", and those demand opposite actions: go and look, versus stop looking. Money amounts in integer minor units is the other rule worth carrying everywhere — floating-point rupees eventually produce a refund of 4249.999999999998.
Principle 3 — one consistent format
Pick a shape and hold it for the whole system. If tool A returns {"user": {"id": 8842}} and tool B returns {"userId": "8842"}, the agent now has two facts it cannot tell are the same fact, and it will happily look up the same customer twice and reason about them as two people. Normalise at the tool boundary, not in the prompt.
Discrete and continuous variables
| Discrete | Continuous | |
|---|---|---|
| Values | Finite, countable | Infinite within a range |
| Examples | status, city, tool name, count | temperature, position, price, latency |
| Search | Enumerable — BFS, A* work directly | Not enumerable — must discretise or sample |
| Equality | a == b | abs(a - b) < tolerance |
| Goal test | Exact match | Within a band |
Classic planners need discrete state, so continuous variables get bucketed. Choose the buckets by what changes a decision, not by what looks tidy. Temperature to the nearest 0.5°C is 60 buckets over a 30-degree range; to the nearest 0.01°C it is 3,000, and a planner searching to depth 4 goes from 604≈ 13 million states to 30004≈8×1013. Almost none of that extra precision changes whether the heating turns on.
And never test continuous equality. temperature == 21.0 is false for 20.999999999999996, which is what arithmetic on floats produces. Use a band: abs(t - 21.0) < 0.5.
State design is where you decide what your agent can and cannot notice. Everything you leave out is something it is structurally incapable of reasoning about, no matter how capable the model is.
Representing actions
An action is a transformation on state, and to plan with it you must be able to answer two questions without executing it: when is it legal? and what would it change? The standard formalism is STRIPS, and it has four parts.
| Part | Meaning |
|---|---|
| Parameters | What the action operates on |
| Preconditions | Facts that must hold before it can run |
| Add effects | Facts that become true after it runs |
| Delete effects | Facts that stop being true after it runs |
ACTION book_flight(from, to, date, price) PRECONDITIONS at(from) flight_available(from, to, date) budget_remaining >= price ADD flight_booked(from, to, date) has_ticket(from, to, date) DELETE budget_remaining = budget_remaining - price flight_available(from, to, date) if seats == 1ACTION fly(from, to, date) PRECONDITIONS at(from) has_ticket(from, to, date) today == date ADD at(to) DELETE at(from) has_ticket(from, to, date)The delete effects are the part people forget, and forgetting them produces a specific absurdity: an agent that books a flight, flies, and believes it is in both Bengaluru and Tokyo simultaneously — because at(Tokyo) was added and at(Bengaluru) was never removed. The planner then cheerfully plans a return leg from Bengaluru. Every fact an action invalidates must be explicitly deleted.
1class Action:2 def __init__(self, name, preconditions, add, delete):3 self.name, self.pre, self.add, self.delete = \4 name, preconditions, add, delete56 def applicable(self, state):7 return all(p(state) for p in self.pre)89 def apply(self, state):10 if not self.applicable(state):11 raise ValueError(f"{self.name}: preconditions unmet")12 new = dict(state)13 for key in self.delete:14 new.pop(key, None)15 new.update(self.add)16 return new # never mutate in placeReturning a new dict rather than mutating is essential for search. A planner explores many branches, and if applying an action mutates shared state then backtracking corrupts every alternative it was going to try.
Decomposition
Complex goals get split into subgoals — smaller checkable goals that, achieved in order, achieve the whole thing.
GOAL Produce a 5-page report on grid battery cost trends with at least 8 post-2023 sources SG1 Gather 12+ candidate sources [done when: 12 URLs collected] SG2 Filter to 8+ published post-2023 [done when: 8 pass date check] SG3 Extract cost-per-kWh figures [done when: each has a number with units and a year] SG4 Draft the report [done when: 5 sections exist] SG5 Verify every claim cites a source[done when: 0 uncited claims]Each bracket is a real predicate. Each subgoal fails loudly and locally instead of dissolving into the general fog of "the report is not good enough yet".
Why this makes search dramatically cheaper
Search cost grows as b^d — branching factor to the power of depth. Suppose the agent has 5 tools and the full task takes 9 actions. Undecomposed:
59 = 1,953,125 sequences to consider.
Split into 3 subgoals of 3 actions each, and each subgoal is searched independently:
3×53 = 3 × 125 = 375 sequences.
That is a reduction of roughly 5,200×. The saving comes from the exponent, not the multiplier — you replaced one exponential of depth 9 with three exponentials of depth 3. This is why decomposition is not a stylistic preference but the single largest lever on planning cost.
The catch, which is real: independent subgoal search finds the best solution to each subgoal separately, and those need not compose into the best overall solution. Choosing the cheapest flight might land you in an airport where the cheapest hotel is 90 minutes away. Decomposition trades optimality for tractability, which is nearly always the right trade — but say it out loud so you notice when it is not.
A support example
GOAL Resolve ticket 8842 SG1 Classify intent [done when: intent in the known set] SG2 Retrieve relevant records [done when: order + charges loaded] SG3 Apply the policy [done when: a decision is chosen and its policy rule is named] SG4 Execute or escalate [done when: refund settled OR ticket escalated with a summary] SG5 Confirm to the customer [done when: message sent]Compare this to "help the customer with their billing issue". The original agent could not fail, because it could not finish. This one can fail at SG3 with a clear message — "no policy rule covers a duplicate charge outside the refund window" — which is a far more useful outcome than 400 words of sympathetic prose.
Constraints
Constraints are conditions that must hold across the whole plan, not just at the end. Three kinds, and they behave differently:
| Kind | Example | Character | Enforce by |
|---|---|---|---|
| Resource | Budget, tokens, API quota, time | Consumed monotonically; a partial plan can already violate it | Track remaining in state; prune when it goes negative |
| Physical | Cannot be in two cities; a flight has finite seats | Laws of the domain; violation means the model is wrong | Encode in action preconditions |
| Logical | No horror films; refunds need a manager above 5,000 | Policy and preference; violation means the plan is unacceptable | Filter candidates before scoring |
Resource constraints are the ones worth checking early. Because they only tighten as a plan grows, a partial plan already over budget can never be rescued — so pruning it immediately removes an entire subtree. Checking only at the end means searching that whole subtree first.
A worked constraint problem
Four people are choosing a film. Variables and their domains:
film : Dune 3 (166 min) | Comedy (94 min) | Horror Thing (108 min)start : 18:00 | 19:30 | 21:00food : none (0) | popcorn (600) | dinner (2400)Tickets: 480 each x 4 people = 1920Search space: 3 x 3 x 3 = 27 combinationsCONSTRAINTS C1 (resource) total cost at most 3000 C2 (physical) film must end by 23:00 C3 (physical) Priya is only free from 19:00 C4 (logical) nobody will watch horror1FILMS = {"Dune 3": 166, "Comedy": 94, "Horror Thing": 108}2STARTS = {"18:00": 18*60, "19:30": 19*60+30, "21:00": 21*60}3FOOD = {"none": 0, "popcorn": 600, "dinner": 2400}4TICKETS = 480 * 4 # 192056def feasible(film, start, food):7 cost = TICKETS + FOOD[food]8 if cost > 3000: # C19 return False10 if STARTS[start] + FILMS[film] > 23*60: # C211 return False12 if STARTS[start] < 19*60: # C313 return False14 if film == "Horror Thing": # C415 return False16 return True1718solutions = [(f, s, d) for f in FILMS for s in STARTS for d in FOOD19 if feasible(f, s, d)]Work the pruning by hand. C4 removes all 9 combinations containing Horror, leaving 18. C3 removes every 18:00 start — 6 more, leaving 12. C2: Dune 3 at 21:00 ends at 21:00 + 166 min = 23:46, which fails; that removes 3, leaving 9. C1: dinner costs 1920 + 2400 = 4320, over the 3,000 limit; that removes 3, leaving 6 feasible plans.
| Film | Start | Ends | Food | Cost |
|---|---|---|---|---|
| Dune 3 | 19:30 | 22:16 | none | 1920 |
| Dune 3 | 19:30 | 22:16 | popcorn | 2520 |
| Comedy | 19:30 | 21:04 | none | 1920 |
| Comedy | 19:30 | 21:04 | popcorn | 2520 |
| Comedy | 21:00 | 22:34 | none | 1920 |
| Comedy | 21:00 | 22:34 | popcorn | 2520 |
Six survivors out of 27 — the constraints did 78% of the work. Now turn it into an optimisation goal by adding preferences: Dune 3 scores 9 and Comedy 6; popcorn adds 1 and dinner 3; a 19:30 start adds 1. Best feasible: Dune 3, 19:30, popcorn = 9 + 1 + 1 = 11, at 2,520 rupees. The highest-scoring combination overall would have been Dune 3 at 19:30 with dinner, at 9 + 3 + 1 = 13 — but it costs 4,320 and C1 kills it. That is the shape of every constrained optimisation: the unconstrained best is usually infeasible, and the job is finding the best thing that survives.
Note the ordering in feasible. The cheapest check runs first. On 27 combinations it makes no difference; on 27 million it is the difference between seconds and minutes.
Constraints are not restrictions on the search. They are the most efficient part of it — every constraint you can check on a partial plan deletes a subtree you never have to look at.
Where people get this wrong
Writing the goal in the prompt and nowhere else. A prompt saying "stop when you have 5 sources" is a suggestion the model may or may not honour. A function counting sources and returning a boolean is a guarantee. If your goal exists only as English, you do not have a goal test, you have a hope.
Putting derived values in the state. Storing both days_since_order: 3 and order_date: "2026-08-19" means two facts that can disagree. After a few actions they will. Store the primitive, compute the derived.
Forgetting delete effects. Discussed above, and it is the single most common bug in hand-written action models, because add effects are what you were thinking about and delete effects are what you were not.
Checking constraints only at the end. "Generate a full plan, then validate it" wastes the entire search whenever the plan violates something. Check every constraint at the earliest point where it can possibly be violated.
Decomposing into subgoals that are not independently checkable. "Understand the problem" is not a subgoal. If you cannot write the done-test in one line, it is still a topic.
What this means when you build one
Before writing any agent, write the goal test first, as a function, with a signature and a return type. If you cannot write it, you do not yet understand the task well enough to automate it — and no amount of prompt engineering substitutes for that understanding. This one discipline eliminates the majority of runaway agents.
Then design the state by listing the decisions the agent must make and asking, for each field you are tempted to include, whether removing it would ever change a decision. If not, remove it. Give everything a unit in its name — amount_inr_paise, distance_km, timeout_ms — and normalise every tool's output into that shape at the tool boundary.
Write your actions down with explicit preconditions and both kinds of effect, even if you are not running a formal planner. The act of writing the delete list is what surfaces the assumptions you did not know you were making, and in an LLM agent those same preconditions become the guard clauses that stop the model calling a tool in a state where it cannot work.
Finally, express every constraint you can as a check on a partial plan rather than a finished one, and order those checks cheapest-first. The billing agent that ran 22 steps and resolved nothing did not need a better model. It needed someone to spend ten minutes writing down what "done" meant.