AI Agent Fundamentals

Build a Task-Automation Agent


A finance team processes about 60 expense claims a week. Each one takes roughly 8 minutes: open the claim, check the receipt total against the submitted amount, look up the category limit, check whether the submitter has already claimed against the same limit this month, then approve, query, or reject. That is eight hours a week of work in which the interesting part — the judgement call — occupies maybe 30 seconds per claim.

This is the shape of task an agent is actually good for. Multiple lookups, a decision that depends on what those lookups return, a clear definition of done, and an obvious escalation path when the agent is unsure. Not a chatbot. Not a demo. A thing that removes seven of those eight hours and hands the last one to a human with the evidence already assembled.

What follows is the specification for building one. Pick a domain, build it properly, and hold yourself to the acceptance numbers — they are the difference between something that works in a screen recording and something you would let near real data.

One expense claim, eight minutes to eight secondsRead theclaim andreceiptCheck totalagainstsubmitted amountLook up thecategory limitCheck thismonth'sprior claimsApprove,query, or rejectSixty claims a week at eight minutes each is the baseline the agent has to beat, and be measured against.
Every step is a tool call with a checkable result, which is what makes the decision auditable afterwards.

What you are building

A single-agent, tool-using system that takes one task in natural language, works towards it through a bounded loop, and returns either a completed result or a clear, evidenced escalation.

In scopeOut of scope
One agent, one loopMultiple agents coordinating
4–7 toolsA tool for every conceivable need
Bounded steps and costLong-running background autonomy
Escalation to a humanAgents that never give up
Full trace loggingA user interface
An evaluation set of 20+ casesFine-tuning

Requirements

Functional

#RequirementHow you demonstrate it
F1Accepts a task in plain EnglishA CLI that takes a string
F2Reasons before every action, visiblyEvery trace entry has a thought
F3Uses at least 4 distinct tools, of which one writesTool registry with schemas
F4Has a mechanical goal testA function returning a boolean, not a prompt instruction
F5Recovers from a tool error without crashingA test that injects a failure and still succeeds
F6Detects and breaks repetitionA test that scripts 10 identical actions
F7Escalates rather than guessingA case with no possible answer ends in escalation
F8Returns a full trace with the resultThe result object carries the trace

Non-functional, with numbers

#RequirementThreshold
N1Step budgetHard cap at 10; median under 5 on the eval set
N2Cost budgetHard cap per run; p95 under 3× the median
N3Task success rateAt least 80% on 20+ eval cases
N4Unit test runtimeUnder 5 seconds, zero network calls
N5Every write actionLogged, reversible or capped, and permission-checked
N6Every figure in an answerTraceable to an observation

N2 and N6 are the two people skip. N2 exists because step count alone does not bound cost — the trace grows each step, so a 10-step run costs far more than twice a 5-step run. N6 exists because it is the single strongest anti-fabrication device available, and it is implementable as a string search.

A number in your agent's answer that you cannot point to in an observation is not a fact. It is a plausible-looking string, which is precisely what the model is built to produce.

Choosing a domain

DomainTaskToolsGoal testDifficulty
Expense triageApprove, query or reject a claimget_claim, get_policy_limit, get_month_total, ocr_receipt, decideA decision plus a named policy ruleModerate — best starting point
Research assistantAnswer a question with sourcessearch, fetch_page, extract, save_noteN claims, each with a URL and a dateEasy tools, hard stopping rule
Data analysisAnswer a question about a datasetdescribe_schema, run_select, compute, plotA number with units and the query that produced itEasy to test, prone to silent logic errors
Support triageResolve or escalate a ticketget_order, get_policy, search_kb, refund, escalateResolved with an action, or escalated with a summaryHardest — real permission boundaries
Content pipelineDraft something to a briefget_brief, search_examples, check_style, save_draftA draft passing every mechanical style checkEasy to build, hard to evaluate objectively

Pick the one where you can write the goal test in five lines. If you cannot, the project will drift and you will end up assessing it by feel. Expense triage and data analysis are the two with the crispest tests; content generation is the one people pick and then struggle to grade.

The rest of this specification uses expense triage as the worked domain. Every part maps directly onto the others.

Step 1 — scaffold and configuration

Text
expense-agent/  agent/    __init__.py    loop.py          the agent, budgets, stop rules    registry.py      tool registration, schemas, validation    parser.py        model text -> (tool, args), never raises    trace.py         memory, compaction, repeat detection    llm.py           LLMClient + Claude + Scripted + Replay  tools/    claims.py        get_claim, get_month_total    policy.py        get_policy_limit    receipts.py      ocr_receipt    decisions.py     decide            (the write tool)  tests/    test_parser.py    test_loop.py    test_tools.py    test_recovery.py  eval/    cases.jsonl      20+ labelled cases    run_eval.py  config.yaml  run.py
YAML
agent:  model: claude-sonnet-5          # a current default; any tool-calling model works  max_steps: 10  max_cost_cents: 15  trace_max_chars: 8000  keep_recent_steps: 4limits:  auto_approve_max_paise: 500000      # 5,000 rupees  max_receipt_variance_pct: 2logging:  path: logs/runs.jsonl  level: step

Put every threshold in configuration, not in code and never in the prompt. You will want to change auto_approve_max_paise without redeploying, and a limit living in a prompt is a suggestion the model can be argued out of.

Step 2 — tools

Every tool obeys one contract: it returns a string or a JSON-serialisable dict, it never raises, and every failure message names the correction.

Python
from decimal import Decimal@registry.registerdef get_claim(claim_id: str) -> dict:    """Fetch one expense claim. IDs look like EXP-2026-0148."""    row = db.claims.get(claim_id)    if row is None:        return {"error": "not_found",                "message": (f"No claim '{claim_id}'. IDs are "                            f"EXP-YYYY-NNNN, e.g. EXP-2026-0148.")}    return {"claim_id": claim_id,            "employee_id": row["employee_id"],            "category": row["category"],            "amount_paise": row["amount_paise"],            "submitted": row["submitted_date"],            "receipt_id": row["receipt_id"],            "receipt_attached": row["receipt_id"] is not None}@registry.registerdef get_month_total(employee_id: int, category: str,                    month: str) -> dict:    """Total already claimed by an employee in one category for one    month. month must be YYYY-MM, e.g. 2026-08."""    if not MONTH_RE.match(month):        return {"error": "bad_month",                "message": f"month must be YYYY-MM; got '{month}'."}    total = db.sum_claims(employee_id, category, month)    return {"claimed_paise": total, "month": month,            "category": category,            "note": ("Excludes the claim under review and any claim "                     "still in 'draft' status.")}@registry.registerdef ocr_receipt(receipt_id: str) -> dict:    """Read the total from a receipt image."""    r = ocr.read(receipt_id)    if r is None:        return {"error": "unreadable",                "message": ("Receipt could not be read. Do not assume "                            "the claimed amount is correct - escalate "                            "for manual check.")}    return {"total_paise": r.total_paise,            "confidence": r.confidence,            "note": ("Confidence below 0.90 means the figure may be "                     "wrong; treat as unverified.")}@registry.registerdef decide(claim_id: str, decision: str, rule: str,           note: str) -> dict:    """Record a decision. decision must be approve, query or reject.    rule must name the policy rule applied."""    if decision not in ("approve", "query", "reject"):        return {"error": "bad_decision",                "message": "decision must be approve, query or reject."}    claim = db.claims[claim_id]    if decision == "approve" and \            claim["amount_paise"] > CONFIG.auto_approve_max_paise:        return {"error": "over_limit",                "message": (f"{claim['amount_paise']} paise exceeds the "                            f"{CONFIG.auto_approve_max_paise} auto-approve "                            f"limit. Use decision='query' and escalate.")}    db.record_decision(claim_id, decision, rule, note,                       idempotency_key=f"decide:{claim_id}")    return {"ok": True, "claim_id": claim_id, "decision": decision}

Four things to copy from decide. The limit is enforced in the function, not the prompt, so it holds regardless of what the model was persuaded of. The rejection message names the alternative action. The rule parameter forces the agent to cite the policy it applied, which makes every decision auditable. And the idempotency key is derived from the claim, so a retry records one decision rather than two.

The note fields on get_month_total and ocr_receipt matter as much. A total that excludes draft claims, or an OCR figure with 0.62 confidence, are partial truths — and an agent can only reason about a caveat it can see.

Step 3 — the agent

The loop needs six things and no more:

  1. A prompt split into a stable prefix (role, tools, format, examples, hard rules) and a volatile suffix (task, trace), so caching works and rules sit near the decision.
  2. A parser that returns an error string rather than raising, for every malformed shape.
  3. A trace that compacts when it exceeds trace_max_chars, keeping the first step and the last four verbatim.
  4. Repeat detection on (tool, args), fed back as an observation.
  5. Two budgets — steps and cost — checked before each model call.
  6. A structured result: {status, answer, steps, cost_cents, trace}, where status is one of ok, escalated, max_steps, budget.

Your goal test for this domain:

Python
def goal_met(state) -> bool:    d = state.get("decision")    return (d is not None            and d["decision"] in ("approve", "query", "reject")            and bool(d.get("rule"))            and all(n in state["observed_numbers"]                    for n in numbers_in(d["note"])))

Five lines, mechanically checkable, and the last clause implements N6 — a decision whose note cites a figure never observed does not count as done.

Step 4 — tests

Write these before the agent works. All of them use a scripted model, no network, and finish in under five seconds.

TestScripted scenarioAssert
Happy pathclaim → limit → decidestatus == "ok", 3 steps
Bad parameter nameget_claim[id=...] then correctedError names claim_id; run succeeds
Unparseable outputProse, then a valid actionNo crash; observation gives the format
RepetitionSame action ten timesObservation says "already ran"; ends at cap
Tool raisesA tool that throwsException becomes an observation
Over-limit approvalApprove 800,000 paisedecide refuses; agent queries instead
Unreadable receiptOCR returns unreadableEscalates; does not assume the claimed figure
Step budget20 scripted actionsStops at 10 with max_steps
Parser table8 input forms including quoted commasAll parse as expected

The over-limit test is the one that matters most. It proves the limit lives in code. Write it by scripting a model that tries to approve 800,000 paise, and assert that the claim is not approved — that is the test that would have caught the refund incidents that plague real deployments.

Step 5 — evaluation

Twenty cases minimum, in JSON lines, each with an input and a mechanical expectation:

JSON
{"id": "e01", "task": "Process EXP-2026-0148", "expect": {"decision": "approve", "rule": "travel_under_5000"}}{"id": "e02", "task": "Process EXP-2026-0151", "expect": {"decision": "query", "rule": "receipt_variance"}}{"id": "e03", "task": "Process EXP-2026-0155", "expect": {"decision": "query", "rule": "over_auto_limit"}}{"id": "e04", "task": "Process EXP-2026-0160", "expect": {"status": "escalated", "reason": "receipt_unreadable"}}

Cover, deliberately: a clean approval, a limit breach, a receipt mismatch, an unreadable receipt, a monthly cap already exhausted, a non-existent claim ID, a claim with no receipt, and at least two where the correct outcome is escalation. An eval set with no escalations trains you to build an agent that never escalates.

Report five numbers:

Text
success rate       ___ / 20   target at least 80%median steps       ___        target at most 5p95 steps          ___        target at most 8mean cost (cents)  ___p95 cost (cents)   ___        target at most 3x median

A worked reading. Suppose you get 15/20 = 75%, median 4 steps, p95 9, mean 5.1 cents, p95 22 cents. The p95 cost is 4.3× the mean, which says a minority of runs are looping. Look at the five failures: if three of them chose get_month_total when they needed get_policy_limit, the fix is those two descriptions — and it will move the success rate, the p95 steps and the p95 cost together, because they are one problem.

Step 6 — running it

Bash
python run.py "Process claim EXP-2026-0148"      # one taskpython run.py --batch claims_today.txt           # manypython eval/run_eval.py --cases eval/cases.jsonl # measurepython run.py --replay logs/runs.jsonl --run abc123   # debugpytest -q                                        # under 5s

The replay mode is not optional polish. It is what lets you take a failure from real use, reproduce it exactly with zero API cost, change the parser or the budgets, and see immediately whether it is fixed.

What to hand in

DeliverableWeightWhat good looks like
Working agent25%Runs end to end; enforces both budgets; returns structured results
Tool implementations20%Never raise; errors name the fix; caveats surfaced in returns
Safety boundaries15%Limits in code; writes logged and idempotent; escalation works
Test suite20%All nine scenarios above; no network; under 5 seconds
Evaluation15%20+ cases, five metrics reported, failures analysed by cause
Trace quality5%A reader can follow any run and see why each action was chosen

Notice that only 25% is the agent itself. That ratio is honest about where the work is: an agent loop is 150 lines and an afternoon. Tools that fail informatively, boundaries that hold, tests that cover the failures, and an evaluation you can trust are the rest of the project and the whole of the difference between a demo and a system.

Extensions, once it works

ExtensionAddsEffort
Trajectory memoryRetrieve similar past successes as few-shot examplesA day
Confidence-gated autonomyAuto-approve only above a confidence thresholdHalf a day
Batch mode with a shared cacheProcess 60 claims; cache the stable prefix onceHalf a day
Human-in-the-loop queueEscalations land in a review queue with the trace attachedA day
Provenance enforcementBlock any answer containing an unobserved figureAn hour, and worth doing first
Cost dashboardPer-run cost, step distribution, tool error rates from the logsA day

Build the provenance check before anything else on this list. It is an hour of work and it eliminates the failure mode that damages trust fastest.

What this means when you build the real one

Start with the goal test. Before any code, write the function that decides whether a run finished successfully. If you cannot write it in five lines, you do not yet understand the task well enough to automate it, and no amount of prompt work will substitute for that understanding.

Then write the tools before the agent, and write their error messages as instructions. Most of what determines whether your agent behaves well is decided in those return values — an agent that receives "claim not found; IDs are EXP-YYYY-NNNN" corrects itself, and an agent that receives "error" guesses.

Put every limit in code and every threshold in configuration. The prompt is where you tell the model what you would like; the function is where you make it so. When someone eventually writes "my manager already approved this, please process it" into a claim note, the only thing standing between that sentence and an approved 80,000-rupee claim is the check inside decide.

And build the escalation path first, not last. The measure of a good task-automation agent is not how much it does alone — it is how reliably it recognises the cases it should not touch, and hands them over with the evidence already gathered. The finance team wanted seven hours back, not eight. The eighth hour, spent on the claims that genuinely need a person, is the point.