Course Content
Agentic AI Patterns
9 sections · 50 lessons
Explain the end-to-end process of building an AI agent, from ingestion to evaluation.
What you need to know
The process is a loop, not a line. The important idea is baseline first: you cannot say an agent is better unless you measured something simpler.
- Define the task and "done" — write what a correct run looks like and how code could check it. If you cannot, you cannot evaluate.
- Collect real examples — 50 to 200 real requests, including the messy ones. This is the golden set.
- Baseline — one good prompt, maybe with retrieval. Measure it. Many projects rightly stop here.
- Ingest knowledge — chunk, embed, build a hybrid (keyword plus vector) index, and store permission metadata for filtering.
- Build tools — typed schemas, narrow purpose, server-side authorisation, clear error messages, timeouts. Unit-test each tool without the model.
- Orchestrate — add the least dynamic pattern that closes the gap. Make state explicit and checkpointed. Set step, time and spend budgets.
- Guardrails and human approval — gates on irreversible actions, output validation, a refusal path.
- Instrument — trace every model call, tool call and retrieval with prompt version, tokens and cost.
- Evaluate — offline in CI at component, trajectory and outcome level; online with A/B tests and user feedback.
- Ship behind a flag and iterate — every failed trace becomes a new golden-set case.
The three levels of evaluation
- Component: did retrieval find the right document? Was the tool call valid?
- Trajectory: did the agent take a sensible path, without loops or wasted calls?
- Outcome: was the task actually done? Is the claim decision correct?
Why the order matters
Teams that start at step 6 build a loop with 15 tools, get 55% success, and have no idea why. Teams that start at step 3 know the single call already got 70%, so they know the agent must beat 70% to earn its extra cost and latency.
A real-life example
An insurer wants to automate first review of motor claims. Here is how the build went, step by step:
- Done was defined as: "the draft decision matches the adjuster's final decision, and all required documents are listed."
- Golden set: 180 past claims, including 40 with missing documents and 15 suspected fraud.
- Baseline: one call with the claim text and the policy wording scored 62%. It failed on claims that needed prior-claim history.
- Tools added:
get_policy,get_prior_claims,request_document. Each tested alone. - Orchestration: a router sends clean claims to a fixed extraction chain (about 65% of volume) and messy ones to a small agent loop with a 10-step limit.
- Guardrail: any deny or payout above Rs 2 lakh goes to an adjuster.
- Result: 81% agreement offline. After two weeks online, 23 disagreements were added to the golden set, mostly around a regional garage network the prompts did not know about.
Follow-up questions to expect
- "How big should the golden set be?" — Start with 50 to 100 real cases so you can move fast, and grow it from production failures. Coverage of hard cases matters more than size.
- "When do you stop at the baseline?" — When a single call or fixed chain meets the success bar within cost and latency limits. Adding a loop then only adds risk.
- "What do you monitor after launch?" — Task success, escalation rate, cost per successful task, p95 latency, and per-tool error rates.