Agentic AI Patterns

Course Content

Agentic AI Patterns

9 sections · 50 lessons

How do AI agents use external knowledge bases and APIs effectively?


One booking, three attempts, one ticketAttempt 1: 503,wait about 0.5 sAttempt 2: 429,wait about 1 sAttempt 3:200,booking HX91Same idempotencykey on all threeA 400 such as a bad date is never retried; it goes back to the model.
Backoff protects the airline, and the reused key protects the traveller from a second ticket.

What you need to know

Enforce in code, not in the prompt

  • Authorisation. Run the call with the user's token, and filter results on the server. Never give the agent an admin key and ask it to behave.
  • Shape. Return 10 useful fields, not 140. Paginate. Cap size.
  • Validation. Check arguments against the schema before calling. Validate the response before it enters the prompt.

Reliability patterns

Python
import random, time, uuidTRANSIENT = {429, 502, 503, 504}def call_with_retry(send, payload, max_attempts=4, base=0.5):    key = str(uuid.uuid4())                        # same key on every retry    for attempt in range(1, max_attempts + 1):        status, body = send(payload, idempotency_key=key)        if status < 400:            return body        if status not in TRANSIENT or attempt == max_attempts:            return {"error": f"{status}: {body}", "retryable": False}        delay = base * 2 ** (attempt - 1)          # 0.5, 1, 2 seconds        time.sleep(delay * random.uniform(0.5, 1.5))   # jitter spreads retries out

Only transient errors (rate limits and gateway errors) are retried; a 400 "bad date" goes straight back to the model as an error it can fix. The idempotency key is created once and reused, so if the first attempt actually succeeded before timing out, the retry does not book twice.

  • Circuit breaker: after repeated failures, stop calling the dependency for a while and return "service unavailable" at once, so the agent does not burn its step budget.
  • Caching: cache read-heavy, slow-changing data with a short TTL.

Tool design drives quality

Precise names, one purpose per tool, enums instead of free text, and actionable error messages do more than prompt tuning. Track per-tool call volume, success rate, latency and how often the model picks the wrong tool.

A real-life example

A travel-booking agent calls an airline booking API. During a sale, the API is overloaded.

Before the fix: the agent's book_flight tool retried any error three times with no key and no delay. A request timed out after the airline had already issued the ticket; the retry issued a second ticket. Over one sale weekend, 37 travellers were double-booked and refunds took two weeks.

After the fix:

  • Each booking gets one idempotency key; retries reuse it and return the existing ticket.
  • Only 429 and 5xx are retried, with backoff and jitter.
  • After 5 failures in 30 seconds, the circuit opens for 60 seconds and the agent tells the traveller "the airline system is busy; I have held your seat request and will retry".
  • Errors like 400: passenger name exceeds 32 characters go to the model, which shortens the middle name and retries correctly.

Follow-up questions to expect

  • "Why add jitter?" — Without it, many clients retry at the same moments and hit the service in waves. Random spread smooths the load.
  • "Where do you validate the response?" — In the tool wrapper, before the result goes back to the model. A malformed or unexpected response becomes a clear error, not confusing input.
  • "How do you handle APIs with no idempotency support?" — Check state before retrying a write ("does a booking with this reference exist?"), or queue writes through a component that de-duplicates.