Course Content
Agentic AI Patterns
9 sections · 50 lessons
What skills are necessary to develop and implement Agentic AI systems?
What you need to know
The skill stack
| Area | What it covers | Why agents need it |
|---|---|---|
| Backend engineering | Async, concurrency, retries, idempotency, timeouts, queues | Tool calls fail, time out, and repeat |
| LLM fundamentals | Tokens, prompting, tool calling, structured output, caching | The model is the decision-maker |
| Context engineering | What goes into the window each step, and what is left out | Long runs overflow and drift |
| Retrieval and data | Chunking, hybrid search, reranking, permission filters | Agents need grounded facts |
| Evaluation | Golden sets, trajectory checks, LLM judges and their calibration | Separates a demo from a product |
| Observability | Traces, spans, cost per task | Debugging non-deterministic runs |
| Security | Injection threat models, sandboxing, least privilege | Agents act on untrusted input |
| Product judgment | Which steps need a model; where humans approve | Most value comes from narrow scope |
Context engineering is a newer term worth using: deciding exactly which instructions, tool results, memories and documents enter the model's window at each step. In long agent runs it matters more than prompt wording.
Why "software engineering first"
An agent calling a payment API that times out has the same problem as any service: did the payment go through? You need idempotency keys so a retry does not charge twice. A run that crashes at step 9 of 12 needs checkpointed state so it can resume. None of this is AI-specific, and all of it breaks agents in production.
A real-life example
A team of four builds an incident-triage agent for their SRE group. Where the time actually went over three months:
- About 35% on tool integration: wrapping the metrics API, log search and deploy history, with timeouts, pagination and output trimming.
- About 25% on evaluation: collecting 120 past incidents with known root causes, and a script to score the agent's hypothesis.
- About 20% on security and permissions: read-only credentials, a sandbox for log queries, and approval for any runbook action.
- About 20% on prompts, model choice and the loop itself.
The engineer who helped most was not the prompt expert. It was the one who noticed that 40% of failures came from one log tool returning 50,000 lines, and capped it at the 200 most relevant.
Follow-up questions to expect
- "Do I need to know how to train models?" — Rarely for agent work. Knowing how models behave (context limits, tool-calling, reasoning effort, cost) matters far more than training them.
- "Which framework should I learn?" — Learn the patterns and one framework well. The loop, tools, state and evals carry across LangGraph, the OpenAI Agents SDK, the Claude Agent SDK and others.
- "How would you show these skills in a portfolio?" — Ship one small agent with a real eval set, traces and a cost-per-task number. That beats five demos.