Agentic AI Patterns

Course Content

Agentic AI Patterns

9 sections · 50 lessons

What is an agent’s context window, and how does it influence design decisions?


Budgeting a 200,000-token windowSystem promptand tools — 20kRetrieved docsand memories — 60kHistory andtool results — 80kOutput andthinking room — 40ktopbottomOne uncapped log search returned 40,000 lines and pushed out the approval rule.
When the window overflows silently, the oldest messages go first — and those are usually the task and the rules.

What you need to know

A context budget

For a model with a 200,000-token window, a triage agent might budget:

PartShareTokens
System prompt and tool definitions10%20,000
Retrieved documents and memories30%60,000
Working history and tool results40%80,000
Headroom for output and reasoning20%40,000

Reasoning models also spend "thinking" tokens inside the output allowance, so leave real headroom.

Enforce it in code

Python
def fit_context(system, messages, limit_tokens, count=lambda s: len(s) // 4):    """Keep the system prompt and the newest turns; replace old tool results with stubs."""    budget = limit_tokens - count(system)    kept, used = [], 0    for msg in reversed(messages):                    # newest first        text = msg["content"]        if used + count(text) > budget and msg["role"] == "tool":            text = f"[tool result {msg['id']} stored; call get_result to reread]"        if used + count(text) > budget:            break                                     # older turns fall out        kept.append({**msg, "content": text})        used += count(text)    return [{"role": "system", "content": system}] + kept[::-1]

The system prompt is always kept. Walking from newest to oldest, a tool result that does not fit is replaced by a stub the agent can use to fetch it again. In practice, use the provider's token counter instead of len // 4, and never let an older message break the pairing between a tool call and its result.

Design decisions it forces

  • Cap every tool output. One raw HTML page or 5,000-row query can fill the window in a single step.
  • Compaction for any long-running agent.
  • Sub-agents for heavy reading. A worker reads 40 pages and returns a summary, so the main context stays small.
  • Stable prefix for caching.
  • Pin the goal and plan, so truncation cannot remove them.

A real-life example

A DevOps triage agent called search_logs with no limit during a big outage. The tool returned 40,000 lines. Its framework silently dropped the oldest messages to fit, and the oldest message was the instruction "never restart production services without approval". Two steps later the agent proposed a restart as if it were routine.

Fixes: search_logs now returns the top 200 lines by relevance plus counts per error type; the full result is stored by ID. A token check runs before every model call and raises an error rather than silently truncating. The approval rule is also enforced in the restart_service tool itself, so it no longer depends on the prompt being present.

Follow-up questions to expect

  • "Why not just use the model with the biggest window?" — Cost and latency rise with every token on every step, and quality on buried facts drops. Focused context usually wins.
  • "What happens when the window overflows?" — The API rejects the call, or a framework truncates silently. Both are bad; check the size yourself before calling.
  • "How do sub-agents help with context?" — Each has its own window. The main agent receives only the result, not the pages the worker read.