Agentic AI Patterns

Course Content

Agentic AI Patterns

9 sections · 50 lessons

How does Agentic AI support cost reduction and operational efficiency?


Engineering the cost of one request12-step run,no caching:about 63 centsCache the stableprefix: about 20 centsRoute 60 percentto a small modelAveragerequest:about 9 centsAssumed prices of 3 and 15 dollars per million input and output tokens.
Routing and caching cut the agent's own bill, but reopened tickets can quietly eat the saving.

What you need to know

Cost of one agent run

A 12-step run, with a 6,000-token system prompt and tool list and about 1,800 tokens added per step, reads about 190,800 input tokens and writes about 3,600. At an assumed $3 per million input and $15 per million output:

Text
No caching:            about $0.63 per runWith prompt caching:   about $0.20 per run   (repeated prefix billed at a tenth)Plus routing 60% of requests to a small model at about $0.02:  0.4 x 0.20 + 0.6 x 0.02 = about $0.09 per request on average

The biggest levers, in usual order: routing easy work to small models, prompt caching, fewer steps, and smaller tool results.

Other levers

  • Budgets per task: steps, time, spend. Stop and escalate at the cap.
  • Cache deterministic tool results.
  • Batch offline work; many providers price batch jobs lower.
  • Retrieve 5 chunks, not 50.

Measuring savings honestly

Include every cost: model and infrastructure, human time on escalations, and rework when the agent gets it wrong.

A real-life example

A food-delivery company's support team handles 10,000 "where is my refund?" and "item missing" tickets a month. Human cost: about Rs 120 per ticket.

Text
Baseline:  10,000 x Rs 120                          = Rs 12,00,000Agent:     10,000 x Rs 8 (model + infra)            = Rs    80,000           3,500 escalated to humans x Rs 120       = Rs  4,20,000           325 reopened (5% of 6,500) x Rs 180      = Rs    58,500Total with agent                                     = Rs  5,58,500Saving                                               = Rs  6,41,500 (about 53%)

The reopen line matters. A version that forced higher containment by resolving borderline cases itself pushed reopens to 15%, and the saving fell. The team also found that 4% of runs hit the 20-step cap, and those runs cost 9 times the median. Fixing one flaky order-lookup tool removed most of them.

Follow-up questions to expect

  • "What is the single biggest cost lever?" — Usually model routing: most requests are easy and do not need the largest model.
  • "How do you set a spend budget?" — From the value of the task. If a ticket is worth Rs 120 of human time, an agent run that has spent Rs 40 should escalate, not continue.
  • "How do you report ROI?" — Cost per successful case versus the baseline, including escalations and rework, over a period long enough to see reopen rates.