Course Content
Agentic AI Patterns
9 sections · 50 lessons
How does Agentic AI support cost reduction and operational efficiency?
What you need to know
Cost of one agent run
A 12-step run, with a 6,000-token system prompt and tool list and about 1,800 tokens added per step, reads about 190,800 input tokens and writes about 3,600. At an assumed $3 per million input and $15 per million output:
No caching: about $0.63 per runWith prompt caching: about $0.20 per run (repeated prefix billed at a tenth)Plus routing 60% of requests to a small model at about $0.02: 0.4 x 0.20 + 0.6 x 0.02 = about $0.09 per request on averageThe biggest levers, in usual order: routing easy work to small models, prompt caching, fewer steps, and smaller tool results.
Other levers
- Budgets per task: steps, time, spend. Stop and escalate at the cap.
- Cache deterministic tool results.
- Batch offline work; many providers price batch jobs lower.
- Retrieve 5 chunks, not 50.
Measuring savings honestly
Include every cost: model and infrastructure, human time on escalations, and rework when the agent gets it wrong.
A real-life example
A food-delivery company's support team handles 10,000 "where is my refund?" and "item missing" tickets a month. Human cost: about Rs 120 per ticket.
Baseline: 10,000 x Rs 120 = Rs 12,00,000Agent: 10,000 x Rs 8 (model + infra) = Rs 80,000 3,500 escalated to humans x Rs 120 = Rs 4,20,000 325 reopened (5% of 6,500) x Rs 180 = Rs 58,500Total with agent = Rs 5,58,500Saving = Rs 6,41,500 (about 53%)The reopen line matters. A version that forced higher containment by resolving borderline cases itself pushed reopens to 15%, and the saving fell. The team also found that 4% of runs hit the 20-step cap, and those runs cost 9 times the median. Fixing one flaky order-lookup tool removed most of them.
Follow-up questions to expect
- "What is the single biggest cost lever?" — Usually model routing: most requests are easy and do not need the largest model.
- "How do you set a spend budget?" — From the value of the task. If a ticket is worth Rs 120 of human time, an agent run that has spent Rs 40 should escalate, not continue.
- "How do you report ROI?" — Cost per successful case versus the baseline, including escalations and rework, over a period long enough to see reopen rates.