Course Content
Agentic AI Patterns
9 sections · 50 lessons
What is Chain-of-Thought (CoT) reasoning? Why is it important for complex tasks?
What you need to know
Why it helps
- More computation. Each generated token is another forward pass. Writing 300 reasoning tokens gives the model 300 more chances to work things out.
- Decomposition. Each small step is easier to get right than one big leap.
- Constraint tracking. Writing constraints down keeps them in view.
How it changed with reasoning models
| Then (2022 to 2023) | Now (2026) |
|---|---|
| Add "Let's think step by step" to the prompt | The model reasons by default |
| Reasoning appears in the visible answer | Reasoning is often hidden or summarised |
| Control by prompt wording | Control by an effort or thinking-budget setting |
| Big gains on maths from the prompt alone | Small extra gains from prompting; big gains from effort |
In agents, reasoning models can also think between tool calls, which improves choosing the next action after a surprising result.
Caveats
- Not faithful. The written reasoning is a plausible story, not a guaranteed record of how the model got there. Do not treat it as an audit trail.
- Cost and latency. Thinking tokens are billed as output. High effort on a simple task wastes money and can even lower accuracy through overthinking.
- Exact maths belongs in tools. CoT reduces arithmetic mistakes; a calculator removes them.
A real-life example
An insurance-claims agent must decide the payable amount on a hospital claim: sum insured Rs 5,00,000, room-rent cap of 1% of sum insured per day, 4 days in a Rs 7,000 room, and 10% co-pay.
A fast model without reasoning answered with a confident wrong total. With reasoning on at medium effort, the model worked through the cap correctly: allowed room rent is Rs 5,000 a day, so proportionate deductions apply. But in 2 out of 40 test cases it still slipped on the proportionate-deduction arithmetic.
The final design: the reasoning model reads the documents and decides which policy rules apply, then calls calculate_payable(rules, bills), which does the arithmetic in code. The reasoning model is used where judgment is needed; the numbers come from deterministic code. Effort is set to low for simple claims with one bill, which halved latency for 60% of traffic.
Follow-up questions to expect
- "Should I still write 'think step by step'?" — For non-reasoning models on multi-step tasks, it can still help. For reasoning models, use the effort setting and spend prompt words on the task and constraints instead.
- "Can I show the reasoning to users?" — Show a short, checked explanation or a summary, not raw reasoning. Raw reasoning can be wrong, messy or reveal internal instructions.
- "How does CoT relate to ReAct?" — ReAct interleaves reasoning with actions: think, call a tool, read the result, think again. CoT is the "think" part.