Course Content
Agents & Tools Interview Prep
6 sections · 40 lessons
What reasoning techniques can agents use for decision-making?
What you need to know
| Technique | Idea | Cost | Use when |
|---|---|---|---|
| Built-in thinking | Model reasons before answering or between tool calls | Thinking tokens | Almost always, tuned by effort |
| ReAct / tool loop | Reason, act, observe, repeat | One call per step | Tool-using tasks with uncertain paths |
| Plan-and-execute | Plan first, then run steps | One planning call plus steps | Long, parallel, approvable tasks |
| Self-consistency | Sample N answers, vote | N× | Single checkable answer, high stakes |
| Tree search | Branch, score, backtrack | Many calls | Partial states can be scored (code compiles) |
| Reflection | Revise with feedback | Extra rounds | A cheap objective check exists |
Built-in thinking in 2026
On current Claude models, you enable adaptive thinking (thinking: {"type": "adaptive"}) and control depth with an effort setting (low to max) instead of writing "think step by step". Lower effort means fewer and more consolidated tool calls — good for routine routes; higher effort for hard agentic work. Thinking can also happen between tool calls, so the model reasons about each result before the next action. OpenAI's reasoning models have a similar reasoning.effort setting.
Self-consistency with tools
For decisions with one right answer — "is this transaction fraud?" — sample 5 independent runs and act only on strong agreement (5/5 or 4/5). Disagreement is a signal to escalate. It multiplies cost, so reserve it for high-value cases.
Tree search
Useful when you can score partial progress: code that compiles, a query that runs, a puzzle state. Expensive and complex to build; most production agents do not need it.
A real-life example
A bank's dispute-handling agent uses different techniques for different steps:
- Collecting facts (transactions, merchant details, past disputes): a tool loop at low effort — mostly lookups.
- Deciding the dispute category (fraud, duplicate charge, merchant dispute): self-consistency with 5 samples at high effort. At 5/5 or 4/5 agreement, it proceeds; at 3/5 or less, it goes to a human analyst. About 11% go to humans.
- Drafting the letter to the merchant: one call, then a validator checks that every amount and date in the letter matches the transaction data (reflection with a real check).
Using high effort and 5 samples for every step would have cost about 6 times more with no gain on the lookups. Matching technique to step kept cost per dispute near ₹9.
Follow-up questions to expect
- "Do reasoning models make ReAct unnecessary?" — No. Thinking improves each decision, but the agent still needs tools and observations for facts it does not have.
- "How do you pick the effort level?" — Measure on real tasks: start at the default, try lower for routine routes and higher for hard ones, and compare quality and cost per completed task.
- "Is chain-of-thought output trustworthy as an explanation?" — Treat it as a debugging aid, not a guaranteed account of why the model decided.