Course Content
Prompt Engineering Mastery
6 sections · 32 lessons
What is Tree of Thoughts (ToT)?
What you need to know
Chain of thought writes one line of reasoning and commits to every step. Tree of Thoughts, introduced in 2023, lets the model explore several options and go back.
The four parts
| Part | Question it answers | Example |
|---|---|---|
| Thought decomposition | What is one step? | One arithmetic operation |
| Thought generator | What are the options? | Propose 5 next operations |
| State evaluator | Is this branch promising? | "sure / maybe / impossible" |
| Search algorithm | Which branch next? | Breadth-first, keep best 5 |
The evaluator can be the model itself (a prompt that rates each branch) or, better, something exact — a unit test, a calculator, a rule checker.
Worked example: the Game of 24
Use 4, 9, 10 and 13 with +, −, × and ÷ to make 24.
Level 1 candidates: 13 - 9 = 4 | 10 - 4 = 6 | 9 + 13 = 22Evaluate: maybe | maybe | unlikely (prune)Level 2 from "10 - 4 = 6", left {6, 9, 13}: 13 - 9 = 4 -> left {6, 4} evaluate: sure (6 x 4)Answer: (10 - 4) x (13 - 9) = 24A single chain of thought might start with 9 + 13 = 22 and never recover. ToT pruned that branch early.
Cost
With 5 candidates per level and 3 levels, you can need dozens of generation and evaluation calls for one problem — seconds to minutes and many times the token cost of one call.
The 2026 view
Reasoning models now explore and backtrack inside their own thinking ("wait, that doesn't work, try..."). So hand-built ToT prompting is rarely needed for everyday tasks. The idea survives in systems that combine several candidates with an exact evaluator: a coding agent that writes several fixes and keeps the one that passes tests, or a planner that checks each plan against constraints in code.
A real-life example
A quick-commerce company plans delivery-slot assignments for 12 riders with rules: maximum 4 orders each, cold items within 20 minutes, no rider crosses the river twice. A single model call produces a plan that breaks the cold-chain rule. The team builds a ToT-style loop:
- The model proposes 4 ways to assign the first batch of orders.
- Code checks each partial plan against the rules and scores total travel time.
- The best 2 branches are expanded with the next batch; rule-breaking branches are dropped.
The final plan satisfies every rule and is 14% shorter than the dispatcher's manual plan. It takes 30 model calls and about 40 seconds — acceptable for a plan made once per hour, not for a customer chat reply.
Follow-up questions to expect
- "Who evaluates the branches?" — Either the model with a rating prompt or external code. External checks are more reliable and cheaper when they exist.
- "How is ToT different from self-consistency?" — Self-consistency samples complete answers and votes at the end; ToT evaluates partial steps and prunes during the search.
- "Would you use ToT in a chatbot?" — Rarely; the latency and cost are too high for interactive replies.