Prompt Engineering Mastery

Course Content

Prompt Engineering Mastery

6 sections · 32 lessons

What is Tree of Thoughts (ToT)?


Game of 24 with 4, 9, 10 and 134, 9, 10, 1310 minus 4 gives 69 plus 13 gives 2213 minus 9gives 4; 6 x 4 = 246 plus 9gives 15 — maybe22 with 4, 10 — prune
A single chain that starts with 9 plus 13 never recovers; the tree prunes that branch and spends its calls elsewhere.

What you need to know

Chain of thought writes one line of reasoning and commits to every step. Tree of Thoughts, introduced in 2023, lets the model explore several options and go back.

The four parts

PartQuestion it answersExample
Thought decompositionWhat is one step?One arithmetic operation
Thought generatorWhat are the options?Propose 5 next operations
State evaluatorIs this branch promising?"sure / maybe / impossible"
Search algorithmWhich branch next?Breadth-first, keep best 5

The evaluator can be the model itself (a prompt that rates each branch) or, better, something exact — a unit test, a calculator, a rule checker.

Worked example: the Game of 24

Use 4, 9, 10 and 13 with +, −, × and ÷ to make 24.

Text
Level 1 candidates:  13 - 9 = 4   |  10 - 4 = 6   |  9 + 13 = 22Evaluate:            maybe        |  maybe        |  unlikely (prune)Level 2 from "10 - 4 = 6", left {6, 9, 13}:                     13 - 9 = 4 -> left {6, 4}   evaluate: sure (6 x 4)Answer: (10 - 4) x (13 - 9) = 24

A single chain of thought might start with 9 + 13 = 22 and never recover. ToT pruned that branch early.

Cost

With 5 candidates per level and 3 levels, you can need dozens of generation and evaluation calls for one problem — seconds to minutes and many times the token cost of one call.

The 2026 view

Reasoning models now explore and backtrack inside their own thinking ("wait, that doesn't work, try..."). So hand-built ToT prompting is rarely needed for everyday tasks. The idea survives in systems that combine several candidates with an exact evaluator: a coding agent that writes several fixes and keeps the one that passes tests, or a planner that checks each plan against constraints in code.

A real-life example

A quick-commerce company plans delivery-slot assignments for 12 riders with rules: maximum 4 orders each, cold items within 20 minutes, no rider crosses the river twice. A single model call produces a plan that breaks the cold-chain rule. The team builds a ToT-style loop:

  • The model proposes 4 ways to assign the first batch of orders.
  • Code checks each partial plan against the rules and scores total travel time.
  • The best 2 branches are expanded with the next batch; rule-breaking branches are dropped.

The final plan satisfies every rule and is 14% shorter than the dispatcher's manual plan. It takes 30 model calls and about 40 seconds — acceptable for a plan made once per hour, not for a customer chat reply.

Follow-up questions to expect

  • "Who evaluates the branches?" — Either the model with a rating prompt or external code. External checks are more reliable and cheaper when they exist.
  • "How is ToT different from self-consistency?" — Self-consistency samples complete answers and votes at the end; ToT evaluates partial steps and prunes during the search.
  • "Would you use ToT in a chatbot?" — Rarely; the latency and cost are too high for interactive replies.