Course Content
Prompt Engineering Mastery
6 sections · 32 lessons
What is Chain of Thought (CoT) prompting?
What you need to know
A model produces each token using everything before it. If it jumps straight to the answer, all the work must happen inside one token prediction. If it writes steps first, each step's result is on the page for the next step to use — like doing long division on paper instead of in your head.
Where it helps
- Arithmetic and totals
- Multi-condition rules ("eligible if A and B, unless C")
- Questions that combine several facts
It helps little for simple lookups, classification with clear labels, or creative writing, and it always costs more tokens and latency.
How reasoning models changed it
By 2026 most frontier models are reasoning models: they produce hidden "thinking" tokens before answering, trained with reinforcement learning to reason well. For these:
- "Think step by step" is redundant. The model already does it.
- You control depth with an API setting — reasoning effort (low, medium, high) — not with prose such as "think harder".
- The raw thinking is usually hidden or returned only as a summary. Prompts that demand the model print its full hidden reasoning can be refused.
When classic CoT is still useful
- Non-reasoning or small models, including cheap models used for high-volume routes.
- Routes where thinking is turned off for latency, but a few steps of visible working still help.
- Audit fields — a short
reasoningfield in JSON that reviewers read. Treat it as a justification, not proof: research shows written reasoning does not always reflect what actually drove the answer.
Faithfulness
The reasoning text can look sensible while the answer came from somewhere else. So validate the answer in code (recompute the total) rather than trusting the explanation.
A real-life example
An invoice-checking tool must confirm that line items plus GST equal the stated total. On a small, fast non-reasoning model, the direct prompt:
Line items: Rs 1,200 x 3, Rs 450 x 2, Rs 2,999 x 1. GST 18%.Stated total: Rs 8,847. Is the total correct? Answer yes or no.The model answers "yes" — wrong. The CoT version:
Line items: Rs 1,200 x 3, Rs 450 x 2, Rs 2,999 x 1. GST 18%.Stated total: Rs 8,847.Work it out step by step: compute each line, the subtotal, the GST,and the grand total. Then give the final answer on its own line as"CORRECT" or "INCORRECT".Output: 3,600 + 900 + 2,999 = 7,499; GST = 1,349.82; total = 8,848.82; INCORRECT. On a reasoning model, the first prompt already gets it right, and the team drops the step-by-step line and sets effort to "low" for this simple check. In both cases the code recomputes the total itself — the model's job is to extract the numbers, and arithmetic belongs in code.
Follow-up questions to expect
- "Why does CoT work?" — Written steps act as working memory and extra computation: each step conditions the next token prediction, so hard problems are broken into easy ones.
- "How do you show CoT output to users?" — Usually you don't. Put reasoning in a separate field and strip it, or on reasoning models use the provider's summary.
- "Is CoT reasoning faithful?" — Not reliably. It can be a plausible story written after the fact, so verify answers independently.