Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Define greedy decoding and discuss its main weaknesses?
What you need to know
Locally best is not globally best
1# a two-step toy language model: P(first), then P(second | first)2p_first = {"A": 0.6, "B": 0.4}3p_second = {"A": {"x": 0.4, "y": 0.3, "z": 0.3},4 "B": {"x": 0.9, "y": 0.05, "z": 0.05}}56g1 = max(p_first, key=p_first.get) # greedy step 1: "A"7g2 = max(p_second[g1], key=p_second[g1].get) # greedy step 2: "x"8print("greedy:", g1 + g2, p_first[g1] * p_second[g1][g2]) # Ax 0.24910seqs = {a + b: p_first[a] * p_second[a][b] for a in p_first for b in p_second[a]}11best = max(seqs, key=seqs.get)12print("best sequence:", best, seqs[best]) # Bx 0.36Greedy picks "A" because 0.6 beats 0.4, but after "A" the model is unsure, while after "B" it is very sure. The best full sequence, "Bx", is never considered. Beam search keeps the top few partial sequences at each step to reduce this.
Repetition
Once a phrase appears, the model often gives it higher probability the next time, so greedy output can loop: "the best way to do this is the best way to do this is...". Holtzman et al. (2019), "The Curious Case of Neural Text Degeneration", documented this and proposed nucleus (top-p) sampling. Beam search, the usual cure for greedy's short-sightedness, makes open-ended text more repetitive, not less.
Other weaknesses
- No diversity. One answer per prompt, so it cannot power self-consistency voting, best-of-N, or several suggestions.
- Bland text. Human writing mixes predictable and surprising words; always taking the most likely one reads flat.
- Reasoning models. Some model providers recommend sampling for long reasoning — for example, DeepSeek-R1's usage guidance suggests a temperature around 0.6 and warns that greedy decoding can cause endless repetition.
When greedy is the right choice
| Use greedy (or temperature near 0) | Use temperature + top-p |
|---|---|
| Extraction, classification, JSON output | Chat, writing, brainstorming |
| Short code completions | Several candidate answers |
| Reproducible evaluation runs | Long reasoning on models that ask for it |
Typical sampling settings: temperature 0.7–1.0 with top-p 0.9–0.95, tuned per task.
A real-life example
A code-completion assistant uses greedy decoding for single-line suggestions: the user wants the most likely completion of for (int i = 0; i , not a creative one, and the same context should give the same suggestion every time. Suggestions are short, so repetition loops rarely have room to start, and a stop rule at the end of the line cuts them off anyway.
The same company's chatbot first shipped with greedy decoding to "be safe". Support staff reported that long answers sometimes repeated the same sentence three times, and every customer got identical, stiff wording. Switching to temperature 0.7 with top-p 0.9 removed the loops in their test set of 500 long conversations, while the JSON tool calls inside the chatbot kept temperature 0, because there any variation is a bug.
Follow-up questions to expect
- "Is greedy the same as temperature 0?" — Yes in effect; many APIs implement temperature 0 as argmax.
- "Does beam search fix greedy's problems?" — It finds higher-probability sequences, which helps translation and summarisation, but for open-ended text it makes output even more repetitive and generic.
- "How do you reduce repetition without sampling?" — Repetition or frequency penalties, or no-repeat-n-gram rules; they help but can also block legitimate repeats such as variable names in code.