Prompt Engineering Mastery

Course Content

Prompt Engineering Mastery

6 sections · 32 lessons

What should you do if the model repeats words or generates loops?


What you need to know

A repetition loop is when the model repeats a phrase or line again and again: "Perfect for daily use. Perfect for daily use. Perfect for...". It happens because each repeated phrase makes the next repetition more likely — the model conditions on its own output.

Common causes

  • Greedy or near-greedy decoding — at temperature 0 the model can fall into a cycle it never leaves.
  • Small models — 1B to 8B open-weight models loop far more than large hosted ones.
  • A repetitive prompt — if your examples all start the same way, the output copies that.
  • No length limit — the model has no reason to stop.
  • Lost signal — a huge, noisy context makes the model fall back on safe, repeated text.

Sampling-side fixes

  • Raise temperature a little (for example 0 to 0.3) where the model accepts it.
  • Add a small frequency penalty (0.3 to 0.5). Large values cause odd word choices and break required repeats, such as field names in JSON.
  • Always set max_tokens and stop sequences so a loop cannot run for thousands of tokens.

Prompt-side fixes (work on every model)

  • State the shape: "exactly 4 bullets, one sentence each".
  • Remove repetition from your own examples.
  • Generate long outputs section by section.
  • Trim irrelevant context.

In code

Detect a loop cheaply: if the same 8-word sequence appears more than three times, discard the output and retry. For JSON, a structured-output mode constrains the shape, so the model cannot repeat keys or ramble.

The 2026 note

Large hosted reasoning models rarely loop in their visible output, and many do not expose penalties. You will still meet loops with small self-hosted models, and a related problem with reasoning models: very long thinking that hits max_tokens before any answer. The fix there is a larger limit or a lower effort setting, not a penalty.

A real-life example

A marketplace runs a 7B open-weight model at temperature 0 to write product bullet points. The prompt:

Text
Write bullet points for this product.Product: Non-stick tawa, 28 cm, induction base.

About 4% of outputs loop: "- Perfect for dosas\n- Perfect for rotis\n- Perfect for parathas\n- Perfect for..." until the 512-token limit. The fixed prompt:

Text
Write exactly 4 bullet points, each under 12 words, each about adifferent feature: size, surface, stove compatibility, cleaning.Product: Non-stick tawa, 28 cm, induction base.

They also set temperature 0.3, repetition_penalty = 1.1, max_tokens = 120, and a loop detector. Loops drop to under 0.1%, and the bullets now cover four different features instead of four dishes.

Follow-up questions to expect

  • "Why not just use a big repetition penalty?" — It punishes legitimate repeats — product names, JSON keys, code keywords — and produces unnatural wording.
  • "What is the difference between frequency and presence penalty?" — Frequency grows with each repeat; presence is a one-time penalty that encourages new topics.
  • "How do you stop loops in structured output?" — Use JSON-schema structured outputs; the shape is enforced, so looping keys cannot happen.