Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Define a Large Language Model (LLM) in your words and outline how it generates text?
What you need to know
What the model is
An LLM is a fixed function: token sequence in, a probability for every vocabulary token out. "Large" means billions of parameters and trillions of training tokens. Predicting the next token well forces the model to learn grammar, facts, code patterns and some reasoning, because all of these help predict text.
Most assistants then add instruction tuning and preference tuning (RLHF or DPO) so the model follows requests. Reasoning models add reinforcement learning on checkable tasks. The generation loop stays the same.
The generation loop
- Tokenise — the prompt becomes ids, for example 40 tokens.
- Forward pass — produce logits of shape
(40, V); keep only the last row. - Turn into probabilities — divide logits by the temperature, then apply softmax.
- Choose — greedy takes the maximum; sampling draws at random, often after top-p keeps only the most likely tokens.
- Append and repeat — the new token becomes input for the next step, using the KV cache so earlier tokens are not recomputed.
- Stop — end-of-sequence token, stop string, or
max_tokens.
A worked step with temperature
Say four candidate tokens have logits [3.0, 2.0, 1.0, 0.5].
probs = softmax(logits / temperature)temperature 1.0 -> [0.631, 0.232, 0.085, 0.052]temperature 0.5 -> [0.862, 0.117, 0.016, 0.006]Lower temperature sharpens the distribution, so output is more predictable. Higher temperature flattens it, so output is more varied. Temperature 0 is treated as greedy.
Why errors compound
Each token is chosen based on earlier tokens, including the model's own choices. One wrong early token (a wrong name, a wrong digit) becomes part of the context, and later tokens are written to be consistent with it. This is one source of hallucination.
A real-life example
An Indian-language news app uses an LLM to write a two-line Hindi summary of each English story. The team notices Hindi summaries take about twice as long as English ones of the same word count.
The reason is the loop above. Their tokenizer splits Devanagari text into more tokens per word than English, so a 40-word Hindi summary needs many more decode steps. Each step is one forward pass, and each costs roughly the same time. The fix options are a model whose tokenizer covers Indic scripts better, a shorter target length, or streaming tokens to the reader so the first words appear quickly. Knowing that generation is one token per step is what lets the engineer find the cause.
Follow-up questions to expect
- "Does the model know facts or look them up?" — Facts are stored implicitly in its weights. It has no lookup at inference unless you add retrieval (RAG) or tools.
- "What is top-p?" — Keep the smallest set of tokens whose probabilities add up to
p(say 0.9), renormalise, and sample only from that set, cutting off the unlikely tail. - "Why is generation slow compared with reading the prompt?" — The prompt is processed in one parallel pass (prefill); output needs one pass per token (decode).