Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
How LLMs work, in the detail an engineer needs
In the first internal pilot, an employee asked how many days of casual leave Harbourline gives in India. The model answered, with complete confidence, "Employees are entitled to 12 days of casual leave per year." The real number has been 8 since the 2024 policy update. Nobody had given the model the policy. It had never seen a single Harbourline document.
So where did "12" come from? It is a common number for casual leave in Indian companies, and the model had read a great deal of text on the public web. It produced the most likely continuation of the sentence, and a likely sentence is not the same as a true one. That single idea explains most of the behaviour you will design around in this course.
You do not need the mathematics of transformers to build good systems. You do need an accurate picture of what happens inside one call, because it tells you why prompts cost what they cost, why answers take the time they take, and why the model must be given your facts instead of being trusted to know them.
Tokens: the unit of everything
A model does not read characters or words. Text is split into tokens, pieces that are often a word, part of a word, or a punctuation mark. "Leave" may be one token; "reimbursement" may be three. For English, a useful rule is that 1 token is about 4 characters, or about 0.75 words. A typical policy page of 450 words is roughly 600 tokens.
Tokens matter for three practical reasons. You pay per token, separately for input and output. The context window, the maximum tokens one call can hold, is counted in tokens. And output speed is measured in tokens per second. Harbourline has another detail to plan for: questions typed in Hindi, Marathi or mixed "Hinglish" usually take two to three times as many tokens per word as English, so the same question can cost more and use more of the window.
One token at a time
At its core, a language model is a function that takes a sequence of tokens and returns a score for every possible next token. Those scores are turned into probabilities, one token is picked, it is added to the sequence, and the whole process repeats. A 250-token answer is 250 of these steps.
The picking step is called sampling. A setting called temperature controls how sharp the probabilities are. This small script shows it with made-up scores for the next token after "Harbourline employees in India get ... days of casual leave:".
1import numpy as np23candidates = ["8", "12", "10", "15"]4logits = np.array([2.1, 2.6, 1.2, 0.3]) # the model's raw scores56def next_token_probs(logits: np.ndarray, temperature: float) -> np.ndarray:7 z = logits / temperature8 z = z - z.max() # keeps exp() from overflowing9 p = np.exp(z)10 return p / p.sum()1112for t in (0.2, 1.0, 1.5):13 probs = next_token_probs(logits, t)14 print(t, {c: round(float(p), 2) for c, p in zip(candidates, probs)})At temperature 1.0 the wrong answer "12" gets 51% and the right answer "8" gets 31%. At 0.2, "12" gets 92%. At 1.5 the probabilities flatten and "12" falls to 43%. Look carefully at what this means: lowering the temperature made the wrong answer more consistent, not less wrong. Temperature controls variety. It does not control truth. Several newer hosted models no longer expose temperature at all, which is fine, because correctness never came from that knob.
Prefill and decode: why there are two latency numbers
A call has two phases. In prefill, the model reads your whole prompt. This work runs in parallel across all input tokens, so it is fast per token but grows with prompt length. In decode, the model writes the answer one token at a time, and each step must wait for the previous one.
This gives you two separate numbers to measure. Time to first token (TTFT) is mostly prefill plus network and queueing; for a 3,000-token prompt on a hosted model it is typically 0.5 to 1 second, and for a 100,000-token prompt it can be several seconds. Output speed is decode; hosted models commonly produce 50 to 100 tokens per second. A 250-token answer at 70 tokens per second takes about 3.6 seconds to write, whatever you do to the prompt.
The practical lessons follow directly. Streaming the answer makes the wait feel short, because the user sees the first words after TTFT instead of after the full answer. Short answers are faster answers. And long prompts hurt twice: they raise TTFT and they cost more. Models also use information in the middle of a very long prompt less reliably than information near the start or end, so a longer prompt is not a safer prompt.
What training gives you, and what it does not
A modern model is built in stages. Pretraining teaches it to predict the next token over a huge amount of public text, which gives it language, reasoning patterns and a broad but frozen picture of the world up to a knowledge cutoff. Instruction tuning teaches it to follow requests. Preference tuning, often called RLHF, teaches it to give answers people rate highly, which makes it helpful and polite, and also makes it lean towards giving some answer rather than saying "I don't know".
None of these stages saw Harbourline's policies, and none of them will see next quarter's update. A model call is also stateless: the model remembers nothing between calls. Everything it can use in one call is its trained weights plus the tokens you send. When a chat looks like it remembers, the application is resending the earlier messages every time.
Comes from training
- Language, grammar, tone
- General reasoning and common patterns
- Public facts up to a cutoff date
- A habit of answering confidently
Must come from your context
- Harbourline's actual policy text
- Which country and grade the user is in
- Today's date and the current policy version
- The user's own leave balance or tickets
This split gives you PolicyPal's basic design rule. Facts come from the context, through retrieval and tools. Behaviour comes from the prompt. Shape comes from a schema. Training is for skills, not for facts that change.
Check your understanding
0 of 3 answered
1.A teammate proposes setting temperature to 0 to stop PolicyPal inventing leave numbers. What is the main problem with this plan?
2.A PolicyPal answer takes 4.5 seconds in total. The prompt is 3,000 tokens and the answer is 250 tokens. Which change will cut the total time the most?
3.Why does a chat with PolicyPal appear to "remember" what the user said two messages ago?