Prompt Engineering for LLMs

Role Prompting and System Messages


Somebody on your team read a prompting thread and now every prompt in the codebase opens the same way.

Text
You are a world-class expert with 30 years of experience and aPhD from MIT. You are extremely intelligent and never makemistakes. Answer the following question.

Ask that prompt to compute a compound-interest figure and it gets it wrong at the same rate as a prompt with no preamble at all. Ask it a question about a library's API and it invents a method name with the same confidence either way. The credentials did nothing, because there is nothing for them to act on. The model has no PhD to invoke. Its parameters are identical whether you address it as a Nobel laureate or as nobody in particular.

And yet — persona instructions clearly do change output. Tell the model it is a paediatric nurse explaining a diagnosis to a worried parent, and you get short sentences, no Latin, an acknowledgement of the worry, and a clear next action. Tell it to write the same content as a clinical handover note and you get abbreviations, numbers, and no reassurance at all. Same facts, entirely different artefact.

So role prompting is neither magic nor useless. It is a real and useful lever attached to a specific mechanism, and almost all the confusion around it comes from misidentifying what the lever controls.

What belongs in the system messageStableacross every turnRules that must never bendThe output contractRefusal andescalation boundariesTone, stated onceNot the user's input
A role earns its tokens only when it changes the objective — 'find what is wrong with this' — never when it just adds a costume.

What a role actually conditions

A language model predicts the next token given everything in its context. Text written by a paediatric nurse to a parent and text written by a consultant to another consultant occupy statistically different regions of the training distribution — different vocabulary, sentence length, hedging, choice of what to explain and what to assume.

When you write "You are a paediatric nurse speaking to a worried parent", you place the conversation in the first region. Continuations typical of that region become more probable. That is the whole effect.

A role instruction selects a style of text, not a level of competence. It changes which continuations are likely; it does not change what the model knows or how carefully it computes.

Once stated plainly, the boundary is obvious.

A role reliably changesA role does not change
Vocabulary and registerFactual knowledge in the weights
Sentence length and structureArithmetic accuracy
How much background is assumedWhether a cited paper exists
Which details are treated as relevantReasoning depth (that comes from generated tokens)
Level of hedging and caveatingKnowledge cutoff
Default format conventions of a fieldAccess to your data

The middle column of failure is the interesting one: "which details are treated as relevant" quietly does a lot of work. Ask "review this contract clause" and you get a general reading. Ask "you are reviewing this clause on behalf of the supplier, whose main concern is uncapped liability" and the model surfaces different clauses as important. That is not a personality change; it is a change in what counts as signal. Framing that shifts the objective is far more valuable than framing that asserts credentials.

Credentials versus audience

Text
BAD:You are a brilliant, world-class senior software engineer with20 years of experience at top tech companies. You have deepexpertise in everything.Explain database indexes.
Text
GOOD:Explain database indexes to a self-taught developer who has beenwriting application code for two years, uses an ORM daily, and hasnever opened a query plan.Assume they know what a table and a WHERE clause are. Do not assumethey know what a B-tree is. Use one concrete example with a table ofabout a million rows and real timings. 300 words or so.

Why the second works. "World-class expert" gives the model nothing it can condition on precisely — expert text spans everything from a beginner's tutorial to a paper on write amplification, so the register stays undetermined. The second version specifies the reader, and the reader determines everything: what can be assumed, what must be defined, how long it should be, what kind of example lands. Every sentence in it constrains an actual decision the model has to make.

The general form of the rule: describing who the output is for beats describing who the model is. The audience is what the writing has to fit; the persona is at best a proxy for it.

System messages: what the API is actually doing

Chat models take a list of messages, each with a role. Roughly:

Python
messages = [    {"role": "system",    "content": "..."},   # standing instructions    {"role": "user",      "content": "..."},   # this turn's input    {"role": "assistant", "content": "..."},   # a previous reply    {"role": "user",      "content": "..."},]

The exact shape varies by provider. Anthropic's Messages API takes standing instructions as a separate top-level system parameter rather than as a message. OpenAI's Responses API uses an instructions parameter, or input messages with the developer role, OpenAI's current name for system-level instructions. The idea is the same in all of them.

Underneath, all of this is flattened into one token sequence using a chat template — special tokens that mark where each role's content begins and ends. The model is not receiving separate channels. It is receiving one stream with markers in it.

So why does the system role behave differently? Because during instruction tuning and preference training, models are trained on conversations where system-role content is followed by responses that obey it, including cases where a later user message asks for something the system content forbids and the trained response declines. The role separation is a learned behaviour reinforced by training data, not an architectural guarantee.

System messages are a trained priority ordering, not a security boundary. They are meaningfully harder to override than ordinary user text, and they are not impossible to override.

Two practical consequences. Put anything that must hold for the whole conversation in the system message — role, tone, format rules, refusal boundaries — because it stays in context for every turn while user messages scroll past. And never rely on a system message alone to enforce something that actually matters. If your assistant must not discuss competitors' pricing, the system message reduces the chance and a check on the output enforces it.

Dividing the content

Put in the system messagePut in the user message
Who the assistant is and who it servesThe actual request or question
Persistent tone and format rulesThe document, ticket or record to work on
Standing boundaries and refusalsAnything that varies per call
What to do when uncertainPer-request overrides you are happy for users to make
Stable domain context (product name, policy)Untrusted, user-supplied content

The last row is worth stating explicitly. Content that came from a user is data. It belongs in a user message, inside delimiters, described as untrusted — never concatenated into the system message, where it inherits the priority you have carefully arranged for your own instructions.

A system prompt that earns its tokens

Text
BAD system message:You are a helpful and friendly customer support assistant for ourcompany. You are knowledgeable, professional, empathetic, patientand always aim to provide excellent service with a positiveattitude. You care deeply about customers.
Text
GOOD system message:You are the support assistant for Northwind, a UK payroll productused by small businesses. You speak to office managers and businessowners, not developers.How you answer:- British English. Two short paragraphs maximum unless the user asks  for steps, in which case use a numbered list.- Never use payroll jargon without a plain-English gloss on first use.- Reference the current UK tax year only if the user has told you  which one they mean; otherwise ask.What you must not do:- Never state a specific tax rate, threshold or deadline. Say the  figures vary and point the user to Settings > Tax year, which shows  the values configured for their account.- Never confirm or deny anything about a specific customer's account.  You have no access to account data.- Never advise on whether someone should be classed as an employee or  a contractor. Reply exactly: "That's a question for an accountant —  I can't advise on employment status."When you don't know:Say "I don't know" and offer to raise a ticket. Do not guess at how afeature works. A wrong answer about payroll costs the user money.

Why the second works. Take the two apart line by line and the difference is not length, it is that every line in the second version changes an output.

  • "Helpful, friendly, knowledgeable, professional, empathetic, patient" are adjectives whose opposites nobody would request. The model is already inclined towards all of them. A word that could not have been otherwise is a word that constrains nothing.
  • "Two short paragraphs maximum unless the user asks for steps" is checkable. Either the output has three paragraphs or it does not.
  • The prohibition on tax figures targets a specific, predictable failure: the model has seen a great many tax rates in training and will happily produce a plausible one, which will be for the wrong jurisdiction or the wrong year. The prohibition comes with a replacement behaviour, which matters — telling a model what not to do without saying what to do instead leaves the gap open.
  • The employment-status rule supplies exact wording. For a boundary that genuinely matters, an exact string is far more reliable than a description of the sentiment you want.
  • The final section makes "I don't know" a legitimate, named option with a follow-up action. Without that, the strongest pressure on the model is to produce a helpful-sounding answer, because that is what support text overwhelmingly looks like.

Roles that change the objective, not the costume

The genuinely powerful uses of role prompting are the ones that change what the model is optimising for. Three that pay for themselves.

Adversarial review

Text
You are reviewing this migration plan as the on-call engineer whowill be paged if it fails at 3am. You are not here to be encouraging.For each step, state what could go wrong and how you would find out.If a step has no rollback, say so explicitly.End with the single most likely cause of an incident, and nothingelse — no summary, no reassurance.

Ask for a plain review and you get a balanced assessment, because balanced assessment is what reviews look like. This framing changes the target: it asks for failure modes specifically and removes the social pressure towards a positive closing note. "You are not here to be encouraging" and "no reassurance" are doing real work, because the default pull towards a constructive ending is strong.

Fixed multi-perspective analysis

Text
Analyse this proposal from exactly these three positions, in order.Use the heading given. Do not blend them, and do not add a fourth.FINANCE: cost, cash flow timing, what happens if revenue is 30%below plan.ENGINEERING: build effort, what it couples us to, ongoing maintenance.CUSTOMER: what changes for an existing user on day one, includinganything that gets worse.Then, under DISAGREEMENT, name the sharpest conflict between thethree positions. Do not resolve it.

A single "analyse this proposal" produces one blended assessment that tends to average over concerns. Forcing separate passes means each perspective is generated in a context primed for that perspective, so cost concerns are not softened by enthusiasm about the architecture. "Do not resolve it" prevents the usual reflex of collapsing tensions into a tidy conclusion, which is precisely where the useful information gets lost.

Withholding the answer

Text
You are tutoring someone learning recursion. They will ask you tofix their code. Do not fix it.Ask one question at a time that helps them locate the problemthemselves. Never write corrected code. Never name the bug. If theyask directly for the answer, ask them what they expect the functionto return for the smallest input, and wait.

Here the role is defined almost entirely by prohibitions, and that is correct: the model's default behaviour on a broken function is to fix it, and this pattern exists to suppress exactly that default. Notice again that every prohibition has a paired positive instruction. "Don't fix it" alone would leave the model with no available action.

Where role prompting goes wrong

Persona mistaken for competence

The most consequential error, and the one that opened this lesson. Published attempts to reproduce persona-based accuracy gains on reasoning and knowledge benchmarks have found the effects small and inconsistent — sometimes positive, sometimes negative, varying by model and task in ways that make them unsafe to assume. Meanwhile the effect on style and format is large and consistent. Use roles where the evidence is strong. If you need accuracy, the levers are examples, decomposition, reasoning tokens, and retrieval of real source material — not job titles.

Adjective stacking

"Insightful, creative, rigorous, thoughtful, precise, imaginative." Several of those pull in opposite directions, and none is measurable. The model resolves the contradiction arbitrarily and differently each run, which shows up as inconsistency you cannot trace to a cause. Replace each adjective with the behaviour you actually want: not "rigorous" but "state your assumptions before your conclusion"; not "creative" but "give three options that use different mechanisms, not three variations of the same one".

Contradictions between the system message and the rest

The system message says "always reply in under 100 words". A few-shot example further down runs to 300. The user asks for a detailed breakdown. Now three parts of the context disagree, and which wins is unpredictable — demonstrations tend to beat descriptions, and recent instructions tend to beat distant ones. Every time you change a rule, re-read the examples and the templates for contradictions. This is the most common bug in mature prompts and it is invisible unless you go looking.

Drift over long conversations

Thirty turns in, the assistant that was terse and formal is chatty. The system message is still in context, but it is now competing with thirty turns of its own output, which forms a much larger and much more recent pattern. If the model has produced friendly, meandering replies for the last twenty turns, that is the strongest available signal for what comes next.

Two fixes that work: restate the critical constraints in the final user message before generation, and periodically summarise old turns instead of carrying them verbatim, so the transcript does not accumulate into a counter-example of your own rules.

Roles used as safety

"You are a helpful assistant that never discusses X" reduces the frequency of discussing X. It does not prevent it. A sufficiently determined user, or a piece of injected text inside a document your system is summarising, can move the conversation somewhere the system message did not anticipate. Where an output genuinely must not occur, the enforcement belongs outside the model: a classifier on the output, a keyword check, a human review step, or simply not exposing that capability.

When you build something with this

Write the system message as a specification, then audit it with one test: cross out every line whose removal would not change any output, and see what is left. If half the message disappears, that half was costing you tokens on every call and buying nothing.

Four questions that produce a system message worth keeping:

  • Who is reading this output, and what do they already know? Answer that concretely and most register decisions make themselves. It is worth more than any number of credentials.
  • What are the three things that must never appear? Name them specifically, and for each one give the replacement behaviour and, where it matters, the exact words. Then enforce it in code as well, because the prompt lowers probability and does not set a rule.
  • What should happen when the model does not know? Uncertainty is the default state on real traffic, and a model with no sanctioned way to express it will produce something confident instead. Make "I don't know" available, phrase it, and attach an action to it.
  • Is anything in here a costume rather than a constraint? If a line asserts expertise, enthusiasm, or personality traits without changing a decision the model has to make, delete it. What remains — audience, format, boundaries, uncertainty handling — is the part that actually steers the output.