Course Content
Generative AI System Design Interview
11 sections · 27 lessons
Text-to-image: conditioning, guidance and the known weaknesses
Three components turn Section 8's unconditional denoiser into a text-to-image model: a text encoder, cross-attention, and a guidance scale. Understanding them is what lets you explain, in the second half of this lesson, why the model still gets a red cube and a blue sphere wrong.
The text encoder
The prompt is tokenised and passed through a text encoder, producing a sequence of embeddings — one per token — not a single vector. That distinction matters: keeping a sequence is what allows the image to attend to different words at different points.
Two families, with different strengths:
- Contrastive text encoders — the text half of an image-text model trained as in the image captioning training lesson. Their representations are already aligned to images, which makes them excellent at nouns, styles, and visual concepts. They are usually trained on short captions and are correspondingly weak on long, syntactically complex prompts.
- Large language-model encoders — a general text model used as a feature extractor. Far better at long prompts, clauses, and compositional structure, and not natively aligned to images, so the model has to learn the mapping.
Many systems use both, concatenating the two sets of embeddings, which is a straightforward way to get the visual grounding of one and the syntactic competence of the other. If asked to choose one, the honest answer is that longer and more compositional prompts push you towards the language-model encoder, and that this choice is an area where practice is still moving.
Cross-attention: how the text reaches the image
The denoising network from How diffusion works, slowly is modified by inserting cross-attention layers at several resolution levels of its U-shaped path. Transformer-based denoisers, now common in recent models, do the same job either with cross-attention in every block or by attending over image and text tokens jointly; the soft-blend behaviour described below applies either way.
At each such layer, the image features query the text embeddings and pull in what is relevant. This is exactly the mechanism from the image captioning model lesson running in the opposite direction: there, a generated word attended over image regions; here, an image region attends over prompt tokens.
Two properties worth stating, because they explain a great deal about the known weaknesses below:
It happens at every step. All 25 or 50 denoising steps consult the text. The prompt is not read once and stored; it is re-queried continuously, which is why guidance strength affects the whole trajectory rather than only the beginning.
It happens at multiple resolutions. Coarse levels, where global composition is decided, and fine levels, where texture is decided, both attend to the text. That is why a prompt can influence both layout and surface detail.
And the property that causes trouble: cross-attention produces a soft weighted blend over tokens for each image position. There is no mechanism that binds red to cube rather than to sphere, and no mechanism that counts. Attention decides how much each token matters at each position; it does not decide which object a word belongs to.
What the scale actually does
Illustrative behaviour for a typical model, described rather than measured:
| Scale | What you get |
|---|---|
| 1 | Beautiful, varied, and frequently ignores the prompt |
| 3 | Loosely follows the prompt; high diversity |
| 7 | The usual default — good adherence, good quality |
| 12 | Strong adherence; colours saturate, contrast hardens, composition simplifies |
| 20 | Prompt-obedient and visibly degraded — burnt highlights, posterised colour, flat scenes |
The curve is not monotone in usefulness. Alignment rises with scale and quality falls, and they cross somewhere around 6–9 for most models. That crossover is a per-model tuning job and it is worth exposing to advanced users while keeping a sensible default for everyone else.
One refinement worth knowing: guidance can be varied across the trajectory. High guidance early, when composition is being decided, and lower guidance late, when texture is being filled in, gives much of the adherence at less of the quality cost — and, as the diffusion serving lesson noted, dropping guidance entirely on late steps also recovers part of its cost.
The known weaknesses
Naming these honestly is what a strong candidate does, and they are the natural deep-dive questions. Each has a mechanism, and the mechanism is what makes the answer good.
Text rendering inside images
Ask for a shop sign reading "OPEN" and you frequently get letter-shaped marks that are not letters.
Three compounding causes:
- The text encoder does not see characters. It sees subword tokens. The token for
OPENcarries the meaning of the word, not the shapes of four specific glyphs in order. - The autoencoder ceiling. As the latent diffusion lesson showed, small text often does not survive an encode-decode round trip. If the decoder cannot render it, the diffusion model cannot produce it. Run the round-trip diagnostic before blaming the generator.
- The captions never transcribed it. Web alt-text almost never says what text appears in the image, so the model was never trained to associate a specific string with specific glyphs.
What has improved it, in order of impact: re-captioning that explicitly transcribes visible text, character-aware or byte-level text encoders, and less aggressive autoencoder compression. This is the weakness where the most progress has been made recently, and it is progress rather than a solution.
Counting
"Five apples" frequently produces four or six.
There is no counting mechanism anywhere in the architecture. Cross-attention distributes attention over tokens; nothing accumulates a tally. The model learns statistical associations between number words and typical quantities in training images, and those associations are weak because captions rarely state counts accurately.
Small numbers — one, two, three — work reasonably because they are common and visually distinct. Beyond about four, performance degrades sharply. Say that; it is more useful than "counting is hard".
Spatial relations
"A cube to the left of a sphere" produces the right objects in an arbitrary arrangement roughly as often as not.
Same root cause: attention says how much a token matters at a position, not where an object should go, and training captions rarely describe spatial layout. The model has learned which objects co-occur, not how they are arranged.
Attribute binding
The clearest single demonstration of the problem. "A red cube and a blue sphere" produces a blue cube and a red sphere often enough to be a standard test case.
The mechanism: the prompt's embeddings carry red, cube, blue, and sphere as a sequence, and cross-attention lets every image region attend to all of them. Nothing enforces that the region rendering the cube should attend to red and not to blue. Attributes leak across objects.
Compositional prompts
Combine the above — several objects, each with attributes, in specified relations, in specified quantities — and reliability falls off a cliff. Two objects with two attributes is often fine; four objects with relations between them is essentially unreliable.