Course Content
Image and Video Generation
4 sections · 7 lessons
Creative Image Series Generator
A designer needs five header images for an article series. Same illustrated style, same colour world, five different scenes. She writes one prompt and runs it five times.
The five images share nothing. One is a flat vector illustration, one is a painterly landscape, one has a photographic depth of field. Different palettes, different line weights. Individually fine, as a set incoherent.
So she fixes the seed and changes only a few words per image. Now the opposite failure: five images that are visibly the same picture with small perturbations — identical composition, identical lighting, the same tree in the same corner. Not a series, a set of near-duplicates.
Both attempts used the same two controls — the prompt and the seed — and both are wrong for the job, because a series is defined by two independent requirements: something must stay constant, and something else must change, and the naive controls move both at once. Building a tool that separates those two axes is the whole project. This brief specifies what to build, explains why each component exists, and names the mistakes that will otherwise cost you a day each.
The three levers, and what each one actually moves
Before any code, get clear on what you are controlling. Latent diffusion generates an image by starting from a grid of random noise and denoising it step by step under text conditioning. Three inputs shape the result, and they are not interchangeable.
| Lever | What it controls | Right job | Wrong job — and the symptom |
|---|---|---|---|
| Initial latent (the seed) | Composition, layout, where things sit in frame | Varying framing while holding subject and style | Using it for style variation: you get a different picture entirely |
| Prompt text | Semantic content — what is present and roughly how it looks | Varying subject or scene across the series | Using it for fine style control: text is a blunt instrument and the model resolves ambiguity per seed |
| Style adapter (LoRA) | Learned visual style, applied to every generation identically | Holding style constant across a whole series | Leaving it off and hoping style words carry the load: they do not, which is failure one above |
A coherent series is one constant axis and one varying axis, chosen deliberately. Change both and you get five unrelated images; change neither and you get five copies.
That gives the design directly. Style is pinned by a LoRA adapter loaded once for the whole series. Content varies through a structured prompt with exactly one field changing. Composition varies through the seed — or is deliberately held, if the series is meant to show the same scene transforming.
What you are building
Four modules with clean responsibilities, a config file, a CLI, and an optional UI. Keep them separable; the tests depend on it.
image-series/ config.yaml # model id, defaults, paths -- no constants in code src/ generator.py # SeriesGenerator: the diffusion pipeline wrapper styles.py # StyleManager: discover, load, weight, unload LoRAs prompts.py # PromptBuilder: templates + one explicit variation axis evaluate.py # coherence + distinctness + adherence metrics io_utils.py # saving, sidecar metadata, contact sheets app/streamlit_app.py # optional UI tests/ # fast tests, plus a slow marker for real generation models/loras/<style>/ # each with config.json + safetensors outputs/<run_id>/ # images + manifest.jsonOne rule that matters more than the layout: every generated image gets a sidecar record with the full prompt, negative prompt, seed, model id, LoRA name and weight, scheduler, steps and guidance scale. Without it you cannot regenerate the one image the client liked, and "regenerate that one but bigger" is the most common request you will receive.
Component 1: the generation core
The pipeline wrapper does four things: load the model once, expose a small parameter surface, drive a per-image generator for reproducibility, and free memory on demand.
1import torch2from diffusers import StableDiffusionPipeline, DPMSolverMultistepScheduler34class SeriesGenerator:5 def __init__(self, model_id, device="cuda", dtype=torch.float16):6 self.device, self.dtype, self.model_id = device, dtype, model_id7 self.pipe = StableDiffusionPipeline.from_pretrained(8 model_id, torch_dtype=dtype, safety_checker=None)9 # Fewer steps for the same quality than the default scheduler.10 self.pipe.scheduler = DPMSolverMultistepScheduler.from_config(11 self.pipe.scheduler.config)12 self.pipe.to(device)13 self.pipe.enable_attention_slicing()14 self.pipe.vae.enable_slicing() # decode a batch one image at a time1516 def generate(self, prompts, seeds, negative=None, steps=25,17 guidance=7.0, height=512, width=512):18 """One image per (prompt, seed) pair. Lists must be equal length."""19 assert len(prompts) == len(seeds)20 gens = [torch.Generator(device=self.device).manual_seed(int(s))21 for s in seeds]22 out = self.pipe(23 prompt=list(prompts),24 negative_prompt=[negative or ""] * len(prompts),25 generator=gens, # one generator PER image26 num_inference_steps=steps, guidance_scale=guidance,27 height=height, width=width)28 return out.images2930 def unload(self):31 del self.pipe32 torch.cuda.empty_cache()Why a list of generators, not one
This is the first trap. Pass a single torch.Generator alongside a batch of four prompts and the four images are drawn from one continuous noise stream. Reproducible as a batch, yes — but image 3 is not reproducible on its own, because its noise depended on how much the generator had already consumed. Ask for image 3 again as a single generation and you get something different. One generator per image, seeded independently, makes every image individually reproducible.
The memory arithmetic behind attention slicing
A 512×512 image has a 64×64 latent, which is 4,096 tokens. Self-attention builds a score matrix of 4,096 × 4,096 = 16.78 million entries per head. In fp16 that is 33.6 MB per head; with 8 heads, 268 MB for one attention operation on one image. Generate a batch of 4 and that single operation peaks at 1.07 GB — on top of weights and all other activations.
Attention slicing computes those heads one at a time, cutting the peak for that operation to roughly 34 MB per image at a small speed cost. VAE slicing does the same for the decoder, which is the other common out-of-memory point because it works at full resolution rather than in latent space.
The weights themselves are modest by comparison: an 860M-parameter UNet at 2 bytes each is 1.72 GB, plus about 0.42 GB for the text encoder and VAE. Almost all memory pressure is activations, which is why batch size pushes you over, not the model.
Steps, guidance and the doubled forward pass
Classifier-free guidance runs the UNet twice per step — once with your prompt, once with the empty prompt — and extrapolates between them. So 50 steps is 100 UNet evaluations, not 50. This is why the step count is the dominant cost term and why the scheduler choice matters so much.
| Setting | Effect | Practical guidance |
|---|---|---|
| Default scheduler, 50 steps | Baseline quality | The usual starting point, and usually wasteful |
| DPM-Solver++, 20–25 steps | Comparable quality, roughly half the compute | Default for series work; 5 images at 25 steps costs what 2.5 cost at 50 |
| Guidance 7–8 | Balanced adherence | Sensible default |
| Guidance above ~12 | Over-saturated, contrast-blown, rigid | Symptom: images look "burnt". Lower it before blaming the prompt |
| Guidance below ~4 | Loose adherence, dreamy | Only when you want the model to wander |
Work an example. At 8 UNet iterations per second, 25 steps with guidance is 50 forward passes ≈ 6.3 seconds per image; a five-image series is about 31 seconds sequentially. The same series at 50 steps is 63 seconds. Halving steps with a better scheduler is the cheapest performance win in the project, and it is a two-line change.
Component 2: style management, and the mistake that destroys your model
A LoRA (low-rank adaptation) does not replace a model's weights. It adds a small learned delta to selected layers:
Here r is the rank — typically 4 to 32 — and s is the strength you set at generation time. The size saving is the point. A 768×768 attention projection has 589,824 parameters. A rank-8 adapter for it stores 768×8 + 8×768 = 12,288 parameters: 2.08% of the layer, a 48× reduction. That is why a style adapter is a few megabytes while the base model is gigabytes, and why you can keep a dozen styles on disk and swap between them.
The anti-pattern
It is tempting to implement "style strength" by scaling the model's own weights:
# WRONG. Do not do this.for param in pipe.unet.parameters(): param.data *= weightThree things are wrong with it, in increasing order of severity. It has nothing to do with LoRA maths — a LoRA adds a low-rank delta, it does not rescale the layer. It is destructive: the base weights are modified in place and there is no undo short of reloading the model. And it compounds silently. Call it three times at weight 0.8 and every parameter in the UNet has been multiplied by 0.8³ = 0.512. The model still runs. It produces washed-out mush, and because there is no error, you will spend an afternoon suspecting your prompts.
If a "strength" control mutates weights, it is a bug wearing a feature's clothes. Adapter strength belongs in the adapter's scale, not in the base model.
The correct implementation
1import json2from pathlib import Path34class StyleManager:5 def __init__(self, root="models/loras"):6 self.root = Path(root)7 self.styles = {}8 for d in sorted(self.root.iterdir()):9 cfg = d / "config.json"10 if d.is_dir() and cfg.exists():11 self.styles[d.name] = json.loads(cfg.read_text())12 self.active = []1314 def list(self):15 return sorted(self.styles)1617 def apply(self, pipe, names, weights=None):18 """Load one or more LoRAs as named adapters and set their strengths."""19 if isinstance(names, str):20 names = [names]21 weights = weights or [1.0] * len(names)22 for n in names:23 if n not in self.styles:24 raise ValueError(f"unknown style {n!r}; have {self.list()}")25 if n not in self.active:26 pipe.load_lora_weights(str(self.root / n), adapter_name=n)27 self.active.append(n)28 # Non-destructive: tells diffusers which adapters are live, and how strong.29 pipe.set_adapters(names, adapter_weights=list(weights))3031 def clear(self, pipe):32 pipe.disable_lora()33 pipe.unload_lora_weights()34 self.active = []Three notes. Load onto the pipeline, not pipe.unet — one call patches the UNet and, where present, the text encoder layers. Always pass adapter_name, or you cannot later address, reweight or remove that adapter. And call clear between series: a leftover adapter is the classic "why is everything watercolour now" bug. Weights above 1.0 exaggerate and destroy detail; blending two styles, keep the sum near 1.0.
Producing your own style, if a downloaded one will not do
| Method | Images needed | Artefact size | Teaches the model | Best for |
|---|---|---|---|---|
| Textual inversion | 3–10 | Kilobytes | A new token pointing at an existing concept | A recurring motif you can already almost describe |
| LoRA | 15–60 | 2–150 MB | A low-rank adjustment to attention layers | Style. The right default for this project |
| DreamBooth | 5–20 | Full checkpoint, GBs | A specific subject bound to a rare token | One exact person, pet or product |
For a series held together by look, LoRA is correct. For a series held together by a specific subject appearing repeatedly, DreamBooth or a LoRA trained on that subject is what you need — style words will not hold a face.
Component 3: prompts with an explicit variation axis
The obvious implementation of "generate variations" is to append a random modifier from a list per image. Do not build that. It fails on three counts: it is not reproducible without capturing the random state, the modifiers frequently contradict each other (an "impressionistic" tag landing in a photorealistic series), and it varies an unnamed mixture of things at once — which is exactly the disease this project exists to cure.
Build variation as a named axis with explicit values instead.
1from dataclasses import dataclass, field23AXES = {4 "camera": ["wide establishing shot", "medium shot", "close-up",5 "overhead view", "low angle"],6 "time": ["dawn", "midday", "golden hour", "dusk", "night"],7 "season": ["early spring", "high summer", "late autumn", "deep winter"],8 "weather": ["clear sky", "low mist", "heavy rain", "falling snow"],9}1011@dataclass12class PromptBuilder:13 template: str # e.g. "{subject}, {axis}, {setting}"14 fields: dict = field(default_factory=dict)15 negative: str = "text, watermark, signature, extra limbs, blurry"1617 def series(self, axis, n):18 if axis not in AXES:19 raise ValueError(f"unknown axis {axis!r}")20 values = AXES[axis][:n]21 if len(values) < n:22 raise ValueError(f"axis {axis!r} has {len(values)} values, need {n}")23 return [self.template.format(axis=v, **self.fields) for v in values]2425pb = PromptBuilder(26 template="{subject} in {setting}, {axis}",27 fields={"subject": "a lone stone lighthouse",28 "setting": "a rocky northern coastline"})2930for p in pb.series("time", 5):31 print(p)32# a lone stone lighthouse in a rocky northern coastline, dawn33# a lone stone lighthouse in a rocky northern coastline, midday34# ...Every prompt in the series is identical except one clause, in a known position, drawn from a fixed ordered list. The series is reproducible from the axis name alone, the progression is legible to a viewer, and when a result is wrong you know exactly which token to blame.
On quality keywords
Tags like "masterpiece, highly detailed, trending on artstation" work by pulling generation toward a cluster of the training distribution that shares those captions. The effect is real, not superstition — but it drags that cluster's visual style along, and that style fights your LoRA. If a series looks subtly like generic concept art despite a strong adapter, remove the quality tags first. Make them a config flag, off by default when a LoRA is active.
Component 4: measuring coherence instead of squinting at it
"The series looks consistent" is the requirement, and until you can compute it you cannot tell whether a change helped. Use CLIP image embeddings: encode each image to a vector, then take the cosine similarity of every pair.
1import itertools, torch23def series_metrics(images, prompts, clip_model, clip_proc, device="cuda"):4 inputs = clip_proc(images=images, return_tensors="pt").to(device)5 with torch.no_grad():6 out = clip_model.get_image_features(**inputs)7 # transformers 5 returns an output object; older versions a plain tensor8 emb = out if torch.is_tensor(out) else out.pooler_output9 emb = emb / emb.norm(dim=-1, keepdim=True)1011 pairs = list(itertools.combinations(range(len(images)), 2))12 sims = [float(emb[i] @ emb[j]) for i, j in pairs]13 return {14 "coherence": sum(sims) / len(sims), # mean pairwise similarity15 "min_pair": min(sims), # the odd one out16 "max_pair": max(sims), # near-duplicate detector17 "duplicates": [(i, j) for (i, j), s in zip(pairs, sims) if s > 0.93],18 }Work through a five-image run. Five images give 5×4/2 = 10 pairs. Suppose the similarities are 0.71, 0.74, 0.69, 0.77, 0.73, 0.70, 0.76, 0.72, 0.75 and 0.68. They sum to 7.25, so coherence = 7.25 / 10 = 0.725, with a minimum pair of 0.68 and a maximum of 0.77.
Now read it. The number alone means nothing; the band is the whole point.
| Mean pairwise similarity | Diagnosis | What to change |
|---|---|---|
| Below ~0.55 | Not a series — the images share no visual identity | Add or strengthen the style LoRA; move more of the prompt into the constant part |
| ~0.60–0.85 | Coherent set with visible variation. The target | Nothing |
| Above ~0.90 | Near-duplicates — you varied nothing meaningful | Vary the seed as well as the prompt, or pick an axis with more visual reach |
| Any single pair above 0.93 | Two images are effectively the same picture | Regenerate one of them with a different seed |
The 0.725 run above sits in the band with no duplicate pairs: a good series. Note that the mean would have looked identical if two images were at 0.95 and the rest scattered low — which is why the minimum and maximum are reported alongside it. A single averaged number hides exactly the failure you care about.
Add prompt adherence the same way, with CLIP text embeddings against each image, so you can catch the case where the series is beautifully coherent and simply not what was asked for.
Performance, configuration and the UI trap
Batching is the main throughput lever. Five sequential generations at 6.3 seconds each is 31.5 seconds; the same five batched is typically 13–15 seconds, roughly 2.2× better, because the GPU stays fed. The cost is peak memory, which scales close to linearly with batch size. On an 8 GB card, batch 4 at 512×512 with slicing is a reasonable ceiling.
| Technique | Memory effect | Speed effect | Use when |
|---|---|---|---|
| fp16 weights | Halves resident weights | Usually faster | Always, on any modern GPU |
| Attention slicing | Large drop in peak | Slightly slower | Under ~12 GB, or large batches |
| VAE slicing | Cuts the decode spike | Negligible | Always; decode is a common OOM point |
| SDPA / flash attention | Lower peak | Faster | Always, where supported |
| Model CPU offload | Very large drop | Noticeably slower | Last resort on small cards |
| Larger batch | Higher peak | Much better throughput | When memory allows |
Keep all of it in config.yaml, including the base model identifier — and check that identifier is current. Model repositories do move and get removed (the original runwayml home of SD 1.5 was taken down in 2024); a hard-coded path inside a class is how a project becomes unrunnable six months later, and a config value is a one-line fix.
This brief uses SD 1.5 because it runs on an 8 GB card and has the largest library of style LoRAs. Newer families — Stable Diffusion 3.5, FLUX.2, Qwen-Image — follow prompts better and render text far better, but they need much more memory, their own pipeline classes, and LoRAs trained for them; an SD 1.5 LoRA will not load onto them. If you switch, the design of the tool does not change, only the generator class and the adapter library.
model: id: "stable-diffusion-v1-5/stable-diffusion-v1-5" dtype: float16 attention_slicing: true vae_slicing: truegeneration: height: 512 width: 512 steps: 25 guidance_scale: 7.0 scheduler: dpmpp_2m negative: "text, watermark, signature, extra limbs, blurry"series: count: 5 axis: time quality_tags: false # off by default when a LoRA is activestyles: dir: models/loras default: null weight: 0.8output: dir: outputs contact_sheet: trueIf you build the optional Streamlit interface, there is one mistake that ruins it. Constructing the generator inside the button handler reloads the whole model on every click — 20 to 40 seconds of dead time per generation, and after a few clicks the old pipelines have not been collected and you are out of memory. Cache the heavy object at module level:
1import streamlit as st23@st.cache_resource # built once per process, reused across reruns4def get_generator(model_id):5 return SeriesGenerator(model_id)Tests worth writing
The instinct is a pytest fixture that builds the generator, with every test generating images. That suite takes minutes, needs a GPU, and will not run in CI — so it will not run at all. Split it: most of your logic needs no model.
| Test | Needs a GPU? | What it protects |
|---|---|---|
| Prompt builder produces N prompts differing in one clause | No | The core variation contract |
| Unknown axis or style raises a clear error | No | Failing loudly instead of silently generating nonsense |
| Style discovery finds every directory with a config | No | Silent "no styles found" on a path typo |
| Config loads and every required key is present | No | Crashes ten minutes into a long run |
| Metrics: identical images score ~1.0, unrelated score low | No (use fixtures) | The evaluator itself being wrong |
| Same seed twice produces byte-identical output | Yes, marked slow | Reproducibility — the guarantee everything else rests on |
| A five-image series returns five images at the requested size | Yes, marked slow, 5 steps | End-to-end wiring |
Mark the GPU tests and keep the fixture session-scoped so the model loads once for the whole run rather than once per test. Drop the step count to 5 in tests — you are checking shapes and determinism, not quality.
Definition of done, and where to take it next
Grade yourself against this before calling it finished.
| Area | Weight | Passing looks like |
|---|---|---|
| Series control | 30% | Constant and varying axes are separately settable; a documented axis produces a legible progression |
| Style system | 20% | LoRAs discovered, loaded by adapter name, weighted non-destructively, cleared between runs |
| Measurement | 20% | Coherence, min/max pair and adherence computed and written into the run manifest |
| Reproducibility | 15% | Any image regenerable from its sidecar record alone |
| Performance | 10% | Batched generation, sensible step count, documented VRAM ceiling |
| Documentation | 5% | A teammate can add a new style and a new axis without asking you anything |
Checklist: five images from one CLI command; a manifest recording prompt, seed, model, LoRA and metrics per image; the same command twice producing identical output; removing the LoRA visibly dropping the coherence score; an automatic contact sheet; fast tests passing without a GPU.
Worthwhile extensions, and what each buys you: ControlNet conditioning turns composition into a hard constraint rather than a soft one; inpainting varies one region while freezing the rest, the cleanest possible separation of axes; frame interpolation between consecutive images turns the series into a short animation; and a REST wrapper makes the generator a service other tools can call.
What this means when you build something
The transferable lesson is the discipline of naming the axes. Most generative tooling is built by someone who wants "variations", implements a random modifier, and then spends weeks adding controls to claw back the determinism they threw away on day one. Deciding up front what is constant, what varies, and which lever moves each is what turns a novelty into a tool someone can use on a deadline.
The second is that generative output needs a scalar you can watch. Not because a similarity score captures taste — it does not — but because without one, every change you make is evaluated by looking at pictures and forming an impression, and impressions do not survive contact with a parameter sweep. Once coherence is a number with a target band, you can sweep LoRA weight from 0.4 to 1.0 and read off the answer instead of arguing about it.
The third is reproducibility as a feature, not a nicety. Sidecar metadata, per-image seeds, pinned model identifiers and a run manifest cost about an hour. The alternative is the moment a client picks image four, asks for it at print resolution, and you find you never recorded its seed.