Image and Video Generation

Creative Image Series Generator


A designer needs five header images for an article series. Same illustrated style, same colour world, five different scenes. She writes one prompt and runs it five times.

The five images share nothing. One is a flat vector illustration, one is a painterly landscape, one has a photographic depth of field. Different palettes, different line weights. Individually fine, as a set incoherent.

So she fixes the seed and changes only a few words per image. Now the opposite failure: five images that are visibly the same picture with small perturbations — identical composition, identical lighting, the same tree in the same corner. Not a series, a set of near-duplicates.

Both attempts used the same two controls — the prompt and the seed — and both are wrong for the job, because a series is defined by two independent requirements: something must stay constant, and something else must change, and the naive controls move both at once. Building a tool that separates those two axes is the whole project. This brief specifies what to build, explains why each component exists, and names the mistakes that will otherwise cost you a day each.

Five images that belong to one seriesScene A:harbourScene B:rooftopScene C:marketScene D:bridgeScene E:stationnullsame style LoRAsame seedand stepsOne variation axis moves; style, seed, sampler, guidance and negative prompt are all held fixed.
Coherence comes from holding everything constant except the one axis you meant to vary — and from measuring it, not squinting at it.

The three levers, and what each one actually moves

Before any code, get clear on what you are controlling. Latent diffusion generates an image by starting from a grid of random noise and denoising it step by step under text conditioning. Three inputs shape the result, and they are not interchangeable.

LeverWhat it controlsRight jobWrong job — and the symptom
Initial latent (the seed)Composition, layout, where things sit in frameVarying framing while holding subject and styleUsing it for style variation: you get a different picture entirely
Prompt textSemantic content — what is present and roughly how it looksVarying subject or scene across the seriesUsing it for fine style control: text is a blunt instrument and the model resolves ambiguity per seed
Style adapter (LoRA)Learned visual style, applied to every generation identicallyHolding style constant across a whole seriesLeaving it off and hoping style words carry the load: they do not, which is failure one above

A coherent series is one constant axis and one varying axis, chosen deliberately. Change both and you get five unrelated images; change neither and you get five copies.

That gives the design directly. Style is pinned by a LoRA adapter loaded once for the whole series. Content varies through a structured prompt with exactly one field changing. Composition varies through the seed — or is deliberately held, if the series is meant to show the same scene transforming.

What you are building

Four modules with clean responsibilities, a config file, a CLI, and an optional UI. Keep them separable; the tests depend on it.

Text
image-series/  config.yaml               # model id, defaults, paths -- no constants in code  src/    generator.py            # SeriesGenerator: the diffusion pipeline wrapper    styles.py               # StyleManager: discover, load, weight, unload LoRAs    prompts.py              # PromptBuilder: templates + one explicit variation axis    evaluate.py             # coherence + distinctness + adherence metrics    io_utils.py             # saving, sidecar metadata, contact sheets  app/streamlit_app.py      # optional UI  tests/                    # fast tests, plus a slow marker for real generation  models/loras/<style>/     # each with config.json + safetensors  outputs/<run_id>/         # images + manifest.json

One rule that matters more than the layout: every generated image gets a sidecar record with the full prompt, negative prompt, seed, model id, LoRA name and weight, scheduler, steps and guidance scale. Without it you cannot regenerate the one image the client liked, and "regenerate that one but bigger" is the most common request you will receive.

Component 1: the generation core

The pipeline wrapper does four things: load the model once, expose a small parameter surface, drive a per-image generator for reproducibility, and free memory on demand.

Python
import torchfrom diffusers import StableDiffusionPipeline, DPMSolverMultistepSchedulerclass SeriesGenerator:    def __init__(self, model_id, device="cuda", dtype=torch.float16):        self.device, self.dtype, self.model_id = device, dtype, model_id        self.pipe = StableDiffusionPipeline.from_pretrained(            model_id, torch_dtype=dtype, safety_checker=None)        # Fewer steps for the same quality than the default scheduler.        self.pipe.scheduler = DPMSolverMultistepScheduler.from_config(            self.pipe.scheduler.config)        self.pipe.to(device)        self.pipe.enable_attention_slicing()        self.pipe.vae.enable_slicing()     # decode a batch one image at a time    def generate(self, prompts, seeds, negative=None, steps=25,                 guidance=7.0, height=512, width=512):        """One image per (prompt, seed) pair. Lists must be equal length."""        assert len(prompts) == len(seeds)        gens = [torch.Generator(device=self.device).manual_seed(int(s))                for s in seeds]        out = self.pipe(            prompt=list(prompts),            negative_prompt=[negative or ""] * len(prompts),            generator=gens,                 # one generator PER image            num_inference_steps=steps, guidance_scale=guidance,            height=height, width=width)        return out.images    def unload(self):        del self.pipe        torch.cuda.empty_cache()

Why a list of generators, not one

This is the first trap. Pass a single torch.Generator alongside a batch of four prompts and the four images are drawn from one continuous noise stream. Reproducible as a batch, yes — but image 3 is not reproducible on its own, because its noise depended on how much the generator had already consumed. Ask for image 3 again as a single generation and you get something different. One generator per image, seeded independently, makes every image individually reproducible.

The memory arithmetic behind attention slicing

A 512×512 image has a 64×64 latent, which is 4,096 tokens. Self-attention builds a score matrix of 4,096 × 4,096 = 16.78 million entries per head. In fp16 that is 33.6 MB per head; with 8 heads, 268 MB for one attention operation on one image. Generate a batch of 4 and that single operation peaks at 1.07 GB — on top of weights and all other activations.

Attention slicing computes those heads one at a time, cutting the peak for that operation to roughly 34 MB per image at a small speed cost. VAE slicing does the same for the decoder, which is the other common out-of-memory point because it works at full resolution rather than in latent space.

The weights themselves are modest by comparison: an 860M-parameter UNet at 2 bytes each is 1.72 GB, plus about 0.42 GB for the text encoder and VAE. Almost all memory pressure is activations, which is why batch size pushes you over, not the model.

Steps, guidance and the doubled forward pass

Classifier-free guidance runs the UNet twice per step — once with your prompt, once with the empty prompt — and extrapolates between them. So 50 steps is 100 UNet evaluations, not 50. This is why the step count is the dominant cost term and why the scheduler choice matters so much.

SettingEffectPractical guidance
Default scheduler, 50 stepsBaseline qualityThe usual starting point, and usually wasteful
DPM-Solver++, 20–25 stepsComparable quality, roughly half the computeDefault for series work; 5 images at 25 steps costs what 2.5 cost at 50
Guidance 7–8Balanced adherenceSensible default
Guidance above ~12Over-saturated, contrast-blown, rigidSymptom: images look "burnt". Lower it before blaming the prompt
Guidance below ~4Loose adherence, dreamyOnly when you want the model to wander

Work an example. At 8 UNet iterations per second, 25 steps with guidance is 50 forward passes ≈ 6.3 seconds per image; a five-image series is about 31 seconds sequentially. The same series at 50 steps is 63 seconds. Halving steps with a better scheduler is the cheapest performance win in the project, and it is a two-line change.

Component 2: style management, and the mistake that destroys your model

A LoRA (low-rank adaptation) does not replace a model's weights. It adds a small learned delta to selected layers:

W′=W+s⋅(BA),B∈Rd×r,  A∈Rr×kW' = W + s \cdot (B A), \qquad B \in \mathbb{R}^{d \times r},\; A \in \mathbb{R}^{r \times k}

Here rr is the rank — typically 4 to 32 — and ss is the strength you set at generation time. The size saving is the point. A 768×768 attention projection has 589,824 parameters. A rank-8 adapter for it stores 768×8 + 8×768 = 12,288 parameters: 2.08% of the layer, a 48× reduction. That is why a style adapter is a few megabytes while the base model is gigabytes, and why you can keep a dozen styles on disk and swap between them.

The anti-pattern

It is tempting to implement "style strength" by scaling the model's own weights:

Python
# WRONG. Do not do this.for param in pipe.unet.parameters():    param.data *= weight

Three things are wrong with it, in increasing order of severity. It has nothing to do with LoRA maths — a LoRA adds a low-rank delta, it does not rescale the layer. It is destructive: the base weights are modified in place and there is no undo short of reloading the model. And it compounds silently. Call it three times at weight 0.8 and every parameter in the UNet has been multiplied by 0.8³ = 0.512. The model still runs. It produces washed-out mush, and because there is no error, you will spend an afternoon suspecting your prompts.

If a "strength" control mutates weights, it is a bug wearing a feature's clothes. Adapter strength belongs in the adapter's scale, not in the base model.

The correct implementation

Python
import jsonfrom pathlib import Pathclass StyleManager:    def __init__(self, root="models/loras"):        self.root = Path(root)        self.styles = {}        for d in sorted(self.root.iterdir()):            cfg = d / "config.json"            if d.is_dir() and cfg.exists():                self.styles[d.name] = json.loads(cfg.read_text())        self.active = []    def list(self):        return sorted(self.styles)    def apply(self, pipe, names, weights=None):        """Load one or more LoRAs as named adapters and set their strengths."""        if isinstance(names, str):            names = [names]        weights = weights or [1.0] * len(names)        for n in names:            if n not in self.styles:                raise ValueError(f"unknown style {n!r}; have {self.list()}")            if n not in self.active:                pipe.load_lora_weights(str(self.root / n), adapter_name=n)                self.active.append(n)        # Non-destructive: tells diffusers which adapters are live, and how strong.        pipe.set_adapters(names, adapter_weights=list(weights))    def clear(self, pipe):        pipe.disable_lora()        pipe.unload_lora_weights()        self.active = []

Three notes. Load onto the pipeline, not pipe.unet — one call patches the UNet and, where present, the text encoder layers. Always pass adapter_name, or you cannot later address, reweight or remove that adapter. And call clear between series: a leftover adapter is the classic "why is everything watercolour now" bug. Weights above 1.0 exaggerate and destroy detail; blending two styles, keep the sum near 1.0.

Producing your own style, if a downloaded one will not do

MethodImages neededArtefact sizeTeaches the modelBest for
Textual inversion3–10KilobytesA new token pointing at an existing conceptA recurring motif you can already almost describe
LoRA15–602–150 MBA low-rank adjustment to attention layersStyle. The right default for this project
DreamBooth5–20Full checkpoint, GBsA specific subject bound to a rare tokenOne exact person, pet or product

For a series held together by look, LoRA is correct. For a series held together by a specific subject appearing repeatedly, DreamBooth or a LoRA trained on that subject is what you need — style words will not hold a face.

Component 3: prompts with an explicit variation axis

The obvious implementation of "generate variations" is to append a random modifier from a list per image. Do not build that. It fails on three counts: it is not reproducible without capturing the random state, the modifiers frequently contradict each other (an "impressionistic" tag landing in a photorealistic series), and it varies an unnamed mixture of things at once — which is exactly the disease this project exists to cure.

Build variation as a named axis with explicit values instead.

Python
from dataclasses import dataclass, fieldAXES = {    "camera":  ["wide establishing shot", "medium shot", "close-up",                "overhead view", "low angle"],    "time":    ["dawn", "midday", "golden hour", "dusk", "night"],    "season":  ["early spring", "high summer", "late autumn", "deep winter"],    "weather": ["clear sky", "low mist", "heavy rain", "falling snow"],}@dataclassclass PromptBuilder:    template: str                  # e.g. "{subject}, {axis}, {setting}"    fields: dict = field(default_factory=dict)    negative: str = "text, watermark, signature, extra limbs, blurry"    def series(self, axis, n):        if axis not in AXES:            raise ValueError(f"unknown axis {axis!r}")        values = AXES[axis][:n]        if len(values) < n:            raise ValueError(f"axis {axis!r} has {len(values)} values, need {n}")        return [self.template.format(axis=v, **self.fields) for v in values]pb = PromptBuilder(    template="{subject} in {setting}, {axis}",    fields={"subject": "a lone stone lighthouse",            "setting": "a rocky northern coastline"})for p in pb.series("time", 5):    print(p)# a lone stone lighthouse in a rocky northern coastline, dawn# a lone stone lighthouse in a rocky northern coastline, midday# ...

Every prompt in the series is identical except one clause, in a known position, drawn from a fixed ordered list. The series is reproducible from the axis name alone, the progression is legible to a viewer, and when a result is wrong you know exactly which token to blame.

On quality keywords

Tags like "masterpiece, highly detailed, trending on artstation" work by pulling generation toward a cluster of the training distribution that shares those captions. The effect is real, not superstition — but it drags that cluster's visual style along, and that style fights your LoRA. If a series looks subtly like generic concept art despite a strong adapter, remove the quality tags first. Make them a config flag, off by default when a LoRA is active.

Component 4: measuring coherence instead of squinting at it

"The series looks consistent" is the requirement, and until you can compute it you cannot tell whether a change helped. Use CLIP image embeddings: encode each image to a vector, then take the cosine similarity of every pair.

Python
import itertools, torchdef series_metrics(images, prompts, clip_model, clip_proc, device="cuda"):    inputs = clip_proc(images=images, return_tensors="pt").to(device)    with torch.no_grad():        out = clip_model.get_image_features(**inputs)    # transformers 5 returns an output object; older versions a plain tensor    emb = out if torch.is_tensor(out) else out.pooler_output    emb = emb / emb.norm(dim=-1, keepdim=True)    pairs = list(itertools.combinations(range(len(images)), 2))    sims = [float(emb[i] @ emb[j]) for i, j in pairs]    return {        "coherence": sum(sims) / len(sims),          # mean pairwise similarity        "min_pair": min(sims),                        # the odd one out        "max_pair": max(sims),                        # near-duplicate detector        "duplicates": [(i, j) for (i, j), s in zip(pairs, sims) if s > 0.93],    }

Work through a five-image run. Five images give 5×4/2 = 10 pairs. Suppose the similarities are 0.71, 0.74, 0.69, 0.77, 0.73, 0.70, 0.76, 0.72, 0.75 and 0.68. They sum to 7.25, so coherence = 7.25 / 10 = 0.725, with a minimum pair of 0.68 and a maximum of 0.77.

Now read it. The number alone means nothing; the band is the whole point.

Mean pairwise similarityDiagnosisWhat to change
Below ~0.55Not a series — the images share no visual identityAdd or strengthen the style LoRA; move more of the prompt into the constant part
~0.60–0.85Coherent set with visible variation. The targetNothing
Above ~0.90Near-duplicates — you varied nothing meaningfulVary the seed as well as the prompt, or pick an axis with more visual reach
Any single pair above 0.93Two images are effectively the same pictureRegenerate one of them with a different seed

The 0.725 run above sits in the band with no duplicate pairs: a good series. Note that the mean would have looked identical if two images were at 0.95 and the rest scattered low — which is why the minimum and maximum are reported alongside it. A single averaged number hides exactly the failure you care about.

Add prompt adherence the same way, with CLIP text embeddings against each image, so you can catch the case where the series is beautifully coherent and simply not what was asked for.

Performance, configuration and the UI trap

Batching is the main throughput lever. Five sequential generations at 6.3 seconds each is 31.5 seconds; the same five batched is typically 13–15 seconds, roughly 2.2× better, because the GPU stays fed. The cost is peak memory, which scales close to linearly with batch size. On an 8 GB card, batch 4 at 512×512 with slicing is a reasonable ceiling.

TechniqueMemory effectSpeed effectUse when
fp16 weightsHalves resident weightsUsually fasterAlways, on any modern GPU
Attention slicingLarge drop in peakSlightly slowerUnder ~12 GB, or large batches
VAE slicingCuts the decode spikeNegligibleAlways; decode is a common OOM point
SDPA / flash attentionLower peakFasterAlways, where supported
Model CPU offloadVery large dropNoticeably slowerLast resort on small cards
Larger batchHigher peakMuch better throughputWhen memory allows

Keep all of it in config.yaml, including the base model identifier — and check that identifier is current. Model repositories do move and get removed (the original runwayml home of SD 1.5 was taken down in 2024); a hard-coded path inside a class is how a project becomes unrunnable six months later, and a config value is a one-line fix.

This brief uses SD 1.5 because it runs on an 8 GB card and has the largest library of style LoRAs. Newer families — Stable Diffusion 3.5, FLUX.2, Qwen-Image — follow prompts better and render text far better, but they need much more memory, their own pipeline classes, and LoRAs trained for them; an SD 1.5 LoRA will not load onto them. If you switch, the design of the tool does not change, only the generator class and the adapter library.

Text
model:  id: "stable-diffusion-v1-5/stable-diffusion-v1-5"  dtype: float16  attention_slicing: true  vae_slicing: truegeneration:  height: 512  width: 512  steps: 25  guidance_scale: 7.0  scheduler: dpmpp_2m  negative: "text, watermark, signature, extra limbs, blurry"series:  count: 5  axis: time  quality_tags: false          # off by default when a LoRA is activestyles:  dir: models/loras  default: null  weight: 0.8output:  dir: outputs  contact_sheet: true

If you build the optional Streamlit interface, there is one mistake that ruins it. Constructing the generator inside the button handler reloads the whole model on every click — 20 to 40 seconds of dead time per generation, and after a few clicks the old pipelines have not been collected and you are out of memory. Cache the heavy object at module level:

Python
import streamlit as st@st.cache_resource      # built once per process, reused across rerunsdef get_generator(model_id):    return SeriesGenerator(model_id)

Tests worth writing

The instinct is a pytest fixture that builds the generator, with every test generating images. That suite takes minutes, needs a GPU, and will not run in CI — so it will not run at all. Split it: most of your logic needs no model.

TestNeeds a GPU?What it protects
Prompt builder produces N prompts differing in one clauseNoThe core variation contract
Unknown axis or style raises a clear errorNoFailing loudly instead of silently generating nonsense
Style discovery finds every directory with a configNoSilent "no styles found" on a path typo
Config loads and every required key is presentNoCrashes ten minutes into a long run
Metrics: identical images score ~1.0, unrelated score lowNo (use fixtures)The evaluator itself being wrong
Same seed twice produces byte-identical outputYes, marked slowReproducibility — the guarantee everything else rests on
A five-image series returns five images at the requested sizeYes, marked slow, 5 stepsEnd-to-end wiring

Mark the GPU tests and keep the fixture session-scoped so the model loads once for the whole run rather than once per test. Drop the step count to 5 in tests — you are checking shapes and determinism, not quality.

Definition of done, and where to take it next

Grade yourself against this before calling it finished.

AreaWeightPassing looks like
Series control30%Constant and varying axes are separately settable; a documented axis produces a legible progression
Style system20%LoRAs discovered, loaded by adapter name, weighted non-destructively, cleared between runs
Measurement20%Coherence, min/max pair and adherence computed and written into the run manifest
Reproducibility15%Any image regenerable from its sidecar record alone
Performance10%Batched generation, sensible step count, documented VRAM ceiling
Documentation5%A teammate can add a new style and a new axis without asking you anything

Checklist: five images from one CLI command; a manifest recording prompt, seed, model, LoRA and metrics per image; the same command twice producing identical output; removing the LoRA visibly dropping the coherence score; an automatic contact sheet; fast tests passing without a GPU.

Worthwhile extensions, and what each buys you: ControlNet conditioning turns composition into a hard constraint rather than a soft one; inpainting varies one region while freezing the rest, the cleanest possible separation of axes; frame interpolation between consecutive images turns the series into a short animation; and a REST wrapper makes the generator a service other tools can call.

What this means when you build something

The transferable lesson is the discipline of naming the axes. Most generative tooling is built by someone who wants "variations", implements a random modifier, and then spends weeks adding controls to claw back the determinism they threw away on day one. Deciding up front what is constant, what varies, and which lever moves each is what turns a novelty into a tool someone can use on a deadline.

The second is that generative output needs a scalar you can watch. Not because a similarity score captures taste — it does not — but because without one, every change you make is evaluated by looking at pictures and forming an impression, and impressions do not survive contact with a parameter sweep. Once coherence is a number with a target band, you can sweep LoRA weight from 0.4 to 1.0 and read off the answer instead of arguing about it.

The third is reproducibility as a feature, not a nicety. Sidecar metadata, per-image seeds, pinned model identifiers and a run manifest cost about an hour. The alternative is the moment a client picks image four, asks for it at print resolution, and you find you never recorded its seed.