Course Content
AI Safety & Guardrails
5 sections · 50 lessons
How do you reduce environmental impact from training large models?
What you need to know
Training
- Don't train from scratch when fine-tuning an existing model works.
- Parameter-efficient fine-tuning: LoRA and QLoRA train small adapter matrices instead of all weights, using a fraction of the memory and compute.
- Right-size: a 7B model that meets the requirement beats a 70B model that exceeds it.
- Efficiency per step: mixed precision (BF16, FP8 on supported hardware), FlashAttention, mixture-of-experts and grouped-query attention.
- Fewer steps: deduplicated, filtered data reaches the same quality with fewer tokens; scaling laws help avoid over-training.
- Smarter search: early-stopping methods (ASHA, Hyperband) or Bayesian optimisation instead of full grid search.
- Checkpointing: a crash near the end should not waste the whole run.
Inference (usually the bigger lifetime cost)
- Quantisation (8-bit or 4-bit weights) and distillation into smaller models.
- Caching: provider prompt caching for repeated prefixes; semantic caching for repeated questions.
- Batching: continuous batching in servers like vLLM raises GPU utilisation.
- Routing: a small model handles easy requests; the big model only the hard ones.
- Output caps: shorter answers use less compute.
Where and when
The same job emits different amounts of CO2 depending on the grid. Choose regions with low carbon intensity, and schedule flexible jobs (training, batch evals, re-embedding) when the grid is cleaner. Water use for data-centre cooling is a growing concern too, and differs by site.
Measuring
- Energy (kWh) and estimated CO2e per training run and per million tokens served.
- Tools like CodeCarbon estimate emissions from hardware power and location.
- Report in the model card, and be honest that these are estimates: data-centre efficiency (PUE), grid mix and hardware utilisation are often assumed. NIST's generative AI profile lists environmental impact as one of its risk areas.
A real-life example
A bank's customer chatbot uses a large model for all 40,000 daily chats. Analysis shows 55% of questions are simple FAQs (branch hours, IFSC codes, card blocking steps). The team adds a semantic cache for top FAQs, routes simple intents to a small model, and caps answers at 250 tokens.
Result: large-model calls fall by 60%, estimated inference energy by about half, and monthly model cost from Rs 18 lakh to Rs 8 lakh, with no drop in customer satisfaction on the answered questions. They also move a monthly re-embedding job to a cloud region with a cleaner grid and run it overnight. Cost and carbon usually point the same way, which makes this an easy case to win internally.
Follow-up questions to expect
- "Is training or inference bigger?" — For a popular product, inference usually dominates over the model's lifetime, because it runs every day for every user; training is a large one-time cost.
- "Does quantisation hurt quality?" — Often a little; 8-bit is usually close to full quality, 4-bit needs testing on your own eval set.
- "How reliable are carbon estimates?" — Rough. They depend on assumed power draw, utilisation and grid data; report ranges and methods, not a single precise figure.