Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
A bank will not let any data leave its network, so the LLM must run on-premises. How do you size and serve it?
What you need to know
Write the workload down first
Say the bank's internal assistant has 40 concurrent users at peak, 3,000-token prompts (policy excerpts plus the question), 400-token answers, and a p95 target under 5 seconds. Every hardware decision follows from those four numbers.
Two memory costs: weights and KV cache
Weights: 8B model at 16-bit ≈ 16 GB 70B model at 8-bit ≈ 70 GB (FP8) 70B model at 4-bit ≈ 35-40 GB (AWQ or GPTQ)KV cache per token (16-bit, grouped-query attention): Llama-3-8B-class: 32 layers x 8 KV heads x 128 dims x 2 (K and V) x 2 bytes ≈ 128 KB Llama-3-70B-class: 80 layers x 8 KV heads x 128 dims x 2 x 2 bytes ≈ 320 KBPeak demand: 40 users x 3,400 tokens = 136,000 tokens 8B: 136,000 x 128 KB ≈ 17 GB -> one 80 GB GPU is comfortable 70B: 136,000 x 320 KB ≈ 44 GB -> plus 70 GB of FP8 weights: needs two 80 GB GPUsA single 80 GB card holding 70 GB of weights has almost nothing left for cache, so it could serve only a handful of users. That is why the 70B tier runs tensor-parallel across two GPUs.
Serving: use an inference server, not a Python loop
A raw transformers generate loop handles one request at a time. vLLM (or SGLang, or TensorRT-LLM) adds:
- Continuous batching — new requests join the running batch as soon as others finish, instead of waiting for a full batch.
- PagedAttention — the KV cache is stored in small pages, like virtual memory, so little memory is wasted on padding.
- An OpenAI-compatible API — application code does not change between the cloud prototype and the on-prem deployment.
vllm serve meta-llama/Llama-3.1-8B-Instruct \ --max-model-len 8192 --gpu-memory-utilization 0.90 \ --max-num-seqs 64--max-num-seqs caps concurrent sequences; --max-model-len bounds the cache each request can claim. Weights must be copied into the bank's registry in advance, since the servers cannot download from the internet.
Operations with no cloud behind you
- Admission control — a token-bucket gate in front of the server queues or rejects requests above capacity, instead of letting latency explode for everyone.
- Pin versions — model checkpoint, quantization and server version live in an internal registry with checksums; nothing is pulled at startup.
- Load test — at 2x expected peak, measuring time-to-first-token and p95 end-to-end latency.
- On-call runbook — there is no vendor to call at 2 AM; the team needs restart, rollback and capacity playbooks.
A real-life example
Scenario (illustrative numbers). A private bank's compliance team wants an assistant over internal circulars. Security rules out any external API. The workload is 40 concurrent users at month-end, 3,000-token prompts and 400-token answers.
The team tests an 8B and a 70B instruct model on 150 labelled compliance questions. The 8B scores 86% and the 70B 92%. They deploy the 8B on one 80 GB GPU and a 70B at FP8 on two GPUs, routing "interpretation" questions to the 70B. A load test at 80 concurrent users shows the 8B at 1.9 s p95, while the 70B reaches 6.5 s, so the gate caps the 70B at 20 concurrent requests and queues the rest with a "busy" message. Month-end runs without an outage.
Follow-up questions to expect
- "Why not just quantize the 70B to 4-bit and use one GPU?" — Possible, but measure quality on your eval first; some reasoning tasks degrade. You also still need cache memory for 40 users.
- "How do you update the model?" — Treat it like a release: new checkpoint into the registry, run the eval suite, canary on a slice of users, keep the old version ready for rollback.
- "What if demand doubles?" — Buy or allocate GPUs ahead of time; on-prem has no autoscaling, so capacity planning replaces it.