Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

A bank will not let any data leave its network, so the LLM must run on-premises. How do you size and serve it?


40 users, 3,400 tokens each: what the GPUs must hold16 GB17 GBone 80 GB GPU70 GB44 GBtwo 80 GB GPUsWeightsKV cache at peakHardware8B at 16-bit70B at FP8About 128 KB per token for the 8B and 320 KB for the 70B.
The weights fit on one card; the cache for 40 concurrent users is what forces the 70B onto two.

What you need to know

Write the workload down first

Say the bank's internal assistant has 40 concurrent users at peak, 3,000-token prompts (policy excerpts plus the question), 400-token answers, and a p95 target under 5 seconds. Every hardware decision follows from those four numbers.

Two memory costs: weights and KV cache

Text
Weights:  8B model at 16-bit   ≈ 16 GB  70B model at 8-bit   ≈ 70 GB   (FP8)  70B model at 4-bit   ≈ 35-40 GB (AWQ or GPTQ)KV cache per token (16-bit, grouped-query attention):  Llama-3-8B-class:  32 layers x 8 KV heads x 128 dims x 2 (K and V) x 2 bytes ≈ 128 KB  Llama-3-70B-class: 80 layers x 8 KV heads x 128 dims x 2 x 2 bytes           ≈ 320 KBPeak demand: 40 users x 3,400 tokens = 136,000 tokens  8B:  136,000 x 128 KB ≈ 17 GB  -> one 80 GB GPU is comfortable  70B: 136,000 x 320 KB ≈ 44 GB  -> plus 70 GB of FP8 weights: needs two 80 GB GPUs

A single 80 GB card holding 70 GB of weights has almost nothing left for cache, so it could serve only a handful of users. That is why the 70B tier runs tensor-parallel across two GPUs.

Serving: use an inference server, not a Python loop

A raw transformers generate loop handles one request at a time. vLLM (or SGLang, or TensorRT-LLM) adds:

  • Continuous batching — new requests join the running batch as soon as others finish, instead of waiting for a full batch.
  • PagedAttention — the KV cache is stored in small pages, like virtual memory, so little memory is wasted on padding.
  • An OpenAI-compatible API — application code does not change between the cloud prototype and the on-prem deployment.
Bash
vllm serve meta-llama/Llama-3.1-8B-Instruct \  --max-model-len 8192 --gpu-memory-utilization 0.90 \  --max-num-seqs 64

--max-num-seqs caps concurrent sequences; --max-model-len bounds the cache each request can claim. Weights must be copied into the bank's registry in advance, since the servers cannot download from the internet.

Operations with no cloud behind you

  1. Admission control — a token-bucket gate in front of the server queues or rejects requests above capacity, instead of letting latency explode for everyone.
  2. Pin versions — model checkpoint, quantization and server version live in an internal registry with checksums; nothing is pulled at startup.
  3. Load test — at 2x expected peak, measuring time-to-first-token and p95 end-to-end latency.
  4. On-call runbook — there is no vendor to call at 2 AM; the team needs restart, rollback and capacity playbooks.

A real-life example

Scenario (illustrative numbers). A private bank's compliance team wants an assistant over internal circulars. Security rules out any external API. The workload is 40 concurrent users at month-end, 3,000-token prompts and 400-token answers.

The team tests an 8B and a 70B instruct model on 150 labelled compliance questions. The 8B scores 86% and the 70B 92%. They deploy the 8B on one 80 GB GPU and a 70B at FP8 on two GPUs, routing "interpretation" questions to the 70B. A load test at 80 concurrent users shows the 8B at 1.9 s p95, while the 70B reaches 6.5 s, so the gate caps the 70B at 20 concurrent requests and queues the rest with a "busy" message. Month-end runs without an outage.

Follow-up questions to expect

  • "Why not just quantize the 70B to 4-bit and use one GPU?" — Possible, but measure quality on your eval first; some reasoning tasks degrade. You also still need cache memory for 40 users.
  • "How do you update the model?" — Treat it like a release: new checkpoint into the registry, run the eval suite, canary on a slice of users, keep the old version ready for rollback.
  • "What if demand doubles?" — Buy or allocate GPUs ahead of time; on-prem has no autoscaling, so capacity planning replaces it.