Course Content
LLMOps & Deployment
6 sections · 40 lessons
How do you deploy and serve large language models in production environments?
What you need to know
The three deployment shapes
Managed API
- Anthropic, OpenAI, Google, or the same models via AWS Bedrock, Azure, Vertex
- No GPUs to run; pay per token
- Best models, fastest start
- Data leaves your network (check retention terms)
Self-hosted inference server
- vLLM, SGLang or TensorRT-LLM on your GPUs
- You pay for GPU hours, busy or idle
- Full control of data, model and latency
- You own scaling, upgrades and on-call
Local runtime
- Ollama, llama.cpp (GGUF files)
- Runs on a laptop, a Mac or one small server
- Great for development and tiny on-prem use
- Not built for many concurrent users
The inference engines, as of 2026
- vLLM — the most widely used open-source server. PagedAttention stores the KV cache in small fixed-size blocks, like pages in an operating system, so memory is not wasted on unused space. Continuous batching adds new requests to the running batch at every decode step instead of waiting for the batch to finish. It also has prefix caching, tensor parallelism, quantization support, multi-LoRA serving and an OpenAI-compatible API.
- SGLang — a strong alternative with similar features. Its RadixAttention reuses shared prompt prefixes very well, which helps agents and multi-turn chat.
- TensorRT-LLM — NVIDIA's engine; usually the fastest on NVIDIA GPUs, with more setup and tied to NVIDIA hardware.
- TGI (Hugging Face Text Generation Inference) — now in maintenance mode; Hugging Face recommends vLLM or SGLang for new deployments. Say this if an interviewer mentions it.
- Ollama / llama.cpp — easy local serving of quantized models; both expose OpenAI-compatible endpoints.
The platform around the engine
On Kubernetes, teams use KServe (which can run vLLM as its LLM runtime) or Ray Serve (its LLM APIs run vLLM underneath), or projects such as llm-d that add KV-cache-aware routing. Around that: autoscaling on queue depth, a gateway (for example a LiteLLM-style proxy) for keys, routing and fallbacks, and SSE streaming to clients.
First sizing question: does the model fit?
Weights need about 2 bytes per parameter in bf16, 1 byte in int8/FP8 and about 0.5 byte in int4. So an 8B model needs about 16 GB in bf16 and a 70B model about 140 GB. The rest of GPU memory holds the KV cache — the saved attention keys and values of every token in flight — and that decides how many requests run at once.
1vllm serve meta-llama/Llama-3.1-8B-Instruct \2 --max-model-len 16384 \3 --gpu-memory-utilization 0.90 \4 --max-num-seqs 64This starts an OpenAI-compatible server. --gpu-memory-utilization is the share of GPU memory vLLM may take (weights plus KV cache), --max-model-len caps prompt plus output length per request, and --max-num-seqs caps how many requests are batched together.
A real-life example
A hospital wants a discharge-note summariser. Patient data cannot go to any cloud, so a managed API is ruled out. The team picks an 8B instruction model and runs vLLM on one server with two 48 GB L40S GPUs, one replica per GPU, behind an internal gateway.
The numbers: 16 GB of weights per GPU leaves roughly 24 GB for KV cache after overheads. At 128 KB of KV cache per token for this model, that holds about 196,000 tokens — around 24 notes of 8,000 tokens each in flight per GPU. With 600 notes a day, that is far more than enough, so they keep the second GPU as a hot spare for failover. The doctors' app talks only to the gateway, so the model can be swapped later without touching the app.
Follow-up questions to expect
- "Why not just use Ollama in production?" — It is great for one user or a small team, but it is not designed for high-concurrency batching, autoscaling and fleet metrics the way vLLM or SGLang are.
- "Why does an OpenAI-compatible endpoint matter?" — The same client code, gateways and observability tools work with your own model and with hosted APIs, so switching is a config change.
- "How would you serve a 70B model?" — About 140 GB in bf16 needs tensor parallelism over two 80 GB GPUs, or int4/FP8 quantization to fit fewer GPUs; I would test quality after quantizing.