LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

How do you handle long context inputs efficiently in production?


KV cache for one request, by context length0.25 GB0.63 GB1 GB2.5 GB4 GB10 GB12.2 GB30.5 GB8B, 128 KB/token70B, 320 KB/token2k tokens8k tokens32k tokens100k tokensLlama 3.1 sizes in bf16; FP8 KV cache halves every cell.
One 100k-token prompt on the 8B model takes the memory of about fifty normal chats, so long inputs need caps and their own pool.

What you need to know

Why long context is expensive

  • Cost — every input token is billed on every call. A 100,000-token prompt at an illustrative $3 per million costs $0.30 per call before any output.
  • Latency — prefill work grows with prompt length, and attention work grows faster than linearly. A 100,000-token prompt can take several seconds before the first token even on a fast GPU.
  • Quality — models often use information at the start and end of a long prompt better than information in the middle.

Memory is the one people forget. The KV cache grows with every token in the prompt. For an 8B model like Llama 3.1 8B:

Text
KV bytes per token = 2 (K and V) × layers × kv_heads × head_dim × bytes                   = 2 × 32 × 8 × 128 × 2 = 131,072 bytes ≈ 128 KB100,000 tokens × 128 KB ≈ 12.5 GB  for ONE request70B model (80 layers): 320 KB per token → about 31 GB for one request

One 100k-token request on the 8B model uses the KV-cache space of about 50 normal 2,000-token chats.

Techniques

  • Retrieve and rerank. Fetch 50 candidates, rerank with a cross-encoder, send the best 3–8.
  • Cache the static part. Put the long, unchanging content first so provider prompt caching or vLLM's prefix caching reuses it; repeated questions about the same document then pay a fraction of the price and get a much faster TTFT.
  • Chunked prefill. vLLM splits a long prefill into pieces and mixes them with other users' decode steps, so their tokens keep flowing.
  • Bound the input. Set a maximum length per route (--max-model-len on vLLM) and reject or summarise larger inputs.
  • Compress history and tool output. Summarise old turns; trim large tool JSON to the fields you need.
  • Map-reduce for whole-document tasks: summarise sections separately, then combine.
  • Separate pool for long-context work so it cannot hurt interactive p95.

A real-life example

A hospital's on-prem summariser receives complete patient records for long stays — up to 300 pages, about 150,000 tokens. The server has an 8B model on a 48 GB GPU with about 24 GB of KV-cache space.

One such record needs about 18 GB of KV cache on its own. When two arrived at once, the other 20 users' short requests were preempted and their latency rose from 4 s to 40 s. The team changes the design:

  • Records are split by section (admission, medication chart, lab results, notes). Each section of 5,000–10,000 tokens is summarised separately, and a final call combines the section summaries.
  • Long jobs run on a separate queue, served by the second GPU.
  • The medication chart, which must be exact, is extracted with a structured-output call rather than summarised.

Peak KV use per request falls from 18 GB to about 1.3 GB, the interactive p95 returns to 4 s, and an eval on 50 long records shows fewer missed medicines than the single-pass version.

Follow-up questions to expect

  • "Why not just use a model with a one-million-token window?" — You can, but you pay for every token on every call, TTFT is high, and recall of mid-document details is not guaranteed. Retrieval is usually cheaper and more accurate.
  • "How does prefix caching help long documents?" — When many questions are asked about the same document, its KV cache is computed once and reused, cutting both cost and TTFT.
  • "What shrinks KV-cache size?" — Models with grouped-query attention (fewer KV heads), FP8 KV cache, and shorter prompts.