Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 10: Offline Enterprise Deployment
Scenario: an enterprise client requires the LangChain application to run entirely inside their network, with no outbound internet. How do you deploy it?
What you need to know
"No outbound internet" breaks more than the model call. Libraries download models at startup, tracing sends data to a cloud service, and pip install reaches the public package index. Each of those must be replaced or switched off.
The component swap
| Component | Cloud version | Offline version |
|---|---|---|
| Chat model | Hosted API | Open-weights model (Llama, Qwen or Mistral family) on vLLM |
| Embeddings | Hosted embeddings API | BGE or E5 on CPU or a shared GPU |
| Vector store | Managed service | pgvector or self-hosted Qdrant |
| Reranker | Hosted rerank API | Local cross-encoder |
| Tracing | LangSmith cloud | Self-hosted LangSmith, OpenTelemetry to their backend, or off |
1from langchain_openai import ChatOpenAI2from langchain_huggingface import HuggingFaceEmbeddings34llm = ChatOpenAI(base_url="http://vllm.internal:8000/v1", api_key="not-used",5 model="Qwen/Qwen3-32B", temperature=0.2)6embeddings = HuggingFaceEmbeddings(model_name="/models/bge-m3") # local path, no downloadvLLM exposes an OpenAI-compatible API, so the same ChatOpenAI class works in both environments; only the base URL and model name change. The embedding model loads from a local directory, not from the Hugging Face Hub.
Close the outbound paths
- Tracing —
LANGSMITH_TRACING=false, or point it at a self-hosted instance. - Model downloads — set
HF_HUB_OFFLINE=1so libraries never try to reach the Hub; load weights from local paths. - Packages — install from an internal PyPI mirror with pinned versions and hashes.
- Artefacts — ship model weights as versioned artefacts with checksums in the client's registry.
- Verify — run the whole stack in a network-isolated test environment; any attempted outbound call should fail loudly in testing, not in production.
Honest trade-offs to state up front
- Quality. A self-hosted 8B to 70B model is usually weaker than a frontier API on hard reasoning. Measure the gap rather than argue about it.
- Capacity. There is no autoscaling. Capacity is fixed by the GPUs installed, so you need admission control, queuing and load shedding.
- Operations. Model updates, security patches and on-call are now the client's or your job, with no vendor to escalate to.
Prove the delta
Run the same eval suite against the cloud stack and the offline stack and report both scores. The client then accepts a known difference, for example 91% versus 85% on their test set, rather than discovering it after go-live.
A real-life example
Scenario (illustrative numbers). A defence-sector manufacturer wants the document assistant a vendor built on a cloud API, but inside an air-gapped network. The vendor's code uses LangChain throughout.
The swap takes three weeks: a 32B open model at FP8 on one 80 GB GPU with vLLM, BGE-M3 embeddings on CPU, pgvector in the client's existing Postgres, and a local cross-encoder. The first test run in an isolated network fails at startup because a tokenizer tried to download from the Hub; HF_HUB_OFFLINE=1 and local paths fix it. On the client's 200-question eval, the offline stack scores 84% against 90% on the cloud stack. The client accepts it in writing, and the gap narrows to 3 points after prompts are tuned for the new model.
Follow-up questions to expect
- "How do you update the model later?" — Deliver a new checksummed artefact, run the eval suite inside their network, canary it, and keep the old version for rollback.
- "Can you still use LangSmith?" — Only self-hosted inside their network, or not at all; OpenTelemetry to their own tracing system is a common alternative.
- "What if they have no GPUs?" — Smaller quantized models on CPU are possible for light use, but latency and quality drop sharply; size the hardware from the workload.