Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 10: Offline Enterprise Deployment


Scenario: an enterprise client requires the LangChain application to run entirely inside their network, with no outbound internet. How do you deploy it?

What you need to know

"No outbound internet" breaks more than the model call. Libraries download models at startup, tracing sends data to a cloud service, and pip install reaches the public package index. Each of those must be replaced or switched off.

The component swap

ComponentCloud versionOffline version
Chat modelHosted APIOpen-weights model (Llama, Qwen or Mistral family) on vLLM
EmbeddingsHosted embeddings APIBGE or E5 on CPU or a shared GPU
Vector storeManaged servicepgvector or self-hosted Qdrant
RerankerHosted rerank APILocal cross-encoder
TracingLangSmith cloudSelf-hosted LangSmith, OpenTelemetry to their backend, or off
Python
from langchain_openai import ChatOpenAIfrom langchain_huggingface import HuggingFaceEmbeddingsllm = ChatOpenAI(base_url="http://vllm.internal:8000/v1", api_key="not-used",                 model="Qwen/Qwen3-32B", temperature=0.2)embeddings = HuggingFaceEmbeddings(model_name="/models/bge-m3")     # local path, no download

vLLM exposes an OpenAI-compatible API, so the same ChatOpenAI class works in both environments; only the base URL and model name change. The embedding model loads from a local directory, not from the Hugging Face Hub.

Close the outbound paths

  1. Tracing — LANGSMITH_TRACING=false, or point it at a self-hosted instance.
  2. Model downloads — set HF_HUB_OFFLINE=1 so libraries never try to reach the Hub; load weights from local paths.
  3. Packages — install from an internal PyPI mirror with pinned versions and hashes.
  4. Artefacts — ship model weights as versioned artefacts with checksums in the client's registry.
  5. Verify — run the whole stack in a network-isolated test environment; any attempted outbound call should fail loudly in testing, not in production.

Honest trade-offs to state up front

  • Quality. A self-hosted 8B to 70B model is usually weaker than a frontier API on hard reasoning. Measure the gap rather than argue about it.
  • Capacity. There is no autoscaling. Capacity is fixed by the GPUs installed, so you need admission control, queuing and load shedding.
  • Operations. Model updates, security patches and on-call are now the client's or your job, with no vendor to escalate to.

Prove the delta

Run the same eval suite against the cloud stack and the offline stack and report both scores. The client then accepts a known difference, for example 91% versus 85% on their test set, rather than discovering it after go-live.

A real-life example

Scenario (illustrative numbers). A defence-sector manufacturer wants the document assistant a vendor built on a cloud API, but inside an air-gapped network. The vendor's code uses LangChain throughout.

The swap takes three weeks: a 32B open model at FP8 on one 80 GB GPU with vLLM, BGE-M3 embeddings on CPU, pgvector in the client's existing Postgres, and a local cross-encoder. The first test run in an isolated network fails at startup because a tokenizer tried to download from the Hub; HF_HUB_OFFLINE=1 and local paths fix it. On the client's 200-question eval, the offline stack scores 84% against 90% on the cloud stack. The client accepts it in writing, and the gap narrows to 3 points after prompts are tuned for the new model.

Follow-up questions to expect

  • "How do you update the model later?" — Deliver a new checksummed artefact, run the eval suite inside their network, canary it, and keep the old version for rollback.
  • "Can you still use LangSmith?" — Only self-hosted inside their network, or not at all; OpenTelemetry to their own tracing system is a common alternative.
  • "What if they have no GPUs?" — Smaller quantized models on CPU are possible for light use, but latency and quality drop sharply; size the hardware from the workload.