Course Content
LangChain Mastery
7 sections · 109 lessons
Implement a custom LangChain LLM wrapper for a local model.
What you need to know
A "wrapper" makes your model look like any other LangChain model, so the rest of the code does not care where it runs.
Try the ready-made options first
| Your set-up | Use |
|---|---|
| Ollama on a laptop or server | ChatOllama from langchain-ollama |
| vLLM, TGI, LM Studio, llama.cpp server (OpenAI-compatible) | ChatOpenAI(base_url="http://host:8000/v1", api_key="EMPTY") |
| A Hugging Face model loaded in Python | HuggingFacePipeline with ChatHuggingFace from langchain-huggingface |
| Anything else (custom C++ runtime, internal RPC) | Write your own class |
Writing your own chat model
Subclass BaseChatModel (message list in, AIMessage out). The older LLM base class (string in, string out) still exists, but chat models are what chains, tool calling and agents expect.
1from typing import Any2from langchain_core.language_models import BaseChatModel3from langchain_core.messages import AIMessage, BaseMessage4from langchain_core.outputs import ChatGeneration, ChatResult56class LocalChatModel(BaseChatModel):7 generate_fn: Any # e.g. your runtime's generate()8 max_new_tokens: int = 256910 @property11 def _llm_type(self) -> str:12 return "local-chat"1314 @property15 def _identifying_params(self) -> dict:16 return {"max_new_tokens": self.max_new_tokens}1718 def _generate(self, messages: list[BaseMessage], stop=None,19 run_manager=None, **kwargs) -> ChatResult:20 prompt = "\n".join(f"{m.type}: {m.text}" for m in messages) + "\nai:"21 text = self.generate_fn(prompt, max_new_tokens=self.max_new_tokens)22 for s in stop or []: # the base class won't do this23 text = text.split(s)[0]24 return ChatResult(generations=[ChatGeneration(message=AIMessage(content=text))])_generateis the only required method. It gets messages, formats them into the prompt your model expects, runs the model and wraps the reply._llm_typenames the model in logs and traces;_identifying_paramsadds its settings.- The message-to-prompt line here is simplified. A real model has a chat template (special tokens around each role); use the tokenizer's
apply_chat_templateso the prompt matches how the model was trained.
Optional methods
_stream: yieldChatGenerationChunkobjects so.stream()sends tokens as they are made. Without it,.stream()returns the whole reply as one chunk._agenerate: a true async version; the default runs_generatein a thread.bind_tools: only if your model supports tool calling and you want to use it in agents.
A real-life example
A hospital network in Kerala cannot send patient notes outside its data centre, so it runs an 8-billion-parameter model on its own GPUs. The first version called the Hugging Face pipeline inside the web process through a hand-written wrapper, and handled 2 requests at a time before memory ran out. The team moved the model behind vLLM, which batches requests on the GPU, and replaced the wrapper with ChatOpenAI(base_url=...). Throughput rose to about 30 concurrent requests on the same hardware, and they deleted 150 lines of custom code. They kept a custom BaseChatModel only for a small de-identification model that runs on an internal RPC service.
Follow-up questions to expect
- "Why subclass
BaseChatModeland notLLM?" — Chat models handle system messages, tool calls and structured output, and agents require them;LLMis the legacy string interface. - "How would you add streaming?" — Implement
_streamto yieldChatGenerationChunk(message=AIMessageChunk(content=token))and callrun_manager.on_llm_new_token(token)so callbacks see each token. - "How do you test it?" — The
langchain-testspackage has standard unit and integration test classes that check a chat model follows the interface.