LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Implement a custom LangChain LLM wrapper for a local model.


Reach for a custom class lastOllama running? Use ChatOllamaOpenAI-compatible server? ChatOpenAI with base_urlHugging Face model in Python? ChatHuggingFaceCustom runtime? SubclassBaseChatModel, write _generate
Moving the model behind vLLM and deleting the wrapper raised throughput from 2 to about 30 concurrent requests on the same GPUs.

What you need to know

A "wrapper" makes your model look like any other LangChain model, so the rest of the code does not care where it runs.

Try the ready-made options first

Your set-upUse
Ollama on a laptop or serverChatOllama from langchain-ollama
vLLM, TGI, LM Studio, llama.cpp server (OpenAI-compatible)ChatOpenAI(base_url="http://host:8000/v1", api_key="EMPTY")
A Hugging Face model loaded in PythonHuggingFacePipeline with ChatHuggingFace from langchain-huggingface
Anything else (custom C++ runtime, internal RPC)Write your own class

Writing your own chat model

Subclass BaseChatModel (message list in, AIMessage out). The older LLM base class (string in, string out) still exists, but chat models are what chains, tool calling and agents expect.

Python
from typing import Anyfrom langchain_core.language_models import BaseChatModelfrom langchain_core.messages import AIMessage, BaseMessagefrom langchain_core.outputs import ChatGeneration, ChatResultclass LocalChatModel(BaseChatModel):    generate_fn: Any                  # e.g. your runtime's generate()    max_new_tokens: int = 256    @property    def _llm_type(self) -> str:        return "local-chat"    @property    def _identifying_params(self) -> dict:        return {"max_new_tokens": self.max_new_tokens}    def _generate(self, messages: list[BaseMessage], stop=None,                  run_manager=None, **kwargs) -> ChatResult:        prompt = "\n".join(f"{m.type}: {m.text}" for m in messages) + "\nai:"        text = self.generate_fn(prompt, max_new_tokens=self.max_new_tokens)        for s in stop or []:                  # the base class won't do this            text = text.split(s)[0]        return ChatResult(generations=[ChatGeneration(message=AIMessage(content=text))])
  • _generate is the only required method. It gets messages, formats them into the prompt your model expects, runs the model and wraps the reply.
  • _llm_type names the model in logs and traces; _identifying_params adds its settings.
  • The message-to-prompt line here is simplified. A real model has a chat template (special tokens around each role); use the tokenizer's apply_chat_template so the prompt matches how the model was trained.

Optional methods

  • _stream: yield ChatGenerationChunk objects so .stream() sends tokens as they are made. Without it, .stream() returns the whole reply as one chunk.
  • _agenerate: a true async version; the default runs _generate in a thread.
  • bind_tools: only if your model supports tool calling and you want to use it in agents.

A real-life example

A hospital network in Kerala cannot send patient notes outside its data centre, so it runs an 8-billion-parameter model on its own GPUs. The first version called the Hugging Face pipeline inside the web process through a hand-written wrapper, and handled 2 requests at a time before memory ran out. The team moved the model behind vLLM, which batches requests on the GPU, and replaced the wrapper with ChatOpenAI(base_url=...). Throughput rose to about 30 concurrent requests on the same hardware, and they deleted 150 lines of custom code. They kept a custom BaseChatModel only for a small de-identification model that runs on an internal RPC service.

Follow-up questions to expect

  • "Why subclass BaseChatModel and not LLM?" — Chat models handle system messages, tool calls and structured output, and agents require them; LLM is the legacy string interface.
  • "How would you add streaming?" — Implement _stream to yield ChatGenerationChunk(message=AIMessageChunk(content=token)) and call run_manager.on_llm_new_token(token) so callbacks see each token.
  • "How do you test it?" — The langchain-tests package has standard unit and integration test classes that check a chat model follows the interface.