LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you configure an LLM in LangChain for text generation?


The same model, three jobs, three settings040030 s0.260030 s0.86015 stemperaturemax_tokenstimeoutInvoice extractionHelp-centre answerPush-notification copymax_tokens sits just above the real need: about 150 tokens of JSON for an invoice.
Settings belong to the task, not the model — extraction wants stable output, copywriting wants variety and a hard length cap.

What you need to know

There are two ways to create a chat model, and they return the same kind of object.

Python
from langchain.chat_models import init_chat_modelfrom langchain_openai import ChatOpenAI# Provider-neutral: the provider comes from the stringllm = init_chat_model(    "openai:gpt-5.4-mini",    temperature=0,     # steady output for extraction/classification    max_tokens=500,    # caps reply length and cost    timeout=30,        # seconds; never wait forever    max_retries=2,     # retries 429s, 5xx and network errors with backoff)# Provider-specific: exposes every provider optionllm2 = ChatOpenAI(model="gpt-5.4-mini", temperature=0, max_tokens=500,                  timeout=30, max_retries=2)msg = llm.invoke("Summarise this release note in two sentences: ...")print(msg.text)                               # the reply as a stringprint(msg.usage_metadata)                     # input/output/total tokensprint(msg.response_metadata.get("finish_reason"))

init_chat_model needs the provider package installed (langchain-openai here). Use the provider class when you need a setting only that provider has, such as reasoning_effort on OpenAI reasoning models.

The settings that matter

SettingWhat it controlsTypical value
modelWhich model; drives quality, speed and priceread from config
temperatureRandomness of word choice0 to 0.2 for facts, 0.7 or more for creative text
max_tokensLongest reply allowedjust above your real need
timeoutSeconds before giving up20 to 60
max_retriesAutomatic retries on 429, 5xx and network errors2 to 6

Reading the reply

invoke returns an AIMessage. In LangChain 1.x, .text gives the reply as one string even when the provider returns a list of content blocks (text plus reasoning or citations); .content is the raw field. finish_reason (OpenAI's name; Anthropic calls it stop_reason) tells you whether the model stopped naturally or hit max_tokens. Use llm.stream(...) to get the reply chunk by chunk for a chat UI.

A real-life example

An invoice extraction service at a Mumbai accounting firm used the provider's default settings. Two problems showed up in the first week: the same invoice gave a different total format on re-runs, and one slow provider response held a worker for 4 minutes. The fix was three arguments. temperature=0 made output stable, max_tokens=400 stopped the occasional rambling reply (the JSON they need is about 150 tokens), and timeout=30 with max_retries=2 meant a stuck call failed in under 2 minutes in the worst case and went to a retry queue. Logging usage_metadata per invoice showed an average cost they could quote to finance.

Follow-up questions to expect

  • "Where should the API key go?" — In an environment variable such as OPENAI_API_KEY, which the class reads automatically; pass api_key= only from a secret manager, never a literal.
  • "How do you know a reply was cut off?" — Check response_metadata: OpenAI sets finish_reason to length, and Anthropic sets stop_reason to max_tokens, when the limit was hit.
  • "Does temperature 0 make output fully deterministic?" — No. It makes it much more stable, but providers do not guarantee identical output, so tests should check shape and meaning, not exact strings.