Course Content
LangChain Mastery
7 sections · 109 lessons
How do you configure an LLM in LangChain for text generation?
What you need to know
There are two ways to create a chat model, and they return the same kind of object.
1from langchain.chat_models import init_chat_model2from langchain_openai import ChatOpenAI34# Provider-neutral: the provider comes from the string5llm = init_chat_model(6 "openai:gpt-5.4-mini",7 temperature=0, # steady output for extraction/classification8 max_tokens=500, # caps reply length and cost9 timeout=30, # seconds; never wait forever10 max_retries=2, # retries 429s, 5xx and network errors with backoff11)1213# Provider-specific: exposes every provider option14llm2 = ChatOpenAI(model="gpt-5.4-mini", temperature=0, max_tokens=500,15 timeout=30, max_retries=2)1617msg = llm.invoke("Summarise this release note in two sentences: ...")18print(msg.text) # the reply as a string19print(msg.usage_metadata) # input/output/total tokens20print(msg.response_metadata.get("finish_reason"))init_chat_model needs the provider package installed (langchain-openai here). Use the provider class when you need a setting only that provider has, such as reasoning_effort on OpenAI reasoning models.
The settings that matter
| Setting | What it controls | Typical value |
|---|---|---|
model | Which model; drives quality, speed and price | read from config |
temperature | Randomness of word choice | 0 to 0.2 for facts, 0.7 or more for creative text |
max_tokens | Longest reply allowed | just above your real need |
timeout | Seconds before giving up | 20 to 60 |
max_retries | Automatic retries on 429, 5xx and network errors | 2 to 6 |
Reading the reply
invoke returns an AIMessage. In LangChain 1.x, .text gives the reply as one string even when the provider returns a list of content blocks (text plus reasoning or citations); .content is the raw field. finish_reason (OpenAI's name; Anthropic calls it stop_reason) tells you whether the model stopped naturally or hit max_tokens. Use llm.stream(...) to get the reply chunk by chunk for a chat UI.
A real-life example
An invoice extraction service at a Mumbai accounting firm used the provider's default settings. Two problems showed up in the first week: the same invoice gave a different total format on re-runs, and one slow provider response held a worker for 4 minutes. The fix was three arguments. temperature=0 made output stable, max_tokens=400 stopped the occasional rambling reply (the JSON they need is about 150 tokens), and timeout=30 with max_retries=2 meant a stuck call failed in under 2 minutes in the worst case and went to a retry queue. Logging usage_metadata per invoice showed an average cost they could quote to finance.
Follow-up questions to expect
- "Where should the API key go?" — In an environment variable such as
OPENAI_API_KEY, which the class reads automatically; passapi_key=only from a secret manager, never a literal. - "How do you know a reply was cut off?" — Check
response_metadata: OpenAI setsfinish_reasontolength, and Anthropic setsstop_reasontomax_tokens, when the limit was hit. - "Does temperature 0 make output fully deterministic?" — No. It makes it much more stable, but providers do not guarantee identical output, so tests should check shape and meaning, not exact strings.