Course Content
LangChain Mastery
7 sections · 109 lessons
Write a function to generate text completions with LangChain.
What you need to know
"Completion" used to mean the old text-in, text-out API. Today you send messages to a chat model and get an AIMessage back, even for a single prompt.
1import logging2from langchain.chat_models import init_chat_model34log = logging.getLogger(__name__)5_llm = init_chat_model("openai:gpt-5.4-mini", temperature=0.7,6 max_tokens=300, timeout=30, max_retries=2)78def complete(prompt: str, system: str = "You are a concise writer.") -> str:9 if not prompt.strip():10 raise ValueError("empty prompt")11 msg = _llm.invoke([("system", system), ("human", prompt)])12 if msg.response_metadata.get("finish_reason") == "length":13 log.warning("reply cut off at max_tokens")14 log.info("tokens: %s", msg.usage_metadata)15 return msg.textThe model is created once at module level, not on every call, which saves set-up time. A list of (role, text) tuples is accepted as messages. .text returns the reply as one string.
Streaming and batching
1def complete_stream(prompt: str):2 for chunk in _llm.stream(prompt):3 yield chunk.text # send each piece to the UI45answers = _llm.batch(prompts, config={"max_concurrency": 8})- Stream when a person is watching. Total time is the same, but the first words appear in under a second instead of after the whole reply.
- Batch when you have many independent prompts. It runs them in parallel and keeps the input order.
- Async (
ainvoke,astream,abatch) inside FastAPI or any async server, so a waiting call does not block a thread.
Edge cases to mention
- Cut-off reply:
finish_reason == "length"meansmax_tokenswas reached. - Empty or refused reply:
.textcan be an empty string; decide what the caller should see. - Different providers name things differently: Anthropic uses
stop_reasoninresponse_metadata, so read both if you switch.
A real-life example
A food-delivery app writes short push-notification copy ("Your biryani is 5 minutes away!") with a completion function. At max_tokens=300 a few replies rambled past the 60-character limit of the notification. The team lowered max_tokens to 40, added a length check that falls back to a fixed template if the text is too long, and used batch with max_concurrency=10 to write copy for 2,000 restaurants at once. Nightly generation went from 35 minutes in a loop to under 4 minutes.
Follow-up questions to expect
- "Why not create the model inside the function?" — It works, but rebuilds the client on every call; create it once and reuse it, since it is safe to share.
- "How do you make the output repeatable for tests?" — Temperature 0 helps, but in tests you swap in a fake model such as
GenericFakeChatModelso results are exact and free. - "How do you stop the user seeing a half reply if streaming fails?" — Catch the error in the stream loop and send an end-of-stream marker with a short apology the UI can show.