LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Write a function to generate text completions with LangChain.


What you need to know

"Completion" used to mean the old text-in, text-out API. Today you send messages to a chat model and get an AIMessage back, even for a single prompt.

Python
import loggingfrom langchain.chat_models import init_chat_modellog = logging.getLogger(__name__)_llm = init_chat_model("openai:gpt-5.4-mini", temperature=0.7,                       max_tokens=300, timeout=30, max_retries=2)def complete(prompt: str, system: str = "You are a concise writer.") -> str:    if not prompt.strip():        raise ValueError("empty prompt")    msg = _llm.invoke([("system", system), ("human", prompt)])    if msg.response_metadata.get("finish_reason") == "length":        log.warning("reply cut off at max_tokens")    log.info("tokens: %s", msg.usage_metadata)    return msg.text

The model is created once at module level, not on every call, which saves set-up time. A list of (role, text) tuples is accepted as messages. .text returns the reply as one string.

Streaming and batching

Python
def complete_stream(prompt: str):    for chunk in _llm.stream(prompt):        yield chunk.text                  # send each piece to the UIanswers = _llm.batch(prompts, config={"max_concurrency": 8})
  • Stream when a person is watching. Total time is the same, but the first words appear in under a second instead of after the whole reply.
  • Batch when you have many independent prompts. It runs them in parallel and keeps the input order.
  • Async (ainvoke, astream, abatch) inside FastAPI or any async server, so a waiting call does not block a thread.

Edge cases to mention

  • Cut-off reply: finish_reason == "length" means max_tokens was reached.
  • Empty or refused reply: .text can be an empty string; decide what the caller should see.
  • Different providers name things differently: Anthropic uses stop_reason in response_metadata, so read both if you switch.

A real-life example

A food-delivery app writes short push-notification copy ("Your biryani is 5 minutes away!") with a completion function. At max_tokens=300 a few replies rambled past the 60-character limit of the notification. The team lowered max_tokens to 40, added a length check that falls back to a fixed template if the text is too long, and used batch with max_concurrency=10 to write copy for 2,000 restaurants at once. Nightly generation went from 35 minutes in a loop to under 4 minutes.

Follow-up questions to expect

  • "Why not create the model inside the function?" — It works, but rebuilds the client on every call; create it once and reuse it, since it is safe to share.
  • "How do you make the output repeatable for tests?" — Temperature 0 helps, but in tests you swap in a fake model such as GenericFakeChatModel so results are exact and free.
  • "How do you stop the user seeing a half reply if streaming fails?" — Catch the error in the stream loop and send an end-of-stream marker with a short apology the UI can show.