LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Implement a chain to handle batch processing in LangChain.


batch with return_exceptions=TrueokokokValueErrorokok012345kept in placeorderpreservedWithout the flag, row 3 raises and rows 0-2 and 4-5 are thrown away.
One malformed invoice should cost one retry, not a rerun of every invoice already paid for.

What you need to know

Python
def run_batch(chain, rows: list[dict], concurrency: int = 8):    results = chain.batch(rows, config={"max_concurrency": concurrency},                          return_exceptions=True)    ok  = [(i, r) for i, r in enumerate(results) if not isinstance(r, Exception)]    bad = [(i, r) for i, r in enumerate(results) if isinstance(r, Exception)]    return ok, badok, bad = run_batch(invoice_chain, [{"text": t} for t in invoice_texts])
  • Order is kept: result i belongs to input i, even though they ran in parallel.
  • max_concurrency caps how many calls run at once. Without it, a large batch can fire hundreds of requests and hit 429 errors.
  • return_exceptions=True puts the exception object in that position instead of raising, so the other results are kept.

Variants

MethodUse it for
batchA list of inputs, sync code, results in order
abatchThe same, in async code
batch_as_completedYields (index, result) as each one finishes
Provider batch APIsVery large offline jobs at lower cost, results within hours

batch_as_completed lets you save each result the moment it is ready, which matters when the job is long:

Python
for i, result in chain.batch_as_completed(rows, config={"max_concurrency": 8},                                          return_exceptions=True):    save_result(ids[i], result)          # write to DB right away

Large jobs

  • Split into chunks of a few hundred rows and record which rows are done.
  • Retry the failed ones separately, after checking why they failed.
  • Watch both requests-per-minute and tokens-per-minute limits; long inputs hit the token limit first.

A real-life example

An accounts team runs an invoice extraction chain over 12,000 invoices at the end of each month. The first script looped with invoke and took over 6 hours; one malformed PDF on invoice 7,431 crashed it, and the rerun paid for the first 7,430 again. The rewrite used batch_as_completed with max_concurrency=12 in chunks of 500, saved each result to Postgres as it arrived, and skipped rows already saved on restart. The run now takes about 40 minutes, and roughly 60 failures a month go to a retry list with the error message attached.

Follow-up questions to expect

  • "Does batch use threads or async?" — Sync batch uses a thread pool; abatch uses asyncio tasks. Both respect max_concurrency.
  • "How do you pick the concurrency?" — Start from the rate limit: requests per minute divided by 60, times the average seconds per call, then leave headroom.
  • "When would you use the provider's batch API instead?" — When nobody needs the result for hours and cost matters more than speed.