Course Content
LangChain Mastery
7 sections · 109 lessons
Implement a chain to handle batch processing in LangChain.
What you need to know
Python
1def run_batch(chain, rows: list[dict], concurrency: int = 8):2 results = chain.batch(rows, config={"max_concurrency": concurrency},3 return_exceptions=True)4 ok = [(i, r) for i, r in enumerate(results) if not isinstance(r, Exception)]5 bad = [(i, r) for i, r in enumerate(results) if isinstance(r, Exception)]6 return ok, bad78ok, bad = run_batch(invoice_chain, [{"text": t} for t in invoice_texts])- Order is kept: result
ibelongs to inputi, even though they ran in parallel. max_concurrencycaps how many calls run at once. Without it, a large batch can fire hundreds of requests and hit 429 errors.return_exceptions=Trueputs the exception object in that position instead of raising, so the other results are kept.
Variants
| Method | Use it for |
|---|---|
batch | A list of inputs, sync code, results in order |
abatch | The same, in async code |
batch_as_completed | Yields (index, result) as each one finishes |
| Provider batch APIs | Very large offline jobs at lower cost, results within hours |
batch_as_completed lets you save each result the moment it is ready, which matters when the job is long:
Python
for i, result in chain.batch_as_completed(rows, config={"max_concurrency": 8}, return_exceptions=True): save_result(ids[i], result) # write to DB right awayLarge jobs
- Split into chunks of a few hundred rows and record which rows are done.
- Retry the failed ones separately, after checking why they failed.
- Watch both requests-per-minute and tokens-per-minute limits; long inputs hit the token limit first.
A real-life example
An accounts team runs an invoice extraction chain over 12,000 invoices at the end of each month. The first script looped with invoke and took over 6 hours; one malformed PDF on invoice 7,431 crashed it, and the rerun paid for the first 7,430 again. The rewrite used batch_as_completed with max_concurrency=12 in chunks of 500, saved each result to Postgres as it arrived, and skipped rows already saved on restart. The run now takes about 40 minutes, and roughly 60 failures a month go to a retry list with the error message attached.
Follow-up questions to expect
- "Does
batchuse threads or async?" — Syncbatchuses a thread pool;abatchuses asyncio tasks. Both respectmax_concurrency. - "How do you pick the concurrency?" — Start from the rate limit: requests per minute divided by 60, times the average seconds per call, then leave headroom.
- "When would you use the provider's batch API instead?" — When nobody needs the result for hours and cost matters more than speed.