Course Content
LangGraph Agents
7 sections · 49 lessons
How do you implement consensus (voting, self-consistency, referee agent) in a graph?
What you need to know
Python
1import operator2from collections import Counter3from typing import Annotated, TypedDict4from langgraph.types import Send56class State(TypedDict):7 claim_text: str8 votes: Annotated[list[str], operator.add]9 label: str10 agreement: float1112def fan_out(state: State) -> list[Send]:13 return [Send("classify", {"claim_text": state["claim_text"], "run": i}) for i in range(5)]1415def classify(inp: dict) -> dict:16 label = classifier.invoke(inp["claim_text"]).label # e.g. "genuine" or "suspicious"17 return {"votes": [label.strip().lower()]}1819def vote(state: State) -> dict:20 label, count = Counter(state["votes"]).most_common(1)[0]21 return {"label": label, "agreement": count / len(state["votes"])}Wire it as START → (fan_out) → classify → vote, with fan_out as a conditional edge. votes needs the appending reducer.
Choosing a pattern
| Pattern | Works for | Cost | Weakness |
|---|---|---|---|
| Self-consistency | Labels, numbers, choices | N calls | Shared model bias |
| Diverse voting | Same, higher stakes | N calls, several models | More setup |
| Referee / judge | Free text, plans | N + 1 calls | Judge has its own bias |
| Verification | Code, SQL, facts with sources | Cheap checks | Only for checkable outputs |
Using agreement
The agreement rate is a confidence signal. Act automatically at 5 of 5 or 4 of 5; send 3 of 5 to a human.
A real-life example
A health insurer flags claims as genuine or suspicious. A single classifier call wrongly flagged 6% of genuine claims, each costing a two-day manual review. With five parallel votes: 5-of-5 and 4-of-5 results (91% of claims) are acted on automatically; the rest go to reviewers. Wrong flags on automatic decisions fell to 2%, at five times the model cost — about Rs 0.60 more per claim, far below the Rs 450 cost of one unnecessary manual review.
Follow-up questions to expect
- "Why not always use a judge?" — It adds a call and its own bias, such as preferring longer answers. For exact outputs, counting is cheaper and more reliable.
- "How do you make samples differ?" — Sampling variation from the model, different prompts, or different models; with temperature zero on the same prompt, the votes are copies.
- "Where does consensus not help?" — When all candidates lack the key information; five guesses without evidence are still guesses.