LangGraph Agents

Course Content

LangGraph Agents

7 sections · 49 lessons

How do you implement consensus (voting, self-consistency, referee agent) in a graph?


What you need to know

Python
import operatorfrom collections import Counterfrom typing import Annotated, TypedDictfrom langgraph.types import Sendclass State(TypedDict):    claim_text: str    votes: Annotated[list[str], operator.add]    label: str    agreement: floatdef fan_out(state: State) -> list[Send]:    return [Send("classify", {"claim_text": state["claim_text"], "run": i}) for i in range(5)]def classify(inp: dict) -> dict:    label = classifier.invoke(inp["claim_text"]).label     # e.g. "genuine" or "suspicious"    return {"votes": [label.strip().lower()]}def vote(state: State) -> dict:    label, count = Counter(state["votes"]).most_common(1)[0]    return {"label": label, "agreement": count / len(state["votes"])}

Wire it as START → (fan_out) → classify → vote, with fan_out as a conditional edge. votes needs the appending reducer.

Choosing a pattern

PatternWorks forCostWeakness
Self-consistencyLabels, numbers, choicesN callsShared model bias
Diverse votingSame, higher stakesN calls, several modelsMore setup
Referee / judgeFree text, plansN + 1 callsJudge has its own bias
VerificationCode, SQL, facts with sourcesCheap checksOnly for checkable outputs

Using agreement

The agreement rate is a confidence signal. Act automatically at 5 of 5 or 4 of 5; send 3 of 5 to a human.

A real-life example

A health insurer flags claims as genuine or suspicious. A single classifier call wrongly flagged 6% of genuine claims, each costing a two-day manual review. With five parallel votes: 5-of-5 and 4-of-5 results (91% of claims) are acted on automatically; the rest go to reviewers. Wrong flags on automatic decisions fell to 2%, at five times the model cost — about Rs 0.60 more per claim, far below the Rs 450 cost of one unnecessary manual review.

Follow-up questions to expect

  • "Why not always use a judge?" — It adds a call and its own bias, such as preferring longer answers. For exact outputs, counting is cheaper and more reliable.
  • "How do you make samples differ?" — Sampling variation from the model, different prompts, or different models; with temperature zero on the same prompt, the votes are copies.
  • "Where does consensus not help?" — When all candidates lack the key information; five guesses without evidence are still guesses.