LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Write a function to implement a self-querying retriever in LangChain?


What you need to know

Users mix meaning and hard constraints in one sentence. Vector search handles the meaning ("good camera") but is bad at numbers and categories ("under 20,000", "phones only"). A self-querying retriever splits the two.

  1. Describe the fields — name, type and meaning of each metadata field the LLM may filter on.
  2. LLM writes a structured query — a search string plus a filter such as and(eq("category", "phone"), lt("price", 20000)).
  3. Translate — a store-specific translator converts that into Chroma, Pinecone, Qdrant or pgvector filter syntax.
  4. Search — the vector store runs a filtered similarity search.

The classic helper

Python
from langchain_classic.retrievers.self_query.base import SelfQueryRetrieverfrom langchain_classic.chains.query_constructor.schema import AttributeInfodef build_self_query(llm, vector_store):    fields = [        AttributeInfo(name="category", description="Product type: phone, laptop or tv", type="string"),        AttributeInfo(name="price", description="Price in Indian rupees", type="integer"),        AttributeInfo(name="rating", description="Average rating from 1 to 5", type="float"),    ]    return SelfQueryRetriever.from_llm(        llm, vector_store,        document_contents="Short product descriptions from an electronics store",        metadata_field_info=fields,        enable_limit=True,          # lets "top 3 ..." set k    )

It needs pip install langchain-classic lark (lark parses the generated query), and a vector store that has a built-in translator (Chroma, Pinecone, Qdrant, pgvector, Elasticsearch and others). In LangChain 1.x these imports come from langchain_classic.

The modern do-it-yourself version

With reliable structured output, many teams skip the helper and ask the model for a Pydantic object:

Python
from typing import Literalfrom pydantic import BaseModel, Fieldclass SearchPlan(BaseModel):    query: str = Field(description="What to search for, without the filter words")    category: Literal["phone", "laptop", "tv"] | None = None    max_price: int | None = Field(None, description="Upper price limit in rupees")planner = llm.with_structured_output(SearchPlan)plan = planner.invoke("phones under 20000 with a good camera")docs = vector_store.similarity_search(    plan.query, k=5,    filter={"$and": [{"category": plan.category}, {"price": {"$lte": plan.max_price}}]})

You control the allowed values (Literal stops invented categories), and you build the filter yourself for your store. In real code, only add the conditions whose values are not None.

Risks

  • Extra latency and cost — one LLM call before every search.
  • Bad filters — the model may filter on a field that doesn't exist or pick an impossible value, returning zero results. Fall back to plain similarity search.
  • Security — never let a self-query filter decide access; add the tenant filter separately in code.

A real-life example

An electronics e-commerce site's chat assistant got "Samsung TV under 40k, 4.5 stars or more". Pure vector search returned a ₹95,000 TV with a matching description. With self-querying, the LLM produced search text "Samsung TV" and the filter price lt 40000 and rating gte 4.5; all 5 results fit the budget.

Monitoring then showed 4% of queries returning zero results, mostly when users wrote "under 40k" and the model produced price lt 40. The team added "prices are in full rupees, e.g. 40k = 40000" to the field description and a fallback: if zero results, drop the filter and search again. Zero-result queries fell below 1%.

Follow-up questions to expect

  • "How is this different from metadata filtering?" — Metadata filtering uses a filter your code writes; self-querying lets the LLM write it from the question.
  • "Which fields should you expose?" — Only fields users actually talk about, with precise descriptions and allowed values; each extra field is a chance to misfilter.
  • "Can you cache it?" — Yes, cache the question-to-plan step for repeated queries.