Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

What is ColBERT, and how does late interaction improve retrieval?


MaxSim: each query token keeps only its best match0.410.970.220.180.300.350.200.930.150.440.120.190.210.880.26elementalQ3DCdorallimitQ3DcadmiumoralScore = 0.97 + 0.93 + 0.88 = 2.78; the other cells are ignored.
Rare terms like an ICH code survive because they are matched token to token instead of being averaged into one vector.

What you need to know

Three ways to score a query against a document

Model typeHow it worksSpeedAccuracy
Bi-encoderOne vector per query, one per passage, one cosineFastest; index offlineLossy: the passage is averaged
Cross-encoderQuery and passage go through the model togetherSlow; nothing precomputedHighest; sees every word pair
Late interaction (ColBERT)One vector per token each side, MaxSim at query timeNear bi-encoder; index offlineClose to cross-encoder

A bi-encoder squeezes a 300-token passage into one vector. If the passage covers pricing, refunds and SLAs, its vector is an average that matches none of the three sharply. A cross-encoder avoids this but must run the model on every (query, passage) pair at query time, so it only works on a shortlist.

MaxSim, in code

Python
import numpy as npdef normalise(m):    return m / np.linalg.norm(m, axis=-1, keepdims=True)def maxsim(q_tokens, d_tokens):    sims = normalise(q_tokens) @ normalise(d_tokens).T   # (q_len, d_len)    return sims.max(axis=1).sum()                         # best match per query tokenrng = np.random.default_rng(0)q = rng.normal(size=(4, 128))                             # 4 query-token vectorsrelevant = np.vstack([q + 0.3 * rng.normal(size=q.shape), # 4 matching tokens...                      rng.normal(size=(60, 128))])        # ...hidden in 60 othersunrelated = rng.normal(size=(64, 128))print(round(float(maxsim(q, relevant)), 2))    # 3.86 (maximum possible is 4)print(round(float(maxsim(q, unrelated)), 2))   # 0.74mean = lambda m: normalise(m.mean(0))print(round(float(mean(q) @ mean(relevant)), 2))   # 0.13 with one averaged vector

The relevant document has only 4 matching tokens among 64. MaxSim finds each one and scores 3.86 out of 4. Mean-pooling the same document into one vector gives a cosine of only 0.13 — the signal is washed out by the other 60 tokens. That is the whole argument for late interaction in ten lines.

The cost: index size

Take a 200-token chunk. A single-vector index stores 1 × 1,024 numbers. ColBERT stores 200 × 128 = 25,600 numbers — about 25 times more before compression. ColBERTv2 reduces this with residual compression (store each token as a nearby centroid ID plus a small, low-bit correction), and the PLAID engine speeds up search over it. It is still larger and more complex to serve than a plain vector index.

How it is used in 2026

  • As a reranker over 100–200 candidates from a cheap first stage. This is the most common and pragmatic use.
  • As a first-stage index in engines with native multi-vector support. Vespa and Qdrant, for example, can store several vectors per point and score with MaxSim.
  • For page images — ColPali applies the same idea to image patches, so scanned pages can be searched without OCR.

A real-life example

A pharma company builds regulatory-document search over guidelines, internal SOPs and past submission letters. Scientists type queries like:

Text
ICH Q3D oral PDE for cadmium

The single-vector retriever returns general pages about "elemental impurities" and "heavy metals", because the passage's vector is dominated by the surrounding prose. The table row mentioning cadmium's permitted daily exposure ranks 30th.

The team keeps its existing hybrid search as the first stage (top 150) and adds a ColBERT-style reranker. Now the query tokens Q3D, cadmium and oral each find their exact partners in the table row, and it moves into the top 3. They check it on a set of 250 queries written by regulatory-affairs staff, comparing the rank of the correct passage before and after, and see the biggest gains on queries containing codes and substance names.

Follow-up questions to expect

  • "Why not always use a cross-encoder instead?" — It cannot precompute document representations, so its cost grows with every candidate at query time. Late interaction precomputes the document side and only does cheap vector maths at query time.
  • "What is PLAID?" — ColBERTv2's search engine: it first uses the compressed centroids to prune candidates quickly, then does full MaxSim only on the survivors.
  • "Does ColBERT replace BM25?" — Not fully. It is much better on exact terms than a single vector, but BM25 is still cheap, transparent and strong on IDs. Many systems run both.