Course Content
Advanced RAG
3 sections · 38 lessons
What is ColBERT, and how does late interaction improve retrieval?
What you need to know
Three ways to score a query against a document
| Model type | How it works | Speed | Accuracy |
|---|---|---|---|
| Bi-encoder | One vector per query, one per passage, one cosine | Fastest; index offline | Lossy: the passage is averaged |
| Cross-encoder | Query and passage go through the model together | Slow; nothing precomputed | Highest; sees every word pair |
| Late interaction (ColBERT) | One vector per token each side, MaxSim at query time | Near bi-encoder; index offline | Close to cross-encoder |
A bi-encoder squeezes a 300-token passage into one vector. If the passage covers pricing, refunds and SLAs, its vector is an average that matches none of the three sharply. A cross-encoder avoids this but must run the model on every (query, passage) pair at query time, so it only works on a shortlist.
MaxSim, in code
1import numpy as np23def normalise(m):4 return m / np.linalg.norm(m, axis=-1, keepdims=True)56def maxsim(q_tokens, d_tokens):7 sims = normalise(q_tokens) @ normalise(d_tokens).T # (q_len, d_len)8 return sims.max(axis=1).sum() # best match per query token910rng = np.random.default_rng(0)11q = rng.normal(size=(4, 128)) # 4 query-token vectors12relevant = np.vstack([q + 0.3 * rng.normal(size=q.shape), # 4 matching tokens...13 rng.normal(size=(60, 128))]) # ...hidden in 60 others14unrelated = rng.normal(size=(64, 128))1516print(round(float(maxsim(q, relevant)), 2)) # 3.86 (maximum possible is 4)17print(round(float(maxsim(q, unrelated)), 2)) # 0.7418mean = lambda m: normalise(m.mean(0))19print(round(float(mean(q) @ mean(relevant)), 2)) # 0.13 with one averaged vectorThe relevant document has only 4 matching tokens among 64. MaxSim finds each one and scores 3.86 out of 4. Mean-pooling the same document into one vector gives a cosine of only 0.13 — the signal is washed out by the other 60 tokens. That is the whole argument for late interaction in ten lines.
The cost: index size
Take a 200-token chunk. A single-vector index stores 1 × 1,024 numbers. ColBERT stores 200 × 128 = 25,600 numbers — about 25 times more before compression. ColBERTv2 reduces this with residual compression (store each token as a nearby centroid ID plus a small, low-bit correction), and the PLAID engine speeds up search over it. It is still larger and more complex to serve than a plain vector index.
How it is used in 2026
- As a reranker over 100–200 candidates from a cheap first stage. This is the most common and pragmatic use.
- As a first-stage index in engines with native multi-vector support. Vespa and Qdrant, for example, can store several vectors per point and score with MaxSim.
- For page images — ColPali applies the same idea to image patches, so scanned pages can be searched without OCR.
A real-life example
A pharma company builds regulatory-document search over guidelines, internal SOPs and past submission letters. Scientists type queries like:
ICH Q3D oral PDE for cadmiumThe single-vector retriever returns general pages about "elemental impurities" and "heavy metals", because the passage's vector is dominated by the surrounding prose. The table row mentioning cadmium's permitted daily exposure ranks 30th.
The team keeps its existing hybrid search as the first stage (top 150) and adds a ColBERT-style reranker. Now the query tokens Q3D, cadmium and oral each find their exact partners in the table row, and it moves into the top 3. They check it on a set of 250 queries written by regulatory-affairs staff, comparing the rank of the correct passage before and after, and see the biggest gains on queries containing codes and substance names.
Follow-up questions to expect
- "Why not always use a cross-encoder instead?" — It cannot precompute document representations, so its cost grows with every candidate at query time. Late interaction precomputes the document side and only does cheap vector maths at query time.
- "What is PLAID?" — ColBERTv2's search engine: it first uses the compressed centroids to prune candidates quickly, then does full MaxSim only on the survivors.
- "Does ColBERT replace BM25?" — Not fully. It is much better on exact terms than a single vector, but BM25 is still cheap, transparent and strong on IDs. Many systems run both.