Course Content
Embeddings and Semantic Search
3 sections · 5 lessons
Mini Project: Build a Semantic Search App
Here is the demo that gets built a thousand times a year. Twenty sentences in a Python list, model.encode(), a cosine similarity call, a print loop. It takes eleven minutes and it genuinely works — type "how do I keep my laptop cool" and it finds the cooling-pad article that shares no words with the query.
Here is what happens when that demo meets ten thousand real documents. Half of them are longer than the model's 256-token window, so three quarters of each is silently thrown away before it is ever indexed. Encoding takes eleven minutes rather than eleven seconds, and nothing is cached, so it happens again on every restart. Somebody searches for a product code and gets nothing useful. Somebody else asks "did search get worse after we changed the model?" and there is no way to answer, because nobody wrote down what "good" was.
This project is the gap between those two things. You are building a semantic search application over a real corpus, with a measured latency budget, a stated quality metric, and an index you can rebuild from source at any time. The interesting work is not the embedding call. It is everything that stops the embedding call from being useless at scale.
What you are building
A command-line application that indexes a document collection once, persists the index, and then answers queries in under half a second with ranked, scored, filterable results.
| Capability | Requirement | Acceptance test |
|---|---|---|
| Indexing | Load documents from JSON, embed them, persist vectors and metadata | Restart the process; search works without re-embedding |
| Query handling | Normalise, encode with the same model used for indexing | A stored manifest names the model; a mismatch refuses to load |
| Retrieval | Top-50 candidates by cosine similarity | Returns 50 distinct ids for a query with 50+ relevant docs |
| Ranking | Cross-encoder rerank of candidates down to top-k | Reranked order differs from retrieval order on at least one test query |
| Filtering | Metadata constraints applied before ranking | A category filter returns a full k results, not a starved list |
| Interface | Interactive CLI plus one-shot --query mode | Both paths produce identical results for the same query |
| Evaluation | A harness reporting recall@50, MRR and MAP | Numbers regenerate from a checked-in query file |
The latency budget, spent in advance
"Under 500 ms" is not a target until you know where the milliseconds go. For 10,000 documents at 384 dimensions on a modest CPU:
| Stage | Work | Budget |
|---|---|---|
| Normalise query | Regex and a token count | < 1 ms |
| Encode query | One transformer pass, ~22M parameters | 8 ms |
| Vector search | 10,000 × 384 = 3.84M multiply-adds | 2 ms |
| Metadata filter | Dictionary lookups over 50 candidates | < 1 ms |
| Cross-encoder rerank | 50 pairs at ~800 pairs/second | 63 ms |
| Format and print | String work | 1 ms |
| Total | ~75 ms |
You have 425 ms of headroom, which is exactly why the budget is worth writing down: it tells you that the reranker — 84% of your spend — is the only stage worth optimising, and that everything else is noise. It also makes two classic mistakes instantly visible. Constructing the CrossEncoder inside the request handler adds 1,800 ms of model loading per query, blowing the budget 24-fold. Encoding documents one at a time instead of in batches turns a 20-second index build into a 6-minute one.
Write the latency budget before the code. It converts "make it fast" into a list of numbers you can test against, and it tells you which nine tenths of the system are not worth optimising.
Memory is the other resource worth predicting. 10,000 vectors at 384 float32 dimensions is 10,000×384×4=15.36 MB — negligible. The models dominate: the bi-encoder is roughly 90 MB and the cross-encoder another 90 MB, so budget around 250 MB resident. Vectors only become the problem at a hundred times this scale.
Layout
semantic-search-app/├── data/│ ├── documents.json 50+ docs across 3+ categories│ └── eval_queries.json queries with known-relevant ids├── src/│ ├── embeddings.py encoding, caching, token-limit guard│ ├── indexing.py FAISS index + manifest + persistence│ ├── search.py retrieve → filter → rerank│ ├── ranking.py cross-encoder and rank fusion│ ├── evaluate.py recall@k, MRR, MAP│ └── app.py CLI, config, result formatting├── tests/├── config.json├── requirements.txt└── main.pypython -m venv venv && source venv/bin/activate # Windows: venv\Scripts\activatepip install sentence-transformers faiss-cpu rank-bm25 numpypip freeze > requirements.txtThe corpus
Fifty documents minimum, spread across at least three categories, with at least a few documents that are near-duplicates of each other. Ten documents cannot distinguish a good ranker from a bad one — with ten documents every query returns something plausible and every metric reads 1.0.
1{2 "documents": [3 {"id": "1", "title": "Introduction to Machine Learning",4 "content": "Machine learning is a subset of artificial intelligence that ...",5 "category": "AI", "year": 2024, "visibility": "public"},6 {"id": "2", "title": "Deep Learning Fundamentals",7 "content": "Deep learning uses neural networks with multiple layers ...",8 "category": "AI", "year": 2023, "visibility": "public"}9 ]10}src/embeddings.py
This module owns one non-negotiable rule: every vector that ever enters the index comes from here, and this module knows which model made it.
1import hashlib2import numpy as np3from sentence_transformers import SentenceTransformer456class EmbeddingManager:7 def __init__(self, model_name="all-MiniLM-L6-v2"):8 self.model_name = model_name9 self.model = SentenceTransformer(model_name)10 self.dimension = self.model.get_embedding_dimension() # get_sentence_embedding_dimension() before v611 self.max_tokens = self.model.max_seq_length # 256 for MiniLM-L612 self._cache = {}1314 def _too_long(self, text):15 # +2 for the [CLS] and [SEP] markers, which also count against the limit16 return len(self.model.tokenizer.tokenize(text)) + 2 > self.max_tokens1718 def encode(self, texts, batch_size=64, use_cache=True):19 """Returns unit-length float32 vectors, shape (len(texts), dimension)."""20 overlong = [i for i, t in enumerate(texts) if self._too_long(t)]21 if overlong:22 raise ValueError(23 f"{len(overlong)} texts exceed {self.max_tokens} tokens "24 f"(first at index {overlong[0]}). Chunk them before indexing."25 )26 if not use_cache:27 return self._encode(texts, batch_size)2829 keys = [hashlib.sha256(f"{self.model_name}|{t}".encode()).hexdigest()30 for t in texts]31 missing = [i for i, k in enumerate(keys) if k not in self._cache]32 if missing:33 fresh = self._encode([texts[i] for i in missing], batch_size)34 for i, vec in zip(missing, fresh):35 self._cache[keys[i]] = vec36 return np.vstack([self._cache[k] for k in keys])3738 def _encode(self, texts, batch_size):39 return self.model.encode(40 texts, batch_size=batch_size,41 normalize_embeddings=True, # unit length: dot product = cosine42 show_progress_bar=len(texts) > 500,43 ).astype("float32")4445 def manifest(self):46 return {"model": self.model_name, "dimension": self.dimension,47 "normalised": True, "max_tokens": self.max_tokens}Three decisions here are the ones markers should look for. Raising on overlong input rather than truncating: silent truncation is the most damaging failure in this whole project because it produces a valid-looking vector representing the first quarter of a document, and no log will ever mention it. Normalising at encode time: unit vectors make inner product identical to cosine similarity, so no later stage can accidentally rank by document length. Including the model name in the cache key: without it, changing models returns stale vectors from the old one, and nothing errors.
src/indexing.py
1import json, os2import faiss3import numpy as np456class VectorIndex:7 def __init__(self, manifest, index_type="flat"):8 self.manifest = dict(manifest, index_type=index_type)9 d = manifest["dimension"]10 if index_type == "flat":11 self.index = faiss.IndexFlatIP(d) # exact; the right default12 elif index_type == "hnsw":13 self.index = faiss.IndexHNSWFlat(d, 32, faiss.METRIC_INNER_PRODUCT)14 self.index.hnsw.efConstruction = 20015 else:16 raise ValueError(f"unknown index_type {index_type!r}")17 self.documents, self.metadata = [], []1819 def add(self, embeddings, documents, metadata):20 assert embeddings.dtype == np.float32, "FAISS requires float32"21 assert len(documents) == len(metadata) == embeddings.shape[0]22 self.index.add(embeddings)23 self.documents.extend(documents)24 self.metadata.extend(metadata)2526 def search(self, query_vec, k=50):27 scores, ids = self.index.search(query_vec.reshape(1, -1), k)28 return [29 {"id": int(i), "score": float(s), "rank": r + 1,30 "doc": self.documents[i], "meta": self.metadata[i]}31 for r, (i, s) in enumerate(zip(ids[0], scores[0])) if i >= 032 ]3334 def save(self, directory):35 os.makedirs(directory, exist_ok=True)36 faiss.write_index(self.index, os.path.join(directory, "vectors.faiss"))37 with open(os.path.join(directory, "store.json"), "w") as f:38 json.dump({"manifest": self.manifest, "documents": self.documents,39 "metadata": self.metadata}, f)4041 @classmethod42 def load(cls, directory, manifest):43 with open(os.path.join(directory, "store.json")) as f:44 saved = json.load(f)45 for key in ("model", "dimension", "normalised"):46 if saved["manifest"][key] != manifest[key]:47 raise RuntimeError(48 f"index built with {key}={saved['manifest'][key]!r}, "49 f"current config uses {manifest[key]!r} - rebuild required"50 )51 obj = cls(manifest, saved["manifest"]["index_type"])52 obj.index = faiss.read_index(os.path.join(directory, "vectors.faiss"))53 obj.documents, obj.metadata = saved["documents"], saved["metadata"]54 return objThe manifest check in load() is the single highest-value defensive line in the project. Two embedding models can both emit 384 dimensions and still occupy completely unrelated spaces, because each learned its own axes from scratch. Build an index with one and query with another and FAISS returns ten results with plausible scores, every one of them meaningless — no exception, no warning, just quietly wrong search forever. Nine lines of comparison turn that into a startup error.
An index is valid only for the exact model that produced its vectors. Store the model name beside the vectors and refuse to load on a mismatch, or you will one day debug a quality regression that has no stack trace.
The if i >= 0 filter matters too: when the index holds fewer than k vectors FAISS pads the id array with -1, and self.documents[-1] cheerfully returns the last document instead of raising.
src/ranking.py
1from sentence_transformers import CrossEncoder234class Reranker:5 """Load ONCE at startup. Constructing this per query costs ~1.8 seconds."""67 def __init__(self, model_name="cross-encoder/ms-marco-MiniLM-L6-v2"):8 self.model = CrossEncoder(model_name)910 def rank(self, query, candidates, top_k=5):11 if not candidates:12 return []13 scores = self.model.predict([[query, c["doc"]] for c in candidates],14 batch_size=32)15 for c, s in zip(candidates, scores):16 c["rerank_score"] = float(s)17 return sorted(candidates, key=lambda c: -c["rerank_score"])[:top_k]181920def reciprocal_rank_fusion(rank_lists, k=60, top_n=50):21 """Merge ranked id lists from several retrievers. No score calibration needed."""22 fused = {}23 for ranking in rank_lists:24 for position, doc_id in enumerate(ranking, start=1):25 fused[doc_id] = fused.get(doc_id, 0.0) + 1.0 / (k + position)26 return sorted(fused, key=fused.get, reverse=True)[:top_n]The retrieval model encodes query and document separately, which is what makes it fast — every document vector is computed once and reused forever — but it also means the model never sees the two together. It cannot tell "refunds after 30 days" from "refunds within 30 days". A cross-encoder concatenates the pair into one input and runs a full transformer over it, so it can attend to "after" against "within". The price is that nothing precomputes: at roughly 800 pairs per second, reranking 50 candidates costs 63 ms and reranking all 10,000 documents would cost 12.5 seconds per query.
Add reciprocal_rank_fusion when your corpus contains identifiers — SKUs, error codes, version numbers — that embeddings blur. Run BM25 alongside the vector search and fuse the two rank lists. RRF needs no tuning because it ignores raw scores entirely: a document at rank 2 in both lists scores 1/62+1/62=0.0323 and beats one that was rank 1 in a single list and rank 30 in the other, at 1/61+1/90=0.0275. Agreement between retrievers outranks a single strong opinion.
src/search.py
1import re234class SearchEngine:5 def __init__(self, embedder, index, reranker=None):6 self.embedder, self.index, self.reranker = embedder, index, reranker78 @staticmethod9 def normalise(query):10 # Collapse whitespace ONLY. Do not lowercase, do not strip punctuation:11 # most encoders are trained on cased, punctuated text, and "XR-4400"12 # must not become "XR 4400".13 return re.sub(r"\s+", " ", query).strip()1415 def search(self, query, top_k=5, filters=None, candidates=50):16 query = self.normalise(query)17 if not query:18 return []19 qvec = self.embedder.encode([query], use_cache=True)[0]20 hits = self.index.search(qvec, k=candidates)21 hits = self._apply_filters(hits, filters) # BEFORE ranking22 if self.reranker:23 return self.reranker.rank(query, hits, top_k=top_k)24 return hits[:top_k]2526 @staticmethod27 def _apply_filters(hits, filters):28 if not filters:29 return hits30 return [h for h in hits31 if all(h["meta"].get(key) == value for key, value in filters.items())]That _apply_filters is deliberately the naive version, and part of the project is understanding why it is only acceptable here. It is a post-filter: it retrieves 50 candidates and then discards ineligible ones. Ask for 5 results with a category filter, retrieve 50, find that only 3 match the category, and show 3 — while hundreds of eligible documents sit further down the ranking, unexamined. The user sees a starved page and concludes search is broken.
Post-filtering is survivable only when the filter is unselective and the candidate pool is generous. It is never acceptable for anything that decides whether a user is permitted to see a document, because by the time the list comprehension runs, the restricted text has already been fetched, ranked and very likely logged. Relevance ranking is not an access control mechanism. To do this properly, either push the constraint into the store (a vector database with a where clause evaluates it during the search) or resolve the eligible id set from a conventional database index first and search only within it.
src/evaluate.py — the part most projects skip
Without this module you cannot answer "did that change help?", and every subsequent decision becomes a matter of taste. Build a file of 30 to 50 real queries, each with the ids you consider relevant:
1[2 {"query": "how do neural networks learn", "relevant": ["2", "17", "31"]},3 {"query": "XR-4400 power supply", "relevant": ["44"]}4]1def recall_at_k(returned_ids, relevant_ids, k):2 hit = len(set(returned_ids[:k]) & set(relevant_ids))3 return hit / len(relevant_ids)45def reciprocal_rank(returned_ids, relevant_ids):6 for position, doc_id in enumerate(returned_ids, start=1):7 if doc_id in relevant_ids:8 return 1.0 / position9 return 0.01011def average_precision(returned_ids, relevant_ids):12 hits, total = 0, 0.013 for position, doc_id in enumerate(returned_ids, start=1):14 if doc_id in relevant_ids:15 hits += 116 total += hits / position17 return total / len(relevant_ids)Work each one by hand once, so the numbers mean something when the harness prints them.
Recall@5. A query has 8 relevant documents in the corpus and your top 5 contains 3 of them. Recall@5 = 3/8 = 0.375. Precision@5 = 3/5 = 0.60. Note they answer different questions: precision asks "is what I showed good?", recall asks "did I find what exists?".
Mean reciprocal rank. Five queries; the first relevant result appears at rank 1, 3, 2, nowhere, and 1. The reciprocal ranks are 1.000, 0.333, 0.500, 0.000, 1.000, so MRR = 2.833/5=0.567. MRR only looks at the first hit, which makes it the right metric when users want one answer.
Mean average precision. Query A has 3 relevant documents, returned at ranks 1, 3 and 6:
Query B has 2 relevant documents at ranks 2 and 5: AP=21(21+52)=0.450. So MAP = (0.722+0.450)/2=0.586. MAP rewards getting all the relevant documents high, which makes it the right metric when the user needs a complete set.
| Metric | Answers | Use it to tune |
|---|---|---|
| Recall@50 | Did retrieval find the answer at all? | Candidate depth, chunking, the embedding model |
| MRR | How high is the first good result? | The reranker |
| MAP | Are all the good results near the top? | Fusion strategy, rerank cut-off |
| P95 latency | Is it fast enough for the worst case? | Batch sizes, candidate depth, caching |
Retrieval recall is the ceiling on every number downstream. Measure it before you tune anything, or you will spend a week improving the order of a list that never contained the right answer.
Measure recall@50 first, because it is a hard ceiling on everything else. A reranker only reorders the list it was handed — if the right document sat at rank 240 and you retrieved 50, no reranker on earth will surface it. If recall@50 comes out at 0.6, stop tuning the ranker: the problem is your chunking, your model, or your data.
src/app.py and the CLI
1import json23class SearchApp:4 def __init__(self, config_path="config.json"):5 cfg = json.load(open(config_path))6 self.embedder = EmbeddingManager(cfg.get("model", "all-MiniLM-L6-v2"))7 self.reranker = Reranker() if cfg.get("use_reranking", True) else None8 self.cfg = cfg9 self.engine = None1011 def build(self, documents_path, index_dir="index/"):12 data = json.load(open(documents_path))["documents"]13 texts = [f"{d['title']}. {d['content']}" for d in data]14 metadata = [{k: v for k, v in d.items() if k != "content"} for d in data]15 vectors = self.embedder.encode(texts, batch_size=self.cfg.get("batch_size", 64))16 index = VectorIndex(self.embedder.manifest(), self.cfg.get("index_type", "flat"))17 index.add(vectors, texts, metadata)18 index.save(index_dir)19 self.engine = SearchEngine(self.embedder, index, self.reranker)2021 def load(self, index_dir="index/"):22 index = VectorIndex.load(index_dir, self.embedder.manifest())23 self.engine = SearchEngine(self.embedder, index, self.reranker)Note f"{title}. {content}" in build(). Embedding title and body together usually beats embedding the body alone, because titles are dense with the topic words a query is likely to use. It is a one-line change that moves MRR measurably, and your evaluation harness is what tells you by how much on your corpus.
Keep configuration in one file so a model swap is a single edit rather than a search-and-replace:
{"model": "all-MiniLM-L6-v2", "index_type": "flat", "use_reranking": true, "top_k": 5, "candidates": 50, "batch_size": 64}The CLI needs an interactive loop and a one-shot --query mode, both routing through the same search() call so they cannot drift apart. When results are empty, print why: how many candidates were retrieved, how many survived filtering. Those two counts distinguish "nothing matched semantically" from "your filter excluded everything", and without them an empty page is undiagnosable.
Tests that would actually have caught something
| Test | Bug it catches |
|---|---|
| Encode a 900-word document | Silent truncation — must raise, not return a vector |
| Save, load, and compare top-5 ids | Persistence that loses metadata or reorders documents |
| Load an index whose manifest names a different model | The cross-model comparison that never errors on its own |
Search an index holding 3 documents with k=50 | The -1 padding that silently returns the last document |
| Filter on a category matching 2 of 50 candidates | Result starvation from post-filtering |
| Query with only whitespace | Empty-string encoding and division-by-zero norms |
| Run the eval harness and assert MRR > a floor | Any regression, from any change, anywhere |
The last one is worth more than the other six combined. A single test asserting that MRR on your checked-in query set stays above a floor turns every future refactor into a decision with evidence behind it.
Extensions, in order of value
- Hybrid retrieval. Index with
BM25Okapialongside the vectors and fuse the two rank lists with RRF. Highest-value addition by a distance if your corpus contains any identifiers. - An HTTP endpoint. A small Flask or FastAPI wrapper:
POST /searchtaking query, top_k and filters, plusGET /health. Load models at import time, never per request. - Per-stage timing. Log milliseconds for retrieve, filter and rerank separately, plus candidate counts in and out of each stage and a zero-result flag. Report P95, not the mean — the mean hides exactly the queries users complain about.
- Query caching. An LRU cache keyed on the normalised query, top_k, and the filters. Omitting filters from the key means one user's filtered results get served to another, which is a correctness bug at best and a data leak at worst.
- Chunking long documents. Split overlong documents into overlapping windows, index each chunk with a parent id, and deduplicate by parent at result time. This is what makes the token-limit error in
encode()actionable rather than merely obstructive.
How the work is judged
| Criterion | Weight | What earns full marks |
|---|---|---|
| Core functionality | 40% | Index, persist, reload, search, filter, rerank — all working end to end from a clean checkout |
| Measurement | 20% | Eval harness runs from a checked-in query file and reports recall@50, MRR, MAP and P95 latency |
| Code quality | 15% | Clear module boundaries; models loaded once; no silent failure paths |
| Documentation | 15% | A README stating the model, the index type, the measured numbers, and the known limitations |
| Testing | 10% | The failure-mode tests above, passing |
Your README should be able to state, in one line: "all-MiniLM-L6-v2, flat index over 10,000 documents, recall@50 = 0.94, MRR = 0.61, P95 latency 82 ms." A project that can say that is finished. One that cannot is a demo, whatever its feature list looks like.
What separates a working build from a good one
Build in the order that surfaces problems earliest, which is not the order the file listing suggests. Get 50 documents indexed and searchable with no reranker, no filters and no cache — then immediately write the evaluation harness and record recall@50. That number is your ceiling, and if it is bad, every hour spent on ranking is wasted. Only once retrieval is measured and sound should the reranker go in, because it is the largest precision gain for the least code.
Make every silent failure loud. Truncation, model mismatch, -1 padding, empty queries, cache keys missing a dimension — the recurring theme across this entire project is that the damaging failures produce plausible output rather than exceptions. A vector for the first quarter of a document looks exactly like a vector for the whole document. Search built with the wrong model returns results with confident scores. Every guard you add converts one of those into a message at the moment it happens rather than a mystery three months later.
Treat the index as derived data. Everything in it can be rebuilt from documents.json and the config, and it should be routine to do so — a command anyone can run, not a recovery procedure someone invents under pressure. That mindset is what lets you change the embedding model, re-chunk the corpus, or switch index types without fear, and it is the difference between a search system you can improve and one you are afraid to touch.