Embeddings and Semantic Search

Mini Project: Build a Semantic Search App


Here is the demo that gets built a thousand times a year. Twenty sentences in a Python list, model.encode(), a cosine similarity call, a print loop. It takes eleven minutes and it genuinely works — type "how do I keep my laptop cool" and it finds the cooling-pad article that shares no words with the query.

Here is what happens when that demo meets ten thousand real documents. Half of them are longer than the model's 256-token window, so three quarters of each is silently thrown away before it is ever indexed. Encoding takes eleven minutes rather than eleven seconds, and nothing is cached, so it happens again on every restart. Somebody searches for a product code and gets nothing useful. Somebody else asks "did search get worse after we changed the model?" and there is no way to answer, because nobody wrote down what "good" was.

This project is the gap between those two things. You are building a semantic search application over a real corpus, with a measured latency budget, a stated quality metric, and an index you can rebuild from source at any time. The interesting work is not the embedding call. It is everything that stops the embedding call from being useless at scale.

A 200 ms budget, spent before you write codeQueryembedding — 15 msANN search overthe index — 8 msCross-encoderre-rank 50 — 120 msFormatting andtransport — 25 msHeadroomleft — 32 mstopbottomRe-ranking is 60 percent of the budget, so the number of candidates it sees is the one dial that matters.
Deciding the budget first turns 'is this fast enough' into arithmetic you can do before the first import, and names the stage to cut when it is not.

What you are building

A command-line application that indexes a document collection once, persists the index, and then answers queries in under half a second with ranked, scored, filterable results.

CapabilityRequirementAcceptance test
IndexingLoad documents from JSON, embed them, persist vectors and metadataRestart the process; search works without re-embedding
Query handlingNormalise, encode with the same model used for indexingA stored manifest names the model; a mismatch refuses to load
RetrievalTop-50 candidates by cosine similarityReturns 50 distinct ids for a query with 50+ relevant docs
RankingCross-encoder rerank of candidates down to top-kReranked order differs from retrieval order on at least one test query
FilteringMetadata constraints applied before rankingA category filter returns a full k results, not a starved list
InterfaceInteractive CLI plus one-shot --query modeBoth paths produce identical results for the same query
EvaluationA harness reporting recall@50, MRR and MAPNumbers regenerate from a checked-in query file

The latency budget, spent in advance

"Under 500 ms" is not a target until you know where the milliseconds go. For 10,000 documents at 384 dimensions on a modest CPU:

StageWorkBudget
Normalise queryRegex and a token count< 1 ms
Encode queryOne transformer pass, ~22M parameters8 ms
Vector search10,000 × 384 = 3.84M multiply-adds2 ms
Metadata filterDictionary lookups over 50 candidates< 1 ms
Cross-encoder rerank50 pairs at ~800 pairs/second63 ms
Format and printString work1 ms
Total~75 ms

You have 425 ms of headroom, which is exactly why the budget is worth writing down: it tells you that the reranker — 84% of your spend — is the only stage worth optimising, and that everything else is noise. It also makes two classic mistakes instantly visible. Constructing the CrossEncoder inside the request handler adds 1,800 ms of model loading per query, blowing the budget 24-fold. Encoding documents one at a time instead of in batches turns a 20-second index build into a 6-minute one.

Write the latency budget before the code. It converts "make it fast" into a list of numbers you can test against, and it tells you which nine tenths of the system are not worth optimising.

Memory is the other resource worth predicting. 10,000 vectors at 384 float32 dimensions is 10,000×384×4=15.3610{,}000 \times 384 \times 4 = 15.36 MB — negligible. The models dominate: the bi-encoder is roughly 90 MB and the cross-encoder another 90 MB, so budget around 250 MB resident. Vectors only become the problem at a hundred times this scale.

Layout

Text
semantic-search-app/├── data/│   ├── documents.json        50+ docs across 3+ categories│   └── eval_queries.json     queries with known-relevant ids├── src/│   ├── embeddings.py         encoding, caching, token-limit guard│   ├── indexing.py           FAISS index + manifest + persistence│   ├── search.py             retrieve → filter → rerank│   ├── ranking.py            cross-encoder and rank fusion│   ├── evaluate.py           recall@k, MRR, MAP│   └── app.py                CLI, config, result formatting├── tests/├── config.json├── requirements.txt└── main.py
Bash
python -m venv venv && source venv/bin/activate      # Windows: venv\Scripts\activatepip install sentence-transformers faiss-cpu rank-bm25 numpypip freeze > requirements.txt

The corpus

Fifty documents minimum, spread across at least three categories, with at least a few documents that are near-duplicates of each other. Ten documents cannot distinguish a good ranker from a bad one — with ten documents every query returns something plausible and every metric reads 1.0.

JSON
{  "documents": [    {"id": "1", "title": "Introduction to Machine Learning",     "content": "Machine learning is a subset of artificial intelligence that ...",     "category": "AI", "year": 2024, "visibility": "public"},    {"id": "2", "title": "Deep Learning Fundamentals",     "content": "Deep learning uses neural networks with multiple layers ...",     "category": "AI", "year": 2023, "visibility": "public"}  ]}

src/embeddings.py

This module owns one non-negotiable rule: every vector that ever enters the index comes from here, and this module knows which model made it.

Python
import hashlibimport numpy as npfrom sentence_transformers import SentenceTransformerclass EmbeddingManager:    def __init__(self, model_name="all-MiniLM-L6-v2"):        self.model_name = model_name        self.model = SentenceTransformer(model_name)        self.dimension = self.model.get_embedding_dimension()   # get_sentence_embedding_dimension() before v6        self.max_tokens = self.model.max_seq_length          # 256 for MiniLM-L6        self._cache = {}    def _too_long(self, text):        # +2 for the [CLS] and [SEP] markers, which also count against the limit        return len(self.model.tokenizer.tokenize(text)) + 2 > self.max_tokens    def encode(self, texts, batch_size=64, use_cache=True):        """Returns unit-length float32 vectors, shape (len(texts), dimension)."""        overlong = [i for i, t in enumerate(texts) if self._too_long(t)]        if overlong:            raise ValueError(                f"{len(overlong)} texts exceed {self.max_tokens} tokens "                f"(first at index {overlong[0]}). Chunk them before indexing."            )        if not use_cache:            return self._encode(texts, batch_size)        keys = [hashlib.sha256(f"{self.model_name}|{t}".encode()).hexdigest()                for t in texts]        missing = [i for i, k in enumerate(keys) if k not in self._cache]        if missing:            fresh = self._encode([texts[i] for i in missing], batch_size)            for i, vec in zip(missing, fresh):                self._cache[keys[i]] = vec        return np.vstack([self._cache[k] for k in keys])    def _encode(self, texts, batch_size):        return self.model.encode(            texts, batch_size=batch_size,            normalize_embeddings=True,        # unit length: dot product = cosine            show_progress_bar=len(texts) > 500,        ).astype("float32")    def manifest(self):        return {"model": self.model_name, "dimension": self.dimension,                "normalised": True, "max_tokens": self.max_tokens}

Three decisions here are the ones markers should look for. Raising on overlong input rather than truncating: silent truncation is the most damaging failure in this whole project because it produces a valid-looking vector representing the first quarter of a document, and no log will ever mention it. Normalising at encode time: unit vectors make inner product identical to cosine similarity, so no later stage can accidentally rank by document length. Including the model name in the cache key: without it, changing models returns stale vectors from the old one, and nothing errors.

src/indexing.py

Python
import json, osimport faissimport numpy as npclass VectorIndex:    def __init__(self, manifest, index_type="flat"):        self.manifest = dict(manifest, index_type=index_type)        d = manifest["dimension"]        if index_type == "flat":            self.index = faiss.IndexFlatIP(d)            # exact; the right default        elif index_type == "hnsw":            self.index = faiss.IndexHNSWFlat(d, 32, faiss.METRIC_INNER_PRODUCT)            self.index.hnsw.efConstruction = 200        else:            raise ValueError(f"unknown index_type {index_type!r}")        self.documents, self.metadata = [], []    def add(self, embeddings, documents, metadata):        assert embeddings.dtype == np.float32, "FAISS requires float32"        assert len(documents) == len(metadata) == embeddings.shape[0]        self.index.add(embeddings)        self.documents.extend(documents)        self.metadata.extend(metadata)    def search(self, query_vec, k=50):        scores, ids = self.index.search(query_vec.reshape(1, -1), k)        return [            {"id": int(i), "score": float(s), "rank": r + 1,             "doc": self.documents[i], "meta": self.metadata[i]}            for r, (i, s) in enumerate(zip(ids[0], scores[0])) if i >= 0        ]    def save(self, directory):        os.makedirs(directory, exist_ok=True)        faiss.write_index(self.index, os.path.join(directory, "vectors.faiss"))        with open(os.path.join(directory, "store.json"), "w") as f:            json.dump({"manifest": self.manifest, "documents": self.documents,                       "metadata": self.metadata}, f)    @classmethod    def load(cls, directory, manifest):        with open(os.path.join(directory, "store.json")) as f:            saved = json.load(f)        for key in ("model", "dimension", "normalised"):            if saved["manifest"][key] != manifest[key]:                raise RuntimeError(                    f"index built with {key}={saved['manifest'][key]!r}, "                    f"current config uses {manifest[key]!r} - rebuild required"                )        obj = cls(manifest, saved["manifest"]["index_type"])        obj.index = faiss.read_index(os.path.join(directory, "vectors.faiss"))        obj.documents, obj.metadata = saved["documents"], saved["metadata"]        return obj

The manifest check in load() is the single highest-value defensive line in the project. Two embedding models can both emit 384 dimensions and still occupy completely unrelated spaces, because each learned its own axes from scratch. Build an index with one and query with another and FAISS returns ten results with plausible scores, every one of them meaningless — no exception, no warning, just quietly wrong search forever. Nine lines of comparison turn that into a startup error.

An index is valid only for the exact model that produced its vectors. Store the model name beside the vectors and refuse to load on a mismatch, or you will one day debug a quality regression that has no stack trace.

The if i >= 0 filter matters too: when the index holds fewer than k vectors FAISS pads the id array with -1, and self.documents[-1] cheerfully returns the last document instead of raising.

src/ranking.py

Python
from sentence_transformers import CrossEncoderclass Reranker:    """Load ONCE at startup. Constructing this per query costs ~1.8 seconds."""    def __init__(self, model_name="cross-encoder/ms-marco-MiniLM-L6-v2"):        self.model = CrossEncoder(model_name)    def rank(self, query, candidates, top_k=5):        if not candidates:            return []        scores = self.model.predict([[query, c["doc"]] for c in candidates],                                    batch_size=32)        for c, s in zip(candidates, scores):            c["rerank_score"] = float(s)        return sorted(candidates, key=lambda c: -c["rerank_score"])[:top_k]def reciprocal_rank_fusion(rank_lists, k=60, top_n=50):    """Merge ranked id lists from several retrievers. No score calibration needed."""    fused = {}    for ranking in rank_lists:        for position, doc_id in enumerate(ranking, start=1):            fused[doc_id] = fused.get(doc_id, 0.0) + 1.0 / (k + position)    return sorted(fused, key=fused.get, reverse=True)[:top_n]

The retrieval model encodes query and document separately, which is what makes it fast — every document vector is computed once and reused forever — but it also means the model never sees the two together. It cannot tell "refunds after 30 days" from "refunds within 30 days". A cross-encoder concatenates the pair into one input and runs a full transformer over it, so it can attend to "after" against "within". The price is that nothing precomputes: at roughly 800 pairs per second, reranking 50 candidates costs 63 ms and reranking all 10,000 documents would cost 12.5 seconds per query.

Add reciprocal_rank_fusion when your corpus contains identifiers — SKUs, error codes, version numbers — that embeddings blur. Run BM25 alongside the vector search and fuse the two rank lists. RRF needs no tuning because it ignores raw scores entirely: a document at rank 2 in both lists scores 1/62+1/62=0.03231/62 + 1/62 = 0.0323 and beats one that was rank 1 in a single list and rank 30 in the other, at 1/61+1/90=0.02751/61 + 1/90 = 0.0275. Agreement between retrievers outranks a single strong opinion.

src/search.py

Python
import reclass SearchEngine:    def __init__(self, embedder, index, reranker=None):        self.embedder, self.index, self.reranker = embedder, index, reranker    @staticmethod    def normalise(query):        # Collapse whitespace ONLY. Do not lowercase, do not strip punctuation:        # most encoders are trained on cased, punctuated text, and "XR-4400"        # must not become "XR 4400".        return re.sub(r"\s+", " ", query).strip()    def search(self, query, top_k=5, filters=None, candidates=50):        query = self.normalise(query)        if not query:            return []        qvec = self.embedder.encode([query], use_cache=True)[0]        hits = self.index.search(qvec, k=candidates)        hits = self._apply_filters(hits, filters)          # BEFORE ranking        if self.reranker:            return self.reranker.rank(query, hits, top_k=top_k)        return hits[:top_k]    @staticmethod    def _apply_filters(hits, filters):        if not filters:            return hits        return [h for h in hits                if all(h["meta"].get(key) == value for key, value in filters.items())]

That _apply_filters is deliberately the naive version, and part of the project is understanding why it is only acceptable here. It is a post-filter: it retrieves 50 candidates and then discards ineligible ones. Ask for 5 results with a category filter, retrieve 50, find that only 3 match the category, and show 3 — while hundreds of eligible documents sit further down the ranking, unexamined. The user sees a starved page and concludes search is broken.

Post-filtering is survivable only when the filter is unselective and the candidate pool is generous. It is never acceptable for anything that decides whether a user is permitted to see a document, because by the time the list comprehension runs, the restricted text has already been fetched, ranked and very likely logged. Relevance ranking is not an access control mechanism. To do this properly, either push the constraint into the store (a vector database with a where clause evaluates it during the search) or resolve the eligible id set from a conventional database index first and search only within it.

src/evaluate.py — the part most projects skip

Without this module you cannot answer "did that change help?", and every subsequent decision becomes a matter of taste. Build a file of 30 to 50 real queries, each with the ids you consider relevant:

JSON
[  {"query": "how do neural networks learn", "relevant": ["2", "17", "31"]},  {"query": "XR-4400 power supply", "relevant": ["44"]}]
Python
def recall_at_k(returned_ids, relevant_ids, k):    hit = len(set(returned_ids[:k]) & set(relevant_ids))    return hit / len(relevant_ids)def reciprocal_rank(returned_ids, relevant_ids):    for position, doc_id in enumerate(returned_ids, start=1):        if doc_id in relevant_ids:            return 1.0 / position    return 0.0def average_precision(returned_ids, relevant_ids):    hits, total = 0, 0.0    for position, doc_id in enumerate(returned_ids, start=1):        if doc_id in relevant_ids:            hits += 1            total += hits / position    return total / len(relevant_ids)

Work each one by hand once, so the numbers mean something when the harness prints them.

Recall@5. A query has 8 relevant documents in the corpus and your top 5 contains 3 of them. Recall@5 = 3/8 = 0.375. Precision@5 = 3/5 = 0.60. Note they answer different questions: precision asks "is what I showed good?", recall asks "did I find what exists?".

Mean reciprocal rank. Five queries; the first relevant result appears at rank 1, 3, 2, nowhere, and 1. The reciprocal ranks are 1.000, 0.333, 0.500, 0.000, 1.000, so MRR = 2.833/5=0.5672.833 / 5 = \mathbf{0.567}. MRR only looks at the first hit, which makes it the right metric when users want one answer.

Mean average precision. Query A has 3 relevant documents, returned at ranks 1, 3 and 6:

AP=13(11+23+36)=2.1673=0.722\text{AP} = \frac{1}{3}\left(\frac{1}{1} + \frac{2}{3} + \frac{3}{6}\right) = \frac{2.167}{3} = 0.722

Query B has 2 relevant documents at ranks 2 and 5: AP=12(12+25)=0.450\text{AP} = \frac{1}{2}(\frac{1}{2} + \frac{2}{5}) = 0.450. So MAP = (0.722+0.450)/2=0.586(0.722 + 0.450)/2 = \mathbf{0.586}. MAP rewards getting all the relevant documents high, which makes it the right metric when the user needs a complete set.

MetricAnswersUse it to tune
Recall@50Did retrieval find the answer at all?Candidate depth, chunking, the embedding model
MRRHow high is the first good result?The reranker
MAPAre all the good results near the top?Fusion strategy, rerank cut-off
P95 latencyIs it fast enough for the worst case?Batch sizes, candidate depth, caching

Retrieval recall is the ceiling on every number downstream. Measure it before you tune anything, or you will spend a week improving the order of a list that never contained the right answer.

Measure recall@50 first, because it is a hard ceiling on everything else. A reranker only reorders the list it was handed — if the right document sat at rank 240 and you retrieved 50, no reranker on earth will surface it. If recall@50 comes out at 0.6, stop tuning the ranker: the problem is your chunking, your model, or your data.

src/app.py and the CLI

Python
import jsonclass SearchApp:    def __init__(self, config_path="config.json"):        cfg = json.load(open(config_path))        self.embedder = EmbeddingManager(cfg.get("model", "all-MiniLM-L6-v2"))        self.reranker = Reranker() if cfg.get("use_reranking", True) else None        self.cfg = cfg        self.engine = None    def build(self, documents_path, index_dir="index/"):        data = json.load(open(documents_path))["documents"]        texts = [f"{d['title']}. {d['content']}" for d in data]        metadata = [{k: v for k, v in d.items() if k != "content"} for d in data]        vectors = self.embedder.encode(texts, batch_size=self.cfg.get("batch_size", 64))        index = VectorIndex(self.embedder.manifest(), self.cfg.get("index_type", "flat"))        index.add(vectors, texts, metadata)        index.save(index_dir)        self.engine = SearchEngine(self.embedder, index, self.reranker)    def load(self, index_dir="index/"):        index = VectorIndex.load(index_dir, self.embedder.manifest())        self.engine = SearchEngine(self.embedder, index, self.reranker)

Note f"{title}. {content}" in build(). Embedding title and body together usually beats embedding the body alone, because titles are dense with the topic words a query is likely to use. It is a one-line change that moves MRR measurably, and your evaluation harness is what tells you by how much on your corpus.

Keep configuration in one file so a model swap is a single edit rather than a search-and-replace:

JSON
{"model": "all-MiniLM-L6-v2", "index_type": "flat", "use_reranking": true, "top_k": 5, "candidates": 50, "batch_size": 64}

The CLI needs an interactive loop and a one-shot --query mode, both routing through the same search() call so they cannot drift apart. When results are empty, print why: how many candidates were retrieved, how many survived filtering. Those two counts distinguish "nothing matched semantically" from "your filter excluded everything", and without them an empty page is undiagnosable.

Tests that would actually have caught something

TestBug it catches
Encode a 900-word documentSilent truncation — must raise, not return a vector
Save, load, and compare top-5 idsPersistence that loses metadata or reorders documents
Load an index whose manifest names a different modelThe cross-model comparison that never errors on its own
Search an index holding 3 documents with k=50The -1 padding that silently returns the last document
Filter on a category matching 2 of 50 candidatesResult starvation from post-filtering
Query with only whitespaceEmpty-string encoding and division-by-zero norms
Run the eval harness and assert MRR > a floorAny regression, from any change, anywhere

The last one is worth more than the other six combined. A single test asserting that MRR on your checked-in query set stays above a floor turns every future refactor into a decision with evidence behind it.

Extensions, in order of value

  1. Hybrid retrieval. Index with BM25Okapi alongside the vectors and fuse the two rank lists with RRF. Highest-value addition by a distance if your corpus contains any identifiers.
  2. An HTTP endpoint. A small Flask or FastAPI wrapper: POST /search taking query, top_k and filters, plus GET /health. Load models at import time, never per request.
  3. Per-stage timing. Log milliseconds for retrieve, filter and rerank separately, plus candidate counts in and out of each stage and a zero-result flag. Report P95, not the mean — the mean hides exactly the queries users complain about.
  4. Query caching. An LRU cache keyed on the normalised query, top_k, and the filters. Omitting filters from the key means one user's filtered results get served to another, which is a correctness bug at best and a data leak at worst.
  5. Chunking long documents. Split overlong documents into overlapping windows, index each chunk with a parent id, and deduplicate by parent at result time. This is what makes the token-limit error in encode() actionable rather than merely obstructive.

How the work is judged

CriterionWeightWhat earns full marks
Core functionality40%Index, persist, reload, search, filter, rerank — all working end to end from a clean checkout
Measurement20%Eval harness runs from a checked-in query file and reports recall@50, MRR, MAP and P95 latency
Code quality15%Clear module boundaries; models loaded once; no silent failure paths
Documentation15%A README stating the model, the index type, the measured numbers, and the known limitations
Testing10%The failure-mode tests above, passing

Your README should be able to state, in one line: "all-MiniLM-L6-v2, flat index over 10,000 documents, recall@50 = 0.94, MRR = 0.61, P95 latency 82 ms." A project that can say that is finished. One that cannot is a demo, whatever its feature list looks like.

What separates a working build from a good one

Build in the order that surfaces problems earliest, which is not the order the file listing suggests. Get 50 documents indexed and searchable with no reranker, no filters and no cache — then immediately write the evaluation harness and record recall@50. That number is your ceiling, and if it is bad, every hour spent on ranking is wasted. Only once retrieval is measured and sound should the reranker go in, because it is the largest precision gain for the least code.

Make every silent failure loud. Truncation, model mismatch, -1 padding, empty queries, cache keys missing a dimension — the recurring theme across this entire project is that the damaging failures produce plausible output rather than exceptions. A vector for the first quarter of a document looks exactly like a vector for the whole document. Search built with the wrong model returns results with confident scores. Every guard you add converts one of those into a message at the moment it happens rather than a mystery three months later.

Treat the index as derived data. Everything in it can be rebuilt from documents.json and the config, and it should be routine to do so — a command anyone can run, not a recovery procedure someone invents under pressure. That mindset is what lets you change the embedding model, re-chunk the corpus, or switch index types without fear, and it is the difference between a search system you can improve and one you are afraid to touch.