Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Implement semantic chunking based on meaning.
What you need to know
Fixed and recursive chunking cut by length. Semantic chunking cuts by meaning: a chunk should hold one topic, and the boundary should be where the text moves to the next one.
The signal is the distance between neighbouring sentences:
distance[i] = 1 − cosine(embedding(sentence i), embedding(sentence i+1))Within a topic, neighbouring sentences are similar (small distance). At a topic change the distance jumps.
Percentile threshold. Some documents are written tightly, others loosely, so an absolute cutoff like 0.3 breaks one document everywhere and another nowhere. Taking the 90th percentile of this document's distances means "break at the biggest 10% of jumps".
Off-by-one. When you look at sentence i + 1 and decide whether it starts a new chunk, the gap to check is distance[i] — between sentence i and sentence i + 1.
1import re2import numpy as np34def split_sentences(text: str) -> list[str]:5 return [s.strip() for s in re.split(r"(?<=[.!?])\s+", text) if s.strip()]67def semantic_chunks(text: str, embed_fn, percentile: float = 90,8 max_chars: int = 1500) -> list[str]:9 """Start a new chunk where neighbouring sentences are unusually far apart."""10 sentences = split_sentences(text)11 if len(sentences) < 3:12 return [" ".join(sentences)] if sentences else []13 v = np.asarray(embed_fn(sentences), dtype=np.float32)14 v /= np.linalg.norm(v, axis=1, keepdims=True) + 1e-1015 distance = 1.0 - np.sum(v[:-1] * v[1:], axis=1) # gap after each sentence16 cutoff = np.percentile(distance, percentile)1718 chunks, current = [], [sentences[0]]19 for i, sentence in enumerate(sentences[1:]): # sentence is sentences[i + 1]20 too_long = sum(map(len, current)) + len(sentence) > max_chars21 if distance[i] > cutoff or too_long:22 chunks.append(" ".join(current))23 current = []24 current.append(sentence)25 chunks.append(" ".join(current))26 return chunksThe tricky parts:
np.sum(v[:-1] * v[1:], axis=1)multiplies each row by the next row element-wise and sums: n − 1 neighbour dot products in one vectorised line. The rows are unit length, so those are cosines.>rather than>=against the cutoff, so the percentile itself is not a boundary; with few sentences this avoids breaking at every other gap.- Fewer than three sentences returns one chunk: there are not enough gaps to have a meaningful percentile.
- The regex sentence splitter is naive: "Dr. Rao" and "e.g. this" split wrongly. Use spaCy or
nltk.sent_tokenizewhen precision matters.
Complexity: one embedding call for n sentences; normalising and the neighbour products are O(n·d); the percentile is O(n log n) at most; the loop is O(n) apart from the length sum, which is O(chunk size). Space O(n·d).
A real-life example
A short news-plus-recipe text, with hand-made 2-dimensional embeddings (first number: sport, second: food):
1VECS = {"Kohli scored a century.": [1.0, 0.0],2 "India won the match by 6 wickets.": [0.9, 0.1],3 "The crowd in Mumbai celebrated.": [0.8, 0.3],4 "For biryani, soak the rice first.": [0.1, 1.0],5 "Add saffron milk before the dum.": [0.0, 1.0],6 "Serve it with raita.": [0.2, 0.9]}7text = " ".join(VECS)8for c in semantic_chunks(text, lambda ss: [VECS[s] for s in ss]):9 print(c)10# Kohli scored a century. India won the match by 6 wickets. The crowd in Mumbai celebrated.11# For biryani, soak the rice first. Add saffron milk before the dum. Serve it with raita.| gap after sentence | distance | above cutoff (0.347)? |
|---|---|---|
| 0 → 1 | 0.006 | no |
| 1 → 2 | 0.031 | no |
| 2 → 3 | 0.557 | yes — new chunk |
| 3 → 4 | 0.005 | no |
| 4 → 5 | 0.024 | no |
The 90th percentile of the five distances is 0.347 (NumPy interpolates between the two largest). One gap stands far above it, and that is exactly where cricket turns into cooking. A fixed 100-character cut would have produced a chunk mixing the celebration with the rice.
News aggregators and long meeting transcripts benefit most, because their topics change without headings.
Follow-up questions to expect
- "Is semantic chunking worth the cost?" — It needs an embedding pass over every sentence at index time. In many benchmarks it beats recursive chunking only modestly; if the documents have headings, splitting on those is usually as good and free. Measure before adopting.
- "What if the whole document is one topic?" — The percentile still marks the top 10% of gaps as breaks, so chunks stay bounded;
max_charscovers the rest. - "Why not compare each sentence with the current chunk's average instead of its neighbour?" — That is a good variant: it is more stable against one odd sentence, at the cost of a little more computation.