Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Implement semantic chunking based on meaning.


Distance between neighbouring sentences0.0060.0310.5570.0050.02401234cricket, cricketover 0.347:new chunkbiryani, biryaniThe cutoff is the 90th percentile of this document's own gaps.
One gap towers over the rest, and that is exactly where the text turns from the match to the recipe.

What you need to know

Fixed and recursive chunking cut by length. Semantic chunking cuts by meaning: a chunk should hold one topic, and the boundary should be where the text moves to the next one.

The signal is the distance between neighbouring sentences:

Text
distance[i] = 1 − cosine(embedding(sentence i), embedding(sentence i+1))

Within a topic, neighbouring sentences are similar (small distance). At a topic change the distance jumps.

Percentile threshold. Some documents are written tightly, others loosely, so an absolute cutoff like 0.3 breaks one document everywhere and another nowhere. Taking the 90th percentile of this document's distances means "break at the biggest 10% of jumps".

Off-by-one. When you look at sentence i + 1 and decide whether it starts a new chunk, the gap to check is distance[i] — between sentence i and sentence i + 1.

Python
import reimport numpy as npdef split_sentences(text: str) -> list[str]:    return [s.strip() for s in re.split(r"(?<=[.!?])\s+", text) if s.strip()]def semantic_chunks(text: str, embed_fn, percentile: float = 90,                    max_chars: int = 1500) -> list[str]:    """Start a new chunk where neighbouring sentences are unusually far apart."""    sentences = split_sentences(text)    if len(sentences) < 3:        return [" ".join(sentences)] if sentences else []    v = np.asarray(embed_fn(sentences), dtype=np.float32)    v /= np.linalg.norm(v, axis=1, keepdims=True) + 1e-10    distance = 1.0 - np.sum(v[:-1] * v[1:], axis=1)       # gap after each sentence    cutoff = np.percentile(distance, percentile)    chunks, current = [], [sentences[0]]    for i, sentence in enumerate(sentences[1:]):          # sentence is sentences[i + 1]        too_long = sum(map(len, current)) + len(sentence) > max_chars        if distance[i] > cutoff or too_long:            chunks.append(" ".join(current))            current = []        current.append(sentence)    chunks.append(" ".join(current))    return chunks

The tricky parts:

  • np.sum(v[:-1] * v[1:], axis=1) multiplies each row by the next row element-wise and sums: n − 1 neighbour dot products in one vectorised line. The rows are unit length, so those are cosines.
  • > rather than >= against the cutoff, so the percentile itself is not a boundary; with few sentences this avoids breaking at every other gap.
  • Fewer than three sentences returns one chunk: there are not enough gaps to have a meaningful percentile.
  • The regex sentence splitter is naive: "Dr. Rao" and "e.g. this" split wrongly. Use spaCy or nltk.sent_tokenize when precision matters.

Complexity: one embedding call for n sentences; normalising and the neighbour products are O(n·d); the percentile is O(n log n) at most; the loop is O(n) apart from the length sum, which is O(chunk size). Space O(n·d).

A real-life example

A short news-plus-recipe text, with hand-made 2-dimensional embeddings (first number: sport, second: food):

Python
VECS = {"Kohli scored a century.": [1.0, 0.0],        "India won the match by 6 wickets.": [0.9, 0.1],        "The crowd in Mumbai celebrated.": [0.8, 0.3],        "For biryani, soak the rice first.": [0.1, 1.0],        "Add saffron milk before the dum.": [0.0, 1.0],        "Serve it with raita.": [0.2, 0.9]}text = " ".join(VECS)for c in semantic_chunks(text, lambda ss: [VECS[s] for s in ss]):    print(c)# Kohli scored a century. India won the match by 6 wickets. The crowd in Mumbai celebrated.# For biryani, soak the rice first. Add saffron milk before the dum. Serve it with raita.
gap after sentencedistanceabove cutoff (0.347)?
0 → 10.006no
1 → 20.031no
2 → 30.557yes — new chunk
3 → 40.005no
4 → 50.024no

The 90th percentile of the five distances is 0.347 (NumPy interpolates between the two largest). One gap stands far above it, and that is exactly where cricket turns into cooking. A fixed 100-character cut would have produced a chunk mixing the celebration with the rice.

News aggregators and long meeting transcripts benefit most, because their topics change without headings.

Follow-up questions to expect

  • "Is semantic chunking worth the cost?" — It needs an embedding pass over every sentence at index time. In many benchmarks it beats recursive chunking only modestly; if the documents have headings, splitting on those is usually as good and free. Measure before adopting.
  • "What if the whole document is one topic?" — The percentile still marks the top 10% of gaps as breaks, so chunks stay bounded; max_chars covers the rest.
  • "Why not compare each sentence with the current chunk's average instead of its neighbour?" — That is a good variant: it is more stable against one odd sentence, at the cost of a little more computation.