Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

What are “hard negatives” in embedding training and how do they improve retrieval quality?


Candidates the retriever ranks for one lawyer's query0.910.890.870.740.710.690.310123456maybe areal answerhard negativeeasy negativeSkip the top ranks, then take negatives from ranks 11 to 50.
The useful negatives sit just below the top: close enough to teach, far enough not to be unlabelled correct answers.

What you need to know

Why easy negatives stop teaching

The contrastive loss is a softmax over similarities. If all negatives are about other topics, the correct passage already wins by a large margin, the softmax gives it probability near 1, and the loss and gradient are near 0. The model looks trained, but it has only learned topic matching. Real retrieval failures happen among 20 passages on the right topic, where only one answers the question.

How to mine them

  1. Retrieve — run your current retriever (BM25 or the current embedding model) for each training query.
  2. Skip the top — ignore the first few ranks, where unlabelled correct answers are most likely.
  3. Filter — drop candidates scoring too close to the positive, or ask a cross-encoder or LLM whether they actually answer the query.
  4. Train — add 1 to 5 hard negatives per query to the in-batch negatives.
  5. Repeat — re-mine with the improved model (the ANCE idea), so negatives get harder as the model improves.
Python
# sentence-transformers 6.xfrom sentence_transformers import SentenceTransformerfrom sentence_transformers.util import mine_hard_negativesmodel = SentenceTransformer("BAAI/bge-m3")with_negs = mine_hard_negatives(    pairs,                    # Dataset with "anchor" and "positive" columns    model,    range_min=10, range_max=50,   # skip ranks 1-10, look at ranks 11-50    relative_margin=0.05,         # negative must score at least 5% below the positive    num_negatives=3,    use_faiss=True,)

The result has anchor, positive and three negatives per row, ready for MultipleNegativesRankingLoss, which uses the extra columns as additional negatives. You can also pass a cross_encoder to rescore candidates before choosing.

The false-negative trap

A corpus often has several passages that answer the same query, but only one is labelled. The top-ranked "negatives" are often these unlabelled answers. Training on them teaches the model to push correct answers away. That is why step 2 and step 3 exist, and why negatives that are too hard can make results worse.

A real-life example

A Mumbai law firm builds a clause search tool over 200,000 clauses from past contracts. A lawyer searches "termination for convenience with 30 days' notice". A typical hard negative is "termination for cause with a 30-day cure period" — nearly the same words, opposite meaning.

Training with in-batch negatives only gives 70% top-5 accuracy on 400 lawyer-labelled queries. The first hard-negative run takes ranks 1–5 as negatives and drops to 66%. A lawyer samples 100 of those negatives and finds that 31 are actually valid answers — near-duplicate clauses from different contracts. The second run skips ranks 1–10, uses a relative margin, and checks candidates with a cross-encoder. Top-5 accuracy rises to 82%. (Made-up numbers for illustration.)

Follow-up questions to expect

  • "How many hard negatives per query?" — Usually 1 to 5, on top of in-batch negatives. More costs memory and raises the false-negative risk.
  • "Do rerankers use hard negatives too?" — Yes, even more so. A cross-encoder reranker only ever sees top candidates in production, so it must be trained on them.
  • "What if you have no labelled positives?" — Generate synthetic queries for each passage with an LLM, then treat (query, source passage) as the positive pair and mine negatives the same way.