Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
What are “hard negatives” in embedding training and how do they improve retrieval quality?
What you need to know
Why easy negatives stop teaching
The contrastive loss is a softmax over similarities. If all negatives are about other topics, the correct passage already wins by a large margin, the softmax gives it probability near 1, and the loss and gradient are near 0. The model looks trained, but it has only learned topic matching. Real retrieval failures happen among 20 passages on the right topic, where only one answers the question.
How to mine them
- Retrieve — run your current retriever (BM25 or the current embedding model) for each training query.
- Skip the top — ignore the first few ranks, where unlabelled correct answers are most likely.
- Filter — drop candidates scoring too close to the positive, or ask a cross-encoder or LLM whether they actually answer the query.
- Train — add 1 to 5 hard negatives per query to the in-batch negatives.
- Repeat — re-mine with the improved model (the ANCE idea), so negatives get harder as the model improves.
1# sentence-transformers 6.x2from sentence_transformers import SentenceTransformer3from sentence_transformers.util import mine_hard_negatives45model = SentenceTransformer("BAAI/bge-m3")6with_negs = mine_hard_negatives(7 pairs, # Dataset with "anchor" and "positive" columns8 model,9 range_min=10, range_max=50, # skip ranks 1-10, look at ranks 11-5010 relative_margin=0.05, # negative must score at least 5% below the positive11 num_negatives=3,12 use_faiss=True,13)The result has anchor, positive and three negatives per row, ready for MultipleNegativesRankingLoss, which uses the extra columns as additional negatives. You can also pass a cross_encoder to rescore candidates before choosing.
The false-negative trap
A corpus often has several passages that answer the same query, but only one is labelled. The top-ranked "negatives" are often these unlabelled answers. Training on them teaches the model to push correct answers away. That is why step 2 and step 3 exist, and why negatives that are too hard can make results worse.
A real-life example
A Mumbai law firm builds a clause search tool over 200,000 clauses from past contracts. A lawyer searches "termination for convenience with 30 days' notice". A typical hard negative is "termination for cause with a 30-day cure period" — nearly the same words, opposite meaning.
Training with in-batch negatives only gives 70% top-5 accuracy on 400 lawyer-labelled queries. The first hard-negative run takes ranks 1–5 as negatives and drops to 66%. A lawyer samples 100 of those negatives and finds that 31 are actually valid answers — near-duplicate clauses from different contracts. The second run skips ranks 1–10, uses a relative margin, and checks candidates with a cross-encoder. Top-5 accuracy rises to 82%. (Made-up numbers for illustration.)
Follow-up questions to expect
- "How many hard negatives per query?" — Usually 1 to 5, on top of in-batch negatives. More costs memory and raises the false-negative risk.
- "Do rerankers use hard negatives too?" — Yes, even more so. A cross-encoder reranker only ever sees top candidates in production, so it must be trained on them.
- "What if you have no labelled positives?" — Generate synthetic queries for each passage with an LLM, then treat (query, source passage) as the positive pair and mine negatives the same way.