Course Content
Fine-Tuning LLMs
6 sections · 52 lessons
What is contrastive learning and why is it central to training embedding models?
What you need to know
The loss, in words
For each query in a batch, compute its similarity to every passage in the batch. Treat this as a classification problem where the correct "class" is its own passage:
loss = -log( exp(sim(q, p_true) / T) / sum over all passages j of exp(sim(q, p_j) / T) )This is InfoNCE, a softmax cross-entropy over similarities. T is the temperature, usually around 0.05. A small temperature sharpens the softmax so the model is pushed hard to separate close candidates.
In-batch negatives
With a batch of 64 pairs, each query has 1 positive and 63 negatives for free: the other 63 passages. More negatives make the task harder and the embeddings better, so batch size is a key hyperparameter for embedding training. Tricks like GradCache (CachedMultipleNegativesRankingLoss in Sentence Transformers) let you use batches of thousands on one GPU by computing the loss in chunks.
Training with Sentence Transformers
1# sentence-transformers 6.x2from datasets import Dataset3from sentence_transformers import (SentenceTransformer, SentenceTransformerTrainer,4 SentenceTransformerTrainingArguments)5from sentence_transformers.losses import MultipleNegativesRankingLoss67model = SentenceTransformer("BAAI/bge-m3")8train = Dataset.from_dict({"anchor": queries, "positive": answers})9loss = MultipleNegativesRankingLoss(model) # in-batch negatives; scale 20 = temperature 0.0510args = SentenceTransformerTrainingArguments(11 output_dir="support-embed", per_device_train_batch_size=64,12 num_train_epochs=1, batch_sampler="no_duplicates")13SentenceTransformerTrainer(model=model, args=args, train_dataset=train, loss=loss).train()batch_sampler="no_duplicates" stops the same passage from appearing twice in one batch, where it would wrongly count as a negative for itself.
Why pairs are the right data
"These two texts mean the same" has no natural label, but pairs are everywhere: question and accepted answer, title and article, search query and clicked result, a sentence and its translation. Strong public embedding models are trained in two stages: first on hundreds of millions of such weak pairs, then on smaller, cleaner sets with hard negatives.
Collapse
If training goes wrong, every text can map to nearly the same vector, which makes every similarity high and retrieval useless. Enough negatives, a sensible temperature and diverse batches prevent it.
A real-life example
An online grocery app's Hindi customer-support bot uses retrieval to find the right help article. Customers write in Hindi, English and Hinglish ("refund kab milega", "paisa wapas nahi aaya"). An off-the-shelf multilingual embedding model finds the right article in the top 5 for 62% of queries on their test set.
They take 40,000 past chats where an agent linked a help article, and use (customer message, linked article) as pairs. One epoch with the code above, batch size 64, takes under an hour on one GPU. Top-5 accuracy rises to 81%, mostly on Hinglish queries, which the original model rarely saw. (Numbers are from this made-up scenario.)
Follow-up questions to expect
- "Why does temperature matter?" — Too high and all candidates look similar, so the gradient is weak; too low and training becomes unstable and overreacts to noisy pairs.
- "Bi-encoder or cross-encoder?" — A bi-encoder embeds query and passage separately, so passages can be indexed in advance; this is what contrastive training produces. A cross-encoder reads both together and is more accurate but too slow for the first search, so it is used to rerank.
- "How do you evaluate an embedding model?" — Recall@k and nDCG@10 on your own labelled queries. Public MTEB scores help with shortlisting only.