Fine-Tuning LLMs

Course Content

Fine-Tuning LLMs

6 sections · 52 lessons

What is contrastive learning and why is it central to training embedding models?


One batch of four query–answer pairs0.820.110.240.090.150.770.310.120.200.280.850.180.070.140.220.79refundfailed paydeliveryaddressrefund kabpaisa wapasorder latechange addressCosine similarities; InfoNCE is a softmax across each row.
Only the diagonal is labelled, yet every other cell in the row is a free negative — which is why batch size drives quality.

What you need to know

The loss, in words

For each query in a batch, compute its similarity to every passage in the batch. Treat this as a classification problem where the correct "class" is its own passage:

Text
loss = -log( exp(sim(q, p_true) / T) / sum over all passages j of exp(sim(q, p_j) / T) )

This is InfoNCE, a softmax cross-entropy over similarities. T is the temperature, usually around 0.05. A small temperature sharpens the softmax so the model is pushed hard to separate close candidates.

In-batch negatives

With a batch of 64 pairs, each query has 1 positive and 63 negatives for free: the other 63 passages. More negatives make the task harder and the embeddings better, so batch size is a key hyperparameter for embedding training. Tricks like GradCache (CachedMultipleNegativesRankingLoss in Sentence Transformers) let you use batches of thousands on one GPU by computing the loss in chunks.

Training with Sentence Transformers

Python
# sentence-transformers 6.xfrom datasets import Datasetfrom sentence_transformers import (SentenceTransformer, SentenceTransformerTrainer,                                   SentenceTransformerTrainingArguments)from sentence_transformers.losses import MultipleNegativesRankingLossmodel = SentenceTransformer("BAAI/bge-m3")train = Dataset.from_dict({"anchor": queries, "positive": answers})loss = MultipleNegativesRankingLoss(model)   # in-batch negatives; scale 20 = temperature 0.05args = SentenceTransformerTrainingArguments(    output_dir="support-embed", per_device_train_batch_size=64,    num_train_epochs=1, batch_sampler="no_duplicates")SentenceTransformerTrainer(model=model, args=args, train_dataset=train, loss=loss).train()

batch_sampler="no_duplicates" stops the same passage from appearing twice in one batch, where it would wrongly count as a negative for itself.

Why pairs are the right data

"These two texts mean the same" has no natural label, but pairs are everywhere: question and accepted answer, title and article, search query and clicked result, a sentence and its translation. Strong public embedding models are trained in two stages: first on hundreds of millions of such weak pairs, then on smaller, cleaner sets with hard negatives.

Collapse

If training goes wrong, every text can map to nearly the same vector, which makes every similarity high and retrieval useless. Enough negatives, a sensible temperature and diverse batches prevent it.

A real-life example

An online grocery app's Hindi customer-support bot uses retrieval to find the right help article. Customers write in Hindi, English and Hinglish ("refund kab milega", "paisa wapas nahi aaya"). An off-the-shelf multilingual embedding model finds the right article in the top 5 for 62% of queries on their test set.

They take 40,000 past chats where an agent linked a help article, and use (customer message, linked article) as pairs. One epoch with the code above, batch size 64, takes under an hour on one GPU. Top-5 accuracy rises to 81%, mostly on Hinglish queries, which the original model rarely saw. (Numbers are from this made-up scenario.)

Follow-up questions to expect

  • "Why does temperature matter?" — Too high and all candidates look similar, so the gradient is weak; too low and training becomes unstable and overreacts to noisy pairs.
  • "Bi-encoder or cross-encoder?" — A bi-encoder embeds query and passage separately, so passages can be indexed in advance; this is what contrastive training produces. A cross-encoder reads both together and is more accurate but too slow for the first search, so it is used to rerank.
  • "How do you evaluate an embedding model?" — Recall@k and nDCG@10 on your own labelled queries. Public MTEB scores help with shortlisting only.