Machine Learning System Design Interview

Course Content

Machine Learning System Design Interview

11 sections · 33 lessons

Visual search: the encoder and contrastive training


Only now, at step four of seven, does the architecture matter — and the answer is deliberately unadventurous. Attic needs one function from image to vector, used over 40 million catalogue images offline and over one query image at request time.

The architecture is the easy half of this lesson. The loss is what actually teaches the encoder what "similar" means, and everything else in training is support for it. The training pairs and hard negatives chosen in step 3 are the raw material; this lesson turns them into an embedding space.

The shape

One encoder. Image in, fixed-length vector out. The same encoder runs over the whole catalogue offline and over the query image at request time, which guarantees both live in the same space.

Two families of image encoder are appropriate, and the choice between them is a trade-off rather than a verdict.

Convolutional encoderVision transformer encoder
How it sees the imageLocal patterns composed into larger ones through stacked filtersImage split into patches, all patches attend to each other
Data appetiteWorks well with moderate dataNeeds more data or heavy pretraining
Inference cost at small sizesLowerHigher
StrengthTexture and local detail — wood grain, fabric weaveGlobal layout and long-range relationships
Reasonable defaultYes, for a first systemYes, if a strong pretrained backbone is available

Deliberately no model names or version numbers here. Both families are actively developed and any specific name will date. What does not date: convolutional encoders are cheaper and texture-biased, transformer encoders are hungrier and better at global structure.

Transfer learning is the right default

Train the encoder from scratch on Attic's 40 million images and you will spend weeks of accelerator time relearning edges, corners, and textures — things every image model needs and that any large-scale pretrained backbone already has.

Instead: take a backbone pretrained on a large general image corpus, replace its final classification layer with a projection to the embedding size, and fine-tune. In practice this gets a usable model in days rather than weeks and reaches better quality, because the pretrained features were learned from far more visual variety than one marketplace contains.

Two fine-tuning schedules, both defensible:

  • Freeze most of the backbone, train the projection head and the last block. Fast, low risk of overfitting, and the right choice with fewer than a few million training pairs.
  • Fine-tune everything at a low learning rate. Better final quality when you have tens of millions of pairs, at more cost and more risk of destroying the pretrained features early in training — which is why a warm-up period matters.

Embedding size

A real decision with a cost attached, not a hyperparameter to shrug at.

DimensionIndex memory for 40M items (float32)Notes
128~20 GBFits comfortably; some fine detail lost
256~41 GBA common sweet spot
512~82 GBNeeds sharding or compression
1024~164 GBRarely worth it for retrieval

Those figures are arithmetic — dimension × 4 bytes × 40 million — not benchmarks. The pattern across most retrieval systems is that recall improves steeply up to a few hundred dimensions and then flattens while memory keeps growing linearly. Start at 256 and measure whether 512 buys anything.

INPUTQuery imageEncodera vision backboneEmbedding256 floatsVector indexHNSW or IVF+PQ40M item vectorsTop-k neighboursforward passnearest neighbour searchthe same encoder must produce the index andthe query vector, or the geometry does not lineupthe index is rebuilt offline; queriesonly read it
Search quality is decided by the encoder, and search latency by the index — they are separate problems with separate fixes.

Contrastive loss, explained from scratch

With the encoder's shape settled, the loss decides what it learns.

Two losses that both teach distanceTriplet loss• Anchor, positive,negative, plus a margin• One negative per step, so mining matters• Collapses if the triplets are too easyIn-batch softmax• Every other item inthe batch is negative• Signal grows with batch size, not mining• Temperature sets howsharp the contrast is
The in-batch version usually wins because a large batch gives free hard negatives that mining has to work for.

Imagine arranging photographs on a very large table. Two photos of the same chair should end up touching. A photo of a chair and a photo of a bicycle should end up at opposite ends. You repeat this millions of times, nudging each photo a little closer to its match and a little further from everything else, until the table's layout encodes similarity as distance.

That is contrastive learning. The loss function is the nudging rule.

The version used in practice works over a batch. Take a batch of 512 image pairs. For each anchor, its own augmented partner is the one positive; the other 1,022 images in the batch are negatives. The loss asks the model to make the anchor–positive similarity high relative to the anchor–negative similarities — technically, it is a cross-entropy over similarity scores, with the positive as the correct class.

A concrete step. Similarities are cosine values between −1 and 1:

  • anchor·positive = 0.62
  • anchor·hardest negative = 0.58
  • everything else ≤ 0.30

The positive wins, but barely. The loss is large, and the gradient pushes the anchor and positive together while pushing the anchor and that hard negative apart. Repeat over millions of batches and the geometry settles.

Where the analogy breaks. On a real table, moving one photo does not move the others. In an embedding space every update moves the shared encoder, so pulling one pair together nudges every image that shares visual features with them. That is why the model generalises to items it never saw — and also why one badly-labelled pair can degrade a whole region of the space.

Triplet loss, and why the batch version usually wins

Triplet loss takes an anchor, one positive, and one negative, and demands that the anchor–positive distance be smaller than the anchor–negative distance by at least a fixed margin — say 0.2. If the gap already exceeds the margin, the loss for that triplet is zero and nothing is learned from it.

That last property is the problem. After early training, most randomly sampled triplets are already satisfied and contribute nothing. You spend most of your compute on examples with zero gradient, which is why triplet training needs aggressive mining to work at all.

Batch contrastive loss sidesteps this by using every other item in the batch as a negative: with 512 pairs you get 1,022 negatives per anchor for the cost of one forward pass. Bigger batches are therefore directly better here, which is an unusual property and a good thing to mention.

Batch construction

Three choices that matter more than the learning rate:

  1. Batch size as large as memory allows. More in-batch negatives means a harder and more informative task. Where memory is the limit, a memory bank of recent embeddings extends the negative pool at low cost.
  2. Do not put two photos of the same listing in one batch as separate anchors. They will be treated as negatives for each other, which teaches the model something false. Deduplicate by listing ID when sampling.
  3. Mix in category-restricted batches. Batches drawn from a single category — all chairs — make every in-batch negative a hard negative at no extra cost.

The temperature parameter

The contrastive loss divides similarities by a small constant before the cross-entropy. Low values (around 0.05) sharpen the distribution and make the model focus hard on the most confusable negatives. High values (around 0.5) flatten it and treat all negatives more equally.

It is one of the few hyperparameters in this system that genuinely changes behaviour rather than convergence speed, and it is worth naming for that reason. Too low and training becomes unstable and over-fixated on a handful of near-duplicates; too high and the space never tightens.

Evaluation split

Split by time, per the training step of the framework: train on listings created before a cut-off date, evaluate on listings created after. A random split lets the model see other photos of the same listing during training, which inflates recall dramatically and measures memorisation.