LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

What are Foundation Models?


What you need to know

Three defining traits

  • Broad pretraining — web text, code, images, audio, not one narrow dataset.
  • General-purpose — no single task is baked in.
  • Adaptable — the base is reused, and effort goes into adapting it.

How you adapt a foundation model (cheapest first)

  1. Prompting — instructions and a few examples. Minutes of work, no training.
  2. Retrieval (RAG) — fetch relevant documents and put them in the prompt. Adds fresh and private knowledge.
  3. Parameter-efficient fine-tuning (LoRA) — train a small adapter for style, format or domain language.
  4. Full fine-tuning — update all weights. Expensive; needed rarely.
  5. Distillation — train a small model to copy the big one for a cheap, fast task model.

The rule of thumb: go down this list only when the step above has been measured and is not good enough.

Why the idea mattered

Before foundation models, a company built a sentiment model, a translation model and an entity extractor separately, each with its own labelled data. Now one pretrained model covers all three with prompts, and the budget moves from data labelling to evaluation and integration.

The risks

  • Single point of failure — a bias or weakness in the base model shows up in every product on top of it.
  • Opacity — training data is often undisclosed, so licensing and contamination are hard to check.
  • Dependence — a small number of providers control the strongest models, and a version change can shift your product's behaviour.

A real-life example

A law firm needs three things: summaries of 40-page contracts, extraction of dates and parties, and tagging of risky clauses. In the old approach, each was a separate model and a separate labelling project.

With a foundation model, the team starts with prompts for all three. The summaries are good. Extraction works after a JSON schema is added to the prompt. Clause tagging is weaker because the firm uses its own risk categories, so they label 2,000 clauses and train a LoRA adapter — a few hours on one GPU. Total: one base model, one adapter, three tasks. When the base model gets a new version, they rerun a 300-case evaluation set before switching, because the firm's work now depends on it.

Follow-up questions to expect

  • "Is every LLM a foundation model?" — Most large pretrained ones are. A small model trained only for one task is not, even if it is a transformer.
  • "When would you fine-tune instead of prompting?" — When prompting plus retrieval has been measured and still misses a consistent format, style or domain skill, or when a smaller fine-tuned model can replace an expensive large one.
  • "What is a multimodal foundation model?" — One pretrained on several kinds of data, such as text, images and audio, so it can take or produce more than one type.