Course Content
LLMs Deep Dive
10 sections · 40 lessons
How does Gemini’s design boost training efficiency and stability vs other multimodal LLMs?
What you need to know
Two ways to build a multimodal model
Bolt-on (modular)
- Take a trained text LLM and a trained image encoder
- Train a small adapter to map image features into the LLM
- Cheap and fast to build; used by LLaVA-style open models
- The parts were never trained together, so alignment is shallow
Native (early fusion)
- One model, all modalities as tokens in one sequence
- Trained jointly from the start of pretraining
- More consistent reasoning across text, image and audio
- Expensive; the data mix must be balanced carefully
Gemini (announced in December 2023) was presented as native from the start. Other labs have moved the same way since — OpenAI described GPT-4o as one model trained end-to-end across text, vision and audio, and Meta described Llama 4 as using early fusion — so native multimodality is now common among frontier models rather than unique to Gemini.
What has been published about efficiency
- Sparse mixture-of-experts — the Gemini 1.5 report described an MoE transformer. Total capacity grows without every parameter running for every token.
- Long context — Gemini 1.5 shipped with a 1-million-token window, and its report showed recall tests at longer lengths in research settings. Later versions kept 1M-token windows and added built-in thinking.
- Hardware co-design — training runs on Google's TPUs with the JAX and Pathways software stack, which lets one job span many TPU pods.
What has been published about stability
Very large runs fail often: a chip dies, a network link drops, or a chip silently computes a wrong number. The Gemini 1.0 report described keeping redundant copies of the model state in memory so a job can resume quickly from a healthy replica rather than reloading a checkpoint from storage, and using deterministic replay to find the source of silent data corruption. Beyond this, standard practice applies: careful initialisation, normalisation and handling of loss spikes.
What is not public
Exact layer counts, number of experts, data mixture and most training recipes are not disclosed. In an interview, say "the report says X" and "I would expect Y", and keep the two apart.
Token cost of media
Media becomes tokens, and tokens cost money and context. The Gemini API documents a fixed token cost per image and per second of audio and video. At a few hundred tokens per second of video, a 10-minute video is on the order of 150,000 tokens before you have asked a question.
A real-life example
A law firm must review a 200-page scanned agreement with handwritten margin notes, plus a 40-minute recorded negotiation call. The old pipeline: OCR for the scan, a speech-to-text model for the call, then a text LLM. OCR dropped the handwritten notes and garbled a table of payment milestones, and the transcript lost who said what, so the summary missed a verbal concession.
With a native multimodal model, the team sends the page images and the audio directly. The model reads the handwriting in context ("Clause 9 — client disagrees, see email") and links a statement in the call to the clause it refers to. Two costs: the scan and audio together use hundreds of thousands of tokens, so each review costs noticeably more than the text pipeline, and the firm still has an associate verify every flagged clause against the original page.
Follow-up questions to expect
- "What is early fusion versus late fusion?" — Early fusion mixes modalities as tokens from the first layer; late fusion processes each modality separately and combines them near the end.
- "Why is native multimodal training harder?" — The data mix must balance modalities so one does not dominate, sequences get long (video), and the whole model must be trained at once at great cost.
- "How would you evaluate a multimodal model for your use case?" — Build a test set of your real documents or media with known answers, including the hard cases (handwriting, tables, accents), and compare against the pipeline you have.