Local LLM Deployment and Quantization

Course Overview
Advanced
Free Course

For engineers who want to run open-weight language models on their own hardware instead of calling a hosted API. You will be able to size a model to your memory, choose a quantisation, fine-tune and serve it locally, and measure its speed and quality honestly.

Instructor: Jaidev
Sections: 3

Course Content

Section 1: Running Models Locally

How local runtimes manage memory, what each GGUF quantisation level trades away, and how memory bandwidth sets your speed on CPU, GPU and Apple Silicon. 3 lessons, about 45 minutes.

Section 2: Deployment Strategies

Fine-tune a model on your own GPU, serve it behind an API that handles concurrency and caching, and measure its latency and resource costs honestly. 3 lessons, about 40 minutes.

Section 3: Mini Project

Size, install, benchmark and write up an 8B model running under Ollama on your own machine. 1 lesson, about 15 minutes.