- MantraMindAI
- Courses
- AI Engineering & MLOps
- Local LLM Deployment and Quantization
Local LLM Deployment and Quantization
Course Overview
Advanced
Free Course
For engineers who want to run open-weight language models on their own hardware instead of calling a hosted API. You will be able to size a model to your memory, choose a quantisation, fine-tune and serve it locally, and measure its speed and quality honestly.
Instructor: Jaidev
Sections: 3
Course Content
Section 1: Running Models Locally
How local runtimes manage memory, what each GGUF quantisation level trades away, and how memory bandwidth sets your speed on CPU, GPU and Apple Silicon. 3 lessons, about 45 minutes.
Section 2: Deployment Strategies
Fine-tune a model on your own GPU, serve it behind an API that handles concurrency and caching, and measure its latency and resource costs honestly. 3 lessons, about 40 minutes.
Section 3: Mini Project
Size, install, benchmark and write up an 8B model running under Ollama on your own machine. 1 lesson, about 15 minutes.