Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your edge-deployed AI assistant must run offline on laptops with no GPU and 8GB RAM. How do you optimize models for constrained edge environments?
What you need to know
The memory maths
weights (bytes) = parameters x bits per weight / 8 3.8B params at ~4.5 bits (Q4_K_M) = about 2.1 GB 8B params at ~4.8 bits = about 4.8 GB (too tight for 8 GB total)KV cache per token = 2 x layers x kv_heads x head_dim x bytes Llama 3.2 3B (28 layers, 8 KV heads, head_dim 128, FP16) = about 112 KB per token at 8,000 tokens of context = about 0.9 GBThe KV cache grows with context length, so a long context can use nearly half as much memory as the weights. Models with grouped-query attention (fewer KV heads) keep it smaller, and capping context at 4K–8K tokens keeps it predictable.
Runtime by hardware
| Laptop | Runtime |
|---|---|
| Any CPU (Windows, Linux, Mac) | llama.cpp, or Ollama built on it |
| Apple silicon | MLX, or llama.cpp with Metal |
| Laptops with an NPU | ONNX Runtime or OpenVINO, where the model is supported |
1from llama_cpp import Llama23llm = Llama(model_path="field-assistant-3b-q4_k_m.gguf",4 n_ctx=4096, n_threads=4, use_mmap=True)5reply = llm.create_chat_completion(6 messages=[{"role": "user", "content": "What is the storage temperature for product X?"}],7 max_tokens=300)8print(reply["choices"][0]["message"]["content"])use_mmap=True maps the file into memory so pages load as needed, and n_ctx=4096 caps context and therefore KV-cache memory.
The plan
- Narrow the task — a small model tuned on one job, for example with LoRA on outputs from your cloud model (distillation), can match a much larger general model on that job.
- Quantise deliberately — 4-bit is usually the sweet spot; 3-bit and below often drop quality sharply. Test your own tasks.
- Control memory — memory-map weights, cap context, prefer grouped-query-attention models.
- Keep local search cheap — SQLite FTS5 for keyword search plus a small embedding model (tens of millions of parameters) so retrieval does not compete with the LLM for RAM.
- Go hybrid — when online, route hard questions to the cloud model; offline, answer locally and say when confidence is lower.
Measure tokens per second on the slowest supported machine, peak memory, first-token latency, and task accuracy against the cloud model.
A real-life example
Scenario, numbers made up. A pharma company gives field representatives an assistant for product questions in villages with no network. The laptops are 8 GB, CPU-only. A generic 8B model at 4-bit loads, but the machines swap memory and produce 1.5 tokens per second.
The team fine-tunes a 3.8B model with LoRA on 20,000 question-answer pairs generated and checked with their cloud model, then quantises it to Q4_K_M (2.3 GB). With a 4K context and a SQLite FTS5 index of product leaflets, peak memory is 3.4 GB and speed is 9 tokens per second on the slowest laptop. On a 300-question test it scores 91% of the cloud model's accuracy, which the team states plainly instead of promising parity. Online, dosage questions still go to the cloud model.
Follow-up questions to expect
- "Why not the biggest model that loads?" — "Loads" is not "runs well": with other apps open the laptop swaps, speed collapses, and the app feels broken.
- "How do you ship updates?" — As versioned model files downloaded when online, often a small LoRA adapter instead of the full model.
- "How do you protect the model file on the device?" — Assume it can be copied. Keep secrets and customer data out of it, and encrypt the local data store.