Course Content
Capstone Project: Multimodal Assistant
1 sections · 6 lessons
Project Setup & Architecture
It is week six of the build. Your assistant is roughly 2,300 lines spread across six Python files. You run it, say "what's in this photo", and get this:
Traceback (most recent call last): File "main.py", line 41, in <module> assistant.run() File "src/assistant.py", line 88, in run return self.process(user_input) File "src/assistant.py", line 112, in process result = self.handle(payload)Exception: 'NoneType' object has no attribute 'strip'Which subsystem failed? The microphone capture, the transcription model, the caption model, the LLM call, or the vector store? You cannot tell. There is no log file, because everything went to print() and your terminal scrollback ended forty minutes ago. You cannot bisect it against a known-good state, because there are no tests. And when you finally do find it — a caption model returning None for a CMYK JPEG — you fix it, and two days later it silently comes back, because nothing checks.
That is not a bad-luck day. That is the predictable output of skipping the boring hour at the start. This stage is that hour. You will write almost no "AI" code here. You will write a directory layout, a pinned dependency set, a configuration loader that refuses to start when something is missing, a logger, an exception hierarchy, an entry point, and five tests. Every later stage of this project reuses all six of those, unchanged.
In a multi-component system the cost of a missing foundation does not appear on day one. It appears in week six, multiplied by every component you added since.
What you are actually building
The finished system is a multimodal assistant: it accepts typed text, an image file, or spoken audio; it reasons about the request, calling real tools when it needs a real answer; it remembers what you told it last week; and it replies in text and optionally in speech. Six subsystems, one orchestrator, two cross-cutting utilities.
INPUT ADAPTERS ORCHESTRATOR SERVICES ------------------ ---------------- --------------------------- voice_handler.py ----> +------------------+ ----> text_processor.py mic / wav -> text | | prompts + LLM calls | MultimodalAssistant vision_handler.py ----> | (assistant.py) | ----> reasoning_engine.py image -> caption | | ReAct loop + JSON schema | routes input, | raw typed text ----> | owns one of | ----> memory_manager.py | each subsystem | history + vector search +------------------+ CROSS-CUTTING (built in this stage, used by everything above): config/settings.py validated configuration, loaded once at import src/utils.py logger, shared helpers exception hierarchy every subsystem raises its own named error typeThe important property of that diagram is the arrows that are not there. vision_handler.py never imports memory_manager.py. voice_handler.py has no idea the reasoning engine exists. Only the orchestrator imports more than one subsystem. That single rule is what lets you test the vision pipeline with no microphone attached, swap the LLM provider without touching a prompt, and — crucially — read a stack trace and know immediately which box broke.
Before you start
- Python 3.12 or 3.13. Check with
python3 --version. The heavy ML wheels in this project (torch, whisper's dependencies, sentence-transformers) have historically lagged a brand-new interpreter by several months, so pick a Python release that has been out for about a year. On a release that is too new you will spend your first afternoon compiling C extensions from source instead of building an assistant. - An LLM API key for whichever provider you use for text generation and embeddings.
- At least 2 GB free disk. The virtual environment plus the models later stages download is not small: a CPU-only torch wheel is roughly 200 MB, Whisper's
basecheckpoint about 139 MB, a sentence-transformer around 90 MB, a caption model a few hundred MB more. - Terminal and git basics. No Docker, Redis, or GPU is needed yet.
Layout first, code second
A flat folder of .py files is genuinely fine at 200 lines. It stops being fine at about 800, which this project passes in its third stage. The layout below is not aesthetic preference; each directory is a boundary that prevents a specific class of mistake.
mkdir -p multimodal-assistant/{src,tests,config,prompts,examples,logs,data}cd multimodal-assistanttouch src/__init__.py tests/__init__.py config/__init__.pymultimodal-assistant/|-- src/| |-- __init__.py| |-- assistant.py orchestrator| |-- voice_handler.py speech-to-text, text-to-speech| |-- vision_handler.py embeddings + captioning| |-- text_processor.py prompt templates + LLM calls| |-- reasoning_engine.py ReAct loop + structured output| |-- memory_manager.py conversation history + vector store| |-- exceptions.py the error hierarchy| `-- utils.py logging and shared helpers|-- config/| |-- __init__.py| `-- settings.py validated configuration|-- tests/| |-- conftest.py shared pytest fixtures| |-- test_settings.py| `-- test_utils.py|-- prompts/ prompt text, kept out of logic|-- examples/ sample images, sample queries|-- logs/ generated, gitignored|-- data/ vector store, sessions; gitignored|-- main.py|-- requirements.txt|-- .env.example|-- .gitignore`-- README.md| Directory | Holds | The mistake it prevents |
|---|---|---|
src/ | One module per subsystem | Vision code quietly growing inside reasoning code, so neither can be tested or replaced alone |
config/ | All configuration | Twenty scattered os.environ.get() calls, each with a different default, none validated |
tests/ | A file mirroring each src/ module | Not noticing which subsystem has zero coverage — the mirror makes gaps visible at a glance |
prompts/ | Prompt strings and tool schemas | Re-testing application logic because you changed the wording of a sentence |
examples/ | Sample inputs | Hunting for a test image every time you want to reproduce a bug |
logs/, data/ | Runtime output | Committing a 40 MB vector index, or a log file containing a user's transcript |
Write .gitignore now, before you create .env and before the first run generates a log:
1venv/2__pycache__/3*.pyc4.env5logs/6data/7*.egg-info/8.pytest_cache/The ordering matters more than it looks. Committing .env once puts the key in git history permanently; deleting the file in a later commit does not remove it. The fix is then key rotation and a history rewrite, which is a bad afternoon. Ignoring the file before it exists costs nothing.
An environment you can reproduce
A virtual environment is a private, disposable copy of Python and its packages that belongs to this project only. Without one, installing torch==2.1.1 here can break a different project on your machine that needed 2.0.0, and you will never be able to state what a fresh clone actually requires.
1python3 -m venv venv2source venv/bin/activate # macOS / Linux3# venv\Scripts\activate # Windows45which python # must print a path inside ./venv/6pip install --upgrade pipNow pin everything with ==, never >=:
# versions current in September 2026 -- bump them deliberately, never by accident# corepython-dotenv==1.2.3pydantic==2.13.5pydantic-settings==2.15.0# language modelsopenai==3.19.2langchain-openai==1.6.5# visionpillow==12.3.0transformers==5.17.0torch==2.14.0torchvision==0.29.0# audioopenai-whisper==20250625pyttsx3==2.99pyaudio==0.2.14# memory and searchchromadb==1.5.9sentence-transformers==6.1.0# api and servingfastapi==0.141.1uvicorn==0.53.0# testing and toolingpytest==9.1.1pytest-cov==7.1.0black==26.5.1Here is the arithmetic that makes == non-negotiable. That file lists 19 direct dependencies, which pull in well over 100 packages in total on Linux. Suppose each package ships a release, on average, once every six weeks — a modest rate for this ecosystem. Over a three-month project that is about two releases each, so more than 200 opportunities for a version to change under you. If just one in fifty of those releases contains a breaking rename, you should expect four or five silent breakages across the project's life, each appearing on a random morning with no code change of your own to blame.
Pinning only the direct dependencies does not get that number to zero, because the packages they pull in still float. A classic case: openai client releases from 2023 declared only httpx<1, so when httpx 0.28 shipped in late 2024, fresh installs of those pinned clients picked it up — and it had removed an argument the old clients passed, so every call failed with a TypeError. So once the install works, freeze the whole resolved set with pip freeze > requirements.lock (or use a lock-file tool such as uv or pip-tools) and install from the lock. Now nothing changes until you decide to bump a version deliberately.
Install and expect it to take a while: pip install -r requirements.txt downloads several hundred megabytes and typically runs 5 to 15 minutes. You pay that once.
Configuration that fails in 0.3 seconds, not 45
The naive approach is os.environ.get("OPENAI_API_KEY") wherever a key is needed. It has three failure modes, all of them quiet. A typo in the variable name returns None rather than raising. A missing value surfaces as an authentication error deep in a library, not as "you forgot to set a key". And a numeric setting read from the environment arrives as the string "0.7", which silently becomes a temperature of nonsense somewhere downstream.
The concrete cost: your assistant loads a Whisper checkpoint and a caption model at startup. That is roughly 45 seconds of model loading before the first LLM call is attempted. With scattered environ.get, a missing key means you wait the full 45 seconds and then get an auth failure. With a validated settings object imported at module load, you fail in about 0.3 seconds with a message naming the exact field. Over a debugging session where you restart thirty times, that is 22 minutes of your life against 9 seconds.
1# config/settings.py2from pathlib import Path3from pydantic import Field, field_validator4from pydantic_settings import BaseSettings, SettingsConfigDict567class Settings(BaseSettings):8 """Every configurable value in the project, validated at import time."""910 model_config = SettingsConfigDict(11 env_file=".env", env_file_encoding="utf-8", extra="ignore"12 )1314 # secrets -- no defaults, so a missing value is a hard error15 openai_api_key: str = Field(..., min_length=20)1617 # model choices -- a small, cheap default; check your provider's current list18 text_model: str = "gpt-6-luna"19 whisper_model: str = "base"20 embedding_model: str = "all-MiniLM-L6-v2"2122 # generation behaviour. None means "not sent": many current reasoning23 # models reject temperature, so only set it for a model that accepts it.24 temperature: float | None = Field(None, ge=0.0, le=2.0)25 # on reasoning models this budget also covers hidden reasoning tokens26 max_tokens: int = Field(4000, gt=0, le=32000)2728 # paths29 data_dir: Path = Path("data")30 log_dir: Path = Path("logs")31 log_level: str = "INFO"3233 @field_validator("log_level")34 @classmethod35 def _valid_level(cls, v: str) -> str:36 allowed = {"DEBUG", "INFO", "WARNING", "ERROR", "CRITICAL"}37 v = v.upper()38 if v not in allowed:39 raise ValueError(f"log_level must be one of {sorted(allowed)}")40 return v414243settings = Settings()44settings.data_dir.mkdir(parents=True, exist_ok=True)45settings.log_dir.mkdir(parents=True, exist_ok=True)Three things are doing real work here. Field(...) with no default makes the key mandatory: import fails immediately if it is absent. ge/le/gt bounds catch a temperature of 7.0 — a plausible typo for 0.7 — at startup rather than as incoherent model output an hour later. And the type annotations coerce: a TEMPERATURE of "0.7" arrives as a real float, data_dir as a real Path, so no caller has to remember to cast.
Temperature defaults to None for a reason. Many current models, including the reasoning models most providers now recommend, reject temperature and top_p or ignore them, so sending 0.7 by habit can turn every call into an HTTP 400. Leave it unset unless your chosen model's documentation says it is supported.
Commit .env.example with the shape but not the values, so a fresh clone knows what to fill in:
1OPENAI_API_KEY=sk-replace-me2TEXT_MODEL=gpt-6-luna3# TEMPERATURE=0.7 only for a model that accepts it4LOG_LEVEL=INFOLogging, because print does not survive contact with reality
print() has no severity, no timestamp, no source, and no persistence. When a transcription fails at 23:14 during an unattended run, a print statement has already scrolled away. A log line is still on disk tomorrow, stamped with the module that emitted it.
1# src/utils.py2import logging3from logging.handlers import RotatingFileHandler4from config.settings import settings567def get_logger(name: str = "assistant") -> logging.Logger:8 log = logging.getLogger(name)9 if log.handlers: # never attach handlers twice10 return log1112 log.setLevel(logging.DEBUG) # handlers decide what actually gets written1314 fmt = logging.Formatter(15 "%(asctime)s | %(levelname)-8s | %(name)s:%(lineno)d | %(message)s",16 datefmt="%Y-%m-%d %H:%M:%S",17 )1819 console = logging.StreamHandler()20 console.setLevel(getattr(logging, settings.log_level))21 console.setFormatter(fmt)2223 disk = RotatingFileHandler(24 settings.log_dir / "assistant.log",25 maxBytes=10_000_000, # 10 MB per file26 backupCount=3, # keep 3 rotations -> 40 MB ceiling27 encoding="utf-8",28 )29 disk.setLevel(logging.DEBUG) # disk keeps everything30 disk.setFormatter(fmt)3132 log.addHandler(console)33 log.addHandler(disk)34 return log353637logger = get_logger()The split levels are the point: the console shows you INFO and above so it stays readable while you work, while the file records DEBUG so the detail is there when you need to reconstruct a failure after the fact. Size the rotation with real arithmetic. One conversational turn through the full pipeline emits roughly 40 DEBUG lines at about 120 bytes each, so about 4.8 KB per turn. At 500 turns of testing per day that is 2.4 MB per day, so a 10 MB file covers about four days and the three backups extend that to roughly sixteen. That is the right ballpark: long enough to cover a weekend of unattended running, small enough that nobody ever notices the disk usage.
| Level | Means | In this project |
|---|---|---|
DEBUG | Detail only useful when diagnosing | Prompt text sent, token counts, similarity scores, audio sample counts |
INFO | Normal progress worth recording | "Whisper base loaded", "indexed 12 documents", "turn completed in 2.4 s" |
WARNING | Degraded but still working | Retrying an API call, falling back to offline TTS, image resized from 4096 px |
ERROR | This operation failed | Transcription raised, JSON validation failed after retries |
CRITICAL | The process cannot continue | Configuration invalid, vector store unreadable |
The failure mode people hit here is logging the wrong thing. Never log an API key, and never log a full audio buffer or base64 image — one logger.debug(f"payload={payload}") on an image request can write 3 MB into a single log line and blow through your rotation in one turn. Log the shape, not the content: logger.debug("image %s, %dx%d, %d bytes", path.name, w, h, size).
An exception hierarchy that names the culprit
The traceback at the top of this lesson ended in a bare Exception. That is the direct consequence of a subsystem that catches everything and re-raises nothing meaningful. The fix is a small tree of typed errors, one per subsystem.
1# src/exceptions.py2class AssistantError(Exception):3 """Base class. Catching this catches everything we raise deliberately."""456class ConfigurationError(AssistantError):7 """Settings missing or invalid."""8910class VoiceError(AssistantError):11 """Microphone capture, transcription, or speech synthesis failed."""121314class VisionError(AssistantError):15 """Image loading, embedding, or captioning failed."""161718class TextProcessingError(AssistantError):19 """LLM call or prompt construction failed."""202122class ReasoningError(AssistantError):23 """The reasoning loop failed or produced unusable output."""242526class MemoryError_(AssistantError):27 """Persistence or vector search failed."""Note the trailing underscore on the last one: Python already has a built-in MemoryError for out-of-memory conditions, and shadowing it would make a genuine allocation failure look like a database problem. Small detail, real bug avoided.
Now compare the two ways of using them. The wrong way:
1try:2 caption = self.captioner(image)3except Exception:4 return None # caller now gets None and has no idea whyEvery downstream 'NoneType' object has no attribute 'strip' in this project traces back to a block shaped like that. Three problems compound: the error type is discarded, the original traceback is discarded, and the failure is converted into a valid-looking return value that travels several frames before exploding somewhere unrelated. The right way:
1try:2 caption = self.captioner(image)3except (OSError, ValueError) as e:4 logger.error("Captioning failed for %s: %s", image_path, e)5 raise VisionError(f"could not caption {image_path}: {e}") from eCatch the specific exceptions you expect, log with context, raise a named error, and keep the cause with from e so the original traceback stays attached. Now the orchestrator can write except VisionError: and degrade gracefully — answer the text part of the question, tell the user the image could not be read — while a KeyboardInterrupt or a genuine bug still propagates instead of being swallowed.
A bare
except Exceptionthat returns a default does not handle an error; it launders one, turning a precise failure into a vague one several frames away.
The entry point
main.py is about how the program is invoked — arguments, the interactive loop, exit codes — and nothing else. Application logic belongs in the orchestrator, which will grow later. For now it needs to prove the foundation works end to end.
1# main.py2import argparse3import sys45from config.settings import settings6from src.exceptions import AssistantError7from src.utils import logger8910def main() -> int:11 parser = argparse.ArgumentParser(description="Multimodal assistant")12 parser.add_argument("--text-only", action="store_true",13 help="skip loading voice and vision models")14 parser.add_argument("--verbose", action="store_true")15 args = parser.parse_args()1617 logger.info("Starting assistant (model=%s, text_only=%s)",18 settings.text_model, args.text_only)1920 try:21 while True:22 user_input = input("\nyou > ").strip()23 if user_input.lower() in {"quit", "exit"}:24 break25 if not user_input:26 continue27 # replaced by the orchestrator in a later stage28 print(f"assistant > echo: {user_input}")29 except KeyboardInterrupt:30 print()31 except AssistantError as e:32 logger.critical("Fatal assistant error: %s", e)33 return 13435 logger.info("Shutting down cleanly")36 return 0373839if __name__ == "__main__":40 sys.exit(main())Returning an exit code rather than calling sys.exit() from inside main() keeps the function testable: a test can call main() and assert on the return value. And a non-zero exit code is what a container orchestrator or CI job actually reads to decide whether the process failed.
Five tests, written before there is anything interesting to test
Writing tests now feels premature — there is barely any behaviour. That is exactly why it works. The habit is cheap to start and expensive to retrofit; a project with zero tests at 2,000 lines almost never gets them. More practically, these particular tests catch the failures that are hardest to diagnose later: a configuration field that silently accepts a bad value, a logger that attaches duplicate handlers and writes every line twice.
1# tests/conftest.py2import pytest3from pathlib import Path456@pytest.fixture7def tmp_env(tmp_path: Path, monkeypatch):8 """Isolated env + throwaway directories, so tests never touch real data."""9 monkeypatch.setenv("OPENAI_API_KEY", "sk-test-key-000000000000")10 monkeypatch.setenv("LOG_LEVEL", "DEBUG")11 monkeypatch.chdir(tmp_path)12 return tmp_path131415@pytest.fixture16def sample_text() -> str:17 return "What is in this image?"1# tests/test_settings.py2import pytest3from pydantic import ValidationError4from config.settings import Settings567def test_missing_key_is_fatal(monkeypatch):8 monkeypatch.delenv("OPENAI_API_KEY", raising=False)9 with pytest.raises(ValidationError):10 Settings(_env_file=None)111213def test_temperature_bounds_are_enforced(tmp_env):14 with pytest.raises(ValidationError):15 Settings(temperature=7.0) # a plausible typo for 0.7161718def test_defaults_have_correct_types(tmp_env):19 s = Settings()20 assert s.temperature is None # not sent unless you set it21 assert isinstance(s.max_tokens, int)22 assert s.data_dir.name == "data"232425def test_bad_log_level_rejected(tmp_env):26 with pytest.raises(ValidationError):27 Settings(log_level="LOUD")1# tests/test_utils.py2from src.utils import get_logger345def test_logger_is_not_double_registered():6 a = get_logger("dup-check")7 n = len(a.handlers)8 b = get_logger("dup-check")9 assert a is b10 assert len(b.handlers) == n # calling twice must not duplicate outputA fixture is just a named piece of setup that pytest injects into any test declaring it as a parameter. tmp_env above does three jobs at once — fake credentials, a fake working directory, and a guaranteed-clean filesystem — and every future test file can request it by name instead of copying that setup. monkeypatch undoes its own changes when the test ends, so an environment variable set in one test cannot leak into the next and produce a pass that depends on test ordering.
Run them: pytest -v, and once you have more code, pytest --cov=src --cov=config.
When things go wrong here
| Symptom | Cause | Fix |
|---|---|---|
ModuleNotFoundError: No module named 'src' when running pytest | pytest was invoked from inside tests/, so the project root is not on the import path | Run pytest from the project root, and make sure src/__init__.py and tests/__init__.py exist |
| Every log line appears twice | get_logger() called more than once and handlers were appended again | The if log.handlers: return log guard — verify test_logger_is_not_double_registered passes |
ValidationError: openai_api_key Field required | No .env in the current working directory, or the process was launched from a different directory | Copy .env.example to .env, fill it in, and launch from the project root |
pip install fails building pyaudio | PortAudio is a C library that pip cannot install | brew install portaudio or sudo apt-get install portaudio19-dev, then reinstall |
| Installed packages are not visible to Python | The virtual environment is not active; pip installed into the system interpreter | source venv/bin/activate, confirm which python points inside venv/ |
| torch install is enormous or times out | pip resolved the CUDA build (over 2 GB) on a machine without a GPU | Install the CPU wheel explicitly from the PyTorch CPU index |
| Settings changes have no effect | A real environment variable is set in the shell and overrides .env | env | grep -i openai, then unset the stale one |
Acceptance criteria for this stage
Do not move on until every one of these is true. These are checkable, not aspirational.
git statusshows no.env, nologs/, nodata/, and novenv/as untracked-but-uncommitted noise.python -c "from config.settings import settings; print(settings.text_model)"prints a model name in under one second.- Temporarily renaming
.envmakes that same command fail in under one second with a message namingopenai_api_key— not with a traceback from inside an HTTP library. python main.py --text-onlystarts, echoes input, and exits cleanly onquitwith exit code 0 (echo $?).logs/assistant.logexists after that run and contains a line at INFO with a timestamp, level, and module name.pytest -vreports 5 passed, 0 failed, in under 2 seconds.grep -rn "sk-" --include="*.py" .returns nothing.
What this buys you on the day something breaks
Return to the opening traceback and imagine it with the foundation in place. The failure now reads VisionError: could not caption examples/fridge.jpg: cannot write mode CMYK as PNG, raised from a named module, with the original OSError attached below it. logs/assistant.log has the preceding forty DEBUG lines including the image dimensions and mode. You add a test that feeds a CMYK JPEG to the vision handler, watch it fail, add a colour-space conversion, and watch it pass — and that test now stands guard so the same bug cannot silently return in week eight.
The diagnosis went from forty minutes of guessing to about ninety seconds of reading. That difference is not a productivity nicety; it is the difference between a project you finish and one you abandon in week six because every fix creates two new problems. The structure, the pins, the validated settings, the levelled logs, the typed exceptions, and the fixtures are all one investment in the same thing: when something breaks — and across four ML subsystems, plenty will — the system tells you exactly where, immediately, and lets you prove the fix.