Capstone Project: Multimodal Assistant

Project Setup & Architecture


It is week six of the build. Your assistant is roughly 2,300 lines spread across six Python files. You run it, say "what's in this photo", and get this:

Text
Traceback (most recent call last):  File "main.py", line 41, in <module>    assistant.run()  File "src/assistant.py", line 88, in run    return self.process(user_input)  File "src/assistant.py", line 112, in process    result = self.handle(payload)Exception: 'NoneType' object has no attribute 'strip'

Which subsystem failed? The microphone capture, the transcription model, the caption model, the LLM call, or the vector store? You cannot tell. There is no log file, because everything went to print() and your terminal scrollback ended forty minutes ago. You cannot bisect it against a known-good state, because there are no tests. And when you finally do find it — a caption model returning None for a CMYK JPEG — you fix it, and two days later it silently comes back, because nothing checks.

That is not a bad-luck day. That is the predictable output of skipping the boring hour at the start. This stage is that hour. You will write almost no "AI" code here. You will write a directory layout, a pinned dependency set, a configuration loader that refuses to start when something is missing, a logger, an exception hierarchy, an entry point, and five tests. Every later stage of this project reuses all six of those, unchanged.

In a multi-component system the cost of a missing foundation does not appear on day one. It appears in week six, multiplied by every component you added since.

The layers you build before anything interestingConfig,validated at importStructuredlogging, not printNamedexception hierarchyHandlers:vision, voice, memoryOrchestratorand entry pointtopbottomConfig that fails in 0.3 seconds beats a stack trace 45 seconds into a model load.
Every layer below the handlers exists so that when 2,300 lines break, the traceback names the culprit instead of the symptom.

What you are actually building

The finished system is a multimodal assistant: it accepts typed text, an image file, or spoken audio; it reasons about the request, calling real tools when it needs a real answer; it remembers what you told it last week; and it replies in text and optionally in speech. Six subsystems, one orchestrator, two cross-cutting utilities.

Text
  INPUT ADAPTERS                ORCHESTRATOR                 SERVICES  ------------------           ----------------              ---------------------------  voice_handler.py   ---->   +------------------+   ---->   text_processor.py    mic / wav -> text        |                  |             prompts + LLM calls                             | MultimodalAssistant  vision_handler.py  ---->   |   (assistant.py) |   ---->   reasoning_engine.py    image -> caption         |                  |             ReAct loop + JSON schema                             |  routes input,   |  raw typed text     ---->   |  owns one of     |   ---->   memory_manager.py                             |  each subsystem  |             history + vector search                             +------------------+  CROSS-CUTTING (built in this stage, used by everything above):    config/settings.py   validated configuration, loaded once at import    src/utils.py         logger, shared helpers    exception hierarchy  every subsystem raises its own named error type

The important property of that diagram is the arrows that are not there. vision_handler.py never imports memory_manager.py. voice_handler.py has no idea the reasoning engine exists. Only the orchestrator imports more than one subsystem. That single rule is what lets you test the vision pipeline with no microphone attached, swap the LLM provider without touching a prompt, and — crucially — read a stack trace and know immediately which box broke.

Before you start

  • Python 3.12 or 3.13. Check with python3 --version. The heavy ML wheels in this project (torch, whisper's dependencies, sentence-transformers) have historically lagged a brand-new interpreter by several months, so pick a Python release that has been out for about a year. On a release that is too new you will spend your first afternoon compiling C extensions from source instead of building an assistant.
  • An LLM API key for whichever provider you use for text generation and embeddings.
  • At least 2 GB free disk. The virtual environment plus the models later stages download is not small: a CPU-only torch wheel is roughly 200 MB, Whisper's base checkpoint about 139 MB, a sentence-transformer around 90 MB, a caption model a few hundred MB more.
  • Terminal and git basics. No Docker, Redis, or GPU is needed yet.

Layout first, code second

A flat folder of .py files is genuinely fine at 200 lines. It stops being fine at about 800, which this project passes in its third stage. The layout below is not aesthetic preference; each directory is a boundary that prevents a specific class of mistake.

Bash
mkdir -p multimodal-assistant/{src,tests,config,prompts,examples,logs,data}cd multimodal-assistanttouch src/__init__.py tests/__init__.py config/__init__.py
Text
multimodal-assistant/|-- src/|   |-- __init__.py|   |-- assistant.py          orchestrator|   |-- voice_handler.py      speech-to-text, text-to-speech|   |-- vision_handler.py     embeddings + captioning|   |-- text_processor.py     prompt templates + LLM calls|   |-- reasoning_engine.py   ReAct loop + structured output|   |-- memory_manager.py     conversation history + vector store|   |-- exceptions.py         the error hierarchy|   `-- utils.py              logging and shared helpers|-- config/|   |-- __init__.py|   `-- settings.py           validated configuration|-- tests/|   |-- conftest.py           shared pytest fixtures|   |-- test_settings.py|   `-- test_utils.py|-- prompts/                  prompt text, kept out of logic|-- examples/                 sample images, sample queries|-- logs/                     generated, gitignored|-- data/                     vector store, sessions; gitignored|-- main.py|-- requirements.txt|-- .env.example|-- .gitignore`-- README.md
DirectoryHoldsThe mistake it prevents
src/One module per subsystemVision code quietly growing inside reasoning code, so neither can be tested or replaced alone
config/All configurationTwenty scattered os.environ.get() calls, each with a different default, none validated
tests/A file mirroring each src/ moduleNot noticing which subsystem has zero coverage — the mirror makes gaps visible at a glance
prompts/Prompt strings and tool schemasRe-testing application logic because you changed the wording of a sentence
examples/Sample inputsHunting for a test image every time you want to reproduce a bug
logs/, data/Runtime outputCommitting a 40 MB vector index, or a log file containing a user's transcript

Write .gitignore now, before you create .env and before the first run generates a log:

Bash
venv/__pycache__/*.pyc.envlogs/data/*.egg-info/.pytest_cache/

The ordering matters more than it looks. Committing .env once puts the key in git history permanently; deleting the file in a later commit does not remove it. The fix is then key rotation and a history rewrite, which is a bad afternoon. Ignoring the file before it exists costs nothing.

An environment you can reproduce

A virtual environment is a private, disposable copy of Python and its packages that belongs to this project only. Without one, installing torch==2.1.1 here can break a different project on your machine that needed 2.0.0, and you will never be able to state what a fresh clone actually requires.

Bash
python3 -m venv venvsource venv/bin/activate        # macOS / Linux# venv\Scripts\activate         # Windowswhich python                    # must print a path inside ./venv/pip install --upgrade pip

Now pin everything with ==, never >=:

Text
# versions current in September 2026 -- bump them deliberately, never by accident# corepython-dotenv==1.2.3pydantic==2.13.5pydantic-settings==2.15.0# language modelsopenai==3.19.2langchain-openai==1.6.5# visionpillow==12.3.0transformers==5.17.0torch==2.14.0torchvision==0.29.0# audioopenai-whisper==20250625pyttsx3==2.99pyaudio==0.2.14# memory and searchchromadb==1.5.9sentence-transformers==6.1.0# api and servingfastapi==0.141.1uvicorn==0.53.0# testing and toolingpytest==9.1.1pytest-cov==7.1.0black==26.5.1

Here is the arithmetic that makes == non-negotiable. That file lists 19 direct dependencies, which pull in well over 100 packages in total on Linux. Suppose each package ships a release, on average, once every six weeks — a modest rate for this ecosystem. Over a three-month project that is about two releases each, so more than 200 opportunities for a version to change under you. If just one in fifty of those releases contains a breaking rename, you should expect four or five silent breakages across the project's life, each appearing on a random morning with no code change of your own to blame.

Pinning only the direct dependencies does not get that number to zero, because the packages they pull in still float. A classic case: openai client releases from 2023 declared only httpx<1, so when httpx 0.28 shipped in late 2024, fresh installs of those pinned clients picked it up — and it had removed an argument the old clients passed, so every call failed with a TypeError. So once the install works, freeze the whole resolved set with pip freeze > requirements.lock (or use a lock-file tool such as uv or pip-tools) and install from the lock. Now nothing changes until you decide to bump a version deliberately.

Install and expect it to take a while: pip install -r requirements.txt downloads several hundred megabytes and typically runs 5 to 15 minutes. You pay that once.

Configuration that fails in 0.3 seconds, not 45

The naive approach is os.environ.get("OPENAI_API_KEY") wherever a key is needed. It has three failure modes, all of them quiet. A typo in the variable name returns None rather than raising. A missing value surfaces as an authentication error deep in a library, not as "you forgot to set a key". And a numeric setting read from the environment arrives as the string "0.7", which silently becomes a temperature of nonsense somewhere downstream.

The concrete cost: your assistant loads a Whisper checkpoint and a caption model at startup. That is roughly 45 seconds of model loading before the first LLM call is attempted. With scattered environ.get, a missing key means you wait the full 45 seconds and then get an auth failure. With a validated settings object imported at module load, you fail in about 0.3 seconds with a message naming the exact field. Over a debugging session where you restart thirty times, that is 22 minutes of your life against 9 seconds.

Python
# config/settings.pyfrom pathlib import Pathfrom pydantic import Field, field_validatorfrom pydantic_settings import BaseSettings, SettingsConfigDictclass Settings(BaseSettings):    """Every configurable value in the project, validated at import time."""    model_config = SettingsConfigDict(        env_file=".env", env_file_encoding="utf-8", extra="ignore"    )    # secrets -- no defaults, so a missing value is a hard error    openai_api_key: str = Field(..., min_length=20)    # model choices -- a small, cheap default; check your provider's current list    text_model: str = "gpt-6-luna"    whisper_model: str = "base"    embedding_model: str = "all-MiniLM-L6-v2"    # generation behaviour. None means "not sent": many current reasoning    # models reject temperature, so only set it for a model that accepts it.    temperature: float | None = Field(None, ge=0.0, le=2.0)    # on reasoning models this budget also covers hidden reasoning tokens    max_tokens: int = Field(4000, gt=0, le=32000)    # paths    data_dir: Path = Path("data")    log_dir: Path = Path("logs")    log_level: str = "INFO"    @field_validator("log_level")    @classmethod    def _valid_level(cls, v: str) -> str:        allowed = {"DEBUG", "INFO", "WARNING", "ERROR", "CRITICAL"}        v = v.upper()        if v not in allowed:            raise ValueError(f"log_level must be one of {sorted(allowed)}")        return vsettings = Settings()settings.data_dir.mkdir(parents=True, exist_ok=True)settings.log_dir.mkdir(parents=True, exist_ok=True)

Three things are doing real work here. Field(...) with no default makes the key mandatory: import fails immediately if it is absent. ge/le/gt bounds catch a temperature of 7.0 — a plausible typo for 0.7 — at startup rather than as incoherent model output an hour later. And the type annotations coerce: a TEMPERATURE of "0.7" arrives as a real float, data_dir as a real Path, so no caller has to remember to cast.

Temperature defaults to None for a reason. Many current models, including the reasoning models most providers now recommend, reject temperature and top_p or ignore them, so sending 0.7 by habit can turn every call into an HTTP 400. Leave it unset unless your chosen model's documentation says it is supported.

Commit .env.example with the shape but not the values, so a fresh clone knows what to fill in:

Bash
OPENAI_API_KEY=sk-replace-meTEXT_MODEL=gpt-6-luna# TEMPERATURE=0.7    only for a model that accepts itLOG_LEVEL=INFO

Logging, because print does not survive contact with reality

print() has no severity, no timestamp, no source, and no persistence. When a transcription fails at 23:14 during an unattended run, a print statement has already scrolled away. A log line is still on disk tomorrow, stamped with the module that emitted it.

Python
# src/utils.pyimport loggingfrom logging.handlers import RotatingFileHandlerfrom config.settings import settingsdef get_logger(name: str = "assistant") -> logging.Logger:    log = logging.getLogger(name)    if log.handlers:            # never attach handlers twice        return log    log.setLevel(logging.DEBUG)  # handlers decide what actually gets written    fmt = logging.Formatter(        "%(asctime)s | %(levelname)-8s | %(name)s:%(lineno)d | %(message)s",        datefmt="%Y-%m-%d %H:%M:%S",    )    console = logging.StreamHandler()    console.setLevel(getattr(logging, settings.log_level))    console.setFormatter(fmt)    disk = RotatingFileHandler(        settings.log_dir / "assistant.log",        maxBytes=10_000_000,     # 10 MB per file        backupCount=3,           # keep 3 rotations -> 40 MB ceiling        encoding="utf-8",    )    disk.setLevel(logging.DEBUG)  # disk keeps everything    disk.setFormatter(fmt)    log.addHandler(console)    log.addHandler(disk)    return loglogger = get_logger()

The split levels are the point: the console shows you INFO and above so it stays readable while you work, while the file records DEBUG so the detail is there when you need to reconstruct a failure after the fact. Size the rotation with real arithmetic. One conversational turn through the full pipeline emits roughly 40 DEBUG lines at about 120 bytes each, so about 4.8 KB per turn. At 500 turns of testing per day that is 2.4 MB per day, so a 10 MB file covers about four days and the three backups extend that to roughly sixteen. That is the right ballpark: long enough to cover a weekend of unattended running, small enough that nobody ever notices the disk usage.

LevelMeansIn this project
DEBUGDetail only useful when diagnosingPrompt text sent, token counts, similarity scores, audio sample counts
INFONormal progress worth recording"Whisper base loaded", "indexed 12 documents", "turn completed in 2.4 s"
WARNINGDegraded but still workingRetrying an API call, falling back to offline TTS, image resized from 4096 px
ERRORThis operation failedTranscription raised, JSON validation failed after retries
CRITICALThe process cannot continueConfiguration invalid, vector store unreadable

The failure mode people hit here is logging the wrong thing. Never log an API key, and never log a full audio buffer or base64 image — one logger.debug(f"payload={payload}") on an image request can write 3 MB into a single log line and blow through your rotation in one turn. Log the shape, not the content: logger.debug("image %s, %dx%d, %d bytes", path.name, w, h, size).

An exception hierarchy that names the culprit

The traceback at the top of this lesson ended in a bare Exception. That is the direct consequence of a subsystem that catches everything and re-raises nothing meaningful. The fix is a small tree of typed errors, one per subsystem.

Python
# src/exceptions.pyclass AssistantError(Exception):    """Base class. Catching this catches everything we raise deliberately."""class ConfigurationError(AssistantError):    """Settings missing or invalid."""class VoiceError(AssistantError):    """Microphone capture, transcription, or speech synthesis failed."""class VisionError(AssistantError):    """Image loading, embedding, or captioning failed."""class TextProcessingError(AssistantError):    """LLM call or prompt construction failed."""class ReasoningError(AssistantError):    """The reasoning loop failed or produced unusable output."""class MemoryError_(AssistantError):    """Persistence or vector search failed."""

Note the trailing underscore on the last one: Python already has a built-in MemoryError for out-of-memory conditions, and shadowing it would make a genuine allocation failure look like a database problem. Small detail, real bug avoided.

Now compare the two ways of using them. The wrong way:

Python
try:    caption = self.captioner(image)except Exception:    return None            # caller now gets None and has no idea why

Every downstream 'NoneType' object has no attribute 'strip' in this project traces back to a block shaped like that. Three problems compound: the error type is discarded, the original traceback is discarded, and the failure is converted into a valid-looking return value that travels several frames before exploding somewhere unrelated. The right way:

Python
try:    caption = self.captioner(image)except (OSError, ValueError) as e:    logger.error("Captioning failed for %s: %s", image_path, e)    raise VisionError(f"could not caption {image_path}: {e}") from e

Catch the specific exceptions you expect, log with context, raise a named error, and keep the cause with from e so the original traceback stays attached. Now the orchestrator can write except VisionError: and degrade gracefully — answer the text part of the question, tell the user the image could not be read — while a KeyboardInterrupt or a genuine bug still propagates instead of being swallowed.

A bare except Exception that returns a default does not handle an error; it launders one, turning a precise failure into a vague one several frames away.

The entry point

main.py is about how the program is invoked — arguments, the interactive loop, exit codes — and nothing else. Application logic belongs in the orchestrator, which will grow later. For now it needs to prove the foundation works end to end.

Python
# main.pyimport argparseimport sysfrom config.settings import settingsfrom src.exceptions import AssistantErrorfrom src.utils import loggerdef main() -> int:    parser = argparse.ArgumentParser(description="Multimodal assistant")    parser.add_argument("--text-only", action="store_true",                        help="skip loading voice and vision models")    parser.add_argument("--verbose", action="store_true")    args = parser.parse_args()    logger.info("Starting assistant (model=%s, text_only=%s)",                settings.text_model, args.text_only)    try:        while True:            user_input = input("\nyou > ").strip()            if user_input.lower() in {"quit", "exit"}:                break            if not user_input:                continue            # replaced by the orchestrator in a later stage            print(f"assistant > echo: {user_input}")    except KeyboardInterrupt:        print()    except AssistantError as e:        logger.critical("Fatal assistant error: %s", e)        return 1    logger.info("Shutting down cleanly")    return 0if __name__ == "__main__":    sys.exit(main())

Returning an exit code rather than calling sys.exit() from inside main() keeps the function testable: a test can call main() and assert on the return value. And a non-zero exit code is what a container orchestrator or CI job actually reads to decide whether the process failed.

Five tests, written before there is anything interesting to test

Writing tests now feels premature — there is barely any behaviour. That is exactly why it works. The habit is cheap to start and expensive to retrofit; a project with zero tests at 2,000 lines almost never gets them. More practically, these particular tests catch the failures that are hardest to diagnose later: a configuration field that silently accepts a bad value, a logger that attaches duplicate handlers and writes every line twice.

Python
# tests/conftest.pyimport pytestfrom pathlib import Path@pytest.fixturedef tmp_env(tmp_path: Path, monkeypatch):    """Isolated env + throwaway directories, so tests never touch real data."""    monkeypatch.setenv("OPENAI_API_KEY", "sk-test-key-000000000000")    monkeypatch.setenv("LOG_LEVEL", "DEBUG")    monkeypatch.chdir(tmp_path)    return tmp_path@pytest.fixturedef sample_text() -> str:    return "What is in this image?"
Python
# tests/test_settings.pyimport pytestfrom pydantic import ValidationErrorfrom config.settings import Settingsdef test_missing_key_is_fatal(monkeypatch):    monkeypatch.delenv("OPENAI_API_KEY", raising=False)    with pytest.raises(ValidationError):        Settings(_env_file=None)def test_temperature_bounds_are_enforced(tmp_env):    with pytest.raises(ValidationError):        Settings(temperature=7.0)          # a plausible typo for 0.7def test_defaults_have_correct_types(tmp_env):    s = Settings()    assert s.temperature is None           # not sent unless you set it    assert isinstance(s.max_tokens, int)    assert s.data_dir.name == "data"def test_bad_log_level_rejected(tmp_env):    with pytest.raises(ValidationError):        Settings(log_level="LOUD")
Python
# tests/test_utils.pyfrom src.utils import get_loggerdef test_logger_is_not_double_registered():    a = get_logger("dup-check")    n = len(a.handlers)    b = get_logger("dup-check")    assert a is b    assert len(b.handlers) == n     # calling twice must not duplicate output

A fixture is just a named piece of setup that pytest injects into any test declaring it as a parameter. tmp_env above does three jobs at once — fake credentials, a fake working directory, and a guaranteed-clean filesystem — and every future test file can request it by name instead of copying that setup. monkeypatch undoes its own changes when the test ends, so an environment variable set in one test cannot leak into the next and produce a pass that depends on test ordering.

Run them: pytest -v, and once you have more code, pytest --cov=src --cov=config.

When things go wrong here

SymptomCauseFix
ModuleNotFoundError: No module named 'src' when running pytestpytest was invoked from inside tests/, so the project root is not on the import pathRun pytest from the project root, and make sure src/__init__.py and tests/__init__.py exist
Every log line appears twiceget_logger() called more than once and handlers were appended againThe if log.handlers: return log guard — verify test_logger_is_not_double_registered passes
ValidationError: openai_api_key Field requiredNo .env in the current working directory, or the process was launched from a different directoryCopy .env.example to .env, fill it in, and launch from the project root
pip install fails building pyaudioPortAudio is a C library that pip cannot installbrew install portaudio or sudo apt-get install portaudio19-dev, then reinstall
Installed packages are not visible to PythonThe virtual environment is not active; pip installed into the system interpretersource venv/bin/activate, confirm which python points inside venv/
torch install is enormous or times outpip resolved the CUDA build (over 2 GB) on a machine without a GPUInstall the CPU wheel explicitly from the PyTorch CPU index
Settings changes have no effectA real environment variable is set in the shell and overrides .envenv | grep -i openai, then unset the stale one

Acceptance criteria for this stage

Do not move on until every one of these is true. These are checkable, not aspirational.

  1. git status shows no .env, no logs/, no data/, and no venv/ as untracked-but-uncommitted noise.
  2. python -c "from config.settings import settings; print(settings.text_model)" prints a model name in under one second.
  3. Temporarily renaming .env makes that same command fail in under one second with a message naming openai_api_key — not with a traceback from inside an HTTP library.
  4. python main.py --text-only starts, echoes input, and exits cleanly on quit with exit code 0 (echo $?).
  5. logs/assistant.log exists after that run and contains a line at INFO with a timestamp, level, and module name.
  6. pytest -v reports 5 passed, 0 failed, in under 2 seconds.
  7. grep -rn "sk-" --include="*.py" . returns nothing.

What this buys you on the day something breaks

Return to the opening traceback and imagine it with the foundation in place. The failure now reads VisionError: could not caption examples/fridge.jpg: cannot write mode CMYK as PNG, raised from a named module, with the original OSError attached below it. logs/assistant.log has the preceding forty DEBUG lines including the image dimensions and mode. You add a test that feeds a CMYK JPEG to the vision handler, watch it fail, add a colour-space conversion, and watch it pass — and that test now stands guard so the same bug cannot silently return in week eight.

The diagnosis went from forty minutes of guessing to about ninety seconds of reading. That difference is not a productivity nicety; it is the difference between a project you finish and one you abandon in week six because every fix creates two new problems. The structure, the pins, the validated settings, the levelled logs, the typed exceptions, and the fixtures are all one investment in the same thing: when something breaks — and across four ML subsystems, plenty will — the system tells you exactly where, immediately, and lets you prove the fix.