Course Content
Capstone Project: Multimodal Assistant
1 sections · 6 lessons
Integration & Deployment
Five subsystems, five green test suites. The vision handler captions images correctly when handed a file path. The voice handler transcribes correctly when handed audio. The reasoning engine loops correctly against a mocked model. The memory manager stores and retrieves correctly when handed strings. Every one of them is right.
You wire them together and run the assistant for a week. Then someone asks it what they said about their trip to Lisbon, and it returns three results, all of which read "Hello! How can I help you today?"
Here is what happened. The microphone occasionally captures a second of room noise and the transcriber, correctly, returns an empty string. The orchestrator passes that empty string to the text path, which builds a prompt with an empty user turn. The model, correctly, responds with a generic greeting. The orchestrator records the exchange, and the memory layer, correctly, indexes it. Over seven days that fired 43 times, and the vector store now holds 43 near-identical greetings that sit close to the centre of the embedding space and therefore rank plausibly for almost any short query.
Nothing crashed. Nothing logged an error. No unit test could have caught it, because every component did exactly what it was specified to do — the bug lives in the sequence, in the fact that nobody decided what should happen when a transcription comes back empty. That decision has no natural home inside any subsystem, because no subsystem knows the others exist.
Giving it a home is what this stage is about: an orchestrator that owns the sequencing, an HTTP surface so something other than a terminal can call it, a container so it runs the same way everywhere, and tests that exercise the seams rather than the parts.
The orchestrator owns the sequence
You could wire everything in main.py. The reason not to is that main.py is about how the program is invoked — argument parsing, the input loop, printing to a terminal. "Voice gets transcribed, then reasoned over, then both spoken and remembered" is not invocation, it is application logic, and it needs to be callable from a second entry point that has no terminal at all.
1# src/assistant.py2from enum import Enum34from src.voice_handler import VoiceHandler5from src.vision_handler import VisionHandler6from src.text_processor import LLMChainManager7from src.reasoning_engine import ReActAgent, registry8from src.memory_manager import MemoryManager9from src.exceptions import (AssistantError, VoiceError, VisionError,10 TextProcessingError)11from src.utils import logger121314class State(Enum):15 IDLE = "idle"; LISTENING = "listening"; PROCESSING = "processing"161718class MultimodalAssistant:19 """The only module permitted to import more than one subsystem."""2021 def __init__(self, text_only: bool = False):22 self.text_only = text_only23 self.state = State.IDLE2425 self.text_processor = LLMChainManager()26 self.tools = registry # already holds calculator, current_time27 self.reasoning = ReActAgent(self.text_processor, self.tools)28 self.memory = MemoryManager()2930 # Optional so the service can run on a box with no microphone,31 # and so a text-only smoke test does not pay to load two vision32 # models and a speech model it will never call.33 self.voice = None if text_only else VoiceHandler()34 self.vision = None if text_only else VisionHandler()3536 self._register_tools()3738 def _register_tools(self) -> None:39 # The reasoning engine knows nothing about MemoryManager. It knows40 # about a name, a callable and a description. Binding the two happens41 # here, at the only point in the codebase that can see both.42 @self.tools.register("search_memory",43 "Search what the user said before, e.g. 'food allergies'")44 def _search(query: str) -> str:45 hits = self.memory.semantic.search(query)46 return "\n".join(h["text"] for h in hits) or "Nothing relevant found."4748 @self.tools.register("remember",49 "Store a durable fact about the user. Input 'key: value', e.g. 'city: Berlin'")50 def _remember(arg: str) -> str:51 key, _, value = arg.partition(":")52 if not value.strip():53 return "Error: expected 'key: value'"54 self.memory.facts.remember(key, value.strip())55 return f"Stored {key.strip()}."5657 def process_text(self, text: str) -> str:58 """The path every other input mode funnels into."""59 if not text.strip(): # GUARD: the opening bug60 raise TextProcessingError("empty input; nothing to process")61 self.state = State.PROCESSING62 try:63 context = self.memory.context_for(text)64 question = f"{context}\n\nUser: {text}" if context else text65 response = self.reasoning.run(question)["answer"]66 self.memory.record("user", text)67 self.memory.record("assistant", response)68 return response69 except AssistantError:70 raise # already typed; let it through71 except Exception as e:72 raise TextProcessingError(str(e)) from e73 finally:74 self.state = State.IDLE7576 def process_voice(self, seconds: int = 5) -> str:77 if self.text_only:78 raise VoiceError("voice is disabled in text-only mode")79 self.state = State.LISTENING80 try:81 text = self.voice.listen(seconds)82 if not text.strip(): # GUARD, again, and earlier83 logger.info("Empty transcription; ignoring")84 return ""85 response = self.process_text(text) # reuse, never duplicate86 self.voice.respond(response)87 return response88 except AssistantError:89 raise90 except Exception as e:91 raise VoiceError(str(e)) from e92 finally:93 self.state = State.IDLE9495 def status(self) -> dict:96 return {"state": self.state.value,97 "components": {"voice": self.voice is not None,98 "vision": self.vision is not None,99 "reasoning": True, "memory": True}}Three things in there are load-bearing.
The import rule. This file imports from every subsystem; no subsystem imports from another. That single constraint is what kept each part independently testable while you built it, and it is what will let you swap the vector store or the speech model later without touching anything else. The moment vision_handler.py imports memory_manager, the rule is gone and it does not come back.
process_voice calls process_text. It does not re-implement reasoning and memory. If it did, the two paths would drift — a fix applied to one, forgotten in the other — and you would eventually get different answers depending on whether the user typed or spoke.
The guards sit at the boundary. The empty-transcription check appears twice on purpose: once where the empty string is produced, so it is handled with context, and once at the entry to the shared path, so no future caller can bypass it. Boundary conditions belong to the orchestrator because it is the only place that can see both sides.
A subsystem can only be correct with respect to its own contract. Deciding what happens when one subsystem's edge case meets another's assumption is a separate job, and if no object owns it, nobody does.
Serving it over HTTP
A terminal loop has exactly one possible client: a person at the same keyboard. An HTTP interface means a web front end, a mobile app, another service, or an automated test can all drive the same logic.
1# src/api_server.py2from contextlib import asynccontextmanager3from fastapi import FastAPI, HTTPException4from pydantic import BaseModel56from src.assistant import MultimodalAssistant7from src.exceptions import AssistantError8from src.utils import logger910STATE = {"assistant": None}111213@asynccontextmanager14async def lifespan(app: FastAPI):15 # Models load ONCE per process, not per request. Doing this inside a16 # request handler would reload roughly 1.2 GB of weights every call.17 logger.info("Loading models...")18 STATE["assistant"] = MultimodalAssistant(text_only=False)19 logger.info("Ready")20 yield21 STATE["assistant"] = None222324app = FastAPI(title="Multimodal Assistant", version="1.0.0", lifespan=lifespan)252627class TextQuery(BaseModel):28 query: str293031@app.get("/health") # LIVENESS: is the process alive?32async def health():33 return {"status": "ok"}343536@app.get("/ready") # READINESS: can it serve traffic?37async def ready():38 if STATE["assistant"] is None:39 raise HTTPException(status_code=503, detail="models still loading")40 return STATE["assistant"].status()414243@app.post("/text")44async def process_text(q: TextQuery):45 if STATE["assistant"] is None:46 raise HTTPException(status_code=503, detail="models still loading")47 try:48 return {"response": STATE["assistant"].process_text(q.query)}49 except AssistantError as e:50 # A known failure mode caused by the input: client error, not 500.51 raise HTTPException(status_code=422, detail=str(e))52 except Exception as e:53 logger.error("unexpected in /text: %s", e, exc_info=True)54 raise HTTPException(status_code=500, detail="internal server error")The split between /health and /ready is not pedantry, and it is the first thing that breaks in production. Loading the caption model, the speech model and the embedding model on a CPU container takes about 40 seconds on a cold start. A container platform that probes a single endpoint every 10 seconds and restarts after three failures will kill the process at 30 seconds — before it has ever finished starting — and then do it again, forever. You get an infinite crash loop from a service that has no bug in it.
Liveness answers "is this process wedged?" and must respond immediately without touching anything heavy. Readiness answers "should traffic be routed here?" and returns 503 until the models are up. Set the platform's startup grace period from the number you actually measured, not from a default.
Error mapping matters for the same reason: a typed AssistantError means the request caused a known failure and the client should see 422 with a readable message. A bare exception means your server is broken and should say 500 and log a traceback. Collapsing both into 500 means every client-side mistake looks like an outage, and your error rate graph stops meaning anything.
One honest limitation to write down now: a /voice endpoint built on voice.listen() records from the microphone attached to the server. On your laptop that is your microphone. In a container it is nothing at all. The real design accepts an uploaded audio file and transcribes the bytes; until you have built that, the endpoint should not be exposed in a deployed configuration.
What "works on my machine" actually costs
The project now depends on a Python version, a set of pinned packages, and two system libraries that pip cannot install:
| Dependency | Why pip cannot supply it | Symptom if missing |
|---|---|---|
ffmpeg | A native binary the speech library shells out to for audio decoding | Transcription raises FileNotFoundError on a file that plays fine locally |
portaudio19-dev | C headers needed to compile the audio capture package from source | pip install fails mid-build with a compiler error about a missing header |
1FROM python:3.12-slim2WORKDIR /app34RUN apt-get update && apt-get install -y --no-install-recommends \5 ffmpeg portaudio19-dev \6 && rm -rf /var/lib/apt/lists/*78# Dependencies BEFORE application code. This ordering is the difference9# between a 3-second rebuild and a 5-minute one; see below.10COPY requirements.txt .11RUN pip install --no-cache-dir -r requirements.txt1213COPY . .1415ENV PYTHONUNBUFFERED=1 LOG_LEVEL=INFO16EXPOSE 800017CMD ["uvicorn", "src.api_server:app", "--host", "0.0.0.0", "--port", "8000"]Do the arithmetic on the layer ordering, because it is the detail people treat as style. Docker caches each instruction and reuses the cache while nothing above it has changed. pip install for this project takes about 4 minutes 50 seconds, most of it downloading a deep learning wheel of roughly 2.5 GB. Copying the source tree takes about 3 seconds.
| Order | Rebuild after a one-line code change | Twenty rebuilds in a day |
|---|---|---|
COPY requirements.txt → pip install → COPY . . | ~3 seconds (only the last layer is invalid) | 1 minute |
COPY . . → pip install | ~4 min 50 s (the code copy invalidates everything below it) | 1 hour 37 minutes |
The second decision is whether to bake the model weights into the image. Downloading them on first use adds about 40 seconds to the first request and requires network access from the container; baking them in adds roughly 1.2 GB to the image, making the final image about 3.9 GB, and makes startup deterministic and offline-capable. For anything that autoscales, bake them in: a new replica that spends its first minute downloading is a replica that is failing health checks while a queue builds behind it.
1# docker-compose.yml2services:3 assistant:4 build: .5 ports: ["8000:8000"]6 environment:7 - OPENAI_API_KEY=${OPENAI_API_KEY} # from the shell or .env, never baked in8 - LOG_LEVEL=INFO9 volumes:10 - ./logs:/app/logs # without these, every restart11 - ./data:/app/data # discards sessions and the indexReproducibility is not finished when the Python dependencies are pinned. It is finished when the operating system libraries, the model weights and the start command are pinned too, because those are the three things that differ between your laptop and the machine that will actually run this.
Two rules survive from development straight into deployment. Secrets come from the environment, never from a file inside the image — an image is copied, pushed to registries, and pulled by people who should not have your key. And anything you want to survive a restart must live on a mounted volume, because a container's own filesystem is discarded when it stops. A vector store written to the container filesystem is a vector store that forgets everything every deploy.
Tests for the seams
Until now, every test mocked the things around the component under test. Integration tests deliberately do the opposite: real objects, real boundaries, no mocks between subsystems. They are slower and less deterministic, and they catch a category of bug that no amount of unit testing can reach.
1# tests/test_integration.py2import pytest3from src.assistant import MultimodalAssistant4from src.exceptions import TextProcessingError, VoiceError567@pytest.fixture(scope="module")8def assistant():9 # Module scope: constructing this loads models. Per-test construction10 # turns a 6-second suite into a 4-minute one.11 return MultimodalAssistant(text_only=True)121314def test_empty_input_is_refused_not_answered(assistant):15 """Regression for the 43 stored greetings."""16 with pytest.raises(TextProcessingError):17 assistant.process_text(" ")181920def test_nothing_is_recorded_when_input_is_refused(assistant):21 before = len(assistant.memory.conversation.turns)22 with pytest.raises(TextProcessingError):23 assistant.process_text("")24 assert len(assistant.memory.conversation.turns) == before252627def test_a_fact_from_one_turn_is_retrievable_in_a_later_one(assistant):28 assistant.process_text("I really enjoy hiking in the mountains at weekends")29 hits = assistant.memory.semantic.search("what do I like doing outdoors?")30 assert hits and "hiking" in hits[0]["text"]313233def test_typed_errors_survive_the_orchestrator(assistant):34 with pytest.raises(VoiceError):35 assistant.process_voice() # text_only=True3637def test_status_reports_disabled_subsystems_honestly(assistant):38 s = assistant.status()39 assert s["components"]["voice"] is False40 assert s["components"]["vision"] is FalseThe second test is the one worth copying into your own habits. It asserts on a side effect that should not have happened. The opening failure was not that something threw; it was that something was silently written. Tests that only check return values are blind to exactly that class of bug.
Write the integration test that reproduces the bug before you fix it, using the real input that caused it. Otherwise you have fixed one instance of a boundary problem and left the boundary unguarded.
Keep the suite runnable: text-only mode for anything that does not genuinely need the vision or speech models, module-scoped fixtures, and a marker such as @pytest.mark.slow on the few tests that must load real weights so the fast suite stays under ten seconds.
Writing down what is not finished
A deployment document that only says how to start the service is half a document. The half that matters lists what is known to be missing, because an undocumented limitation is indistinguishable from a bug to whoever inherits this.
| Limitation | Why it is there | What it needs before production |
|---|---|---|
| The voice endpoint records from the server's microphone | The capture method was written for a local terminal, where the assumption held | Accept an uploaded audio file and transcribe from bytes |
| No authentication on any endpoint | Nothing needed it while the only client was localhost | An API key check or auth middleware; anyone who can reach the port can spend your model budget |
| No rate limiting | A single human typing is self-limiting | A per-client limit; one loop in a broken client can exhaust a daily quota in minutes |
| Plain HTTP | Local development never left the machine | TLS at a reverse proxy or load balancer |
| Unbounded conversation context | Sessions were short during development | A token budget on assembled context; long sessions get slower and dearer every turn |
| No backup of the data volume | The data was disposable while building | A tested restore, not just a backup job — an untested backup is a hope |
When things go wrong here
| Symptom | Cause | Fix |
|---|---|---|
| Every subsystem's tests pass; the assembled assistant misbehaves on real input | A boundary case nobody owns — an empty, oversized, or oddly punctuated value crossing between subsystems | Guard at the orchestrator, add a regression test using the exact input, and assert on side effects as well as return values |
| Container restarts forever, logs stop partway through model loading | The health probe fires before startup finishes and the platform kills the process | Separate liveness from readiness; set the grace period above the measured load time |
| Docker build fails compiling the audio package | portaudio19-dev missing from the image | Install system libraries before pip install, not after |
| Every rebuild reinstalls all dependencies | COPY . . placed above pip install, invalidating the cache | Copy requirements.txt alone, install, then copy the source |
| Sessions and the vector index vanish on every deploy | Writes went to the container's own filesystem | Mount volumes for data/ and logs/; verify by restarting and re-querying |
| Client mistakes appear as 500s and swamp the error rate | All exceptions collapsed into one handler | Typed errors to 422, unexpected ones to 500 with a logged traceback |
| Two entry points give different answers to the same question | Logic duplicated in the loop and in a route instead of shared | Both call the identical orchestrator method; the route should be four lines |
| First request after deploy times out at the load balancer | Weights downloaded on first use rather than baked into the image | Bake the models in, and warm the process during startup rather than on request |
| Responses correct but slow under any concurrency | One process, CPU inference, blocking calls inside async handlers | Measure which call dominates; run blocking work in a thread pool; scale replicas rather than threads for model work |
Acceptance criteria for this stage
MultimodalAssistantconstructs in both text-only and full modes, and text-only mode does not load the vision or speech models — check by timing it.- A grep of the codebase shows exactly one file importing from more than one subsystem.
- Empty or whitespace-only input raises a typed error, and nothing is written to memory when it does.
- A deliberately bad image path produces a typed vision error at the API boundary, returned as 422 with a readable message.
/healthanswers within milliseconds while the models are still loading;/readyreturns 503 until they are up, then reports subsystem status.- A text request through the API returns exactly what calling the orchestrator method directly returns, for the same input.
- The image builds from a clean checkout, and a one-line source change rebuilds in seconds rather than minutes.
- The container survives a stop and start with sessions and the vector index intact.
- The fast test suite runs in under ten seconds; the slow suite passes with real credentials.
- You can state, without looking, two documented limitations and why each exists.
What you have when this runs
The assistant now understands typed text, spoken input and images, reasons over multi-step problems with tools, remembers across sessions, searches its own history by meaning, and serves all of it over HTTP from a reproducible container. That is the visible result. The structural result is more useful.
Every subsystem is still replaceable. Because only one file knows the others exist, swapping the vector store, the speech model, or the whole reasoning strategy is a change to one constructor and its tests, not an archaeology project across five modules. Because both entry points call the same methods, adding a third — a scheduled job, a message queue consumer, a chat integration — costs a few lines rather than a reimplementation.
The habit worth carrying forward is the one the opening bug teaches. Systems assembled from correct parts fail at the joins, and they usually fail quietly: an empty string that should have been rejected, a value stored that should have been dropped, a probe that answered before it was true. Loud failures find themselves. Quiet ones are found by someone deciding, explicitly, what should happen when one component's edge case meets another's assumption — writing that decision into the object that owns the sequence, and pinning it with a test that asserts on what did not happen.