Capstone Project: Multimodal Assistant

Integration & Deployment


Five subsystems, five green test suites. The vision handler captions images correctly when handed a file path. The voice handler transcribes correctly when handed audio. The reasoning engine loops correctly against a mocked model. The memory manager stores and retrieves correctly when handed strings. Every one of them is right.

You wire them together and run the assistant for a week. Then someone asks it what they said about their trip to Lisbon, and it returns three results, all of which read "Hello! How can I help you today?"

Here is what happened. The microphone occasionally captures a second of room noise and the transcriber, correctly, returns an empty string. The orchestrator passes that empty string to the text path, which builds a prompt with an empty user turn. The model, correctly, responds with a generic greeting. The orchestrator records the exchange, and the memory layer, correctly, indexes it. Over seven days that fired 43 times, and the vector store now holds 43 near-identical greetings that sit close to the centre of the embedding space and therefore rank plausibly for almost any short query.

Nothing crashed. Nothing logged an error. No unit test could have caught it, because every component did exactly what it was specified to do — the bug lives in the sequence, in the fact that nobody decided what should happen when a transcription comes back empty. That decision has no natural home inside any subsystem, because no subsystem knows the others exist.

Giving it a home is what this stage is about: an orchestrator that owns the sequencing, an HTTP surface so something other than a terminal can call it, a container so it runs the same way everywhere, and tests that exercise the seams rather than the parts.

Five green test suites, one thing that owns the sequenceOrchestratorVision handler:file to captionVoice handler: audio to textReasoning engine: ReAct loopMemory manager:store and recallHTTP API and health check
Each subsystem passed its tests in isolation; the bugs that remain all live in the seams between them.

The orchestrator owns the sequence

You could wire everything in main.py. The reason not to is that main.py is about how the program is invoked — argument parsing, the input loop, printing to a terminal. "Voice gets transcribed, then reasoned over, then both spoken and remembered" is not invocation, it is application logic, and it needs to be callable from a second entry point that has no terminal at all.

Python
# src/assistant.pyfrom enum import Enumfrom src.voice_handler import VoiceHandlerfrom src.vision_handler import VisionHandlerfrom src.text_processor import LLMChainManagerfrom src.reasoning_engine import ReActAgent, registryfrom src.memory_manager import MemoryManagerfrom src.exceptions import (AssistantError, VoiceError, VisionError,                            TextProcessingError)from src.utils import loggerclass State(Enum):    IDLE = "idle"; LISTENING = "listening"; PROCESSING = "processing"class MultimodalAssistant:    """The only module permitted to import more than one subsystem."""    def __init__(self, text_only: bool = False):        self.text_only = text_only        self.state = State.IDLE        self.text_processor = LLMChainManager()        self.tools = registry              # already holds calculator, current_time        self.reasoning = ReActAgent(self.text_processor, self.tools)        self.memory = MemoryManager()        # Optional so the service can run on a box with no microphone,        # and so a text-only smoke test does not pay to load two vision        # models and a speech model it will never call.        self.voice = None if text_only else VoiceHandler()        self.vision = None if text_only else VisionHandler()        self._register_tools()    def _register_tools(self) -> None:        # The reasoning engine knows nothing about MemoryManager. It knows        # about a name, a callable and a description. Binding the two happens        # here, at the only point in the codebase that can see both.        @self.tools.register("search_memory",            "Search what the user said before, e.g. 'food allergies'")        def _search(query: str) -> str:            hits = self.memory.semantic.search(query)            return "\n".join(h["text"] for h in hits) or "Nothing relevant found."        @self.tools.register("remember",            "Store a durable fact about the user. Input 'key: value', e.g. 'city: Berlin'")        def _remember(arg: str) -> str:            key, _, value = arg.partition(":")            if not value.strip():                return "Error: expected 'key: value'"            self.memory.facts.remember(key, value.strip())            return f"Stored {key.strip()}."    def process_text(self, text: str) -> str:        """The path every other input mode funnels into."""        if not text.strip():                      # GUARD: the opening bug            raise TextProcessingError("empty input; nothing to process")        self.state = State.PROCESSING        try:            context = self.memory.context_for(text)            question = f"{context}\n\nUser: {text}" if context else text            response = self.reasoning.run(question)["answer"]            self.memory.record("user", text)            self.memory.record("assistant", response)            return response        except AssistantError:            raise                                  # already typed; let it through        except Exception as e:            raise TextProcessingError(str(e)) from e        finally:            self.state = State.IDLE    def process_voice(self, seconds: int = 5) -> str:        if self.text_only:            raise VoiceError("voice is disabled in text-only mode")        self.state = State.LISTENING        try:            text = self.voice.listen(seconds)            if not text.strip():                   # GUARD, again, and earlier                logger.info("Empty transcription; ignoring")                return ""            response = self.process_text(text)     # reuse, never duplicate            self.voice.respond(response)            return response        except AssistantError:            raise        except Exception as e:            raise VoiceError(str(e)) from e        finally:            self.state = State.IDLE    def status(self) -> dict:        return {"state": self.state.value,                "components": {"voice": self.voice is not None,                               "vision": self.vision is not None,                               "reasoning": True, "memory": True}}

Three things in there are load-bearing.

The import rule. This file imports from every subsystem; no subsystem imports from another. That single constraint is what kept each part independently testable while you built it, and it is what will let you swap the vector store or the speech model later without touching anything else. The moment vision_handler.py imports memory_manager, the rule is gone and it does not come back.

process_voice calls process_text. It does not re-implement reasoning and memory. If it did, the two paths would drift — a fix applied to one, forgotten in the other — and you would eventually get different answers depending on whether the user typed or spoke.

The guards sit at the boundary. The empty-transcription check appears twice on purpose: once where the empty string is produced, so it is handled with context, and once at the entry to the shared path, so no future caller can bypass it. Boundary conditions belong to the orchestrator because it is the only place that can see both sides.

A subsystem can only be correct with respect to its own contract. Deciding what happens when one subsystem's edge case meets another's assumption is a separate job, and if no object owns it, nobody does.

Serving it over HTTP

A terminal loop has exactly one possible client: a person at the same keyboard. An HTTP interface means a web front end, a mobile app, another service, or an automated test can all drive the same logic.

Python
# src/api_server.pyfrom contextlib import asynccontextmanagerfrom fastapi import FastAPI, HTTPExceptionfrom pydantic import BaseModelfrom src.assistant import MultimodalAssistantfrom src.exceptions import AssistantErrorfrom src.utils import loggerSTATE = {"assistant": None}@asynccontextmanagerasync def lifespan(app: FastAPI):    # Models load ONCE per process, not per request. Doing this inside a    # request handler would reload roughly 1.2 GB of weights every call.    logger.info("Loading models...")    STATE["assistant"] = MultimodalAssistant(text_only=False)    logger.info("Ready")    yield    STATE["assistant"] = Noneapp = FastAPI(title="Multimodal Assistant", version="1.0.0", lifespan=lifespan)class TextQuery(BaseModel):    query: str@app.get("/health")                       # LIVENESS: is the process alive?async def health():    return {"status": "ok"}@app.get("/ready")                        # READINESS: can it serve traffic?async def ready():    if STATE["assistant"] is None:        raise HTTPException(status_code=503, detail="models still loading")    return STATE["assistant"].status()@app.post("/text")async def process_text(q: TextQuery):    if STATE["assistant"] is None:        raise HTTPException(status_code=503, detail="models still loading")    try:        return {"response": STATE["assistant"].process_text(q.query)}    except AssistantError as e:        # A known failure mode caused by the input: client error, not 500.        raise HTTPException(status_code=422, detail=str(e))    except Exception as e:        logger.error("unexpected in /text: %s", e, exc_info=True)        raise HTTPException(status_code=500, detail="internal server error")

The split between /health and /ready is not pedantry, and it is the first thing that breaks in production. Loading the caption model, the speech model and the embedding model on a CPU container takes about 40 seconds on a cold start. A container platform that probes a single endpoint every 10 seconds and restarts after three failures will kill the process at 30 seconds — before it has ever finished starting — and then do it again, forever. You get an infinite crash loop from a service that has no bug in it.

Liveness answers "is this process wedged?" and must respond immediately without touching anything heavy. Readiness answers "should traffic be routed here?" and returns 503 until the models are up. Set the platform's startup grace period from the number you actually measured, not from a default.

Error mapping matters for the same reason: a typed AssistantError means the request caused a known failure and the client should see 422 with a readable message. A bare exception means your server is broken and should say 500 and log a traceback. Collapsing both into 500 means every client-side mistake looks like an outage, and your error rate graph stops meaning anything.

One honest limitation to write down now: a /voice endpoint built on voice.listen() records from the microphone attached to the server. On your laptop that is your microphone. In a container it is nothing at all. The real design accepts an uploaded audio file and transcribes the bytes; until you have built that, the endpoint should not be exposed in a deployed configuration.

What "works on my machine" actually costs

The project now depends on a Python version, a set of pinned packages, and two system libraries that pip cannot install:

DependencyWhy pip cannot supply itSymptom if missing
ffmpegA native binary the speech library shells out to for audio decodingTranscription raises FileNotFoundError on a file that plays fine locally
portaudio19-devC headers needed to compile the audio capture package from sourcepip install fails mid-build with a compiler error about a missing header
Dockerfile
FROM python:3.12-slimWORKDIR /appRUN apt-get update && apt-get install -y --no-install-recommends \        ffmpeg portaudio19-dev \    && rm -rf /var/lib/apt/lists/*# Dependencies BEFORE application code. This ordering is the difference# between a 3-second rebuild and a 5-minute one; see below.COPY requirements.txt .RUN pip install --no-cache-dir -r requirements.txtCOPY . .ENV PYTHONUNBUFFERED=1 LOG_LEVEL=INFOEXPOSE 8000CMD ["uvicorn", "src.api_server:app", "--host", "0.0.0.0", "--port", "8000"]

Do the arithmetic on the layer ordering, because it is the detail people treat as style. Docker caches each instruction and reuses the cache while nothing above it has changed. pip install for this project takes about 4 minutes 50 seconds, most of it downloading a deep learning wheel of roughly 2.5 GB. Copying the source tree takes about 3 seconds.

OrderRebuild after a one-line code changeTwenty rebuilds in a day
COPY requirements.txt → pip install → COPY . .~3 seconds (only the last layer is invalid)1 minute
COPY . . → pip install~4 min 50 s (the code copy invalidates everything below it)1 hour 37 minutes

The second decision is whether to bake the model weights into the image. Downloading them on first use adds about 40 seconds to the first request and requires network access from the container; baking them in adds roughly 1.2 GB to the image, making the final image about 3.9 GB, and makes startup deterministic and offline-capable. For anything that autoscales, bake them in: a new replica that spends its first minute downloading is a replica that is failing health checks while a queue builds behind it.

YAML
# docker-compose.ymlservices:  assistant:    build: .    ports: ["8000:8000"]    environment:      - OPENAI_API_KEY=${OPENAI_API_KEY}     # from the shell or .env, never baked in      - LOG_LEVEL=INFO    volumes:      - ./logs:/app/logs                     # without these, every restart      - ./data:/app/data                     # discards sessions and the index

Reproducibility is not finished when the Python dependencies are pinned. It is finished when the operating system libraries, the model weights and the start command are pinned too, because those are the three things that differ between your laptop and the machine that will actually run this.

Two rules survive from development straight into deployment. Secrets come from the environment, never from a file inside the image — an image is copied, pushed to registries, and pulled by people who should not have your key. And anything you want to survive a restart must live on a mounted volume, because a container's own filesystem is discarded when it stops. A vector store written to the container filesystem is a vector store that forgets everything every deploy.

Tests for the seams

Until now, every test mocked the things around the component under test. Integration tests deliberately do the opposite: real objects, real boundaries, no mocks between subsystems. They are slower and less deterministic, and they catch a category of bug that no amount of unit testing can reach.

Python
# tests/test_integration.pyimport pytestfrom src.assistant import MultimodalAssistantfrom src.exceptions import TextProcessingError, VoiceError@pytest.fixture(scope="module")def assistant():    # Module scope: constructing this loads models. Per-test construction    # turns a 6-second suite into a 4-minute one.    return MultimodalAssistant(text_only=True)def test_empty_input_is_refused_not_answered(assistant):    """Regression for the 43 stored greetings."""    with pytest.raises(TextProcessingError):        assistant.process_text("   ")def test_nothing_is_recorded_when_input_is_refused(assistant):    before = len(assistant.memory.conversation.turns)    with pytest.raises(TextProcessingError):        assistant.process_text("")    assert len(assistant.memory.conversation.turns) == beforedef test_a_fact_from_one_turn_is_retrievable_in_a_later_one(assistant):    assistant.process_text("I really enjoy hiking in the mountains at weekends")    hits = assistant.memory.semantic.search("what do I like doing outdoors?")    assert hits and "hiking" in hits[0]["text"]def test_typed_errors_survive_the_orchestrator(assistant):    with pytest.raises(VoiceError):        assistant.process_voice()          # text_only=Truedef test_status_reports_disabled_subsystems_honestly(assistant):    s = assistant.status()    assert s["components"]["voice"] is False    assert s["components"]["vision"] is False

The second test is the one worth copying into your own habits. It asserts on a side effect that should not have happened. The opening failure was not that something threw; it was that something was silently written. Tests that only check return values are blind to exactly that class of bug.

Write the integration test that reproduces the bug before you fix it, using the real input that caused it. Otherwise you have fixed one instance of a boundary problem and left the boundary unguarded.

Keep the suite runnable: text-only mode for anything that does not genuinely need the vision or speech models, module-scoped fixtures, and a marker such as @pytest.mark.slow on the few tests that must load real weights so the fast suite stays under ten seconds.

Writing down what is not finished

A deployment document that only says how to start the service is half a document. The half that matters lists what is known to be missing, because an undocumented limitation is indistinguishable from a bug to whoever inherits this.

LimitationWhy it is thereWhat it needs before production
The voice endpoint records from the server's microphoneThe capture method was written for a local terminal, where the assumption heldAccept an uploaded audio file and transcribe from bytes
No authentication on any endpointNothing needed it while the only client was localhostAn API key check or auth middleware; anyone who can reach the port can spend your model budget
No rate limitingA single human typing is self-limitingA per-client limit; one loop in a broken client can exhaust a daily quota in minutes
Plain HTTPLocal development never left the machineTLS at a reverse proxy or load balancer
Unbounded conversation contextSessions were short during developmentA token budget on assembled context; long sessions get slower and dearer every turn
No backup of the data volumeThe data was disposable while buildingA tested restore, not just a backup job — an untested backup is a hope

When things go wrong here

SymptomCauseFix
Every subsystem's tests pass; the assembled assistant misbehaves on real inputA boundary case nobody owns — an empty, oversized, or oddly punctuated value crossing between subsystemsGuard at the orchestrator, add a regression test using the exact input, and assert on side effects as well as return values
Container restarts forever, logs stop partway through model loadingThe health probe fires before startup finishes and the platform kills the processSeparate liveness from readiness; set the grace period above the measured load time
Docker build fails compiling the audio packageportaudio19-dev missing from the imageInstall system libraries before pip install, not after
Every rebuild reinstalls all dependenciesCOPY . . placed above pip install, invalidating the cacheCopy requirements.txt alone, install, then copy the source
Sessions and the vector index vanish on every deployWrites went to the container's own filesystemMount volumes for data/ and logs/; verify by restarting and re-querying
Client mistakes appear as 500s and swamp the error rateAll exceptions collapsed into one handlerTyped errors to 422, unexpected ones to 500 with a logged traceback
Two entry points give different answers to the same questionLogic duplicated in the loop and in a route instead of sharedBoth call the identical orchestrator method; the route should be four lines
First request after deploy times out at the load balancerWeights downloaded on first use rather than baked into the imageBake the models in, and warm the process during startup rather than on request
Responses correct but slow under any concurrencyOne process, CPU inference, blocking calls inside async handlersMeasure which call dominates; run blocking work in a thread pool; scale replicas rather than threads for model work

Acceptance criteria for this stage

  1. MultimodalAssistant constructs in both text-only and full modes, and text-only mode does not load the vision or speech models — check by timing it.
  2. A grep of the codebase shows exactly one file importing from more than one subsystem.
  3. Empty or whitespace-only input raises a typed error, and nothing is written to memory when it does.
  4. A deliberately bad image path produces a typed vision error at the API boundary, returned as 422 with a readable message.
  5. /health answers within milliseconds while the models are still loading; /ready returns 503 until they are up, then reports subsystem status.
  6. A text request through the API returns exactly what calling the orchestrator method directly returns, for the same input.
  7. The image builds from a clean checkout, and a one-line source change rebuilds in seconds rather than minutes.
  8. The container survives a stop and start with sessions and the vector index intact.
  9. The fast test suite runs in under ten seconds; the slow suite passes with real credentials.
  10. You can state, without looking, two documented limitations and why each exists.

What you have when this runs

The assistant now understands typed text, spoken input and images, reasons over multi-step problems with tools, remembers across sessions, searches its own history by meaning, and serves all of it over HTTP from a reproducible container. That is the visible result. The structural result is more useful.

Every subsystem is still replaceable. Because only one file knows the others exist, swapping the vector store, the speech model, or the whole reasoning strategy is a change to one constructor and its tests, not an archaeology project across five modules. Because both entry points call the same methods, adding a third — a scheduled job, a message queue consumer, a chat integration — costs a few lines rather than a reimplementation.

The habit worth carrying forward is the one the opening bug teaches. Systems assembled from correct parts fail at the joins, and they usually fail quietly: an empty string that should have been rejected, a value stored that should have been dropped, a probe that answered before it was true. Loud failures find themselves. Quiet ones are found by someone deciding, explicitly, what should happen when one component's edge case meets another's assumption — writing that decision into the object that owns the sequence, and pinning it with a test that asserts on what did not happen.