Course Content
Machine Learning Essentials
6 sections · 16 lessons
Deployment with Streamlit & Flask
Your spam classifier gets 96% on held-out data. The product manager wants to try it on some real emails.
So you send her the notebook and the model file. She does not have Python. You spend a morning walking her through installing Anaconda; the environment resolves to a different scikit-learn version and the model fails to load. You give up and offer to run emails yourself, pasting results into a spreadsheet. Two days later the backend engineer asks how his service should call your model. You have no answer, because there is nothing to call.
The model was finished. The work was not. A trained model is a file that only helps people who can run Python, load the right dependencies, and shape their data into the exact format it expects — which is you, and possibly nobody else.
Deployment is the step that turns that file into something other people can use, and there are two fundamentally different destinations depending on who "other people" are.
Two audiences, two answers
| A human needs to try it | A program needs to call it | |
|---|---|---|
| What you build | A web app with inputs and a visible result | An API returning JSON |
| Interface | Buttons, sliders, text boxes | An HTTP endpoint |
| Tool | Streamlit | Flask |
| Typical volume | Tens of requests a day | Thousands a minute |
| Purpose | Demos, internal tools, exploration | Production integration |
These are not competitors. Most projects eventually want both — an API carrying the production traffic, and a small app so that the people who requested the model can see it working.
Client and server, in plain terms
Both approaches rest on the same arrangement, and it is worth stating plainly because the vocabulary trips people up.
A server is a program that starts up, loads your model into memory, and waits. A client — a browser, a mobile app, another service — sends a request. The server does the work and sends back a response.
CLIENT SERVER | | (started earlier, model in memory) |--- POST /predict --------------->| | {"income": 52000, "age": 34} | validate input | | model.predict_proba(...) |<-- 200 OK -----------------------| | {"probability": 0.83} |Four pieces of vocabulary carry most of the weight:
- Method —
GETretrieves something,POSTsends data for processing. Predictions use POST because the input is a body of data, not a name in a URL. - Path —
/predict,/health. Each is a separate function on the server. - Body — the data, almost always JSON.
- Status code — 200 succeeded, 400 the client sent something invalid, 422 the values failed validation, 500 the server broke. Getting these right is how the caller's code knows whether to retry.
Streamlit: an interface without writing any front-end code
Streamlit turns a Python script into a web app. There is no HTML, no JavaScript, no callbacks — you write top-to-bottom Python and Streamlit renders it.
The mental model that explains everything
Streamlit's execution model surprises everyone once, so learn it before writing anything: every time the user interacts with any widget, Streamlit reruns your entire script from the first line.
Not a callback. Not the affected section. The whole file, top to bottom, with the widget now returning its new value.
This is why the framework feels so simple — there is no event wiring, because state changes just rerun everything. It is also why the naive version is unusably slow:
1import streamlit as st2import joblib34model = joblib.load("pipeline.joblib") # WRONG: reloads on every keystrokeMove a slider and that 200 MB file is deserialised again. The fix is a decorator that tells Streamlit to compute something once and reuse it:
1@st.cache_resource2def load_model():3 return joblib.load("pipeline.joblib")45model = load_model() # loaded once per server processTwo decorators, used for different things:
| Decorator | For | Behaviour |
|---|---|---|
@st.cache_resource | Models, database connections | One shared object across all sessions; not copied |
@st.cache_data | DataFrames, computed results | Cached per argument value; returns a copy so callers cannot corrupt it |
Getting these the wrong way round is the most common Streamlit bug. Caching a model with cache_data copies it on every access; caching a DataFrame with cache_resource lets one user's mutation leak into another user's session.
A complete app
1import streamlit as st2import pandas as pd3import joblib45st.set_page_config(page_title="Churn risk", page_icon="📉")67@st.cache_resource8def load_model():9 return joblib.load("pipeline.joblib")1011model = load_model()1213st.title("Customer churn risk")14st.caption("Enter customer details to estimate the probability of churn.")1516col1, col2 = st.columns(2)17with col1:18 income = st.number_input("Annual income (£)", 0, 500_000, 45_000, step=1_000)19 age = st.slider("Age", 18, 90, 38)20 tenure = st.slider("Years as customer", 0.0, 30.0, 3.5, step=0.5)21with col2:22 products = st.selectbox("Products held", [1, 2, 3, 4, 5])23 region = st.selectbox("Region", ["north", "south", "east", "west"])24 contract = st.checkbox("On a contract", value=True)2526threshold = st.sidebar.slider("Decision threshold", 0.05, 0.95, 0.34, 0.01)2728if st.button("Predict", type="primary"):29 row = pd.DataFrame([{30 "income": income, "age": age, "tenure_years": tenure,31 "num_products": products, "region": region,32 "has_contract": int(contract),33 }])3435 prob = float(model.predict_proba(row)[0, 1])3637 st.metric("Churn probability", f"{prob:.1%}")38 st.progress(prob)3940 if prob >= threshold:41 st.error(f"Flagged for retention outreach (threshold {threshold:.0%})")42 else:43 st.success("No action needed")4445 with st.expander("What the model received"):46 st.dataframe(row)Run it with streamlit run app.py. That is roughly fifty lines for a working internal tool, which is why Streamlit has become the default for demos.
Two details worth copying. The threshold is a sidebar control rather than a constant, so the person using the tool can see the trade-off between catching churners and generating false alarms rather than being handed a fixed answer. And the expander showing the exact input frame turns "the model gave a weird answer" into a debuggable report.
Where Streamlit stops
Streamlit is a UI framework, not a service platform. It has no clean way to be called by another program, its reruns make anything stateful awkward, and it holds one script execution per connected user, which limits concurrency sharply. Use it for people; use something else for machines.
Flask: an API other programs can call
1from flask import Flask, request, jsonify2import pandas as pd3import joblib4import logging56app = Flask(__name__)7log = logging.getLogger(__name__)89MODEL = joblib.load("pipeline.joblib") # once, at import10MODEL_VERSION = "2.3.0"11FEATURES = ["income", "age", "tenure_years", "num_products",12 "region", "has_contract"]1314@app.get("/health")15def health():16 return jsonify(status="ok", model_version=MODEL_VERSION)1718@app.post("/predict")19def predict():20 payload = request.get_json(silent=True)21 if payload is None:22 return jsonify(error="body must be valid JSON"), 4002324 missing = [f for f in FEATURES if f not in payload]25 if missing:26 return jsonify(error="missing fields", fields=missing), 4222728 try:29 row = pd.DataFrame([{f: payload[f] for f in FEATURES}])30 prob = float(MODEL.predict_proba(row)[0, 1])31 except Exception:32 log.exception("prediction failed")33 return jsonify(error="prediction failed"), 5003435 return jsonify(36 probability=round(prob, 4),37 prediction=int(prob >= 0.34),38 model_version=MODEL_VERSION,39 )Calling it:
1curl -X POST http://localhost:5000/predict \2 -H "Content-Type: application/json" \3 -d '{"income": 52000, "age": 34, "tenure_years": 2.5,4 "num_products": 2, "region": "north", "has_contract": 1}'56# {"model_version":"2.3.0","prediction":1,"probability":0.8312}Three things in that handler are not decoration.
The health endpoint is how load balancers and orchestrators decide whether this instance should receive traffic. Without one, a process whose model failed to load keeps receiving requests.
The model version in the response means every logged prediction can later be traced to the artefact that produced it. When results shift, that field answers "did the model change?" in seconds.
The catch-all except that logs and returns 500 ensures an unexpected input produces a clean error rather than a stack trace in the response body. Returning internal tracebacks to callers leaks file paths and library versions.
Validate before the model ever sees the data
A model cannot tell you its input was nonsense. Give it an age of 900 and it will return a confident probability, because 900 is a perfectly good float.
| Check | Catches | Response |
|---|---|---|
| Field present | Typos, schema drift in the caller | 422 with the missing field names |
| Correct type | "thirty-four" where a number is expected | 422 |
| Plausible range | age 900, income −5,000 | 422 |
| Category is known | region: "atlantis" | 422, or rely on handle_unknown="ignore" |
| Inside training range | income of £40m when training topped out at £500k | 200, but flag the response as low-confidence |
| Not all defaults | A caller sending an empty template | 422 |
1RANGES = {"income": (0, 1_000_000), "age": (18, 100),2 "tenure_years": (0, 60), "num_products": (1, 10)}3CATEGORIES = {"region": {"north", "south", "east", "west"}}45def validate(payload):6 errors = []7 for field, (lo, hi) in RANGES.items():8 value = payload.get(field)9 if not isinstance(value, (int, float)) or isinstance(value, bool):10 errors.append(f"{field} must be a number")11 elif not (lo <= value <= hi):12 errors.append(f"{field} must be between {lo} and {hi}, got {value}")13 for field, allowed in CATEGORIES.items():14 if payload.get(field) not in allowed:15 errors.append(f"{field} must be one of {sorted(allowed)}")16 return errorsThe isinstance(value, bool) exclusion is not pedantry: in Python True is an instance of int, so a caller sending "age": true would otherwise sail through as the number 1.
An unvalidated API does not fail when it receives bad data. It returns a confident number, and someone acts on it.
Choosing between them
| Streamlit | Flask | |
|---|---|---|
| Audience | People | Programs |
| Lines for a working version | ~30 | ~40, plus a client |
| Front-end code required | None | None — but there is no UI either |
| Callable from another service | Not practically | Yes, that is the point |
| Concurrency | Poor — a script run per user | Good, with a proper WSGI server |
| Input validation | Constrained by the widgets | You must write it |
| Authentication | Awkward | Standard middleware |
| Best for | Demos, internal tools, exploration | Production integration |
A common and sensible arrangement is to build the Flask API first and then write the Streamlit app as a client of it. The app calls POST /predict over HTTP rather than holding its own copy of the model. One model, one validation path, two front doors.
Getting it off your laptop
When Flask starts it prints a warning that people ignore:
WARNING: This is a development server. Do not use it in a production deployment.It means it. The built-in server handles one request at a time, has no process management, and no protection against slow clients. In production, run the app under a proper WSGI server:
gunicorn --workers 4 --bind 0.0.0.0:8000 --timeout 60 app:appEach worker is a separate process with its own copy of the model in memory — worth remembering when the model is 200 MB and you asked for eight workers. A reasonable starting point is 2 × cores + 1, adjusted downwards if memory is the constraint.
Then make the environment reproducible with a container:
1# Dockerfile2FROM python:3.12-slim3WORKDIR /app4COPY requirements.txt .5RUN pip install --no-cache-dir -r requirements.txt6COPY app.py pipeline.joblib ./7EXPOSE 80008CMD ["gunicorn", "--workers", "4", "--bind", "0.0.0.0:8000", "app:app"]with exact pins in requirements.txt:
flask==3.1.3gunicorn==26.2.0scikit-learn==1.9.1pandas==3.0.6joblib==1.6.0The versions above were current when this was written; pin whatever you trained with. Those pins are the whole point. A model file saved under scikit-learn 1.9.1 and loaded under a different version may fail, or may load and behave slightly differently. Pinning turns "works on my machine" into "works wherever this image runs".
Finally, keep configuration out of the code. Model paths, thresholds, and log levels belong in environment variables so that staging and production can differ without a code change:
1import os23MODEL_PATH = os.environ.get("MODEL_PATH", "pipeline.joblib")4THRESHOLD = float(os.environ.get("DECISION_THRESHOLD", "0.34"))Log what you predicted
One habit separates a service you can improve from one you can only restart. Log every prediction — inputs, output, model version, timestamp.
1log.info("prediction", extra={2 "model_version": MODEL_VERSION,3 "features": {f: payload[f] for f in FEATURES},4 "probability": prob,5})Those logs are what let you detect that the average incoming income has shifted 30% since training, or compare predictions against outcomes once they arrive, or reproduce exactly what happened when a customer complains. Without them a deployed model is a black box that you can only observe by its silence.
What this means when you build something
Decide who is calling before you choose a tool. If the answer is "a colleague who wants to click something", write the Streamlit app — it will take an afternoon and the conversation it enables is worth more than another point of accuracy. If the answer is "our backend, on every checkout", write the API, and accept that validation, logging, health checks, and pinned dependencies are part of the job rather than polish to add later.
Whichever you build, load the model exactly once when the process starts, keep every preprocessing step inside the saved pipeline so the serving code never reimplements it, and validate every field before it reaches predict. Those three things prevent the great majority of production incidents involving machine learning models, and none of them takes more than twenty lines.