Python for AI and Data Science

Customizing Plots & Building Dashboards


You put a chart in front of the finance director. It is technically correct: the data is right, the aggregation is right, the trend is real. The axis labels read col_3 and value. There are eight lines in eight default colours, and the legend sits on top of the most important one. She looks at it for four seconds and says, "So what am I meant to take from this?"

That is the failure this lesson is about. The chart was not wrong; it was unread. Every second the viewer spends decoding your encoding is a second they are not spending on your finding, and past about ten seconds most people stop trying. Customisation is not decoration. It is the work of removing everything between the viewer and the point.

Styling in the order that actually helps the readerName the decision the figure has to supportSet one baseline: fonts, sizes, figure sizeMute every series except the one that mattersAnnotate the single point you want seenArrange panels so the eye reads them in order
The finance director spent four seconds on it — everything above buys back some of those seconds.

The one question to ask before styling anything

What single sentence should the viewer leave with? Write it down. Then make that sentence the title, and make every styling decision serve it.

Weak titleTitle that does the work
Monthly RevenueRevenue fell 23% after the June price change
Model ComparisonGradient boosting beats logistic regression on recall, at twice the training cost
Feature CorrelationsOnly three of eleven features carry independent signal

A chart title should state the finding, not name the axes. If the title merely repeats what the axis labels already say, it is wasted space at the top of your most valuable real estate.

Once you have the sentence, the rest follows mechanically. Everything that supports it gets emphasis; everything else gets muted or deleted.

Setting a consistent baseline

Do not restyle every chart individually. Set defaults once, at the top of the file, and every plot inherits them.

Python
import matplotlib.pyplot as pltimport seaborn as snssns.set_theme(style="whitegrid", context="talk", palette="colorblind")plt.rcParams.update({    "figure.figsize": (10, 6),    "figure.dpi": 110,    "axes.titlesize": 15,    "axes.titleweight": "bold",    "axes.labelsize": 12,    "axes.spines.top": False,      # remove the box: two spines are enough    "axes.spines.right": False,    "grid.alpha": 0.3,             # gridlines should whisper, not shout    "legend.frameon": False,    "savefig.bbox": "tight",    "savefig.dpi": 300,})

The context argument scales every font and line at once: "paper", "notebook", "talk", "poster". A chart that reads perfectly in a notebook has unreadable 8-point labels when projected on a wall, and switching context is a one-word fix instead of a dozen font-size arguments.

Colour, chosen deliberately

Palette typeForExamplesFails when
QualitativeUnordered categoriescolorblind, tab10, Set2More than about 7 categories
SequentialLow to highviridis, Blues, magmaUsed for categories — implies false ordering
DivergingDeviation from a midpointcoolwarm, RdBuThe midpoint is not set to zero

Two rules do most of the work. First, use a colourblind-safe palette by default — roughly one man in twelve cannot reliably distinguish red from green, and a red/green traffic-light chart is illegible to them. viridis and Seaborn's colorblind palette are safe and also survive being photocopied in greyscale, because they vary in lightness as well as hue.

Second, when using a diverging palette, always pin the centre:

Python
sns.heatmap(corr, cmap="coolwarm", center=0, vmin=-1, vmax=1, annot=True)

Without center=0, the colour scale stretches to fit whatever range your data happens to have, so a correlation of +0.05 can render as deep red and look like a strong relationship. The colours must mean the same thing in every chart you produce, or the reader learns to distrust them.

Emphasis by muting

The most effective single technique in chart design is to grey out everything except the thing you are talking about.

Python
fig, ax = plt.subplots()for name, group in df.groupby("region"):    highlight = name == "North"    ax.plot(group.month, group.revenue,            color="#c0392b" if highlight else "#cccccc",            linewidth=2.5 if highlight else 1,            zorder=3 if highlight else 1)    if highlight:        last = group.iloc[-1]        ax.text(last.month, last.revenue, f"  {name}",                color="#c0392b", va="center", fontweight="bold")

Eight lines in eight colours means the viewer must consult the legend eight times. One red line among seven grey ones means they see the point immediately. Labelling the line directly at its end, rather than in a legend, removes the lookup entirely — the eye never has to leave the data.

Annotation: telling the viewer where to look

Python
fig, ax = plt.subplots()ax.plot(dates, revenue, color="#2c3e50", linewidth=2)peak = revenue.idxmax()ax.annotate(f"Peak: £{revenue.max():,.0f}",            xy=(dates[peak], revenue[peak]),            # the point            xytext=(dates[peak], revenue[peak] * 1.15), # where the text goes            arrowprops=dict(arrowstyle="->", color="gray"),            ha="center")ax.axhline(revenue.mean(), color="gray", linestyle="--", linewidth=1)ax.text(dates[0], revenue.mean(), " average", va="bottom", color="gray", fontsize=9)ax.axvspan(pd.Timestamp("2024-06-01"), pd.Timestamp("2024-08-31"),           alpha=0.12, color="orange")ax.text(pd.Timestamp("2024-07-15"), ax.get_ylim()[1] * 0.95,        "price change", ha="center", fontsize=9, color="#b9770e")

A shaded band with a label saying "price change" turns a line chart into an explanation. Without it, the viewer sees a dip and has to ask what happened; with it, the chart has already answered.

Multi-panel layouts

plt.subplots(rows, cols) gives you a uniform grid. Real figures are rarely uniform — you usually want one large panel and several small supporting ones. subplot_mosaic is the readable way to express that:

Python
fig, axd = plt.subplot_mosaic(    """    AAB    AAC    DDD    """,    figsize=(13, 9),    gridspec_kw={"hspace": 0.35, "wspace": 0.3},)axd["A"].scatter(df.income, df.spend, alpha=0.5)   # the headline, 2x2axd["B"].hist(df.income, bins=30)                  # supportingaxd["C"].hist(df.spend, bins=30)axd["D"].plot(monthly.index, monthly.values)       # full-width strip

The ASCII layout string is the actual layout. Each letter is a panel, and repeating a letter makes that panel span those cells. It is far easier to read a month later than the equivalent GridSpec slicing.

Layout rules that keep a multi-panel figure readable

  • Share axes when panels are comparable. sharey=True means two bars of equal height mean equal values; without it, independent scales make wildly different quantities look identical.
  • Call fig.tight_layout() or constrained_layout=True. Otherwise labels from one panel overlap the next, and the figure looks careless.
  • Cap the panel count. Beyond about six, nobody reads them all; they scan for the biggest one. Split into two figures instead.

Interactivity: when it earns its place

A static image is better for a report, a paper or a slide — it makes one point and cannot be misread. Interactivity is worth the extra weight when the viewer needs to ask their own questions: zoom into a spike, read the exact value of a point, or hide a series to see behind it.

Python
import plotly.express as pxfig = px.scatter(df, x="income", y="spend",                 color="region", size="orders",                 hover_data=["customer_id", "signup_date"],                 title="Spend against income by region")fig.update_layout(template="plotly_white", height=520,                  legend=dict(orientation="h", y=-0.2))fig.write_html("scatter.html")     # self-contained, opens in any browser

The hover_data argument is the honest reason to use Plotly. With ten thousand points, a static scatter can show the pattern but not identify the offender; hovering over the outlier and reading its customer ID turns "there is a strange point" into "customer 88213 is strange", which is an action rather than an observation.

UseWhenWatch out for
Matplotlib / SeabornReports, papers, slides, anything printedNothing — this is the default
PlotlyExploration, sharing an HTML file, dashboardsFile size: 50k points makes a browser crawl

Dashboards

A dashboard is not a pile of charts. It is a page that answers a specific recurring question for a specific person, and its hardest constraint is that they will glance at it for fifteen seconds.

Python
import streamlit as stimport pandas as pd, plotly.express as pxst.set_page_config(page_title="Sales", layout="wide")@st.cache_data                      # re-runs only when the file changesdef load():    return pd.read_parquet("sales.parquet")df = load()region = st.sidebar.multiselect("Region", sorted(df.region.unique()),                                default=sorted(df.region.unique()))start, end = st.sidebar.date_input("Period", [df.date.min(), df.date.max()])view = df[df.region.isin(region) & df.date.between(pd.Timestamp(start), pd.Timestamp(end))]c1, c2, c3 = st.columns(3)c1.metric("Revenue", f"£{view.revenue.sum():,.0f}", f"{view.revenue.sum()/df.revenue.sum()-1:.1%}")c2.metric("Orders", f"{len(view):,}")c3.metric("Average order", f"£{view.revenue.mean():,.2f}")st.plotly_chart(    px.line(view.groupby("date", as_index=False).revenue.sum(), x="date", y="revenue"),    width="stretch",                # fill the column's width)

Run it with streamlit run app.py. The mechanism is that Streamlit re-executes the whole script top to bottom on every interaction — which is why @st.cache_data is not optional. Without it, every click on a filter re-reads the file from disk, and a dashboard over a large dataset becomes unusable.

Layout rules for a dashboard

RuleWhy
Headline numbers at the topThe 15-second reader gets the answer without scrolling
Filters in a sidebar, not inlineControls and content stay visually separate
Four to six charts maximumBeyond that, people stop reading and start scanning
Every number carries a comparison"£2.4m" means nothing; "£2.4m, up 12%" means something
State the data's freshnessPrevents decisions being made on last week's figures

The comparison rule is the one most often skipped and the one that matters most. A metric with no baseline cannot be acted upon, because the viewer has no way to know whether it is good.

Building one figure that carries a decision

Here is the whole approach applied at once: a headline panel with the finding in the title, supporting context, a muted palette with a single highlight, and direct annotation.

Python
import matplotlib.pyplot as plt, seaborn as snssns.set_theme(style="whitegrid", context="talk", palette="colorblind")fig, axd = plt.subplot_mosaic("AAB\nAAC", figsize=(14, 7), constrained_layout=True)# A: the headlinefor name, g in monthly.groupby("region"):    focus = name == "North"    axd["A"].plot(g.month, g.revenue,                  color="#c0392b" if focus else "#d5d8dc",                  linewidth=3 if focus else 1.5, zorder=3 if focus else 1)    if focus:        axd["A"].text(g.month.iloc[-1], g.revenue.iloc[-1], "  North",                      color="#c0392b", fontweight="bold", va="center")axd["A"].axvspan(5.5, 8.5, color="orange", alpha=0.12)axd["A"].text(7, axd["A"].get_ylim()[1] * 0.96, "price change",              ha="center", fontsize=10, color="#b9770e")axd["A"].set_title("North revenue fell 23% after the June price change",                   loc="left", fontweight="bold")axd["A"].set_ylabel("Revenue (£000)")axd["A"].set_xlabel("")# B and C: supporting contextsns.barplot(data=by_region, x="revenue", y="region", ax=axd["B"],            color="#d5d8dc", order=by_region.sort_values("revenue").region)axd["B"].set_title("Total by region", loc="left", fontsize=11)axd["B"].set_xlabel(""); axd["B"].set_ylabel("")sns.histplot(data=orders, x="order_value", bins=30, ax=axd["C"], color="#5d6d7e")axd["C"].set_title("Order value distribution", loc="left", fontsize=11)axd["C"].set_xlabel("£"); axd["C"].set_ylabel("")fig.savefig("north_revenue.png", dpi=300, bbox_inches="tight")

Notice how little of that code is about drawing data. Three lines plot the numbers; the rest decides what the viewer notices first, second and third. That ratio is normal, and it is the difference between a chart that gets a decision made and one that gets the question "so what am I meant to take from this?"

When you are unsure whether a styling change is worth it, use the ten-second test: show the figure to someone who has not seen the data, count to ten, take it away, and ask what it said. If they cannot tell you, the problem is almost never the data — it is that the finding was never made visible.