Course Content
Python for AI and Data Science
5 sections · 13 lessons
Customizing Plots & Building Dashboards
You put a chart in front of the finance director. It is technically correct: the data is right, the aggregation is right, the trend is real. The axis labels read col_3 and value. There are eight lines in eight default colours, and the legend sits on top of the most important one. She looks at it for four seconds and says, "So what am I meant to take from this?"
That is the failure this lesson is about. The chart was not wrong; it was unread. Every second the viewer spends decoding your encoding is a second they are not spending on your finding, and past about ten seconds most people stop trying. Customisation is not decoration. It is the work of removing everything between the viewer and the point.
The one question to ask before styling anything
What single sentence should the viewer leave with? Write it down. Then make that sentence the title, and make every styling decision serve it.
| Weak title | Title that does the work |
|---|---|
| Monthly Revenue | Revenue fell 23% after the June price change |
| Model Comparison | Gradient boosting beats logistic regression on recall, at twice the training cost |
| Feature Correlations | Only three of eleven features carry independent signal |
A chart title should state the finding, not name the axes. If the title merely repeats what the axis labels already say, it is wasted space at the top of your most valuable real estate.
Once you have the sentence, the rest follows mechanically. Everything that supports it gets emphasis; everything else gets muted or deleted.
Setting a consistent baseline
Do not restyle every chart individually. Set defaults once, at the top of the file, and every plot inherits them.
1import matplotlib.pyplot as plt2import seaborn as sns34sns.set_theme(style="whitegrid", context="talk", palette="colorblind")56plt.rcParams.update({7 "figure.figsize": (10, 6),8 "figure.dpi": 110,9 "axes.titlesize": 15,10 "axes.titleweight": "bold",11 "axes.labelsize": 12,12 "axes.spines.top": False, # remove the box: two spines are enough13 "axes.spines.right": False,14 "grid.alpha": 0.3, # gridlines should whisper, not shout15 "legend.frameon": False,16 "savefig.bbox": "tight",17 "savefig.dpi": 300,18})The context argument scales every font and line at once: "paper", "notebook", "talk", "poster". A chart that reads perfectly in a notebook has unreadable 8-point labels when projected on a wall, and switching context is a one-word fix instead of a dozen font-size arguments.
Colour, chosen deliberately
| Palette type | For | Examples | Fails when |
|---|---|---|---|
| Qualitative | Unordered categories | colorblind, tab10, Set2 | More than about 7 categories |
| Sequential | Low to high | viridis, Blues, magma | Used for categories — implies false ordering |
| Diverging | Deviation from a midpoint | coolwarm, RdBu | The midpoint is not set to zero |
Two rules do most of the work. First, use a colourblind-safe palette by default — roughly one man in twelve cannot reliably distinguish red from green, and a red/green traffic-light chart is illegible to them. viridis and Seaborn's colorblind palette are safe and also survive being photocopied in greyscale, because they vary in lightness as well as hue.
Second, when using a diverging palette, always pin the centre:
sns.heatmap(corr, cmap="coolwarm", center=0, vmin=-1, vmax=1, annot=True)Without center=0, the colour scale stretches to fit whatever range your data happens to have, so a correlation of +0.05 can render as deep red and look like a strong relationship. The colours must mean the same thing in every chart you produce, or the reader learns to distrust them.
Emphasis by muting
The most effective single technique in chart design is to grey out everything except the thing you are talking about.
1fig, ax = plt.subplots()2for name, group in df.groupby("region"):3 highlight = name == "North"4 ax.plot(group.month, group.revenue,5 color="#c0392b" if highlight else "#cccccc",6 linewidth=2.5 if highlight else 1,7 zorder=3 if highlight else 1)8 if highlight:9 last = group.iloc[-1]10 ax.text(last.month, last.revenue, f" {name}",11 color="#c0392b", va="center", fontweight="bold")Eight lines in eight colours means the viewer must consult the legend eight times. One red line among seven grey ones means they see the point immediately. Labelling the line directly at its end, rather than in a legend, removes the lookup entirely — the eye never has to leave the data.
Annotation: telling the viewer where to look
1fig, ax = plt.subplots()2ax.plot(dates, revenue, color="#2c3e50", linewidth=2)34peak = revenue.idxmax()5ax.annotate(f"Peak: £{revenue.max():,.0f}",6 xy=(dates[peak], revenue[peak]), # the point7 xytext=(dates[peak], revenue[peak] * 1.15), # where the text goes8 arrowprops=dict(arrowstyle="->", color="gray"),9 ha="center")1011ax.axhline(revenue.mean(), color="gray", linestyle="--", linewidth=1)12ax.text(dates[0], revenue.mean(), " average", va="bottom", color="gray", fontsize=9)1314ax.axvspan(pd.Timestamp("2024-06-01"), pd.Timestamp("2024-08-31"),15 alpha=0.12, color="orange")16ax.text(pd.Timestamp("2024-07-15"), ax.get_ylim()[1] * 0.95,17 "price change", ha="center", fontsize=9, color="#b9770e")A shaded band with a label saying "price change" turns a line chart into an explanation. Without it, the viewer sees a dip and has to ask what happened; with it, the chart has already answered.
Multi-panel layouts
plt.subplots(rows, cols) gives you a uniform grid. Real figures are rarely uniform — you usually want one large panel and several small supporting ones. subplot_mosaic is the readable way to express that:
1fig, axd = plt.subplot_mosaic(2 """3 AAB4 AAC5 DDD6 """,7 figsize=(13, 9),8 gridspec_kw={"hspace": 0.35, "wspace": 0.3},9)1011axd["A"].scatter(df.income, df.spend, alpha=0.5) # the headline, 2x212axd["B"].hist(df.income, bins=30) # supporting13axd["C"].hist(df.spend, bins=30)14axd["D"].plot(monthly.index, monthly.values) # full-width stripThe ASCII layout string is the actual layout. Each letter is a panel, and repeating a letter makes that panel span those cells. It is far easier to read a month later than the equivalent GridSpec slicing.
Layout rules that keep a multi-panel figure readable
- Share axes when panels are comparable.
sharey=Truemeans two bars of equal height mean equal values; without it, independent scales make wildly different quantities look identical. - Call
fig.tight_layout()orconstrained_layout=True. Otherwise labels from one panel overlap the next, and the figure looks careless. - Cap the panel count. Beyond about six, nobody reads them all; they scan for the biggest one. Split into two figures instead.
Interactivity: when it earns its place
A static image is better for a report, a paper or a slide — it makes one point and cannot be misread. Interactivity is worth the extra weight when the viewer needs to ask their own questions: zoom into a spike, read the exact value of a point, or hide a series to see behind it.
1import plotly.express as px23fig = px.scatter(df, x="income", y="spend",4 color="region", size="orders",5 hover_data=["customer_id", "signup_date"],6 title="Spend against income by region")7fig.update_layout(template="plotly_white", height=520,8 legend=dict(orientation="h", y=-0.2))9fig.write_html("scatter.html") # self-contained, opens in any browserThe hover_data argument is the honest reason to use Plotly. With ten thousand points, a static scatter can show the pattern but not identify the offender; hovering over the outlier and reading its customer ID turns "there is a strange point" into "customer 88213 is strange", which is an action rather than an observation.
| Use | When | Watch out for |
|---|---|---|
| Matplotlib / Seaborn | Reports, papers, slides, anything printed | Nothing — this is the default |
| Plotly | Exploration, sharing an HTML file, dashboards | File size: 50k points makes a browser crawl |
Dashboards
A dashboard is not a pile of charts. It is a page that answers a specific recurring question for a specific person, and its hardest constraint is that they will glance at it for fifteen seconds.
1import streamlit as st2import pandas as pd, plotly.express as px34st.set_page_config(page_title="Sales", layout="wide")56@st.cache_data # re-runs only when the file changes7def load():8 return pd.read_parquet("sales.parquet")910df = load()1112region = st.sidebar.multiselect("Region", sorted(df.region.unique()),13 default=sorted(df.region.unique()))14start, end = st.sidebar.date_input("Period", [df.date.min(), df.date.max()])1516view = df[df.region.isin(region) & df.date.between(pd.Timestamp(start), pd.Timestamp(end))]1718c1, c2, c3 = st.columns(3)19c1.metric("Revenue", f"£{view.revenue.sum():,.0f}", f"{view.revenue.sum()/df.revenue.sum()-1:.1%}")20c2.metric("Orders", f"{len(view):,}")21c3.metric("Average order", f"£{view.revenue.mean():,.2f}")2223st.plotly_chart(24 px.line(view.groupby("date", as_index=False).revenue.sum(), x="date", y="revenue"),25 width="stretch", # fill the column's width26)Run it with streamlit run app.py. The mechanism is that Streamlit re-executes the whole script top to bottom on every interaction — which is why @st.cache_data is not optional. Without it, every click on a filter re-reads the file from disk, and a dashboard over a large dataset becomes unusable.
Layout rules for a dashboard
| Rule | Why |
|---|---|
| Headline numbers at the top | The 15-second reader gets the answer without scrolling |
| Filters in a sidebar, not inline | Controls and content stay visually separate |
| Four to six charts maximum | Beyond that, people stop reading and start scanning |
| Every number carries a comparison | "£2.4m" means nothing; "£2.4m, up 12%" means something |
| State the data's freshness | Prevents decisions being made on last week's figures |
The comparison rule is the one most often skipped and the one that matters most. A metric with no baseline cannot be acted upon, because the viewer has no way to know whether it is good.
Building one figure that carries a decision
Here is the whole approach applied at once: a headline panel with the finding in the title, supporting context, a muted palette with a single highlight, and direct annotation.
1import matplotlib.pyplot as plt, seaborn as sns23sns.set_theme(style="whitegrid", context="talk", palette="colorblind")45fig, axd = plt.subplot_mosaic("AAB\nAAC", figsize=(14, 7), constrained_layout=True)67# A: the headline8for name, g in monthly.groupby("region"):9 focus = name == "North"10 axd["A"].plot(g.month, g.revenue,11 color="#c0392b" if focus else "#d5d8dc",12 linewidth=3 if focus else 1.5, zorder=3 if focus else 1)13 if focus:14 axd["A"].text(g.month.iloc[-1], g.revenue.iloc[-1], " North",15 color="#c0392b", fontweight="bold", va="center")1617axd["A"].axvspan(5.5, 8.5, color="orange", alpha=0.12)18axd["A"].text(7, axd["A"].get_ylim()[1] * 0.96, "price change",19 ha="center", fontsize=10, color="#b9770e")20axd["A"].set_title("North revenue fell 23% after the June price change",21 loc="left", fontweight="bold")22axd["A"].set_ylabel("Revenue (£000)")23axd["A"].set_xlabel("")2425# B and C: supporting context26sns.barplot(data=by_region, x="revenue", y="region", ax=axd["B"],27 color="#d5d8dc", order=by_region.sort_values("revenue").region)28axd["B"].set_title("Total by region", loc="left", fontsize=11)29axd["B"].set_xlabel(""); axd["B"].set_ylabel("")3031sns.histplot(data=orders, x="order_value", bins=30, ax=axd["C"], color="#5d6d7e")32axd["C"].set_title("Order value distribution", loc="left", fontsize=11)33axd["C"].set_xlabel("£"); axd["C"].set_ylabel("")3435fig.savefig("north_revenue.png", dpi=300, bbox_inches="tight")Notice how little of that code is about drawing data. Three lines plot the numbers; the rest decides what the viewer notices first, second and third. That ratio is normal, and it is the difference between a chart that gets a decision made and one that gets the question "so what am I meant to take from this?"
When you are unsure whether a styling change is worth it, use the ten-second test: show the figure to someone who has not seen the data, count to ten, take it away, and ask what it said. If they cannot tell you, the problem is almost never the data — it is that the finding was never made visible.