Course Content
Data Science Fundamentals
4 sections · 10 lessons
Visual Analytics
A support team reports that the average ticket resolution time is 30 minutes. Management is satisfied — 30 minutes is the target. Then somebody plots the actual distribution.
count 60 | ### ### 40 | ##### ##### 20 | ####### ####### 0 +---------------------------------------------- 5 10 15 20 25 30 35 40 45 50 minutesThere are two humps and almost nothing in the middle. Roughly half the tickets resolve in about 10 minutes and half take about 45. Nobody experiences a 30-minute resolution. The mean sits in the empty valley between two entirely different processes — a self-service path and a path that requires a human specialist — and it describes neither.
No summary statistic could have revealed that. The mean, median, standard deviation and skewness of this distribution are all unremarkable. It took thirty seconds of plotting to see something that months of reporting had hidden. That is the case for visual analytics: not to decorate findings after the fact, but to see structure that numbers compress away.
Choosing the Right Picture
Almost every exploratory chart answers one of five questions. Start from the question and the chart type follows.
| Question | Variables | Chart |
|---|---|---|
| What does this variable look like? | 1 numeric | Histogram, density plot, boxplot |
| How do these two move together? | 2 numeric | Scatter, hexbin, regression plot |
| How do groups compare? | 1 categorical + 1 numeric | Bar with error bars, boxplot, violin |
| How do categories relate? | 2 categorical | Grouped/stacked bar, heatmap of counts |
| How does this change over time? | Time + numeric | Line plot |
Seeing a Single Variable
Histograms and the bin-count problem
1import matplotlib.pyplot as plt2import seaborn as sns34fig, axes = plt.subplots(1, 3, figsize=(15, 4))5for ax, bins in zip(axes, [8, 30, 120]):6 ax.hist(df["resolution_minutes"], bins=bins, edgecolor="white")7 ax.set_title(f"{bins} bins")8plt.tight_layout()Run that on the support data and the point becomes obvious: with 8 bins the two humps merge into one broad lump and the bimodality vanishes; with 120 bins the shape dissolves into noise; around 30 the structure is unmistakable. The bin count is not a cosmetic setting. It decides what you are able to see.
Since no single choice is right, always look at two or three. The Freedman–Diaconis rule gives a sensible starting width, scaling with the interquartile range and shrinking as the sample grows:
1import numpy as np2x = df["resolution_minutes"].dropna()3iqr = x.quantile(0.75) - x.quantile(0.25)4width = 2 * iqr / len(x) ** (1/3)5print("suggested bins:", int((x.max() - x.min()) / width))Density plots
sns.kdeplot(data=df, x="resolution_minutes", hue="channel", fill=True, alpha=0.35, common_norm=False)A kernel density estimate smooths the histogram into a curve, which makes overlaying several groups readable where overlapping histograms become mud. Two cautions. It smooths, so a small bump may be an artefact of the bandwidth rather than a feature of the data — check it against a histogram before believing it. And it will happily draw density below zero for a strictly positive quantity like duration, because the smoother does not know about the boundary.
common_norm=False matters when groups differ in size: it normalises each group separately, so you compare shapes rather than comparing "how many rows are in this group", which the bar heights already told you.
Boxplots and violins
A boxplot draws five numbers: the median as the centre line, the first and third quartiles as the box edges, whiskers extending to the furthest points within 1.5 × IQR of the box, and everything beyond as individual dots.
1fig, axes = plt.subplots(1, 2, figsize=(13, 5), sharey=True)2sns.boxplot(data=df, x="team", y="resolution_minutes", ax=axes[0])3sns.violinplot(data=df, x="team", y="resolution_minutes", ax=axes[1], inner="quartile")4axes[0].set_title("Boxplot: five numbers")5axes[1].set_title("Violin: full shape")Here is the trap that the opening example illustrates: a boxplot of bimodal data looks completely ordinary. The median sits in the empty valley, the box spans both humps, and nothing signals that there are two populations. Violin plots draw the density along each side, so both humps appear. When comparing groups where distribution shape might differ, use violins.
Boxplots compare centres and spreads. They cannot show shape, so they hide exactly the structure that makes a summary statistic misleading.
| Chart | Shows | Best when | Hides |
|---|---|---|---|
| Histogram | Full shape, one group | Exploring a single variable | Sensitive to bin choice |
| Density (KDE) | Smoothed shape, several groups | Comparing 2–4 distributions | Small features; boundaries |
| Boxplot | Median, IQR, outliers | Comparing many groups compactly | Multimodality, sample size |
| Violin | Shape per group | Groups may differ in shape | Needs enough data per group |
| Strip / swarm | Every individual point | Fewer than ~200 points | Becomes unreadable when dense |
Seeing Two Variables Together
sns.scatterplot(data=df, x="tenure_months", y="lifetime_value", hue="segment", alpha=0.5, s=25)Once you pass a few thousand points, a scatter plot becomes a solid block and the eye can no longer judge where the density is. Three fixes, in order of how much data they handle:
1fig, axes = plt.subplots(1, 3, figsize=(16, 4.5))23# 1. Transparency - works to roughly 20k points4axes[0].scatter(df.x, df.y, alpha=0.05, s=8)56# 2. Hexagonal binning - counts per hexagon7hb = axes[1].hexbin(df.x, df.y, gridsize=45, cmap="Blues", mincnt=1)8plt.colorbar(hb, ax=axes[1], label="count")910# 3. 2D density contours11sns.kdeplot(data=df, x="x", y="y", fill=True, levels=8, ax=axes[2])Hexbin is the workhorse for large scatter data. It shows where the mass is instead of where the ink is, and unlike transparency it does not saturate.
Adding a trend without lying about it
1fig, axes = plt.subplots(1, 2, figsize=(13, 5))2sns.regplot(data=df, x="tenure_months", y="lifetime_value", ax=axes[0],3 scatter_kws={"alpha": 0.3}, line_kws={"color": "crimson"})4sns.regplot(data=df, x="tenure_months", y="lifetime_value", ax=axes[1],5 lowess=True, scatter_kws={"alpha": 0.3}, line_kws={"color": "crimson"})6axes[0].set_title("Straight-line fit")7axes[1].set_title("LOWESS: follows the data")Draw both. If the straight line and the LOWESS curve diverge, the relationship is not linear and any correlation coefficient you computed is understating it. This is the visual version of the check that saves you from dropping a genuinely predictive feature because a linear measure scored it near zero.
Correlation heatmaps
1corr = df.select_dtypes("number").corr()2mask = np.triu(np.ones_like(corr, dtype=bool)) # hide the mirror half34sns.heatmap(corr, mask=mask, cmap="RdBu_r", center=0, vmin=-1, vmax=1,5 annot=True, fmt=".2f", square=True, linewidths=0.5,6 cbar_kws={"shrink": 0.7})Three settings do the real work. center=0 with a diverging colour map puts white at zero, so positive and negative are visually distinct rather than being two shades of the same colour. vmin and vmax fix the scale to the full range, so heatmaps from different datasets are comparable — without them, a matrix whose strongest correlation is 0.3 renders that 0.3 in the same deep colour another matrix uses for 0.95. Masking the upper triangle removes redundant information, since the matrix is symmetric.
Categorical Comparisons
order = df.groupby("region")["revenue"].mean().sort_values().indexsns.barplot(data=df, y="region", x="revenue", order=order, errorbar=("ci", 95), color="#4C72B0")Three deliberate choices there. Sorting by value rather than alphabetically means the ranking is readable at a glance — alphabetical order encodes nothing. Horizontal orientation keeps category labels readable without rotation. And the confidence interval prevents the most common misreading of a bar chart: two bars of different heights whose intervals overlap heavily are not evidence of a real difference, and a plain bar chart gives no way to tell.
Bars must start at zero. Truncating the axis at 90 to make a difference between 92 and 95 look dramatic is the classic misleading chart, because a bar's meaning is carried by its length. Line and scatter plots have no such constraint — they encode position, not length — so zooming their axes is legitimate.
Pie charts
Use them for at most three or four slices where the point is "roughly half" or "about a quarter". Human beings judge angles poorly, so a pie with nine slices takes longer to read than a sorted bar chart and conveys less. If the slices need labelling with their percentages to be legible, the chart is doing no work and a table would serve better.
Two categoricals at once
ct = pd.crosstab(df["region"], df["product"], normalize="index") * 100sns.heatmap(ct, annot=True, fmt=".1f", cmap="YlGnBu", cbar_kws={"label": "% of region"})normalize="index" converts counts into row percentages, which is usually the question you mean. Raw counts mostly show which regions are large; percentages show which products are disproportionately popular where.
More Than Two Variables
1# Every pairwise relationship plus the diagonal distributions2sns.pairplot(df[["age", "income", "spend", "tenure", "segment"]],3 hue="segment", diag_kind="kde", plot_kws={"alpha": 0.4, "s": 18})45# The same relationship, repeated per subgroup6g = sns.FacetGrid(df, col="region", row="segment", height=3, margin_titles=True)7g.map_dataframe(sns.scatterplot, x="tenure", y="spend", alpha=0.5)Pairplots are the fastest way to survey a small numeric dataset, but the number of panels grows as the square of the columns — ten columns produce a hundred panels. Restrict to eight or fewer, and sample rows if there are many.
Facet grids ("small multiples") are the most underrated technique here. Repeating one simple chart across subgroups, on shared axes, lets the eye compare directly. It reveals Simpson's paradox situations, where an overall trend reverses within every subgroup — something no single aggregate chart can show.
Time
1ts = df.set_index("date")["revenue"].resample("D").sum()23fig, ax = plt.subplots(figsize=(13, 5))4ax.plot(ts.index, ts, alpha=0.3, lw=0.8, label="daily")5ax.plot(ts.index, ts.rolling(7, center=True).mean(), lw=2, label="7-day average")6ax.fill_between(ts.index,7 ts.rolling(7).mean() - ts.rolling(7).std(),8 ts.rolling(7).mean() + ts.rolling(7).std(),9 alpha=0.15)10ax.legend()11ax.set_ylabel("Revenue (GBP)")Plotting the raw series and a rolling average together is the standard pattern: the faint line preserves the real variation, the bold line shows the trend. A 7-day window on daily data removes the day-of-week cycle specifically, which is why weekly business data almost always wants exactly 7 rather than a rounder number.
For strong seasonality, decomposition separates the components explicitly:
from statsmodels.tsa.seasonal import seasonal_decomposeseasonal_decompose(ts.dropna(), model="additive", period=7).plot()Making Figures Someone Else Can Read
An exploratory chart has an audience of one and can be scruffy. A chart that leaves your machine needs work that has nothing to do with the data.
1plt.rcParams.update({2 "figure.figsize": (10, 6),3 "figure.dpi": 110,4 "axes.spines.top": False,5 "axes.spines.right": False,6 "axes.grid": True,7 "grid.alpha": 0.25,8 "font.size": 11,9 "axes.titlesize": 14,10})1112fig, ax = plt.subplots()13ax.plot(ts.index, ts.rolling(7).mean(), lw=2)14ax.set_title("Revenue recovered after the March pricing change") # the finding15ax.set_xlabel("")16ax.set_ylabel("Revenue, 7-day average (GBP)") # units17ax.annotate("New pricing", xy=(pd.Timestamp("2024-03-15"), 42000),18 xytext=(pd.Timestamp("2024-02-01"), 55000),19 arrowprops={"arrowstyle": "->"})20fig.savefig("revenue.png", dpi=200, bbox_inches="tight")The title carries the finding, not the variable name. "Revenue over time" makes the reader work out the point; "Revenue recovered after the March pricing change" states it, and the chart becomes the evidence. Axis labels carry units. Removing the top and right frame lines and fading the grid puts the ink on the data rather than the furniture.
The title should carry the finding and the axis labels should carry the units. A chart whose title names the variables makes every reader rediscover the point for themselves.
Colour
| Data type | Palette kind | Examples |
|---|---|---|
| Unordered categories | Qualitative | tab10, Set2 |
| Low to high | Sequential | viridis, Blues |
| Diverging around a midpoint | Diverging | RdBu_r, coolwarm |
Avoid the rainbow (jet) map for continuous data. Its brightness does not increase monotonically, so it invents visual boundaries where the data changes smoothly and flattens real differences elsewhere. viridis was designed to be perceptually uniform and to remain readable in greyscale and to viewers with common forms of colour blindness — roughly 8% of men. Never encode a critical distinction in red versus green alone; add a shape, a line style or a direct label.
What This Means When You Build Something
Make a single function that produces the standard first look at any dataset, and run it before you form a single opinion:
1def eda_overview(df, target=None, max_cols=8):2 num = df.select_dtypes("number").columns[:max_cols]3 cat = [c for c in df.select_dtypes(["object", "string", "category"]).columns4 if df[c].nunique() <= 12][:4]56 # 1. Every numeric distribution7 df[num].hist(bins=40, figsize=(14, 3 * ((len(num) + 2) // 3)),8 edgecolor="white", layout=((len(num) + 2) // 3, 3))9 plt.suptitle("Distributions", y=1.01)1011 # 2. Missingness pattern12 plt.figure(figsize=(12, 4))13 sns.heatmap(df.isna(), cbar=False, cmap="viridis")14 plt.title("Missing values (light = missing)")1516 # 3. Correlations17 plt.figure(figsize=(9, 7))18 corr = df[num].corr()19 sns.heatmap(corr, mask=np.triu(np.ones_like(corr, dtype=bool)),20 cmap="RdBu_r", center=0, vmin=-1, vmax=1, annot=True, fmt=".2f")2122 # 4. Each category against the target23 if target:24 for c in cat:25 plt.figure(figsize=(8, 4))26 sns.violinplot(data=df, x=c, y=target, inner="quartile")27 plt.xticks(rotation=30, ha="right")28 plt.show()The missingness heatmap in step two is the one people leave out and then regret. Rendered as an image, a block of missing values that spans one contiguous stretch of rows immediately tells you a data feed was down for a period — information that a table of missing-value percentages cannot convey at all.
Plot first, summarise second. Every misleading number in this lesson — a mean in an empty valley, a boxplot concealing two populations, a correlation that a single point created — was visible in a chart and invisible in a table. The chart takes thirty seconds. Being wrong takes considerably longer to undo.