Data Science Fundamentals

Visual Analytics


A support team reports that the average ticket resolution time is 30 minutes. Management is satisfied — 30 minutes is the target. Then somebody plots the actual distribution.

Text
count 60 |     ###                              ### 40 |    #####                            ##### 20 |   #######                          #######  0 +----------------------------------------------      5   10   15   20   25   30   35   40   45   50   minutes

There are two humps and almost nothing in the middle. Roughly half the tickets resolve in about 10 minutes and half take about 45. Nobody experiences a 30-minute resolution. The mean sits in the empty valley between two entirely different processes — a self-service path and a path that requires a human specialist — and it describes neither.

No summary statistic could have revealed that. The mean, median, standard deviation and skewness of this distribution are all unremarkable. It took thirty seconds of plotting to see something that months of reporting had hidden. That is the case for visual analytics: not to decorate findings after the fact, but to see structure that numbers compress away.

Match the picture to the questionWhat areyou asking?One numericcolumn: histogramTwo numeric:scatter with a trendCategory againstnumeric: box or violinTwocategoricals: grouped barsSomething overtime: line, never bars
The support team's 30-minute average sat between a 10-minute peak and a 45-minute peak, and no bar chart would have shown it.

Choosing the Right Picture

Almost every exploratory chart answers one of five questions. Start from the question and the chart type follows.

QuestionVariablesChart
What does this variable look like?1 numericHistogram, density plot, boxplot
How do these two move together?2 numericScatter, hexbin, regression plot
How do groups compare?1 categorical + 1 numericBar with error bars, boxplot, violin
How do categories relate?2 categoricalGrouped/stacked bar, heatmap of counts
How does this change over time?Time + numericLine plot

Seeing a Single Variable

Histograms and the bin-count problem

Python
import matplotlib.pyplot as pltimport seaborn as snsfig, axes = plt.subplots(1, 3, figsize=(15, 4))for ax, bins in zip(axes, [8, 30, 120]):    ax.hist(df["resolution_minutes"], bins=bins, edgecolor="white")    ax.set_title(f"{bins} bins")plt.tight_layout()

Run that on the support data and the point becomes obvious: with 8 bins the two humps merge into one broad lump and the bimodality vanishes; with 120 bins the shape dissolves into noise; around 30 the structure is unmistakable. The bin count is not a cosmetic setting. It decides what you are able to see.

Since no single choice is right, always look at two or three. The Freedman–Diaconis rule gives a sensible starting width, scaling with the interquartile range and shrinking as the sample grows:

h=2×IQRn1/3h = 2 \times \frac{IQR}{n^{1/3}}

Python
import numpy as npx = df["resolution_minutes"].dropna()iqr = x.quantile(0.75) - x.quantile(0.25)width = 2 * iqr / len(x) ** (1/3)print("suggested bins:", int((x.max() - x.min()) / width))

Density plots

Python
sns.kdeplot(data=df, x="resolution_minutes", hue="channel",            fill=True, alpha=0.35, common_norm=False)

A kernel density estimate smooths the histogram into a curve, which makes overlaying several groups readable where overlapping histograms become mud. Two cautions. It smooths, so a small bump may be an artefact of the bandwidth rather than a feature of the data — check it against a histogram before believing it. And it will happily draw density below zero for a strictly positive quantity like duration, because the smoother does not know about the boundary.

common_norm=False matters when groups differ in size: it normalises each group separately, so you compare shapes rather than comparing "how many rows are in this group", which the bar heights already told you.

Boxplots and violins

A boxplot draws five numbers: the median as the centre line, the first and third quartiles as the box edges, whiskers extending to the furthest points within 1.5 × IQR of the box, and everything beyond as individual dots.

Python
fig, axes = plt.subplots(1, 2, figsize=(13, 5), sharey=True)sns.boxplot(data=df, x="team", y="resolution_minutes", ax=axes[0])sns.violinplot(data=df, x="team", y="resolution_minutes", ax=axes[1], inner="quartile")axes[0].set_title("Boxplot: five numbers")axes[1].set_title("Violin: full shape")

Here is the trap that the opening example illustrates: a boxplot of bimodal data looks completely ordinary. The median sits in the empty valley, the box spans both humps, and nothing signals that there are two populations. Violin plots draw the density along each side, so both humps appear. When comparing groups where distribution shape might differ, use violins.

Boxplots compare centres and spreads. They cannot show shape, so they hide exactly the structure that makes a summary statistic misleading.

ChartShowsBest whenHides
HistogramFull shape, one groupExploring a single variableSensitive to bin choice
Density (KDE)Smoothed shape, several groupsComparing 2–4 distributionsSmall features; boundaries
BoxplotMedian, IQR, outliersComparing many groups compactlyMultimodality, sample size
ViolinShape per groupGroups may differ in shapeNeeds enough data per group
Strip / swarmEvery individual pointFewer than ~200 pointsBecomes unreadable when dense

Seeing Two Variables Together

Python
sns.scatterplot(data=df, x="tenure_months", y="lifetime_value",                hue="segment", alpha=0.5, s=25)

Once you pass a few thousand points, a scatter plot becomes a solid block and the eye can no longer judge where the density is. Three fixes, in order of how much data they handle:

Python
fig, axes = plt.subplots(1, 3, figsize=(16, 4.5))# 1. Transparency - works to roughly 20k pointsaxes[0].scatter(df.x, df.y, alpha=0.05, s=8)# 2. Hexagonal binning - counts per hexagonhb = axes[1].hexbin(df.x, df.y, gridsize=45, cmap="Blues", mincnt=1)plt.colorbar(hb, ax=axes[1], label="count")# 3. 2D density contourssns.kdeplot(data=df, x="x", y="y", fill=True, levels=8, ax=axes[2])

Hexbin is the workhorse for large scatter data. It shows where the mass is instead of where the ink is, and unlike transparency it does not saturate.

Adding a trend without lying about it

Python
fig, axes = plt.subplots(1, 2, figsize=(13, 5))sns.regplot(data=df, x="tenure_months", y="lifetime_value", ax=axes[0],            scatter_kws={"alpha": 0.3}, line_kws={"color": "crimson"})sns.regplot(data=df, x="tenure_months", y="lifetime_value", ax=axes[1],            lowess=True, scatter_kws={"alpha": 0.3}, line_kws={"color": "crimson"})axes[0].set_title("Straight-line fit")axes[1].set_title("LOWESS: follows the data")

Draw both. If the straight line and the LOWESS curve diverge, the relationship is not linear and any correlation coefficient you computed is understating it. This is the visual version of the check that saves you from dropping a genuinely predictive feature because a linear measure scored it near zero.

Correlation heatmaps

Python
corr = df.select_dtypes("number").corr()mask = np.triu(np.ones_like(corr, dtype=bool))     # hide the mirror halfsns.heatmap(corr, mask=mask, cmap="RdBu_r", center=0, vmin=-1, vmax=1,            annot=True, fmt=".2f", square=True, linewidths=0.5,            cbar_kws={"shrink": 0.7})

Three settings do the real work. center=0 with a diverging colour map puts white at zero, so positive and negative are visually distinct rather than being two shades of the same colour. vmin and vmax fix the scale to the full range, so heatmaps from different datasets are comparable — without them, a matrix whose strongest correlation is 0.3 renders that 0.3 in the same deep colour another matrix uses for 0.95. Masking the upper triangle removes redundant information, since the matrix is symmetric.

Categorical Comparisons

Python
order = df.groupby("region")["revenue"].mean().sort_values().indexsns.barplot(data=df, y="region", x="revenue", order=order,            errorbar=("ci", 95), color="#4C72B0")

Three deliberate choices there. Sorting by value rather than alphabetically means the ranking is readable at a glance — alphabetical order encodes nothing. Horizontal orientation keeps category labels readable without rotation. And the confidence interval prevents the most common misreading of a bar chart: two bars of different heights whose intervals overlap heavily are not evidence of a real difference, and a plain bar chart gives no way to tell.

Bars must start at zero. Truncating the axis at 90 to make a difference between 92 and 95 look dramatic is the classic misleading chart, because a bar's meaning is carried by its length. Line and scatter plots have no such constraint — they encode position, not length — so zooming their axes is legitimate.

Pie charts

Use them for at most three or four slices where the point is "roughly half" or "about a quarter". Human beings judge angles poorly, so a pie with nine slices takes longer to read than a sorted bar chart and conveys less. If the slices need labelling with their percentages to be legible, the chart is doing no work and a table would serve better.

Two categoricals at once

Python
ct = pd.crosstab(df["region"], df["product"], normalize="index") * 100sns.heatmap(ct, annot=True, fmt=".1f", cmap="YlGnBu", cbar_kws={"label": "% of region"})

normalize="index" converts counts into row percentages, which is usually the question you mean. Raw counts mostly show which regions are large; percentages show which products are disproportionately popular where.

More Than Two Variables

Python
# Every pairwise relationship plus the diagonal distributionssns.pairplot(df[["age", "income", "spend", "tenure", "segment"]],             hue="segment", diag_kind="kde", plot_kws={"alpha": 0.4, "s": 18})# The same relationship, repeated per subgroupg = sns.FacetGrid(df, col="region", row="segment", height=3, margin_titles=True)g.map_dataframe(sns.scatterplot, x="tenure", y="spend", alpha=0.5)

Pairplots are the fastest way to survey a small numeric dataset, but the number of panels grows as the square of the columns — ten columns produce a hundred panels. Restrict to eight or fewer, and sample rows if there are many.

Facet grids ("small multiples") are the most underrated technique here. Repeating one simple chart across subgroups, on shared axes, lets the eye compare directly. It reveals Simpson's paradox situations, where an overall trend reverses within every subgroup — something no single aggregate chart can show.

Time

Python
ts = df.set_index("date")["revenue"].resample("D").sum()fig, ax = plt.subplots(figsize=(13, 5))ax.plot(ts.index, ts, alpha=0.3, lw=0.8, label="daily")ax.plot(ts.index, ts.rolling(7, center=True).mean(), lw=2, label="7-day average")ax.fill_between(ts.index,                ts.rolling(7).mean() - ts.rolling(7).std(),                ts.rolling(7).mean() + ts.rolling(7).std(),                alpha=0.15)ax.legend()ax.set_ylabel("Revenue (GBP)")

Plotting the raw series and a rolling average together is the standard pattern: the faint line preserves the real variation, the bold line shows the trend. A 7-day window on daily data removes the day-of-week cycle specifically, which is why weekly business data almost always wants exactly 7 rather than a rounder number.

For strong seasonality, decomposition separates the components explicitly:

Python
from statsmodels.tsa.seasonal import seasonal_decomposeseasonal_decompose(ts.dropna(), model="additive", period=7).plot()

Making Figures Someone Else Can Read

An exploratory chart has an audience of one and can be scruffy. A chart that leaves your machine needs work that has nothing to do with the data.

Python
plt.rcParams.update({    "figure.figsize": (10, 6),    "figure.dpi": 110,    "axes.spines.top": False,    "axes.spines.right": False,    "axes.grid": True,    "grid.alpha": 0.25,    "font.size": 11,    "axes.titlesize": 14,})fig, ax = plt.subplots()ax.plot(ts.index, ts.rolling(7).mean(), lw=2)ax.set_title("Revenue recovered after the March pricing change")   # the findingax.set_xlabel("")ax.set_ylabel("Revenue, 7-day average (GBP)")                       # unitsax.annotate("New pricing", xy=(pd.Timestamp("2024-03-15"), 42000),            xytext=(pd.Timestamp("2024-02-01"), 55000),            arrowprops={"arrowstyle": "->"})fig.savefig("revenue.png", dpi=200, bbox_inches="tight")

The title carries the finding, not the variable name. "Revenue over time" makes the reader work out the point; "Revenue recovered after the March pricing change" states it, and the chart becomes the evidence. Axis labels carry units. Removing the top and right frame lines and fading the grid puts the ink on the data rather than the furniture.

The title should carry the finding and the axis labels should carry the units. A chart whose title names the variables makes every reader rediscover the point for themselves.

Colour

Data typePalette kindExamples
Unordered categoriesQualitativetab10, Set2
Low to highSequentialviridis, Blues
Diverging around a midpointDivergingRdBu_r, coolwarm

Avoid the rainbow (jet) map for continuous data. Its brightness does not increase monotonically, so it invents visual boundaries where the data changes smoothly and flattens real differences elsewhere. viridis was designed to be perceptually uniform and to remain readable in greyscale and to viewers with common forms of colour blindness — roughly 8% of men. Never encode a critical distinction in red versus green alone; add a shape, a line style or a direct label.

What This Means When You Build Something

Make a single function that produces the standard first look at any dataset, and run it before you form a single opinion:

Python
def eda_overview(df, target=None, max_cols=8):    num = df.select_dtypes("number").columns[:max_cols]    cat = [c for c in df.select_dtypes(["object", "string", "category"]).columns           if df[c].nunique() <= 12][:4]    # 1. Every numeric distribution    df[num].hist(bins=40, figsize=(14, 3 * ((len(num) + 2) // 3)),                 edgecolor="white", layout=((len(num) + 2) // 3, 3))    plt.suptitle("Distributions", y=1.01)    # 2. Missingness pattern    plt.figure(figsize=(12, 4))    sns.heatmap(df.isna(), cbar=False, cmap="viridis")    plt.title("Missing values (light = missing)")    # 3. Correlations    plt.figure(figsize=(9, 7))    corr = df[num].corr()    sns.heatmap(corr, mask=np.triu(np.ones_like(corr, dtype=bool)),                cmap="RdBu_r", center=0, vmin=-1, vmax=1, annot=True, fmt=".2f")    # 4. Each category against the target    if target:        for c in cat:            plt.figure(figsize=(8, 4))            sns.violinplot(data=df, x=c, y=target, inner="quartile")            plt.xticks(rotation=30, ha="right")    plt.show()

The missingness heatmap in step two is the one people leave out and then regret. Rendered as an image, a block of missing values that spans one contiguous stretch of rows immediately tells you a data feed was down for a period — information that a table of missing-value percentages cannot convey at all.

Plot first, summarise second. Every misleading number in this lesson — a mean in an empty valley, a boxplot concealing two populations, a correlation that a single point created — was visible in a chart and invisible in a table. The chart takes thirty seconds. Being wrong takes considerably longer to undo.