Data Science Fundamentals

Descriptive Statistics


A 12-person startup advertises an "average salary of £91,000". Here are the actual salaries, in thousands:

Text
28  31  32  34  35  36  38  40  42  45  48  683

The mean really is £91,000. Ten of the twelve people earn less than half of it. The twelfth is the founder, and one number has dragged the summary of an entire company away from anything that describes a single employee.

The median is £37,000. Nobody was lied to about the arithmetic, and yet the advertised figure conveys almost the opposite of the truth. That gap — between a number that is correct and a number that is informative — is what descriptive statistics is about. The job is not to compute summaries. It is to choose summaries that survive contact with the actual shape of your data, and to know which ones fall apart.

Twelve salaries behind an average of 91283132343536384042454868301234567891011medianis 37the founderFigures in thousands of pounds; ten of the twelve earn less than half the advertised average.
The mean is a balance point and one heavy value moves it; the median counts positions, so it cannot be dragged.

Where the Middle Is

Three ways to say "typical", and they disagree exactly when it matters most.

MeasureDefinitionBreaks down whenBest for
MeanSum divided by countSkew or outliersSymmetric data; anything you need to add up
MedianThe middle value when sortedRarely — very robustIncome, prices, durations, response times
ModeThe most frequent valueContinuous data; tiesCategories, discrete counts

xˉ=1n∑i=1nxi\bar{x} = \frac{1}{n}\sum_{i=1}^{n} x_i

The formula shows why the mean is fragile: every observation enters with equal weight, so one value of 683 contributes as much as 683 values of one. The median only cares about ordering, so replacing 683 with 6,830 does not move it at all.

Python
import pandas as pdimport numpy as npsalaries = pd.Series([28, 31, 32, 34, 35, 36, 38, 40, 42, 45, 48, 683])print(salaries.mean(), salaries.median())    # 91.0  37.0# A quick skew detector: how far apart are they, relative to spread?gap = (salaries.mean() - salaries.median()) / salaries.std()print(round(gap, 2))                          # 0.29

When the mean and the median disagree substantially, report both — and lead with the median. The disagreement is itself a finding about the distribution.

Weighted and grouped means

A plain mean of means is wrong whenever the groups differ in size. Suppose three shops have average baskets of £12 (2,000 customers), £45 (100 customers) and £30 (500 customers). The mean of 12, 45 and 30 is £29. The true average basket is:

xˉw=∑wixi∑wi=2000(12)+100(45)+500(30)2600=43,5002600≈16.73\bar{x}_w = \frac{\sum w_i x_i}{\sum w_i} = \frac{2000(12) + 100(45) + 500(30)}{2600} = \frac{43{,}500}{2600} \approx 16.73
Python
shops = pd.DataFrame({    "avg_basket": [12, 45, 30],    "customers": [2000, 100, 500],})print(np.average(shops["avg_basket"], weights=shops["customers"]).round(2))  # 16.73

£29 versus £16.73 is not a rounding difference; it is a wrong answer produced by a correct-looking calculation. This is one of the most common mistakes in dashboard work, because averaging a column of averages is exactly what a naive query does.

How Spread Out It Is

Two datasets can share a mean and describe completely different worlds. Delivery times averaging 30 minutes could mean every order arrives between 28 and 32 minutes, or that half arrive in 10 and half in 50. Only the spread distinguishes them, and the spread is usually what the customer experiences.

Python
x = df["delivery_minutes"]print("range:", x.max() - x.min())print("IQR:", x.quantile(0.75) - x.quantile(0.25))print("std:", x.std())print("CV:", x.std() / x.mean())            # unitless, comparable across columnsprint(x.quantile([0.05, 0.25, 0.5, 0.75, 0.95]))
MeasureWhat it saysRobust to outliers?
RangeDistance between the two most extreme pointsNo — defined entirely by them
IQRWidth of the middle 50%Yes
Variance / standard deviationTypical squared / absolute distance from the meanNo
Coefficient of variationSpread relative to sizeNo, but comparable across units

Standard deviation is the square root of variance:

s2=1n−1∑i=1n(xi−xˉ)2,s=s2s^2 = \frac{1}{n-1}\sum_{i=1}^{n}(x_i - \bar{x})^2, \qquad s = \sqrt{s^2}

The n − 1 instead of n is Bessel's correction. Dividing by n would systematically underestimate the true population variance, because the sample mean is itself pulled towards the sample points. pandas uses n − 1 by default; NumPy uses n. That difference is the reason df["x"].std() and np.std(df["x"]) give different answers, which surprises people every time. Pass ddof=1 to NumPy to match.

The coefficient of variation earns its place when comparing columns with different units. A standard deviation of 3 means something quite different for ages than for house prices; a CV of 0.05 versus 0.6 is directly comparable.

The Shape of the Distribution

Skewness — which way the tail runs

Python
print(df["income"].skew())      # positive: long right tailprint(df["exam_score"].skew())  # negative: long left tail
SkewnessInterpretationConsequence
≈ 0 (|skew| < 0.5)Roughly symmetricMean is a fair summary
0.5 to 1Moderately skewedReport the median too
> 1Strongly skewedMean is misleading; consider a log transform

Positive skew means the tail stretches right and the mean sits above the median — income, wealth, city sizes, waiting times. Negative skew means the tail stretches left, typical of scores capped at a maximum where most people cluster near the top.

Kurtosis — how heavy the tails are

Kurtosis measures how much of the variance comes from rare extreme values rather than from moderate ones. Both pandas and SciPy report excess kurtosis, where a normal distribution scores 0.

Python
print(df["daily_return"].kurtosis())   # financial returns: often 3 to 10

High kurtosis is a warning that "three standard deviations" is not as rare as normal-distribution intuition suggests. Financial returns are the standard example: models built assuming normal tails badly underestimate the frequency of large losses, because the real distribution has far more mass out there than the bell curve allows.

Is It Normal? Test Before You Assume

Many procedures assume approximate normality. Checking is cheap.

Python
from scipy import statsx = df["value"].dropna()if len(x) <= 5000:    stat, p = stats.shapiro(x)    print(f"Shapiro-Wilk: W={stat:.4f}, p={p:.4g}")else:    stat, p = stats.normaltest(x)          # D'Agostino-Pearson    print(f"D'Agostino: stat={stat:.2f}, p={p:.4g}")print("normal-ish" if p > 0.05 else "not normal")

Shapiro–Wilk is powerful on small samples but becomes almost useless on very large ones — with 100,000 rows it will report a significant departure from normality for data that is normal enough for every practical purpose. A quantile-quantile plot is often the better check on big data: plot your sorted values against the values a normal distribution would produce, and a straight line means normal. Curvature at the ends is exactly the heavy tails kurtosis was measuring.

What a p-value actually means

This is worth stating precisely, because the wrong version is repeated constantly.

A p-value is the probability of seeing a result at least as extreme as yours if the null hypothesis were true. It is not the probability that the null hypothesis is true. It is not the probability your finding is a fluke. It is not a measure of effect size.

ClaimCorrect?Why
"p = 0.03, so there's a 3% chance the null is true"NoReverses the conditional. p is P(data | null), not P(null | data)
"p = 0.03, so the effect is large"NoWith a big enough sample, trivial effects reach significance
"p = 0.21, so there is no effect"NoAbsence of evidence is not evidence of absence
"If the null were true, data this extreme would appear 3% of the time"YesThat is the definition

Always report an effect size alongside a p-value. Significance tells you whether an effect is detectable; effect size tells you whether anyone should care.

Comparing Groups

Two groups: the t-test

Python
a = df.loc[df["variant"] == "A", "conversion_time"]b = df.loc[df["variant"] == "B", "conversion_time"]t, p = stats.ttest_ind(a, b, equal_var=False)     # Welch's t-testprint(f"t={t:.3f}, p={p:.4g}")# Cohen's d - the effect sizepooled = np.sqrt(((len(a)-1)*a.var(ddof=1) + (len(b)-1)*b.var(ddof=1))                 / (len(a)+len(b)-2))d = (a.mean() - b.mean()) / pooledprint(f"Cohen's d = {d:.3f}")

Use equal_var=False as the default. Welch's version does not assume the two groups have equal variance, and it costs almost nothing in power when they do. Interpret Cohen's d as roughly: 0.2 small, 0.5 medium, 0.8 large — a difference of 0.8 standard deviations is one you could see in a chart without a test.

Use ttest_rel instead when the same subjects are measured twice (before and after). Treating paired data as independent throws away the pairing and makes the test far less sensitive.

Three or more groups: ANOVA

Python
groups = [g["score"].values for _, g in df.groupby("branch")]f, p = stats.f_oneway(*groups)print(f"F={f:.3f}, p={p:.4g}")

Why not just run a t-test on every pair? Because with five groups there are ten pairs, and at a 5% threshold the probability of at least one false positive is 1 − 0.95¹⁰ ≈ 40%. ANOVA asks one question — "are these groups all the same?" — at one significance level. If it rejects, follow up with post-hoc pairwise tests that correct for the multiple comparisons, such as Tukey's HSD.

Python
from statsmodels.stats.multicomp import pairwise_tukeyhsdprint(pairwise_tukeyhsd(df["score"], df["branch"], alpha=0.05))

Two categorical variables: chi-square

Python
table = pd.crosstab(df["region"], df["churned"])chi2, p, dof, expected = stats.chi2_contingency(table)print(f"chi2={chi2:.2f}, p={p:.4g}")# Cramer's V - effect size for categorical associationn = table.values.sum()v = np.sqrt(chi2 / (n * (min(table.shape) - 1)))print(f"Cramer's V = {v:.3f}")

Chi-square becomes unreliable when expected cell counts drop below about five. Check expected.min(); if it is too small, merge sparse categories or use Fisher's exact test.

Choosing the test

QuestionDataTestNon-parametric alternative
Do two independent groups differ?NumericWelch's t-testMann–Whitney U
Did the same subjects change?Numeric, pairedPaired t-testWilcoxon signed-rank
Do 3+ groups differ?NumericOne-way ANOVAKruskal–Wallis
Are two categories associated?CategoricalChi-squareFisher's exact (small counts)
Do two numeric variables move together?NumericPearson correlationSpearman correlation

Reach for the non-parametric column when the data is strongly skewed, the sample is small, or outliers are present and real. These tests work on ranks rather than values, so they give up a little power in exchange for not assuming a distribution shape.

One Function That Tells You What You Have

Python
def describe_numeric(df, cols=None):    cols = cols or df.select_dtypes(include=np.number).columns    rows = []    for c in cols:        s = df[c].dropna()        if len(s) < 3:            continue        q1, q3 = s.quantile([0.25, 0.75])        rows.append({            "column": c,            "n": len(s),            "missing_pct": round(df[c].isna().mean() * 100, 1),            "mean": round(s.mean(), 2),            "median": round(s.median(), 2),            "std": round(s.std(), 2),            "cv": round(s.std() / s.mean(), 2) if s.mean() else np.nan,            "iqr": round(q3 - q1, 2),            "skew": round(s.skew(), 2),            "kurtosis": round(s.kurtosis(), 2),            "p05": round(s.quantile(0.05), 2),            "p95": round(s.quantile(0.95), 2),            "outliers_iqr": int(((s < q1 - 1.5*(q3-q1)) | (s > q3 + 1.5*(q3-q1))).sum()),        })    return pd.DataFrame(rows).set_index("column")summary = describe_numeric(df)print(summary)# Columns that need attentionprint(summary.query("abs(skew) > 1 or missing_pct > 10 or outliers_iqr > 0.05 * n"))

That final query is the part worth keeping. A summary table is only useful if it points at something; filtering it down to the columns that violate your assumptions turns a wall of numbers into a short list of decisions.

What This Means When You Build Something

Before any modelling or reporting, run three checks on every numeric column and let the answers dictate what you do next.

CheckIf it failsDo this
Mean ≈ median?They divergeReport the median; consider a log transform for modelling
|Skew| < 1?Strongly skewedTransform, or use rank-based methods
Excess kurtosis near 0?Heavy tailsDo not trust normal-based intervals; use bootstrapping

And when you report a comparison, give three numbers, never one: the effect size, a confidence interval, and the sample size. "Variant B converted 2.1 percentage points higher (95% CI: 0.4 to 3.8, n = 12,400)" tells a reader how large the effect is, how certain you are, and how much data stands behind it. "p < 0.05" tells them almost nothing — and on a large enough sample, it tells them nothing at all, because with enough rows every difference becomes statistically significant regardless of whether it means anything.