Course Content
Data Science Fundamentals
4 sections · 10 lessons
Descriptive Statistics
A 12-person startup advertises an "average salary of £91,000". Here are the actual salaries, in thousands:
28 31 32 34 35 36 38 40 42 45 48 683The mean really is £91,000. Ten of the twelve people earn less than half of it. The twelfth is the founder, and one number has dragged the summary of an entire company away from anything that describes a single employee.
The median is £37,000. Nobody was lied to about the arithmetic, and yet the advertised figure conveys almost the opposite of the truth. That gap — between a number that is correct and a number that is informative — is what descriptive statistics is about. The job is not to compute summaries. It is to choose summaries that survive contact with the actual shape of your data, and to know which ones fall apart.
Where the Middle Is
Three ways to say "typical", and they disagree exactly when it matters most.
| Measure | Definition | Breaks down when | Best for |
|---|---|---|---|
| Mean | Sum divided by count | Skew or outliers | Symmetric data; anything you need to add up |
| Median | The middle value when sorted | Rarely — very robust | Income, prices, durations, response times |
| Mode | The most frequent value | Continuous data; ties | Categories, discrete counts |
The formula shows why the mean is fragile: every observation enters with equal weight, so one value of 683 contributes as much as 683 values of one. The median only cares about ordering, so replacing 683 with 6,830 does not move it at all.
1import pandas as pd2import numpy as np34salaries = pd.Series([28, 31, 32, 34, 35, 36, 38, 40, 42, 45, 48, 683])5print(salaries.mean(), salaries.median()) # 91.0 37.067# A quick skew detector: how far apart are they, relative to spread?8gap = (salaries.mean() - salaries.median()) / salaries.std()9print(round(gap, 2)) # 0.29When the mean and the median disagree substantially, report both — and lead with the median. The disagreement is itself a finding about the distribution.
Weighted and grouped means
A plain mean of means is wrong whenever the groups differ in size. Suppose three shops have average baskets of £12 (2,000 customers), £45 (100 customers) and £30 (500 customers). The mean of 12, 45 and 30 is £29. The true average basket is:
Python1shops = pd.DataFrame({2 "avg_basket": [12, 45, 30],3 "customers": [2000, 100, 500],4})5print(np.average(shops["avg_basket"], weights=shops["customers"]).round(2)) # 16.73£29 versus £16.73 is not a rounding difference; it is a wrong answer produced by a correct-looking calculation. This is one of the most common mistakes in dashboard work, because averaging a column of averages is exactly what a naive query does.
How Spread Out It Is
Two datasets can share a mean and describe completely different worlds. Delivery times averaging 30 minutes could mean every order arrives between 28 and 32 minutes, or that half arrive in 10 and half in 50. Only the spread distinguishes them, and the spread is usually what the customer experiences.
Python1x = df["delivery_minutes"]23print("range:", x.max() - x.min())4print("IQR:", x.quantile(0.75) - x.quantile(0.25))5print("std:", x.std())6print("CV:", x.std() / x.mean()) # unitless, comparable across columns7print(x.quantile([0.05, 0.25, 0.5, 0.75, 0.95]))
Measure What it says Robust to outliers? Range Distance between the two most extreme points No — defined entirely by them IQR Width of the middle 50% Yes Variance / standard deviation Typical squared / absolute distance from the mean No Coefficient of variation Spread relative to size No, but comparable across units Standard deviation is the square root of variance:
s2=n−11i=1∑n(xi−xˉ)2,s=s2The n − 1 instead of n is Bessel's correction. Dividing by n would systematically underestimate the true population variance, because the sample mean is itself pulled towards the sample points. pandas uses n − 1 by default; NumPy uses n. That difference is the reason
df["x"].std()andnp.std(df["x"])give different answers, which surprises people every time. Passddof=1to NumPy to match.The coefficient of variation earns its place when comparing columns with different units. A standard deviation of 3 means something quite different for ages than for house prices; a CV of 0.05 versus 0.6 is directly comparable.
The Shape of the Distribution
Skewness — which way the tail runs
Pythonprint(df["income"].skew()) # positive: long right tailprint(df["exam_score"].skew()) # negative: long left tail
Skewness Interpretation Consequence ≈ 0 (|skew| < 0.5) Roughly symmetric Mean is a fair summary 0.5 to 1 Moderately skewed Report the median too > 1 Strongly skewed Mean is misleading; consider a log transform Positive skew means the tail stretches right and the mean sits above the median — income, wealth, city sizes, waiting times. Negative skew means the tail stretches left, typical of scores capped at a maximum where most people cluster near the top.
Kurtosis — how heavy the tails are
Kurtosis measures how much of the variance comes from rare extreme values rather than from moderate ones. Both pandas and SciPy report excess kurtosis, where a normal distribution scores 0.
Pythonprint(df["daily_return"].kurtosis()) # financial returns: often 3 to 10High kurtosis is a warning that "three standard deviations" is not as rare as normal-distribution intuition suggests. Financial returns are the standard example: models built assuming normal tails badly underestimate the frequency of large losses, because the real distribution has far more mass out there than the bell curve allows.
Is It Normal? Test Before You Assume
Many procedures assume approximate normality. Checking is cheap.
Python1from scipy import stats23x = df["value"].dropna()45if len(x) <= 5000:6 stat, p = stats.shapiro(x)7 print(f"Shapiro-Wilk: W={stat:.4f}, p={p:.4g}")8else:9 stat, p = stats.normaltest(x) # D'Agostino-Pearson10 print(f"D'Agostino: stat={stat:.2f}, p={p:.4g}")1112print("normal-ish" if p > 0.05 else "not normal")Shapiro–Wilk is powerful on small samples but becomes almost useless on very large ones — with 100,000 rows it will report a significant departure from normality for data that is normal enough for every practical purpose. A quantile-quantile plot is often the better check on big data: plot your sorted values against the values a normal distribution would produce, and a straight line means normal. Curvature at the ends is exactly the heavy tails kurtosis was measuring.
What a p-value actually means
This is worth stating precisely, because the wrong version is repeated constantly.
A p-value is the probability of seeing a result at least as extreme as yours if the null hypothesis were true. It is not the probability that the null hypothesis is true. It is not the probability your finding is a fluke. It is not a measure of effect size.
Claim Correct? Why "p = 0.03, so there's a 3% chance the null is true" No Reverses the conditional. p is P(data | null), not P(null | data) "p = 0.03, so the effect is large" No With a big enough sample, trivial effects reach significance "p = 0.21, so there is no effect" No Absence of evidence is not evidence of absence "If the null were true, data this extreme would appear 3% of the time" Yes That is the definition Always report an effect size alongside a p-value. Significance tells you whether an effect is detectable; effect size tells you whether anyone should care.
Comparing Groups
Two groups: the t-test
Python1a = df.loc[df["variant"] == "A", "conversion_time"]2b = df.loc[df["variant"] == "B", "conversion_time"]34t, p = stats.ttest_ind(a, b, equal_var=False) # Welch's t-test5print(f"t={t:.3f}, p={p:.4g}")67# Cohen's d - the effect size8pooled = np.sqrt(((len(a)-1)*a.var(ddof=1) + (len(b)-1)*b.var(ddof=1))9 / (len(a)+len(b)-2))10d = (a.mean() - b.mean()) / pooled11print(f"Cohen's d = {d:.3f}")Use
equal_var=Falseas the default. Welch's version does not assume the two groups have equal variance, and it costs almost nothing in power when they do. Interpret Cohen's d as roughly: 0.2 small, 0.5 medium, 0.8 large — a difference of 0.8 standard deviations is one you could see in a chart without a test.Use
ttest_relinstead when the same subjects are measured twice (before and after). Treating paired data as independent throws away the pairing and makes the test far less sensitive.Three or more groups: ANOVA
Pythongroups = [g["score"].values for _, g in df.groupby("branch")]f, p = stats.f_oneway(*groups)print(f"F={f:.3f}, p={p:.4g}")Why not just run a t-test on every pair? Because with five groups there are ten pairs, and at a 5% threshold the probability of at least one false positive is 1 − 0.95¹⁰ ≈ 40%. ANOVA asks one question — "are these groups all the same?" — at one significance level. If it rejects, follow up with post-hoc pairwise tests that correct for the multiple comparisons, such as Tukey's HSD.
Pythonfrom statsmodels.stats.multicomp import pairwise_tukeyhsdprint(pairwise_tukeyhsd(df["score"], df["branch"], alpha=0.05))Two categorical variables: chi-square
Python1table = pd.crosstab(df["region"], df["churned"])2chi2, p, dof, expected = stats.chi2_contingency(table)3print(f"chi2={chi2:.2f}, p={p:.4g}")45# Cramer's V - effect size for categorical association6n = table.values.sum()7v = np.sqrt(chi2 / (n * (min(table.shape) - 1)))8print(f"Cramer's V = {v:.3f}")Chi-square becomes unreliable when expected cell counts drop below about five. Check
expected.min(); if it is too small, merge sparse categories or use Fisher's exact test.Choosing the test
Question Data Test Non-parametric alternative Do two independent groups differ? Numeric Welch's t-test Mann–Whitney U Did the same subjects change? Numeric, paired Paired t-test Wilcoxon signed-rank Do 3+ groups differ? Numeric One-way ANOVA Kruskal–Wallis Are two categories associated? Categorical Chi-square Fisher's exact (small counts) Do two numeric variables move together? Numeric Pearson correlation Spearman correlation Reach for the non-parametric column when the data is strongly skewed, the sample is small, or outliers are present and real. These tests work on ranks rather than values, so they give up a little power in exchange for not assuming a distribution shape.
One Function That Tells You What You Have
Python1def describe_numeric(df, cols=None):2 cols = cols or df.select_dtypes(include=np.number).columns3 rows = []4 for c in cols:5 s = df[c].dropna()6 if len(s) < 3:7 continue8 q1, q3 = s.quantile([0.25, 0.75])9 rows.append({10 "column": c,11 "n": len(s),12 "missing_pct": round(df[c].isna().mean() * 100, 1),13 "mean": round(s.mean(), 2),14 "median": round(s.median(), 2),15 "std": round(s.std(), 2),16 "cv": round(s.std() / s.mean(), 2) if s.mean() else np.nan,17 "iqr": round(q3 - q1, 2),18 "skew": round(s.skew(), 2),19 "kurtosis": round(s.kurtosis(), 2),20 "p05": round(s.quantile(0.05), 2),21 "p95": round(s.quantile(0.95), 2),22 "outliers_iqr": int(((s < q1 - 1.5*(q3-q1)) | (s > q3 + 1.5*(q3-q1))).sum()),23 })24 return pd.DataFrame(rows).set_index("column")2526summary = describe_numeric(df)27print(summary)2829# Columns that need attention30print(summary.query("abs(skew) > 1 or missing_pct > 10 or outliers_iqr > 0.05 * n"))That final query is the part worth keeping. A summary table is only useful if it points at something; filtering it down to the columns that violate your assumptions turns a wall of numbers into a short list of decisions.
What This Means When You Build Something
Before any modelling or reporting, run three checks on every numeric column and let the answers dictate what you do next.
Check If it fails Do this Mean ≈ median? They diverge Report the median; consider a log transform for modelling |Skew| < 1? Strongly skewed Transform, or use rank-based methods Excess kurtosis near 0? Heavy tails Do not trust normal-based intervals; use bootstrapping And when you report a comparison, give three numbers, never one: the effect size, a confidence interval, and the sample size. "Variant B converted 2.1 percentage points higher (95% CI: 0.4 to 3.8, n = 12,400)" tells a reader how large the effect is, how certain you are, and how much data stands behind it. "p < 0.05" tells them almost nothing — and on a large enough sample, it tells them nothing at all, because with enough rows every difference becomes statistically significant regardless of whether it means anything.