Python for AI and Data Science

Exploratory Data Analysis (EDA)


A team builds a customer churn model. They load the CSV, split it, fit a random forest, and get 99.7% accuracy on the held-out data. Everyone is delighted. It goes live, and it predicts almost nothing correctly.

The dataset contained a column called cancellation_reason. It was blank for every customer who stayed and filled in for every customer who left. The model had not learned to predict churn; it had learned to read the answer. Thirty seconds of looking at the data — one groupby on the target — would have exposed it.

Python
df.groupby("churned")["cancellation_reason"].apply(lambda s: s.notna().mean())# churned# False    0.00# True     1.00      <- this column is the answer, not a feature

Exploratory data analysis is the discipline of looking before you model. It is not a preliminary formality; it is where you find the leaked column, the duplicated ingest, the category that means something different after March, and the relationship that reverses inside every subgroup. All of these are invisible to a training script, and all of them are fatal.

The four questions EDA is actually askingBefore any modelWhat shape andtypes is this?How is eachcolumn distributed?What moves with the target?Which column isleaking the answer?
The 99.7% churn model failed on the fourth question, which is the only one a fitted model cannot answer for you.

The four questions EDA answers

StageQuestionTypical finding
StructureWhat have I actually got?Wrong dtypes, duplicates, 40% missing in one column
UnivariateWhat does each variable look like on its own?Heavy skew, impossible values, a category with 5,000 levels
BivariateWhat relates to what?A leaked feature; a strong predictor; two identical columns
MultivariateWhat happens when variables interact?Multicollinearity; a relationship that reverses by group

Work in that order. Each stage narrows what you need to look at in the next, and skipping to the correlation heatmap — which is what everyone wants to do — means you interpret correlations on columns whose dtypes and missing-value patterns you have not checked.

Stage 1: structure

Python
import pandas as pd, numpy as npdf = pd.read_csv("customers.csv")print(df.shape)df.info(memory_usage="deep")print(df.head(10))overview = pd.DataFrame({    "dtype": df.dtypes,    "missing_pct": (df.isna().mean() * 100).round(1),    "n_unique": df.nunique(),    "example": df.iloc[0],}).sort_values("missing_pct", ascending=False)print(overview)

That single table catches most structural problems at once. Read it looking for four specific things:

  • A numeric column typed as text (str in pandas 3, object in older versions) — something non-numeric is hiding in it, usually a placeholder like "N/A" or a currency symbol.
  • n_unique equal to the row count on anything other than an ID — that column is an identifier and has no predictive content, but a tree model will happily memorise it.
  • n_unique equal to 1 — a constant column, which carries zero information and should be dropped.
  • Very high missingness — above roughly 60%, filling is mostly invention.
Python
print("duplicate rows:", df.duplicated().sum())print("duplicate IDs:", df.customer_id.duplicated().sum())print(df.churned.value_counts(normalize=True))# False    0.94# True     0.06     <- 6% positive class: accuracy is now a useless metric

Checking the target's balance early changes decisions you make much later. At 6% positive, a model that predicts "no churn" for everybody scores 94% accuracy while being completely worthless, so you know now to judge by precision and recall instead.

Stage 2: one variable at a time

Numeric columns

Python
num = df.select_dtypes(include=np.number)stats = num.describe().Tstats["skew"] = num.skew()stats["zeros_pct"] = (num == 0).mean() * 100stats["negatives"] = (num < 0).sum()print(stats.round(2))

The extra three columns earn their place. Skew above about 1 means a long right tail — income, order value, page views — which distorts any model that assumes symmetry and usually wants a log transform. A high zero percentage often signals a hidden placeholder rather than a genuine zero. And negatives in a column that cannot be negative is a data error you want to know about now, not after training.

Python
import matplotlib.pyplot as plt, seaborn as snscols = num.columns[:9]fig, axes = plt.subplots(3, 3, figsize=(13, 9))for ax, col in zip(axes.ravel(), cols):    sns.histplot(df[col].dropna(), bins=40, kde=True, ax=ax)    ax.set_title(col, fontsize=10)fig.tight_layout()

Look for three shapes: a second hump, which means two populations mixed together and often means a missing grouping column; a spike at one value, which is a default or placeholder; and a hard wall at zero or a round number, which is a truncation or a data-entry limit.

Categorical columns

Python
for col in df.select_dtypes(include=["object", "str"]).columns:    vc = df[col].value_counts(dropna=False)    print(f"\n{col}: {df[col].nunique()} levels")    print(vc.head(5))    print("levels appearing once:", (vc == 1).sum())

High cardinality is the thing to watch. A category with 3,000 levels cannot be one-hot encoded into 3,000 columns without destroying your model, and levels that appear once are noise — group them into an "Other" bucket. This is also where you spot "UK", "uk" and "U.K." living as three separate categories.

Stage 3: two variables together

Numeric against numeric, and the correlation trap

Python
print(df[["income", "spend"]].corr(method="pearson"))    # linear relationshipsprint(df[["income", "spend"]].corr(method="spearman"))   # any monotonic relationship

Pearson's coefficient measures linear association only. Consider y=x2y = x^2 over xx from −5-5 to 55: the relationship is perfect and deterministic, and Pearson reports approximately zero, because for every rise on the right there is a matching fall on the left.

Python
x = np.linspace(-5, 5, 100)y = x ** 2print(np.corrcoef(x, y)[0, 1].round(3))          # 0.0    -- "no relationship"

A large gap between Pearson and Spearman is itself informative: it means the relationship is real but curved, and a linear model will underfit it. This is why you plot rather than only tabulate.

A correlation of zero means "no straight-line relationship", not "no relationship". Scatter the pair before you conclude a feature is useless.

Categorical against numeric

Python
print(df.groupby("plan")["monthly_spend"].agg(["count", "mean", "median", "std"]))fig, ax = plt.subplots(figsize=(9, 5))sns.boxplot(data=df, x="plan", y="monthly_spend", ax=ax, showfliers=False)sns.stripplot(data=df, x="plan", y="monthly_spend", ax=ax,              color="black", alpha=0.25, size=2)

Always print the count alongside the mean. A group of 7 rows with a striking average is not a finding — it is noise, and it will not reproduce.

Categorical against categorical

Python
ct = pd.crosstab(df.plan, df.churned, normalize="index")print((ct * 100).round(1))#          False  True# basic     88.4  11.6# standard  95.1   4.9# premium   97.8   2.2

normalize="index" converts counts to row percentages, which is what makes the comparison meaningful — raw counts mostly tell you how big each group is, not how the groups differ.

Screening for leakage

The failure that opened this lesson has a systematic check. Any feature suspiciously predictive of the target on its own deserves suspicion, not celebration:

Python
for col in num.columns:    if col == "churned":        continue    diff = df.groupby("churned")[col].mean()    pooled = df[col].std()    if pooled and abs(diff.diff().iloc[-1]) / pooled > 2:        print(f"{col}: suspiciously separated by target — check it")for col in df.select_dtypes(include=["object", "str"]):    rate = df.groupby("churned")[col].apply(lambda s: s.notna().mean())    if abs(rate.diff().iloc[-1]) > 0.5:        print(f"{col}: missingness itself predicts the target — likely leakage")

Then ask the question a statistic cannot answer: would this value have existed at the moment I need to make the prediction? A cancellation reason, a refund date, a final invoice total — all of these exist only after the outcome they are supposed to predict. Timing, not correlation, is what defines leakage.

Stage 4: many variables at once

The correlation matrix, read properly

Python
corr = num.corr()mask = np.triu(np.ones_like(corr, dtype=bool))     # hide the mirrored halffig, ax = plt.subplots(figsize=(10, 8))sns.heatmap(corr, mask=mask, cmap="coolwarm", center=0, vmin=-1, vmax=1,            annot=True, fmt=".2f", square=True, ax=ax)# find the redundant pairs explicitlypairs = corr.where(~mask).stack()print(pairs[pairs.abs() > 0.85].sort_values(ascending=False))

Two features correlated above about 0.9 are effectively the same feature measured twice. Keeping both does not add information; it splits the credit between them, which makes coefficients unstable and importance scores meaningless. Keep the one that is cheaper to collect or easier to explain.

Pair plots and dimensionality reduction

Python
sns.pairplot(df[["income", "spend", "tenure", "churned"]].sample(1500),             hue="churned", diag_kind="kde", plot_kws={"alpha": 0.4, "s": 12})

Sample first. A pair plot on 200,000 rows takes minutes and renders as solid blocks of colour; 1,500 rows shows the same structure in seconds.

When you have thirty numeric features, no grid of scatter plots is readable. Principal component analysis compresses them into two axes that capture as much of the variation as possible, purely so you can look at the cloud:

Python
from sklearn.decomposition import PCAfrom sklearn.preprocessing import StandardScalerX = StandardScaler().fit_transform(num.dropna())     # PCA requires scaling firstpcs = PCA(n_components=2).fit(X)proj = pcs.transform(X)print(pcs.explained_variance_ratio_.round(3))   # e.g. [0.41 0.19] -> 60% of variancesns.scatterplot(x=proj[:, 0], y=proj[:, 1], hue=df.dropna(subset=num.columns).churned,                alpha=0.5, s=12)

Scaling before PCA is not optional. PCA finds directions of maximum variance, so an unscaled column measured in thousands will dominate one measured in units purely because of its scale, and the components describe your units rather than your data.

The reversal that catches everyone

Here is admissions data for two departments, and it is the most important pattern in this lesson:

DepartmentMen appliedMen admittedMen rateWomen appliedWomen admittedWomen rate
A (easy)40024060%1006565%
B (hard)1002020%40010025%
Combined50026052%50016533%

Women were admitted at a higher rate in both departments, and at a much lower rate overall. Nothing is miscalculated. The aggregate reverses because the two groups applied to different departments: most men applied to the easy one, most women to the hard one, and the combined figure mostly measures where people applied rather than how they were treated.

Any aggregate can reverse when you split by a variable you did not include. Before trusting an overall rate, check it within the groups that make up the population.

Python
# the systematic check: does the headline hold inside every segment?overall = df.churned.mean()by_segment = df.groupby("plan").churned.mean()print(f"overall {overall:.1%}")print(by_segment.apply(lambda v: f"{v:.1%}"))

What you should have at the end

EDA produces an artefact, not just a feeling. Write down what you found, as decisions:

Python
findings = {    "rows": len(df),    "duplicates_removed": int(df.duplicated().sum()),    "target_positive_rate": round(df.churned.mean(), 4),    "dropped_leakage": ["cancellation_reason", "final_invoice_date"],    "dropped_constant": ["region_code_v1"],    "dropped_redundant": ["total_spend (0.97 with monthly_spend x tenure)"],    "log_transform": ["income", "order_value"],    "high_cardinality_to_bucket": ["postcode (3,214 levels)"],    "impute": {"income": "median within job_title", "tenure": "median"},    "metric_choice": "precision/recall — 6% positive class makes accuracy meaningless",    "open_questions": ["churn rate doubles after March — process change or data change?"],}

That dictionary is worth more than any chart you produced, because it is the set of decisions the modelling work will inherit, with reasons attached. Six weeks later, when someone asks why total_spend is not in the model, the answer is written down.

The habit underneath all of this is simple and hard to keep: be suspicious of good news. A feature that predicts beautifully, a correlation of 0.98, an accuracy of 99.7% — in real data these are far more often mistakes than discoveries. EDA is the half hour you spend trying to prove your own result wrong, before you spend three weeks building on top of it.