Mathew K Analytics

Lesson 9 · Statistics for data analysts

Non-Parametric Tests: When Assumptions Fail | Statistics #9

Video nine of the 15-part series: rank-based tests for when the normal-distribution assumptions behind a t-test or ANOVA don't hold. Real Sample Superstore…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Statistics for Data Analysts, Video 9: Non-Parametric Tests#

  • Video nine of the 15-part series: rank-based tests for when the normal-distribution assumptions behind a t-test or ANOVA don't hold.
  • Real Sample Superstore order data throughout, including a real case where a parametric test and a non-parametric test genuinely disagree.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • You'll need pandas, NumPy, and SciPy.
  • Place superstore_sales.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from scipy import stats

df = pd.read_csv('superstore_sales.csv')
df['Order Date'] = pd.to_datetime(df['Order Date'])
print(f'Real Sales skewness: {stats.skew(df["Sales"]):.2f}')
Real Sales skewness: 12.97

Part 1: Why Non-Parametric Tests#

Converting real values to ranks throws away their exact magnitude but keeps their order, which makes rank-based tests far less sensitive to a handful of extreme real outliers, exactly the kind Sales is full of. The tradeoff: slightly less statistical power than a t-test when the data genuinely is close to normal, in exchange for real robustness when it isn't.

Part 2: The Mann-Whitney U Test#

furniture = df[df['Category'] == 'Furniture']['Sales']
office = df[df['Category'] == 'Office Supplies']['Sales']
print(f'Real Furniture: n={len(furniture)}, median={furniture.median():.2f}, mean={furniture.mean():.2f}')
print(f'Real Office Supplies: n={len(office)}, median={office.median():.2f}, mean={office.mean():.2f}')
Real Furniture: n=2121, median=182.22, mean=349.83
Real Office Supplies: n=6026, median=27.42, mean=119.32
u_stat, p_mwu = stats.mannwhitneyu(furniture, office, alternative='two-sided')
t_stat, p_ttest = stats.ttest_ind(furniture, office, equal_var=False)
print(f'Real Mann-Whitney U: statistic={u_stat:.0f}, p-value={p_mwu:.2e}')
print(f'Real Welch t-test: statistic={t_stat:.2f}, p-value={p_ttest:.2e}')
Real Mann-Whitney U: statistic=9727490, p-value=5.39e-281
Real Welch t-test: statistic=19.24, p-value=7.01e-78

Part 3: The Wilcoxon Signed-Rank Test#

df['Year'] = df['Order Date'].dt.year
yearly = df[df['Year'].isin([2016, 2017])].groupby(['Sub-Category', 'Year'])['Sales'].mean().unstack().dropna()
print(f'Real sub-categories with data in both years: {len(yearly)}')
print(yearly.head())
Real sub-categories with data in both years: 17
Year                2016        2017
Sub-Category                        
Accessories   225.246527  217.986298
Appliances    228.511535  260.163224
Art            32.573268   31.429319
Binders       119.718855  145.576090
Bookcases     486.582713  395.056312
w_stat, p_wilcoxon = stats.wilcoxon(yearly[2016], yearly[2017])
t_paired, p_paired = stats.ttest_rel(yearly[2016], yearly[2017])
print(f'Real Wilcoxon signed-rank: statistic={w_stat:.1f}, p-value={p_wilcoxon:.4f}')
print(f'Real paired t-test: statistic={t_paired:.3f}, p-value={p_paired:.4f}')
Real Wilcoxon signed-rank: statistic=46.0, p-value=0.1594
Real paired t-test: statistic=1.791, p-value=0.0922

Part 4: Kruskal-Wallis, and a Real Disagreement#

regions = df['Region'].unique()
region_groups = [df[df['Region'] == r]['Sales'] for r in regions]
h_stat, p_kw = stats.kruskal(*region_groups)
f_stat, p_anova = stats.f_oneway(*region_groups)
print(f'Real Kruskal-Wallis: statistic={h_stat:.2f}, p-value={p_kw:.2e}')
print(f'Real one-way ANOVA: statistic={f_stat:.3f}, p-value={p_anova:.3f}')
Real Kruskal-Wallis: statistic=26.10, p-value=9.07e-06
Real one-way ANOVA: statistic=0.801, p-value=0.493
print('Real median Sales by region:')
print(df.groupby('Region')['Sales'].median().sort_values())
print()
print('Real mean Sales by region:')
print(df.groupby('Region')['Sales'].mean().sort_values())
Real median Sales by region:
Region
Central    45.980
South      54.594
East       54.900
West       60.840
Name: Sales, dtype: float64

Real mean Sales by region:
Region
Central    215.772661
West       226.493233
East       238.336110
South      241.803645
Name: Sales, dtype: float64

This is the honest reason the two tests disagree: ANOVA compares real group means, and real means here are dominated by a small number of very large orders that happen to land unevenly across regions, essentially noise relative to the question 'do regions behave differently.' Kruskal-Wallis compares real rank distributions, which is far less swayed by a few extreme values and picks up the genuine, consistent regional pattern visible in the medians. When a parametric and a non-parametric test disagree on real skewed data like this, the non-parametric result is usually the more trustworthy one.

Part 5: Choosing the Right Test#

  • Two independent groups, roughly normal: two-sample t-test.
  • Two independent groups, skewed or with real outliers: Mann-Whitney U.
  • Two paired/matched measurements, roughly normal differences: paired t-test.
  • Two paired/matched measurements, skewed differences: Wilcoxon signed-rank.
  • Three or more independent groups, roughly normal: one-way ANOVA.
  • Three or more independent groups, skewed or with real outliers: Kruskal-Wallis.
  • When in doubt on real business data, especially anything involving dollar amounts, revenue, or order sizes, which are almost always right-skewed, lean non-parametric, or run both and treat a disagreement as a signal to look closer, exactly as part four just demonstrated.

Wrap-Up: What You Learned#

  • Why rank-based tests exist: robustness to skew and real outliers, at a small cost in power when data is genuinely normal.
  • Mann-Whitney U, for comparing two independent real groups, agreeing with a t-test on a clear real Furniture-versus-Office-Supplies difference.
  • Wilcoxon signed-rank, for paired real comparisons, applied to sub-category Sales across two real years.
  • Kruskal-Wallis, for three or more groups, including a genuine, honest disagreement with one-way ANOVA on real regional Sales, and why the rank-based result deserved more trust here.
  • A practical decision guide for choosing between parametric and non-parametric tests on real data.
  • Video ten builds directly on the idea of running several tests at once, covering the multiple comparisons problem and how to correct for it. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.