Mathew K Analytics

Lesson 18 · Probability and Statistics in python

Understanding Permutation Tests: A Flexible Approach to Hypothesis Testing

Welcome! In this lesson, we will explore permutation tests for hypothesis testing. Permutation tests are a flexible, non-parametric approach for comparing…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Permutation Tests for Hypothesis Testing#

Welcome! In this lesson, we will explore permutation tests for hypothesis testing. Permutation tests are a flexible, non-parametric approach for comparing groups when you are unsure about the underlying distribution.

  • We will start by understanding the basics of probability and statistics.*
  • Then we will learn about descriptive statistics and hypothesis testing.*
  • We will use Python and real datasets to practice each step.*

By the end, you will be ready to apply permutation tests to your own data.

# Importing all the libraries we will need
import warnings; warnings.filterwarnings('ignore')
import numpy as np
import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt
# Data setup
# We will use the Tips dataset for our examples
tips = sns.load_dataset('tips')
print('Dataset shape:', tips.shape)
tips.head()
Dataset shape: (244, 7)
total_bill tip sex smoker day time size
0 16.99 1.01 Female No Sun Dinner 2
1 10.34 1.66 Male No Sun Dinner 3
2 21.01 3.50 Male No Sun Dinner 3
3 23.68 3.31 Male No Sun Dinner 2
4 24.59 3.61 Female No Sun Dinner 4

Basics of Probability#

Probability tells us how likely an event is to happen. The value is always between 0 and 1.

Example:

Flipping a fair coin: Probability of getting heads is 0.5 (50%).

# Simulate flipping a coin 10 times
np.random.seed(42)
flips = np.random.choice(['heads', 'tails'], size=10)
print('Results:', flips)
heads_count = np.sum(flips == 'heads')
print('Number of heads:', heads_count)
Results: ['heads' 'tails' 'heads' 'heads' 'heads' 'tails' 'heads' 'heads' 'heads'
 'tails']
Number of heads: 7
# Calculate the probability of getting heads in our simulation
prob_heads = heads_count / len(flips)
print('Simulated probability of heads:', prob_heads)
Simulated probability of heads: 0.7

Descriptive Statistics Overview#

Descriptive statistics help us understand data quickly. Measures like mean (average) and median show central tendency.

Example: What is the average tip amount in our dataset?

# Calculate mean and median tip
mean_tip = tips['tip'].mean()
median_tip = tips['tip'].median()
print('Mean tip:', mean_tip)
print('Median tip:', median_tip)
Mean tip: 2.99827868852459
Median tip: 2.9
# Check for missing data
missing = tips.isnull().sum()
print('Missing values in each column:')
print(missing)
Missing values in each column:
total_bill    0
tip           0
sex           0
smoker        0
day           0
time          0
size          0
dtype: int64
# Remove rows with missing values
tips_clean = tips.dropna()
print('Shape after removing missing data:', tips_clean.shape)
Shape after removing missing data: (244, 7)

What is a Permutation Test?#

A permutation test helps us compare groups by shuffling data many times. If two groups are truly similar, randomly mixing them should not change the average difference much.

Permutation tests do not require normal data or equal group sizes.

# Visualize tip amounts by gender
sns.boxplot(x='sex', y='tip', data=tips_clean)
plt.title('Tip Amounts by Gender')
plt.show()
No description has been provided for this image
# State the hypothesis
print('Null hypothesis: Average tip is the same for men and women.')
print('Alternative hypothesis: Average tip is different for men and women.')
Null hypothesis: Average tip is the same for men and women.
Alternative hypothesis: Average tip is different for men and women.
# Observed mean difference in tip by gender
mean_male = tips_clean[tips_clean['sex'] == 'Male']['tip'].mean()
mean_female = tips_clean[tips_clean['sex'] == 'Female']['tip'].mean()
obs_diff = mean_male - mean_female
print('Observed difference in mean tip (Male - Female):', obs_diff)
Observed difference in mean tip (Male - Female): 0.25616955853283585
# Run a permutation test
n_permutations = 1000
combined = tips_clean['tip'].values
labels = tips_clean['sex'].values
diffs = []
for _ in range(n_permutations):
    shuffled_labels = np.random.permutation(labels)
    mean_male = combined[shuffled_labels == 'Male'].mean()
    mean_female = combined[shuffled_labels == 'Female'].mean()
    diffs.append(mean_male - mean_female)
diffs = np.array(diffs)
# Plot permutation differences and observed value
plt.hist(diffs, bins=30, color='skyblue', edgecolor='black')
plt.axvline(obs_diff, color='red', linestyle='--', label='Observed diff')
plt.xlabel('Mean difference (Male - Female)')
plt.ylabel('Count')
plt.title('Permutation Test Results')
plt.legend()
plt.show()
No description has been provided for this image
# Calculate the p-value
p_value = np.mean(np.abs(diffs) >= np.abs(obs_diff))
print('Estimated p-value:', p_value)
Estimated p-value: 0.158

Interpreting Results#

If the p-value is below your threshold (often 0.05), you can reject the null hypothesis.

Permutation tests provide a way to assess significance without strict requirements about the data.

# Try a two-group comparison for smokers vs non-smokers
mean_smoker = tips_clean[tips_clean['smoker'] == 'Yes']['tip'].mean()
mean_nonsmoker = tips_clean[tips_clean['smoker'] == 'No']['tip'].mean()
obs_diff_smoker = mean_smoker - mean_nonsmoker
print('Observed difference in mean tip (Smoker - Non-Smoker):', obs_diff_smoker)
Observed difference in mean tip (Smoker - Non-Smoker): 0.016855372783593392

Mini-project: Simulate your own permutation experiment!#

  1. Choose two groups in the Tips data (for example, lunch vs dinner).
  2. Compute the observed mean tip difference.
  3. Shuffle the group labels and store the mean differences.
  4. Visualize results and calculate your own p-value.

Share what you learned in the comments!

Recap#

  • Permutation tests help compare groups even with unusual data.
  • You learned to write and run your own in Python.
  • Visualization and p-values help you decide if a difference is meaningful.

Thanks for learning with us today!

What to try next?#

  • Practice on the Titanic dataset: does class or gender change survival rates?
  • Try more advanced tests, like comparing several groups.
  • See how permutation tests connect with bootstrapping.

If you found this helpful, please like and subscribe for more tutorials!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.