Mathew K Analytics

Lesson 15 · Probability and Statistics in python

Hypothesis Testing Explained: Concepts, Errors, and How to Interpret P-Values in Python

Welcome! Today we are learning how to compare groups, check if our results are real or by chance, and use statistics to make decisions. We will cover key…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Introduction to Hypothesis Testing in Python#

Welcome! Today we are learning how to compare groups, check if our results are real or by chance, and use statistics to make decisions.

We will cover key ideas, learn what errors and p-values are, and work hands-on in Python.

Let us get started!

# Always suppress warnings to keep everything neat
import warnings
from statsmodels.tools.sm_exceptions import ConvergenceWarning, ValueWarning
warnings.filterwarnings('ignore', category=ValueWarning)
warnings.filterwarnings('ignore', category=ConvergenceWarning)
warnings.filterwarnings('ignore', category=RuntimeWarning)  # e.g., from log/exp on edge values

# Set up key libraries for data, stats, and plots
import pandas as pd
import numpy as np
import seaborn as sns
import matplotlib.pyplot as plt
from scipy import stats

sns.set(style='whitegrid')

What is Hypothesis Testing?#

Hypothesis testing helps us decide: Are our results real or due to randomness?

We start with a 'null hypothesis.' This is usually a statement of 'no difference' or 'nothing's going on.'

We collect data, run a test, and based on the result, we either reject or fail to reject the null hypothesis.

Common examples include comparing two averages, checking if a medicine works, or seeing if a new teaching method helps.

# Let us simulate flipping a coin 20 times
np.random.seed(42)
flips = np.random.choice(['Heads', 'Tails'], size=20)
print('Coin flips:', flips)
Coin flips: ['Heads' 'Tails' 'Heads' 'Heads' 'Heads' 'Tails' 'Heads' 'Heads' 'Heads'
 'Tails' 'Heads' 'Heads' 'Heads' 'Heads' 'Tails' 'Heads' 'Tails' 'Tails'
 'Tails' 'Heads']
# Let us count the outcomes
heads = np.sum(flips == 'Heads')
tails = np.sum(flips == 'Tails')
print('Heads:', heads)
print('Tails:', tails)
Heads: 13
Tails: 7

Types of Hypotheses#

Null Hypothesis (H0): There is no real change or effect.

Alternative Hypothesis (H1): There is a real change or effect.

For our coin, H0 says the coin is fair. H1 says the coin is not fair.

# How unusual is our outcome? Let us compute a p-value.
# We use a binomial test: for 20 flips, expected heads = 10
try:
    p_value = stats.binomtest(heads, n=20, p=0.5, alternative='two-sided')
    # Newer scipy: returns a BinomTestResult object with .pvalue
    print('p-value:', p_value.pvalue)
except AttributeError:
    # Legacy: use stats.binom_test
    pvalue_legacy = stats.binom_test(heads, n=20, p=0.5, alternative='two-sided')
    print('p-value:', pvalue_legacy)
    # for lesson, make sure p_value is defined & compatible
    class Dummy:
        def __init__(self, p): self.pvalue = p
    p_value = Dummy(pvalue_legacy)
    
p-value: 0.26317596435546875
# What does the p-value mean?
if p_value.pvalue < 0.05:
    print('We reject the null hypothesis. The coin may not be fair!')
else:
    print('We fail to reject the null. The coin seems fair enough.')
    
We fail to reject the null. The coin seems fair enough.

Introducing Type I and Type II Errors#

  • Type I Error: Rejecting the null when it is actually true (false alarm).
  • Type II Error: Not rejecting the null when it is actually false (missed it).

If we set our cutoff for p-value as 0.05, there is a 5% chance of a Type I error.

# See errors in action: let us simulate many experiments
n_exp = 1000
false_alarms = 0
for i in range(n_exp):
    flips_sim = np.random.choice(['Heads', 'Tails'], size=20)
    heads_sim = np.sum(flips_sim == 'Heads')
    try:
        p_sim = stats.binomtest(heads_sim, n=20, p=0.5, alternative='two-sided')
        pval = p_sim.pvalue
    except AttributeError:
        pval = stats.binom_test(heads_sim, n=20, p=0.5, alternative='two-sided')
    if pval < 0.05:
        false_alarms += 1
print('False alarms (Type I errors) in 1000 simulated fair coins:', false_alarms)
False alarms (Type I errors) in 1000 simulated fair coins: 37

Switching to Real Data: Restaurant Tips Example#

We will use the 'tips' dataset to learn how to test if people tip differently on Sundays.

Real datasets let us practice making discoveries with actual observations.

# Data setup: Load the tips dataset
tips = sns.load_dataset('tips')
print('Shape:', tips.shape)
tips.head()
Shape: (244, 7)
total_bill tip sex smoker day time size
0 16.99 1.01 Female No Sun Dinner 2
1 10.34 1.66 Male No Sun Dinner 3
2 21.01 3.50 Male No Sun Dinner 3
3 23.68 3.31 Male No Sun Dinner 2
4 24.59 3.61 Female No Sun Dinner 4
# Compare tips on Sunday vs. other days: plot and summary
sunday = tips[tips['day'] == 'Sun']['tip']
other = tips[tips['day'] != 'Sun']['tip']
plt.figure(figsize=(7,4))
sns.histplot(sunday, color='red', kde=True, label='Sunday', bins=12)
sns.histplot(other, color='blue', kde=True, label='Other Days', bins=12)
plt.legend()
plt.title('Tip distributions: Sunday vs. Other Days')
plt.xlabel('Tip Amount ($)')
plt.ylabel('Frequency')
plt.show()
No description has been provided for this image
# Hypothesis test: Are Sunday tips different?
t_stat, p_tip = stats.ttest_ind(sunday, other, equal_var=False)
print('t-statistic:', t_stat)
print('p-value:', p_tip)
t-statistic: 2.0753622376964542
p-value: 0.03948957581757897
# Interpreting our result
if p_tip < 0.05:
    print('Tips on Sundays are likely different from other days!')
else:
    print('No significant difference found in average tips.')
    
Tips on Sundays are likely different from other days!

Key Takeaways: Errors, p-values, and Real Decisions#

  • P-values tell us how likely our result is just from chance.
  • Type I errors are false alarms: finding a difference that is not really there.
  • Type II errors are missed discoveries: not finding a real difference.

Final step: Let us practice with one more example!

# Short practice: Enter tip amounts for two groups and test
group1 = []
group2 = []
print('Enter 5 tip amounts for Group 1:')
for i in range(5):
    group1.append(float(input(f'Tip {i+1} for Group 1: ')))
print('Enter 5 tip amounts for Group 2:')
for i in range(5):
    group2.append(float(input(f'Tip {i+1} for Group 2: ')))
t_stat2, p_value2 = stats.ttest_ind(group1, group2, equal_var=False)
print('t-statistic:', t_stat2)
print('p-value:', p_value2)
Enter 5 tip amounts for Group 1:
Enter 5 tip amounts for Group 2:
t-statistic: 3.642141316611115
p-value: 0.015409766797827351
# Practice: Interpret your result
if p_value2 < 0.05:
    print('Your groups have a real difference in tips!')
else:
    print('No major difference in tips between your groups.')
    
Your groups have a real difference in tips!

Challenge Exercise#

Pick two other days from the tips dataset.

Test if the tip amounts are different between them.

Try plotting the results and explaining your findings.

Recap: What Have You Learned?#

  • Null and alternative hypotheses help us check for real differences.
  • P-values tell us if our findings could just be luck.
  • We practiced with coin flips and real restaurant data.
  • Errors remind us to be careful drawing conclusions.

Keep exploring to grow your data skills!

Thanks for Learning! Watch More and Subscribe#

If you enjoyed this lesson, check out the next video for deeper statistics.

Like and subscribe for new tutorials each week!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.