Lesson 16 · Probability and Statistics in python
Understanding t-Tests, Chi-Square Tests, and ANOVA in Python for Statistical Analysis
In this lesson, we will explore t-Tests, Chi-Square Tests, and ANOVA. We will use real datasets to learn how these methods work and when to use them. No…
- CourseProbability and Statistics in python
- Lesson16 of 35
- Video11 min
- FormatJupyter notebook · 13 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWelcome to Hypothesis Testing in Python!#
In this lesson, we will explore t-Tests, Chi-Square Tests, and ANOVA.
We will use real datasets to learn how these methods work and when to use them.
No prior experience is needed. Let us start our journey into statistical testing!
What are Hypothesis Tests?#
Hypothesis tests help us decide if we see real differences between groups or just random chance.
We often use these tests in science, business, and medicine to make decisions.
# Suppress warnings for a cleaner notebook
import warnings; warnings.filterwarnings("ignore")
# Import key packages
import pandas as pd
import numpy as np
from scipy import stats
import seaborn as sns
import matplotlib.pyplot as plt
# Data setup
tips = sns.load_dataset("tips")
print("Shape:", tips.shape)
tips.head()
Why do we use t-Tests?#
A t-Test checks if the average value from two groups is really different, or just by chance.
For example: Do people tip more on weekends?
# Quick explore: average tip on weekend vs weekday
tips["day_type"] = tips["day"].apply(lambda x: "Weekend" if x in ["Sat", "Sun"] else "Weekday")
weekend_mean = tips.query('day_type == "Weekend"')["tip"].mean()
weekday_mean = tips.query('day_type == "Weekday"')["tip"].mean()
print("Average tip on weekends:", round(weekend_mean, 2))
print("Average tip on weekdays:", round(weekday_mean, 2))
# Doing the independent t-Test
weekend_tips = tips.query('day_type == "Weekend"')["tip"]
weekday_tips = tips.query('day_type == "Weekday"')["tip"]
t_stat, p_val = stats.ttest_ind(weekend_tips, weekday_tips, equal_var=False)
print("t-statistic:", round(t_stat, 3))
print("p-value:", p_val)
# Visualize the difference in average tips
sns.boxplot(x="day_type", y="tip", data=tips, palette="Set2");
plt.title("Tip Distribution: Weekend vs Weekday")
plt.ylabel("Tip Amount ($)")
plt.show()
What about more than two groups?#
To compare more than two groups (for example, tips on different days), we use ANOVA.
ANOVA asks if there are any group means that are different from the others.
# One-way ANOVA: average tips by each day
days = [day for day in tips["day"].unique()]
groups = [tips.query('day == @d')["tip"] for d in days]
anova_stat, anova_p = stats.f_oneway(*groups)
print("ANOVA F-statistic:", round(anova_stat, 3))
print("ANOVA p-value:", anova_p)
# Visualize tips for each day
sns.boxplot(x="day", y="tip", data=tips, palette="Set3");
plt.title("Tip Distribution by Day")
plt.ylabel("Tip Amount ($)")
plt.show()
What is a Chi-Square Test?#
A Chi-Square Test checks if two categories are related or independent.
Example: Are men and women equally likely to tip above $3?
# Chi-Square test: Are males/females equally likely to tip above $3?
tips["tip_above_3"] = tips["tip"] > 3
contingency = pd.crosstab(tips["sex"], tips["tip_above_3"])
chi2, p, dof, ex = stats.chi2_contingency(contingency)
print("Chi-square statistic:", round(chi2, 3))
print("p-value:", p)
contingency
# Visualize the rate of tips above $3 by gender
tip_rates = contingency.div(contingency.sum(axis=1), axis=0)
tip_rates[True].plot(kind="bar", color=["skyblue", "salmon"]);
plt.ylabel("Fraction tipping above $3")
plt.title("Tip above $3: Male vs Female")
plt.ylim(0,1)
plt.show()
# Let us confirm: What is our sample size for the main tests?
print("Number of rows:", len(tips))
tips["sex"].value_counts()
Handling missing data safely#
Sometimes real data has missing values.
Let us check and clean if needed for these tests.
# Check for missing data
print(tips.isnull().sum())
# Drop rows with missing values for our columns of interest
clean_tips = tips.dropna(subset=["tip", "day", "sex"]).copy()
print("Rows after cleaning:", len(clean_tips))
Practice: t-Test for smokers vs non-smokers#
Try testing if smokers tip more than non-smokers.
You can use: stats.ttest_ind() just like before.
# Solution: t-Test for smokers vs non-smokers
smoker_tips = clean_tips[clean_tips["smoker"]=="Yes"]["tip"]
nonsmoker_tips = clean_tips[clean_tips["smoker"]=="No"]["tip"]
t_stat_smoker, p_val_smoker = stats.ttest_ind(smoker_tips, nonsmoker_tips, equal_var=False)
print("Smoker avg:", round(smoker_tips.mean(),2))
print("Non-smoker avg:", round(nonsmoker_tips.mean(),2))
print("t-statistic:", round(t_stat_smoker, 3))
print("p-value:", p_val_smoker)
# Ask user: Would you expect smokers or non-smokers to tip more? Why?
user_guess = input("Who do you think tips more on average: smokers or non-smokers? Why? ")
print("You guessed:", user_guess)
Extra: Try ANOVA on lunch versus dinner tips by day#
Can you run an ANOVA test but split by meal time?
What story might the results tell for a restaurant owner?
Recap: What have we learned?#
You can now:
- Compare two group means with t-Tests
- Compare many groups with ANOVA
- Test if two categories are related with Chi-Square
- Check and handle missing data
These skills unlock many insights in real-world data!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



