Mathew K Analytics

Lesson 54 · Python for Data Science

4 Hypothesis Testing in Python: Step-by-Step Guide for Data Science

Have you ever wondered how scientists make decisions from data? Hypothesis testing is a set of powerful tools that helps us do just that. Today we will…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Introduction to Hypothesis Testing in Python#

Have you ever wondered how scientists make decisions from data? Hypothesis testing is a set of powerful tools that helps us do just that. Today we will learn how to perform hypothesis testing step-by-step in Python.

# Let us import the tools we will use for hypothesis testing
import numpy as np
import scipy.stats as stats

What is Hypothesis Testing?#

Hypothesis testing is a process to decide if a belief about data is correct.

  • We start with a claim, or hypothesis.
  • We gather data and run a test.
  • The test tells us if our claim is supported or rejected.

Many fields use hypothesis testing: science, business, psychology, and more.

# Let us set up some pretend data for our test
np.random.seed(42)  # This makes our random numbers repeatable
sample_A = np.random.normal(loc=50, scale=5, size=30)
sample_B = np.random.normal(loc=53, scale=5, size=30)

Step 1: State the Hypotheses#

  • Null hypothesis (H0): There is no difference between the groups.
  • Alternative hypothesis (H1): There is a difference.

We use hypothesis testing to see if the difference is likely real or just due to chance.

# Let us look at our sample data to get a feel for the values
print('Sample A scores:', sample_A)
print('Sample B scores:', sample_B)
Sample A scores: [52.48357077 49.30867849 53.23844269 57.61514928 48.82923313 48.82931522
 57.89606408 53.83717365 47.65262807 52.71280022 47.68291154 47.67135123
 51.20981136 40.43359878 41.37541084 47.18856235 44.9358444  51.57123666
 45.45987962 42.93848149 57.32824384 48.8711185  50.33764102 42.87625907
 47.27808638 50.55461295 44.24503211 51.87849009 46.99680655 48.54153125]
Sample B scores: [49.99146694 62.26139092 52.93251388 47.71144536 57.11272456 46.89578175
 54.04431798 43.20164938 46.35906976 53.98430618 56.6923329  53.85684141
 52.42175859 51.49448152 45.60739005 49.40077896 50.69680615 58.28561113
 54.71809145 44.18479922 54.62041985 51.0745886  49.61539    56.05838144
 58.15499761 57.6564006  48.80391238 51.45393812 54.65631716 57.87772564]
# Let us calculate and print the mean of each group
mean_A = np.mean(sample_A)
mean_B = np.mean(sample_B)
print('Group A average:', mean_A)
print('Group B average:', mean_B)
Group A average: 49.05926552074481
Group B average: 52.39418764855029

Step 2: Choose a Significance Level#

Most of the time, we use a significance level of 0.05. This means we are willing to accept a 5% chance of a false positive. If our p-value is lower than this, we consider the result statistically significant.

# Let us set our significance level
alpha = 0.05
print('Significance level alpha:', alpha)
Significance level alpha: 0.05

Step 3: Pick the Test#

Because we are comparing means of two groups, we will use the t-test. A t-test is a tool to see if the averages of two groups are different. There are different kinds of t-tests, but let us start with the most common.

# Let us run an independent two-sample t-test
t_stat, p_value = stats.ttest_ind(sample_A, sample_B)
print('t statistic:', t_stat)
print('p-value:', p_value)
t statistic: -2.821074768517716
p-value: 0.006542565964576734
# Let us interpret the results
if p_value < alpha:
    print('Reject the null hypothesis. There is a statistically significant difference!')
else:
    print('Fail to reject the null hypothesis. We do not have enough evidence of a difference.')
    
Reject the null hypothesis. There is a statistically significant difference!

Step 4: Assumptions and Safe Practices#

T-tests assume that data is nearly normal and groups have similar spread (variance). Always check your data before trusting the results.

Let us check for normality with a simple test.

# Let us check normality with the Shapiro-Wilk test
shapiro_A = stats.shapiro(sample_A)
shapiro_B = stats.shapiro(sample_B)
print('Sample A normality p-value:', shapiro_A.pvalue)
print('Sample B normality p-value:', shapiro_B.pvalue)
Sample A normality p-value: 0.6868054942916989
Sample B normality p-value: 0.9129582559088167
# Handling user input: Let us let the user set a significance level
alpha_input = input('Please type your desired significance level (for example, 0.01 or 0.05): ')
alpha = float(alpha_input)
print('Updated significance level:', alpha)
Updated significance level: 0.01
 
# Test again with the updated significance level
if p_value < alpha:
    print('With your alpha, we reject the null hypothesis!')
else:
    print('With your alpha, we do not reject the null hypothesis.')
    
With your alpha, we reject the null hypothesis!

What if Data is not Normal?#

If our data is not normal, we can use a nonparametric test. Mann-Whitney U test is a good choice when the t-test is not suitable.

Let us try it.

# Let us run the Mann-Whitney U test
u_stat, u_p_value = stats.mannwhitneyu(sample_A, sample_B)
print('Mann-Whitney U statistic:', u_stat)
print('p-value:', u_p_value)
Mann-Whitney U statistic: 269.0
p-value: 0.007617064306244217

Challenge: Mini Project#

Can you test if people who sleep more, on average, score higher in a quiz?

  1. Make up two groups: 'less_sleep' and 'more_sleep', each with at least 20 scores.
  2. Calculate measures, compare means, and pick a test.
  3. Decide if the difference is meaningful.

Try using both t-test and Mann-Whitney!

# Example for the mini project:
less_sleep = np.random.normal(55, 6, 25)
more_sleep = np.random.normal(60, 6, 25)
print('Average score (less sleep):', np.mean(less_sleep))
print('Average score (more sleep):', np.mean(more_sleep))

# Let us perform a t-test
mp_t_stat, mp_p_val = stats.ttest_ind(less_sleep, more_sleep)
print('p-value (t-test):', mp_p_val)
Average score (less sleep): 54.91855354265035
Average score (more sleep): 59.69750759216644
p-value (t-test): 0.003132224197586407
# Try the Mann-Whitney U test for the same data
mp_u_stat, mp_u_p_val = stats.mannwhitneyu(less_sleep, more_sleep)
print('p-value (Mann-Whitney):', mp_u_p_val)
p-value (Mann-Whitney): 0.004901861038274268

Troubleshooting: Common Mistakes#

  1. Define your hypotheses before testing.
  2. Check if data is normal and variances are similar.
  3. Be careful not to run tests just to get a low p-value.
  4. Always look at the real-world importance, not just the numbers.
  5. Do not forget to check your sample sizes.
# Quick tips: Always label your variables clearly and document what each test is for
# Use comments to remind your future self why you did things
 

Summary#

  • Hypothesis testing helps us make decisions from data.
  • Always state your hypotheses and check assumptions first.
  • Use t-tests for normal data, Mann-Whitney for non-normal data.
  • Look beyond p-values to what the results mean for the real world.

Great work today!

Thank you for joining!#

If you enjoyed this lesson:

  • Like the video
  • Subscribe for more beginner Python tutorials
  • Share with a friend who wants to learn data science

Let us keep learning together. See you in the next video!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.