Lesson 52 · Python for Data Science
2 Probability and Normal Distribution in Python
Welcome! Today we will learn about probability and the normal distributiontwo important ideas for data science and many real-world tasks. This lesson will…
- CoursePython for Data Science
- Lesson52 of 38
- Video1 min
- FormatJupyter notebook · 19 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Probability and Normal Distribution in Python#
Welcome! Today we will learn about probability and the normal distributiontwo important ideas for data science and many real-world tasks.
This lesson will help you see how these ideas work in Python.
By the end, you will know how to make random numbers, visualize probability, and use normal distributions for real-life problems.
Let us get started!
What is Probability?#
Probability measures how likely something is to happen. It is a number between 0 and 1.
For example, 0 means impossible, 1 means certain, and 0.5 means a 50-50 chance.
# Let us use Python to flip a coin.
import random
coin = random.choice(['Heads', 'Tails'])
print('Coin flip result:', coin)
# Let us repeat the coin flip many times to see the probability in action.
flips = 100
heads_count = 0
for i in range(flips):
if random.choice(['Heads', 'Tails']) == 'Heads':
heads_count += 1
prob_heads = heads_count / flips
print('Proportion of heads:', prob_heads)
What is a Normal Distribution?#
The normal distribution is a famous pattern in statistics. Many real-world things, like heights or scores, follow this pattern.
It looks like a bell curvemost values are near the middle, and fewer are far away.
# Let us make some random numbers that follow a normal distribution.
import numpy as np
numbers = np.random.normal(loc=0, scale=1, size=1000)
print('First 10 numbers:', numbers[:10])
# Let us see what shape our numbers make.
import matplotlib.pyplot as plt
plt.hist(numbers, bins=30, color='skyblue')
plt.title('Histogram of Normally Distributed Numbers')
plt.xlabel('Value')
plt.ylabel('Frequency')
plt.show()
# What do mean and standard deviation mean?
mean = np.mean(numbers)
std = np.std(numbers)
print('Mean:', mean)
print('Standard deviation:', std)
Why Does Normal Distribution Matter?#
Many real-world measurements (like exam scores, peoples heights, or measurement errors) closely follow the normal distribution.
Knowing this helps us make predictions and decisions.
# How likely is a number in our group to be close to zero?
close_count = sum(abs(x) < 1 for x in numbers)
prob_close = close_count / len(numbers)
print('Fraction between -1 and 1:', prob_close)
# Let the user try picking their own mean and standard deviation.
user_mean = float(input('Pick a mean (try 5): '))
user_std = float(input('Pick a standard deviation (try 2): '))
user_numbers = np.random.normal(loc=user_mean, scale=user_std, size=1000)
plt.hist(user_numbers, bins=30, color='lightgreen')
plt.title('Your Normal Distribution')
plt.xlabel('Value')
plt.ylabel('Frequency')
plt.show()
# Let us calculate the probability that a number is greater than 7 in your data.
count_above_7 = sum(x > 7 for x in user_numbers)
prob_above_7 = count_above_7 / len(user_numbers)
print('Fraction above 7:', prob_above_7)
# What if we want to know about two conditions at once?
between_count = sum(3 < x < 7 for x in user_numbers)
prob_between = between_count / len(user_numbers)
print('Fraction between 3 and 7:', prob_between)
Working with Real-World Data#
Let us use height data as an example.
Suppose we have heights of some people, and want to know if they follow a normal distribution.
# Here is some pretend height data (in centimeters).
heights = [172, 160, 168, 174, 181, 165, 177, 170, 169, 173, 175, 162, 180]
plt.hist(heights, bins=5, color='plum')
plt.title('Histogram of Heights')
plt.xlabel('Height (cm)')
plt.ylabel('Number of People')
plt.show()
# Let us check mean and standard deviation of the heights.
height_mean = np.mean(heights)
height_std = np.std(heights)
print('Mean height:', height_mean)
print('Standard deviation:', height_std)
# Let us standardize the heights (z-scores).
z_scores = [(h - height_mean) / height_std for h in heights]
print('Z-scores:', z_scores)
# Let us sort the heights to see the range.
sorted_heights = sorted(heights)
print('Sorted heights:', sorted_heights)
# Let us combine two groups of heights and see how statistics change.
more_heights = [178, 166, 164, 171]
combined_heights = heights + more_heights
combined_mean = np.mean(combined_heights)
print('New mean:', combined_mean)
Mini-Project: Simple Grade Simulator#
Suppose test scores for a big class are normally distributed with an average of 70 and a standard deviation of 10.
We want to simulate some scores and count how many students scored above 85.
# Make 200 random test scores and count high scores.
scores = np.random.normal(loc=70, scale=10, size=200)
num_high = sum(s > 85 for s in scores)
print('Scores above 85:', num_high)
# What is the highest and lowest score in this class?
highest = np.max(scores)
lowest = np.min(scores)
print('Highest:', highest)
print('Lowest:', lowest)
# Let us draw the score distribution for the class.
plt.hist(scores, bins=20, color='salmon')
plt.title('Class Test Scores')
plt.xlabel('Score')
plt.ylabel('Number of Students')
plt.show()
Common Pitfalls#
- Not importing numpy or matplotlib before using them
- Using a wrong size or type for input
- Mistaking mean and standard deviation
- Misunderstanding how probabilities work
Checking these will help avoid errors!
# Final challenge: Ask the user to simulate their own test results.
user_avg = float(input('Pick a class mean (try 75): '))
user_sd = float(input('Pick a class standard deviation (try 12): '))
how_many = int(input('How many students? (try 40): '))
test_results = np.random.normal(user_avg, user_sd, how_many)
plt.hist(test_results, bins=10, color='orange')
plt.title('Your Simulated Test Results')
plt.xlabel('Score')
plt.ylabel('Number of Students')
plt.show()
print('Mean:', np.mean(test_results))
print('Standard deviation:', np.std(test_results))
Summary#
- Probability helps us understand chance in life and data
- The normal distribution is everywhere in nature, tests, and science
- Python makes simulating and visualizing these concepts much easier
- Now you can explore random events, analyze data, and predict outcomes
Great job finishing the lesson!
Want more lessons? Please like, comment, subscribe, and share with friends!#
Thanks for learning with usyou are ready to try Python and statistics on your own!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



