Lesson 73 · Python Fundamentals
Fundamentals of Descriptive Statistics in Python for Data Analysis
Welcome! In this lesson, we explore how Python helps us summarize and understand our data. Descriptive statistics let us describe and interpret numbers…
- CoursePython Fundamentals
- Lesson73 of 22
- Video10 min
- FormatJupyter notebook · 14 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Descriptive Statistics in Python#
Welcome! In this lesson, we explore how Python helps us summarize and understand our data.
Descriptive statistics let us describe and interpret numbers quickly and clearly. This helps us make sense of the information around us.
No advanced math needed! We will use simple Python code and relatable examples.
What Are Descriptive Statistics?#
Descriptive statistics are numbers that help describe a set of data. They answer questions like:
- What is typical in my data?
- How spread out are the values?
- Are there any interesting patterns?
Let us jump in and see how Python makes this easy!
# First, let us create a simple list of numbers
ages = [23, 25, 31, 18, 24, 29, 21]
print(ages) # Output the list so we can see it
What is the Mean?#
The mean is what many people call the 'average.' It adds all the numbers together and divides by how many numbers there are.
# Calculate the mean age
mean_age = sum(ages) / len(ages)
print('The mean age is:', mean_age)
What is the Median?#
The median is the number that is right in the middle when all values are put in order. If there is an even number, it averages the two middle ones.
# Let us find the median age
sorted_ages = sorted(ages)
n = len(sorted_ages)
if n % 2 == 1:
median_age = sorted_ages[n // 2] # Odd case, one middle value
else:
middle1 = sorted_ages[n // 2 - 1]
middle2 = sorted_ages[n // 2]
median_age = (middle1 + middle2) / 2 # Even case, average two middle values
print('The median age is:', median_age)
What is the Mode?#
The mode is the value that appears most often in the list. It shows what is most common.
# Find the mode of the ages
from collections import Counter
counts = Counter(ages)
max_count = max(counts.values())
modes = [age for age, count in counts.items() if count == max_count]
print('The mode(s) of ages is/are:', modes)
Range: How Spread Out?#
The range is the difference between the largest and smallest values. It tells us how wide our set of numbers is.
# Let us calculate the range of ages
range_ages = max(ages) - min(ages)
print('The range of ages is:', range_ages)
Standard Deviation: How Much Do Values Vary?#
Standard deviation shows how much the numbers are spread out around the mean. A small standard deviation means most values are close to the average. A big one means values are more widely spread.
# Compute the standard deviation without using libraries
mean = sum(ages) / len(ages)
variance = sum((x - mean) ** 2 for x in ages) / len(ages)
std_dev = variance ** 0.5
print('Standard deviation is:', std_dev)
Using Built-in Python Statistics#
Python has a standard library called statistics that makes summary calculations easy. Let us see how simple it can be.
import statistics
# Quick summary using the statistics library
print('Mean using statistics:', statistics.mean(ages))
print('Median using statistics:', statistics.median(ages))
print('Standard deviation:', statistics.stdev(ages))
Mini-Project: Survey Results Analyzer (Part 1)#
Let us analyze data from a pretend survey: people's scores on a quiz. We will ask you to enter your scores below:
Try it with numbers like 70, 80, 85, 60, 95, 100, 88.
# Accept user input for survey scores
scores_input = input('Please enter your scores separated by commas: ')
scores = [int(x.strip()) for x in scores_input.split(',')]
print('Your scores are:', scores)
# Show all summary statistics on scores
print('Mean score:', statistics.mean(scores))
print('Median score:', statistics.median(scores))
print('Mode(s):', statistics.multimode(scores))
print('Range:', max(scores) - min(scores))
print('Standard deviation:', statistics.stdev(scores))
Mini-Project (Part 2): What is Typical?#
Sometimes a very high or very low value can change the mean. Try adding a score like 1000 to see what happens to the statistics.
Do you notice which statistic stays the same and which one changes a lot?
# Add 1000 to scores and see the new mean and median
scores_with_outlier = scores + [1000]
print('New mean:', statistics.mean(scores_with_outlier))
print('New median:', statistics.median(scores_with_outlier))
Challenge: Your Turn!#
Try making your own list of numbersfor example, daily steps, test scores, or hours slept each night.
Calculate the mean, median, and standard deviation for your list.
Which value best describes your data?
# Ask the user for a new list to analyze
user_input = input('Type numbers, separated by commas: ')
user_numbers = [float(num.strip()) for num in user_input.split(',')]
print('Mean:', statistics.mean(user_numbers))
print('Median:', statistics.median(user_numbers))
print('Standard deviation:', statistics.stdev(user_numbers))
Recap & Next Steps#
Descriptive statistics help us quickly sum up what is going on in any set of numbers. We learned how to get the mean, median, mode, range, and standard deviation with Python.
What else would you like to analyze? The next lesson will reveal even more tools!
Thanks for following along!
Thanks for Watching!#
If you enjoyed this beginner Python lesson, please:
- Like the video
- Subscribe for more tutorials
- Comment your favorite statistic
- Share with a friend
See you next time!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



