Mathew K Analytics

Lesson 46 · Python for Data Science

Distribution & Relationship Analysis in Python | Exploratory Data Analysis Tutorial

Welcome! Today we will learn how to look at and analyze data: how things are spread out, and how they connect. Knowing these skills helps you understand…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Exploring Distributions and Relationships in Python#

Welcome! Today we will learn how to look at and analyze data: how things are spread out, and how they connect.

Knowing these skills helps you understand patterns, spot outliers, and make better decisions with data.

We will use easy Python tools, beginner tips, and real data examples.

Let us get started!

# Let us start by making a simple list of numbers to represent test scores.
scores = [90, 85, 78, 82, 95, 88, 76, 91, 87, 85]
print('Test scores:', scores)
Test scores: [90, 85, 78, 82, 95, 88, 76, 91, 87, 85]
# To understand how the scores are spread out, let us check their minimum and maximum values.
minimum = min(scores)
maximum = max(scores)
print('Lowest score:', minimum)
print('Highest score:', maximum)
Lowest score: 76
Highest score: 95

What is a Distribution?#

In data, 'distribution' means how values are grouped or spread out.

For example, are most test scores high, low, or in the middle?

Understanding distribution helps you find trends and unusual points.

# Let us calculate the average score, also called the mean.
average = sum(scores) / len(scores)
print('Average score:', average)
Average score: 85.7
# Median helps us know the middle value in our sorted data.
sorted_scores = sorted(scores)
n = len(sorted_scores)
if n % 2 == 1:
    median = sorted_scores[n//2]
else:
    median = (sorted_scores[n//2 - 1] + sorted_scores[n//2]) / 2
print('Median score:', median)
Median score: 86.0
# Let us check how far scores are from the average using standard deviation.
import math
mean = average
variance = sum((x - mean)**2 for x in scores) / len(scores)
std_dev = math.sqrt(variance)
print('Standard deviation:', std_dev)
Standard deviation: 5.550675634551167

Visualizing Distributions without Graphs#

Sometimes, just printing data in order helps.

If you understand the spread in numbers, you can make smart guessesand you do not need fancy graphs for that!

# Print the sorted scores to see the shape of our data.
print('Sorted scores:', sorted_scores)
Sorted scores: [76, 78, 82, 85, 85, 87, 88, 90, 91, 95]
# Let us find the most common score using a frequency count.
from collections import Counter
counts = Counter(scores)
most_common = counts.most_common(1)[0]
print('Most common score:', most_common[0], 'appears', most_common[1], 'times')
Most common score: 85 appears 2 times

What about Relationships?#

A 'relationship' in data means one value changes as another does.

For example, maybe higher hours studied leads to higher test scores.

We look for patterns between two sets of data to see if they might connect.

# Let us make two lists: hours studied and scores received.
hours_studied = [5, 4, 3, 2, 6, 4, 1, 5, 3, 2]
print('Hours studied:', hours_studied)
print('Test scores:', scores)
Hours studied: [5, 4, 3, 2, 6, 4, 1, 5, 3, 2]
Test scores: [90, 85, 78, 82, 95, 88, 76, 91, 87, 85]
# Let us pair up hours and scores to look for patterns.
paired = list(zip(hours_studied, scores))
print('Hours and scores:', paired)
Hours and scores: [(5, 90), (4, 85), (3, 78), (2, 82), (6, 95), (4, 88), (1, 76), (5, 91), (3, 87), (2, 85)]
# We can measure relationship strength using correlation.
def mean(lst):
    return sum(lst) / len(lst)

def correlation(x, y):
    avg_x = mean(x)
    avg_y = mean(y)
    num = sum((xi - avg_x) * (yi - avg_y) for xi, yi in zip(x, y))
    den_x = math.sqrt(sum((xi - avg_x) ** 2 for xi in x))
    den_y = math.sqrt(sum((yi - avg_y) ** 2 for yi in y))
    return num / (den_x * den_y)

corr = correlation(hours_studied, scores)
print('Correlation between hours studied and scores:', corr)
Correlation between hours studied and scores: 0.870764867478004
# Safe data access: Let us try reading a value and handle possible errors.
try:
    idx = int(input('Enter a student index (0-9): '))
    score = scores[idx]
    print('Student', idx, 'score:', score)
except (ValueError, IndexError):
    print('Please enter a valid index between 0 and 9.')
    
Student 2 score: 78
 
# Let us update a score to see how data changes.
print('Original scores:', scores)
scores[0] = 99
print('Updated scores:', scores)
Original scores: [90, 85, 78, 82, 95, 88, 76, 91, 87, 85]
Updated scores: [99, 85, 78, 82, 95, 88, 76, 91, 87, 85]
# Remove a scorelet us see what happens to the distribution.
removed = scores.pop()
print('Removed last score:', removed)
print('Scores now:', scores)
Removed last score: 85
Scores now: [99, 85, 78, 82, 95, 88, 76, 91, 87]
# Quick look at built-in stats: Use statistics.mean for easy averages.
import statistics
avg2 = statistics.mean(scores)
print('Easy mean:', avg2)
Easy mean: 86.77777777777777
# Let us loop through all scores: What if you want to know which are above 90?
high_scores = []
for score in scores:
    if score > 90:
        high_scores.append(score)
print('High scores over 90:', high_scores)
High scores over 90: [99, 95, 91]
# List comprehensions give us a shorter way to filter data.
eighty_and_up = [s for s in scores if s >= 80]
print('Scores 80 or higher:', eighty_and_up)
Scores 80 or higher: [99, 85, 82, 95, 88, 91, 87]
# Let us sort the paired hours and scores by hours studied.
sorted_pairs = sorted(zip(hours_studied, scores), key=lambda x: x[0])
print('Sorted by hours studied:', sorted_pairs)
Sorted by hours studied: [(1, 76), (2, 82), (3, 78), (3, 87), (4, 85), (4, 88), (5, 99), (5, 91), (6, 95)]
# Now, let us look at a small real-world-like set: survey ages and satisfaction.
ages = [21, 34, 19, 45, 28, 30, 40, 23, 25, 37]
satisfaction = [6, 8, 5, 7, 9, 6, 7, 8, 5, 9]
print('Ages:', ages)
print('Satisfaction (1-10):', satisfaction)
Ages: [21, 34, 19, 45, 28, 30, 40, 23, 25, 37]
Satisfaction (1-10): [6, 8, 5, 7, 9, 6, 7, 8, 5, 9]
# Analyze the relationship between age and satisfaction.
corr_age_sat = correlation(ages, satisfaction)
print('Correlation between age and satisfaction:', corr_age_sat)
Correlation between age and satisfaction: 0.4147806778921701
# Mini-project: Let us make a function that shows full summary statistics.
def summarize(data):
    print('Min:', min(data))
    print('Max:', max(data))
    print('Mean:', statistics.mean(data))
    print('Median:', statistics.median(data))
    print('Standard deviation:', statistics.stdev(data))
    counts = Counter(data)
    mode_val = counts.most_common(1)[0][0]
    print('Mode:', mode_val)

print('Summary for ages:')
summarize(ages)
print()
print('Summary for satisfaction:')
summarize(satisfaction)
Summary for ages:
Min: 19
Max: 45
Mean: 30.2
Median: 29.0
Standard deviation: 8.625543461139129
Mode: 21

Summary for satisfaction:
Min: 5
Max: 9
Mean: 7
Median: 7.0
Standard deviation: 1.4907119849998598
Mode: 6
# Mini-project, Part 2: Using user input to analyze your own data.
raw = input('Enter a list of numbers, separated by spaces: ')
user_data = [float(num) for num in raw.strip().split()]
print('Your data:', user_data)
summarize(user_data)
Your data: [20.0, 30.0, 25.0, 40.0, 35.0, 50.0]
Min: 20.0
Max: 50.0
Mean: 33.333333333333336
Median: 32.5
Standard deviation: 10.801234497346433
Mode: 20.0
 
# Optimization tip: Use NumPy for even faster math on big lists.
import numpy as np
biglist = np.array([1, 2, 3, 4, 5, 6, 7, 8, 9, 10])
print('Mean:', biglist.mean())
print('Standard deviation:', biglist.std())
Mean: 5.5
Standard deviation: 2.8722813232690143
# Troubleshooting: Catching division by zero and empty data errors.
try:
    empty_list = []
    avg_empty = sum(empty_list) / len(empty_list)
except ZeroDivisionError:
    print('You cannot divide by zero. The list is empty.')
    
You cannot divide by zero. The list is empty.
# Extra trick: Remove outliers by filtering numbers far from the mean.
mean_val = statistics.mean(scores)
std_val = statistics.stdev(scores)
filtered = [x for x in scores if abs(x - mean_val) <= 2 * std_val]
print('Scores with outliers removed:', filtered)
Scores with outliers removed: [99, 85, 78, 82, 95, 88, 76, 91, 87]
# Challenge: Write your own function to get the range (max minus min) of a list.
def data_range(data):
    return max(data) - min(data)

print('Range for satisfaction:', data_range(satisfaction))
Range for satisfaction: 4

Recap: What Did We Learn?#

  • Distributions show how data is spread.
  • Relationships reveal how two things move together.
  • We calculated, filtered, and combined data in Python.
  • Safe coding protects against errors and bad inputs.
  • Mini-projects help connect all the skills together.

With these basics, you can understand almost any data set you run into!

Thanks for learning with us!#

If you liked this lesson on analyzing distributions and relationships in Python, please:

  • Like the video
  • Comment and tell us what you built
  • Subscribe for more Python basics
  • Share with your friends who want to learn data skills

Keep exploringthere is so much you can do with data and Python!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.