Mathew K Analytics

Lesson 15 · Data visualisation in python

Understanding Distribution Plots in Python: Histogram, KDE, Violin, and Box Plots Explained

Today we are learning how to visualize data distributions. We will explore histograms, KDEs, violin plots, and box plots. Let us see how these plots help us…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Introduction to Distribution Plots in Python#

Today we are learning how to visualize data distributions.

We will explore histograms, KDEs, violin plots, and box plots.

Let us see how these plots help us understand time series data!

import warnings; warnings.filterwarnings("ignore")  # Ignore warnings for cleaner output

# Let us import necessary libraries
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns

Data setup#

Let us load actual time series data to explore its distribution.

We will use the Airline Passengers dataset.

This dataset counts the monthly number of airline passengers from 1949 to 1960.

url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/airline-passengers.csv"
data = pd.read_csv(url)
print("Shape of the data:", data.shape)
data.head()
Shape of the data: (144, 2)
Month Passengers
0 1949-01 112
1 1949-02 118
2 1949-03 132
3 1949-04 129
4 1949-05 121
# Let us plot a simple line chart of the time series
plt.figure(figsize=(10,4))
plt.plot(pd.to_datetime(data["Month"]), data["Passengers"], marker="o")
plt.title("Monthly Airline Passengers Over Time")
plt.xlabel("Month")
plt.ylabel("Number of Passengers")
plt.show()
No description has been provided for this image

What is a distribution?#

A distribution shows how often values appear in our data.

Some values are common, some are rare.

Plots like histograms help us see the shape and spread of the data.

# Let us start with a histogram
plt.figure(figsize=(8,4))
sns.histplot(data["Passengers"], bins=12, color="skyblue", kde=False)
plt.title("Histogram of Monthly Airline Passengers")
plt.xlabel("Number of Passengers")
plt.ylabel("Count")
plt.show()
No description has been provided for this image

What can you see in the histogram?#

Are most months busy? Or do quiet months appear often?

Notice where the bars are tallest or shortest.

# Plotting a KDE (Kernel Density Estimate)
plt.figure(figsize=(8,4))
sns.kdeplot(data["Passengers"], fill=True, color="orange")
plt.title("KDE: Smooth Estimate of Passenger Distribution")
plt.xlabel("Number of Passengers")
plt.ylabel("Density")
plt.show()
No description has been provided for this image

KDE plots help us see patterns#

A KDE can show multiple bumps or peaks.

These peaks suggest values that happen more often.

Use KDEs for exploring patterns in your data.

# Let us see both histogram and KDE together
plt.figure(figsize=(8,4))
sns.histplot(data["Passengers"], bins=12, color="skyblue", kde=True)
plt.title("Histogram and KDE of Passenger Numbers")
plt.xlabel("Number of Passengers")
plt.ylabel("Count / Density")
plt.show()
No description has been provided for this image
# Box plots are useful for spotting outliers
plt.figure(figsize=(6,3))
sns.boxplot(y=data["Passengers"], color="lightgreen")
plt.title("Box Plot of Monthly Airline Passengers")
plt.ylabel("Number of Passengers")
plt.show()
No description has been provided for this image

Understanding box plot parts#

The box shows the middle half of the data.

The line inside is the median: the value in the center.

Dots beyond the lines are called outliers.

Box plots quickly show differences between groups.

# Violin plots mix KDE and box ideas
plt.figure(figsize=(6,3))
sns.violinplot(y=data["Passengers"], color="violet")
plt.title("Violin Plot of Monthly Airline Passengers")
plt.ylabel("Number of Passengers")
plt.show()
No description has been provided for this image
# Compare box and violin plot side by side
fig, axes = plt.subplots(1, 2, figsize=(10,3))
sns.boxplot(y=data["Passengers"], color="lightgreen", ax=axes[0])
axes[0].set_title("Box Plot")
sns.violinplot(y=data["Passengers"], color="violet", ax=axes[1])
axes[1].set_title("Violin Plot")
plt.tight_layout()
plt.show()
No description has been provided for this image
# What if there are missing values?
data_missing = data.copy()
data_missing.loc[::12, "Passengers"] = None  # Remove one value per year
print("Missing data count:", data_missing["Passengers"].isnull().sum())
Missing data count: 12
# Plotting with missing values (will skip missing data)
plt.figure(figsize=(8,3))
sns.histplot(data_missing["Passengers"], bins=12, color="coral")
plt.title("Histogram with Missing Values")
plt.xlabel("Number of Passengers")
plt.ylabel("Count")
plt.show()
No description has been provided for this image
# You can fill missing values if you want
filled_data = data_missing["Passengers"].fillna(data_missing["Passengers"].mean())
plt.figure(figsize=(8,3))
sns.histplot(filled_data, bins=12, color="navy")
plt.title("Histogram with Filled Missing Values")
plt.xlabel("Number of Passengers")
plt.ylabel("Count")
plt.show()
No description has been provided for this image

Challenge: Compare different plot types#

Try making a histogram, KDE, box, and violin plot for another dataset.

Which plot helps you see patterns best? Why?

# Mini-project: Explore temperature data!
temp_url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/daily-min-temperatures.csv"
temps = pd.read_csv(temp_url)
print("Temperature data shape:", temps.shape)
temps.head()
Temperature data shape: (3650, 2)
Date Temp
0 1981-01-01 20.7
1 1981-01-02 17.9
2 1981-01-03 18.8
3 1981-01-04 14.6
4 1981-01-05 15.8
# Make all four distribution plots for temperatures
plt.figure(figsize=(16,3))
plt.subplot(1,4,1)
sns.histplot(temps["Temp"], bins=15, color="skyblue")
plt.title("Histogram")
plt.subplot(1,4,2)
sns.kdeplot(temps["Temp"], fill=True, color="orange")
plt.title("KDE")
plt.subplot(1,4,3)
sns.boxplot(y=temps["Temp"], color="lightgreen")
plt.title("Box")
plt.subplot(1,4,4)
sns.violinplot(y=temps["Temp"], color="violet")
plt.title("Violin")
plt.tight_layout()
plt.show()
No description has been provided for this image
# Best practices
# - Check for missing values
# - Use the right plot for your question
# - Combine plots for deeper insight
# - Always add clear labels and titles
# Troubleshooting: What if a plot looks empty or wrong?
try:
    sns.histplot([], bins=10)
    plt.show()
except Exception as e:
    print("Error:", e)
    
No description has been provided for this image
# Extra tip: Plot single value many times
ones_data = [1]*20
plt.figure(figsize=(4,2))
sns.histplot(ones_data, bins=5, color="red")
plt.title("When All Values Are the Same")
plt.show()
No description has been provided for this image
# Ready to practice? Make your own customized plot.
user_vals = input("Enter some numbers (comma-separated): ")
my_vals = [float(x.strip()) for x in user_vals.split(",") if x.strip()]
plt.figure(figsize=(5,2))
sns.histplot(my_vals, bins=min(len(my_vals), 10), color="teal")
plt.title("Your Own Data Histogram")
plt.show()
No description has been provided for this image

Recap: What did we learn?#

We learned how to see distributions with histograms, KDEs, box, and violin plots.

These tools help us discover patterns, spot outliers, and share our story with data.

Remember: experiment with real datasets and try each plot for yourself!

Thanks for joining! Please like, comment, and subscribe for more beginner Python lessons.#

Let us keep exploring the world of data together!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.