Lesson 15 · Data visualisation in python
Understanding Distribution Plots in Python: Histogram, KDE, Violin, and Box Plots Explained
Today we are learning how to visualize data distributions. We will explore histograms, KDEs, violin plots, and box plots. Let us see how these plots help us…
- CourseData visualisation in python
- Lesson15 of 34
- Video11 min
- FormatJupyter notebook · 18 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIntroduction to Distribution Plots in Python#
Today we are learning how to visualize data distributions.
We will explore histograms, KDEs, violin plots, and box plots.
Let us see how these plots help us understand time series data!
import warnings; warnings.filterwarnings("ignore") # Ignore warnings for cleaner output
# Let us import necessary libraries
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
Data setup#
Let us load actual time series data to explore its distribution.
We will use the Airline Passengers dataset.
This dataset counts the monthly number of airline passengers from 1949 to 1960.
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/airline-passengers.csv"
data = pd.read_csv(url)
print("Shape of the data:", data.shape)
data.head()
# Let us plot a simple line chart of the time series
plt.figure(figsize=(10,4))
plt.plot(pd.to_datetime(data["Month"]), data["Passengers"], marker="o")
plt.title("Monthly Airline Passengers Over Time")
plt.xlabel("Month")
plt.ylabel("Number of Passengers")
plt.show()
What is a distribution?#
A distribution shows how often values appear in our data.
Some values are common, some are rare.
Plots like histograms help us see the shape and spread of the data.
# Let us start with a histogram
plt.figure(figsize=(8,4))
sns.histplot(data["Passengers"], bins=12, color="skyblue", kde=False)
plt.title("Histogram of Monthly Airline Passengers")
plt.xlabel("Number of Passengers")
plt.ylabel("Count")
plt.show()
What can you see in the histogram?#
Are most months busy? Or do quiet months appear often?
Notice where the bars are tallest or shortest.
# Plotting a KDE (Kernel Density Estimate)
plt.figure(figsize=(8,4))
sns.kdeplot(data["Passengers"], fill=True, color="orange")
plt.title("KDE: Smooth Estimate of Passenger Distribution")
plt.xlabel("Number of Passengers")
plt.ylabel("Density")
plt.show()
KDE plots help us see patterns#
A KDE can show multiple bumps or peaks.
These peaks suggest values that happen more often.
Use KDEs for exploring patterns in your data.
# Let us see both histogram and KDE together
plt.figure(figsize=(8,4))
sns.histplot(data["Passengers"], bins=12, color="skyblue", kde=True)
plt.title("Histogram and KDE of Passenger Numbers")
plt.xlabel("Number of Passengers")
plt.ylabel("Count / Density")
plt.show()
# Box plots are useful for spotting outliers
plt.figure(figsize=(6,3))
sns.boxplot(y=data["Passengers"], color="lightgreen")
plt.title("Box Plot of Monthly Airline Passengers")
plt.ylabel("Number of Passengers")
plt.show()
Understanding box plot parts#
The box shows the middle half of the data.
The line inside is the median: the value in the center.
Dots beyond the lines are called outliers.
Box plots quickly show differences between groups.
# Violin plots mix KDE and box ideas
plt.figure(figsize=(6,3))
sns.violinplot(y=data["Passengers"], color="violet")
plt.title("Violin Plot of Monthly Airline Passengers")
plt.ylabel("Number of Passengers")
plt.show()
# Compare box and violin plot side by side
fig, axes = plt.subplots(1, 2, figsize=(10,3))
sns.boxplot(y=data["Passengers"], color="lightgreen", ax=axes[0])
axes[0].set_title("Box Plot")
sns.violinplot(y=data["Passengers"], color="violet", ax=axes[1])
axes[1].set_title("Violin Plot")
plt.tight_layout()
plt.show()
# What if there are missing values?
data_missing = data.copy()
data_missing.loc[::12, "Passengers"] = None # Remove one value per year
print("Missing data count:", data_missing["Passengers"].isnull().sum())
# Plotting with missing values (will skip missing data)
plt.figure(figsize=(8,3))
sns.histplot(data_missing["Passengers"], bins=12, color="coral")
plt.title("Histogram with Missing Values")
plt.xlabel("Number of Passengers")
plt.ylabel("Count")
plt.show()
# You can fill missing values if you want
filled_data = data_missing["Passengers"].fillna(data_missing["Passengers"].mean())
plt.figure(figsize=(8,3))
sns.histplot(filled_data, bins=12, color="navy")
plt.title("Histogram with Filled Missing Values")
plt.xlabel("Number of Passengers")
plt.ylabel("Count")
plt.show()
Challenge: Compare different plot types#
Try making a histogram, KDE, box, and violin plot for another dataset.
Which plot helps you see patterns best? Why?
# Mini-project: Explore temperature data!
temp_url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/daily-min-temperatures.csv"
temps = pd.read_csv(temp_url)
print("Temperature data shape:", temps.shape)
temps.head()
# Make all four distribution plots for temperatures
plt.figure(figsize=(16,3))
plt.subplot(1,4,1)
sns.histplot(temps["Temp"], bins=15, color="skyblue")
plt.title("Histogram")
plt.subplot(1,4,2)
sns.kdeplot(temps["Temp"], fill=True, color="orange")
plt.title("KDE")
plt.subplot(1,4,3)
sns.boxplot(y=temps["Temp"], color="lightgreen")
plt.title("Box")
plt.subplot(1,4,4)
sns.violinplot(y=temps["Temp"], color="violet")
plt.title("Violin")
plt.tight_layout()
plt.show()
# Best practices
# - Check for missing values
# - Use the right plot for your question
# - Combine plots for deeper insight
# - Always add clear labels and titles
# Troubleshooting: What if a plot looks empty or wrong?
try:
sns.histplot([], bins=10)
plt.show()
except Exception as e:
print("Error:", e)
# Extra tip: Plot single value many times
ones_data = [1]*20
plt.figure(figsize=(4,2))
sns.histplot(ones_data, bins=5, color="red")
plt.title("When All Values Are the Same")
plt.show()
# Ready to practice? Make your own customized plot.
user_vals = input("Enter some numbers (comma-separated): ")
my_vals = [float(x.strip()) for x in user_vals.split(",") if x.strip()]
plt.figure(figsize=(5,2))
sns.histplot(my_vals, bins=min(len(my_vals), 10), color="teal")
plt.title("Your Own Data Histogram")
plt.show()
Recap: What did we learn?#
We learned how to see distributions with histograms, KDEs, box, and violin plots.
These tools help us discover patterns, spot outliers, and share our story with data.
Remember: experiment with real datasets and try each plot for yourself!
Thanks for joining! Please like, comment, and subscribe for more beginner Python lessons.#
Let us keep exploring the world of data together!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



