Lesson 47 · Mastering Pandas
How to Use Pandas with Matplotlib and Seaborn for Effective Data Visualization
In this lesson, you will learn how to use pandas together with Python's most popular plotting libraries. You will explore real datasets using pandas, create…
- CourseMastering Pandas
- Lesson47 of 44
- Video25 min
- FormatJupyter notebook · 17 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIntegrating Pandas with Matplotlib and Seaborn#
In this lesson, you will learn how to use pandas together with Python's most popular plotting libraries. You will explore real datasets using pandas, create clean visualizations with Matplotlib and Seaborn, and learn some tips for combining them in your workflow.
By the end, you can build more informative charts and get insights quickly!
import warnings
warnings.filterwarnings("ignore")
# We will use the Tips dataset for demonstration
import seaborn as sns
import matplotlib.pyplot as plt
import numpy as np
np.random.seed(42)
df = sns.load_dataset('tips')
print(df.shape)
print(df.head(3))
Why visualize with pandas + Matplotlib/Seaborn?#
Pandas makes data cleaning and selection fast. Matplotlib and Seaborn make it easy to plot and compare. They work together to help you explore, communicate, and share insights.
# Plotting the total bill amounts as a histogram
df['total_bill'].plot(kind='hist', bins=20, edgecolor='black', figsize=(8,4))
plt.title('Distribution of Total Bill Amounts')
plt.xlabel('Total Bill')
plt.ylabel('Frequency')
plt.show()
# Seaborn: a prettier histogram with a kde curve
sns.histplot(df['total_bill'], bins=20, kde=True, color='skyblue', edgecolor='black')
plt.title('Total Bill Distribution (Seaborn)')
plt.xlabel('Total Bill')
plt.ylabel('Count')
plt.show()
Quick comparison: Pandas vs. Seaborn vs. Matplotlib#
- Pandas: Easy for quick, basic charts built right into your DataFrame.
- Matplotlib: The most flexible (and lower-level) for custom figures, axes, and styles.
- Seaborn: Best for statistical plots with nice colors, legends, and built-in statistics.
# Scatter plot: Tip vs. Total Bill
plt.figure(figsize=(8,5))
plt.scatter(df['total_bill'], df['tip'], alpha=0.7)
plt.title('Scatter Plot of Total Bill vs Tip Amount')
plt.xlabel('Total Bill')
plt.ylabel('Tip')
plt.show()
# Seaborn scatter plot with automatic regression line
sns.lmplot(x='total_bill', y='tip', data=df, height=5, aspect=1.5)
plt.title('Seaborn: Tips vs. Total Bill with Trend Line')
plt.xlabel('Total Bill')
plt.ylabel('Tip')
plt.show()
# Plotting by category: Average tip by day
avg_tip_per_day = df.groupby('day')['tip'].mean()
avg_tip_per_day.plot(kind='bar', color='orange', figsize=(6,4))
plt.title('Average Tip by Day of Week')
plt.xlabel('Day of Week')
plt.ylabel('Average Tip')
plt.show()
# Seaborn barplot: Average tip by sex and smoker status
plt.figure(figsize=(6,4))
sns.barplot(x='sex', y='tip', hue='smoker', data=df, ci=None, palette='Set2')
plt.title('Average Tip by Gender and Smoking Status')
plt.xlabel('Gender')
plt.ylabel('Average Tip')
plt.show()
# Heatmap: Correlation between numeric columns
corr = df.corr(numeric_only=True)
sns.heatmap(corr, annot=True, cmap='coolwarm', fmt='.2f')
plt.title('Heatmap of Feature Correlations')
plt.show()
# Pairplot: Visualize all feature relationships at once
sns.pairplot(df, diag_kind='kde', hue='sex', palette='muted', plot_kws={'alpha':0.7})
plt.suptitle('Pairplot of Tips Dataset (Colored by Gender)', y=1.02)
plt.show()
Customizing your plots#
You can easily style, annotate, and tweak your charts using options in Matplotlib and Seaborn. This control helps you make your visualizations clear and beautiful for reports or presentations.
# Changing style and adding annotations
sns.set_style('whitegrid')
plt.figure(figsize=(7,5))
sns.boxplot(x='day', y='total_bill', data=df, palette='pastel')
plt.title('Total Bill by Day of Week')
plt.xlabel('Day of Week')
plt.ylabel('Total Bill')
plt.annotate('Outlier?', xy=(1, 50), xytext=(1.6, 45), arrowprops=dict(arrowstyle='->'))
plt.show()
# Saving your plots to disk
fig, ax = plt.subplots(figsize=(6,4))
sns.histplot(df['tip'], bins=15, ax=ax, color='purple')
ax.set_title('Histogram of Tips')
ax.set_xlabel('Tip Amount')
ax.set_ylabel('Count')
fig.tight_layout()
fig.savefig('tips_histogram.png')
# EXERCISE: Customize a plot yourself
plt.figure(figsize=(8,4))
sns.violinplot(x='sex', y='tip', data=df, palette='husl')
plt.title('Tip Distribution by Gender')
plt.xlabel('Gender')
plt.ylabel('Tip Amount')
plt.show()
MINI-PROJECT PART 1: Tiny EDA on the Tips Dataset#
Let us see if tips are higher during weekends or weekdays. We will make a grouped barplot and check the averages.
# Create a column labeling each entry as 'Weekday' or 'Weekend'
df['day_type'] = df['day'].apply(lambda x: 'Weekend' if x in ['Sat', 'Sun'] else 'Weekday')
# Group by day_type and calculate mean tip
weekday_vs_weekend = df.groupby('day_type')['tip'].mean().reset_index()
sns.barplot(x='day_type', y='tip', data=weekday_vs_weekend, palette='viridis')
plt.title('Average Tip: Weekday vs Weekend')
plt.xlabel('Day Type')
plt.ylabel('Average Tip')
plt.show()
# MINI-PROJECT PART 2: Visualizing tip percentage by party size
df['tip_pct'] = 100 * df['tip'] / df['total_bill']
sns.boxplot(x='size', y='tip_pct', data=df, palette='rocket')
plt.title('Tip Percentage by Party Size')
plt.xlabel('Party Size')
plt.ylabel('Tip % of Total Bill')
plt.ylim(0,50)
plt.show()
# Best practice: Use style context manager for publication-ready charts
with sns.axes_style('darkgrid'):
plt.figure(figsize=(7,4))
sns.countplot(x='day', hue='smoker', data=df, palette='deep')
plt.title('Smoker Count by Day of Week')
plt.xlabel('Day of Week')
plt.ylabel('Number of Smokers')
plt.show()
# Troubleshooting: What if you get a confusing plot or an error?
# Example: Misspelling a column name
try:
sns.boxplot(x='day', y='full_bill', data=df)
except Exception as e:
print('Error:', e)
# CHALLENGE: Plot tip percentage by time and smoker status
plt.figure(figsize=(8,5))
sns.violinplot(x='time', y='tip_pct', hue='smoker', split=True, data=df, palette='Set3')
plt.title('Tip Percentage by Meal Time and Smoker')
plt.xlabel('Meal Time')
plt.ylabel('Tip % of Total Bill')
plt.ylim(0,50)
plt.show()
Recap: Takeaways#
- Pandas lets you prepare, clean, and filter your data for charting.
- Matplotlib is the backbone that all these visualizations use.
- Seaborn supercharges plotting with color, statistics, and easy grouping.
- Combining them lets you quickly explore and tell data stories.
Experiment, practice, and do not be afraid to try new chart types!
Thank you! Next steps#
Try making plots with your own data, or revisit this lesson to dig deeper. If you enjoyed this, please like and subscribe for more pandas and Python tutorials, and leave a comment with your favorite new chart!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



