Mathew K Analytics

Lesson 1 · Full-Length Courses

Data Visualization for Beginners, with Matplotlib

Matplotlib is the foundational Python library for drawing charts: line plots, scatter plots, histograms, bar charts, and more. This lesson runs entirely…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Data Visualization for Beginners, with Matplotlib#

  • Matplotlib is the foundational Python library for drawing charts: line plots, scatter plots, histograms, bar charts, and more.
  • This lesson runs entirely inside a Jupyter Notebook in VS Code, one cell at a time.
  • No prior charting experience is required, only the basic Python and pandas skills from earlier lessons.
  • We will build every chart step by step using a small real-world restaurant tips dataset, ending with a mini dashboard that combines several charts in one figure.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • If matplotlib, pandas, or seaborn are not installed yet, open a terminal in VS Code and run: pip install matplotlib pandas seaborn
  • Place tips.csv in the same folder as your notebook so pandas can find it with just its file name.
import matplotlib.pyplot as plt
import pandas as pd
import seaborn as sns
print('Ready to plot!')
Ready to plot!

Your First Chart: A Bare-Bones Line Plot#

  • Before touching real data, let's draw the simplest possible chart so you see the basic pattern: give matplotlib some numbers, then call plt.show().
plt.plot([1, 2, 3, 4], [10, 20, 15, 30])
plt.show()
No description has been provided for this image

Working With a Real Dataset: tips.csv#

  • tips.csv contains 244 real restaurant bills, and is a classic dataset for learning both pandas and visualization.
  • total_bill: the bill amount in dollars
  • tip: the tip amount in dollars
  • sex: the sex of the person who paid, Male or Female
  • smoker: whether the party included a smoker, Yes or No
  • day: the day of the week, Thur, Fri, Sat, or Sun
  • time: Lunch or Dinner
  • size: the number of people in the party
tips = pd.read_csv('tips.csv')
tips.head()
total_bill tip sex smoker day time size
0 21.34 4.01 Male No Sun Dinner 2
1 13.14 2.63 Male No Sat Dinner 2
2 19.50 2.84 Female No Sun Dinner 2
3 23.59 2.41 Female Yes Sun Dinner 2
4 33.88 4.54 Female No Thur Lunch 6
tips.shape
(244, 7)

Line Plots: Showing a Trend#

  • A line plot connects points in order, and is the natural choice whenever your x-axis has a clear sequence, like days of the week or dates.
  • Unlike our very first bare-bones example, this one uses a real summary computed from our dataset.
avg_tip_by_day = tips.groupby('day', observed=True)['tip'].mean()
plt.plot(avg_tip_by_day.index.astype(str), avg_tip_by_day.values, marker='o')
plt.xlabel('Day')
plt.ylabel('Average Tip ($)')
plt.title('Average Tip by Day')
plt.show()
No description has been provided for this image

Scatter Plots: Comparing Two Numbers#

  • A scatter plot places one dot per row, using one column for the x position and another column for the y position.
  • It is the go-to chart for asking: does one number tend to go up when another number goes up?
plt.scatter(tips['total_bill'], tips['tip'])
plt.show()
No description has been provided for this image
plt.scatter(tips['total_bill'], tips['tip'], edgecolors='k')
plt.xlabel('Total Bill ($)')
plt.ylabel('Tip ($)')
plt.title('Total Bill vs. Tip')
plt.show()
No description has been provided for this image

Customizing With Color and Shape#

colors_by_sex = {'Male': 'blue', 'Female': 'red'}
plt.scatter(tips['total_bill'], tips['tip'], c=tips['sex'].map(colors_by_sex), edgecolors='k')
plt.xlabel('Total Bill ($)')
plt.ylabel('Tip ($)')
plt.title('Total Bill vs. Tip, Colored by Sex')
plt.show()
No description has been provided for this image
markers_by_smoker = {'No': 'o', 'Yes': '^'}
for smoker_status, group in tips.groupby('smoker'):
    plt.scatter(group['total_bill'], group['tip'], marker=markers_by_smoker[smoker_status], label=smoker_status, edgecolors='k')
plt.xlabel('Total Bill ($)')
plt.ylabel('Tip ($)')
plt.title('Total Bill vs. Tip, Shaped by Smoker Status')
plt.legend(title='Smoker')
plt.show()
No description has been provided for this image

Histograms: Seeing How Values Are Distributed#

  • A histogram groups numbers into buckets, called bins, and shows how many values fall into each bucket.
  • It answers a different question than a scatter plot: not 'how do two numbers relate', but 'what does this one number's spread look like'.
plt.hist(tips['total_bill'], bins=20)
plt.xlabel('Total Bill ($)')
plt.ylabel('Frequency')
plt.title('Distribution of Total Bill Amounts')
plt.show()
No description has been provided for this image
plt.hist(tips['total_bill'], bins=20, alpha=0.5, label='Total Bill')
plt.hist(tips['tip'], bins=20, alpha=0.5, label='Tip')
plt.xlabel('Amount ($)')
plt.ylabel('Frequency')
plt.title('Distribution of Total Bill and Tip Amounts')
plt.legend()
plt.show()
No description has been provided for this image

Subplots: Multiple Charts in One Figure#

fig, (ax1, ax2) = plt.subplots(nrows=1, ncols=2, figsize=(12, 4))
ax1.hist(tips['total_bill'], bins=20, color='green')
ax1.set_xlabel('Total Bill ($)')
ax1.set_ylabel('Frequency')
ax1.set_title('Total Bill Amounts')
ax2.hist(tips['tip'], bins=20, color='red')
ax2.set_xlabel('Tip ($)')
ax2.set_ylabel('Frequency')
ax2.set_title('Tip Amounts')
plt.tight_layout()
plt.show()
No description has been provided for this image

Controlling Figure Size and Style#

  • plt.figure(figsize=...) controls how large a chart appears, in inches.
  • plt.style.use lets you switch matplotlib's entire look, colors, gridlines, fonts, and all, with a single line.
plt.figure(figsize=(8, 5))
plt.scatter(tips['total_bill'], tips['tip'], edgecolors='k')
plt.xlabel('Total Bill ($)')
plt.ylabel('Tip ($)')
plt.title('Total Bill vs. Tip (Wider Figure)')
plt.show()
No description has been provided for this image
print(plt.style.available[:8])
['Solarize_Light2', '_classic_test_patch', '_mpl-gallery', '_mpl-gallery-nogrid', 'bmh', 'classic', 'dark_background', 'fast']
plt.style.use('ggplot')
plt.scatter(tips['total_bill'], tips['tip'], edgecolors='k')
plt.xlabel('Total Bill ($)')
plt.ylabel('Tip ($)')
plt.title('Total Bill vs. Tip (ggplot Style)')
plt.show()
No description has been provided for this image
plt.style.use('default')

Box Plots: Comparing Groups at a Glance#

  • A box plot summarizes a number's distribution with five landmarks: the minimum, the 25th percentile, the median, the 75th percentile, and the maximum, plus any outlier points.
  • Building one from scratch in pure matplotlib takes real effort, so here we borrow one ready-made helper from seaborn, which is designed to work hand-in-hand with matplotlib.
sns.boxplot(x='day', y='total_bill', data=tips)
plt.title('Total Bill Distribution by Day')
plt.show()
No description has been provided for this image

Axis Limits, Gridlines, and Annotations#

  • xlim and ylim let you control exactly which range of values is visible on each axis.
  • grid adds light gridlines, making it easier to read values off a chart.
  • annotate lets you point directly at one specific spot on a chart and label it with text.
plt.scatter(tips['total_bill'], tips['tip'], edgecolors='k')
plt.xlabel('Total Bill ($)')
plt.ylabel('Tip ($)')
plt.title('Total Bill vs. Tip, With Fixed Limits and Gridlines')
plt.xlim(0, 60)
plt.ylim(0, 12)
plt.grid(True)
plt.show()
No description has been provided for this image
biggest_tip_row = tips.loc[tips['tip'].idxmax()]
plt.scatter(tips['total_bill'], tips['tip'], edgecolors='k')
plt.annotate('Biggest tip!', xy=(biggest_tip_row['total_bill'], biggest_tip_row['tip']),
             xytext=(10, 10), textcoords='offset points', arrowprops=dict(arrowstyle='->'))
plt.xlabel('Total Bill ($)')
plt.ylabel('Tip ($)')
plt.title('Total Bill vs. Tip, With the Biggest Tip Annotated')
plt.show()
No description has been provided for this image

Bar Charts: Comparing Category Averages#

  • A bar chart compares one summary number, often a mean, across categories.
avg_by_smoker = tips.groupby('smoker')[['total_bill', 'tip']].mean()
avg_by_smoker.index = avg_by_smoker.index.map({'No': 'Non-Smoker', 'Yes': 'Smoker'})
avg_by_smoker.plot(kind='bar')
plt.ylabel('Amount ($)')
plt.title('Average Total Bill and Tip, Smoker vs. Non-Smoker')
plt.xticks(rotation=0)
plt.show()
No description has been provided for this image

Horizontal Bar Charts and Error Bars#

  • Switching a bar chart to horizontal often makes long category names much easier to read.
  • Error bars show how uncertain an average is, using the standard deviation or another spread measure.
avg_bill_by_day = tips.groupby('day', observed=True)['total_bill'].mean()
avg_bill_by_day.plot(kind='barh')
plt.xlabel('Average Total Bill ($)')
plt.ylabel('Day')
plt.title('Average Total Bill by Day (Horizontal)')
plt.show()
No description has been provided for this image
avg_bill = tips.groupby('day', observed=True)['total_bill'].mean()
std_bill = tips.groupby('day', observed=True)['total_bill'].std()
plt.bar(avg_bill.index.astype(str), avg_bill.values, yerr=std_bill.values, capsize=5)
plt.xlabel('Day')
plt.ylabel('Average Total Bill ($)')
plt.title('Average Total Bill by Day, With Error Bars')
plt.show()
No description has been provided for this image

Stacked Bar Charts: Comparing Category Breakdowns#

bills_by_day_sex = tips.groupby(['day', 'sex']).size().unstack()
bills_by_day_sex.plot(kind='bar', stacked=True)
plt.ylabel('Number of Bills')
plt.title('Bills Paid by Sex, Across Days')
plt.show()
No description has been provided for this image
bar_colors = ['#1f77b4', '#ff7f0e']
patterns = ['//', 'xx']
fig, ax = plt.subplots()
for i, (sex_label, pattern) in enumerate(zip(bills_by_day_sex.columns, patterns)):
    ax.bar(bills_by_day_sex.index, bills_by_day_sex[sex_label], bottom=bills_by_day_sex.iloc[:, :i].sum(axis=1), color=bar_colors[i], hatch=pattern, label=sex_label)
ax.set_xlabel('Day')
ax.set_ylabel('Number of Bills')
ax.set_title('Bills Paid by Sex, Across Days (with Patterns)')
ax.legend(title='Sex')
plt.show()
No description has been provided for this image

Pie Charts: Showing Proportions of a Whole#

  • A pie chart shows how a total splits into parts. It works best with only a few categories.
sex_counts = tips['sex'].value_counts()
plt.pie(sex_counts, labels=sex_counts.index, autopct='%.1f%%')
plt.title('Bills Paid by Sex')
plt.show()
No description has been provided for this image

Coloring Points by a Continuous Number#

  • So far, color has represented a category, like sex. Color can also represent a continuous number, like party size, using a colormap and a colorbar.
sc = plt.scatter(tips['total_bill'], tips['tip'], c=tips['size'], cmap='viridis', edgecolors='k')
plt.xlabel('Total Bill ($)')
plt.ylabel('Tip ($)')
plt.title('Total Bill vs. Tip, Colored by Party Size')
cb = plt.colorbar(sc)
cb.set_label('Party Size')
plt.show()
No description has been provided for this image

Bonus: Density Plots for Crowded Data#

  • When a scatter plot has too many overlapping points to read clearly, density-style plots can show where points cluster most heavily.
sns.kdeplot(data=tips, x='total_bill', y='tip', hue='sex')
plt.xlabel('Total Bill ($)')
plt.ylabel('Tip ($)')
plt.title('Density of Total Bill vs. Tip, by Sex')
plt.show()
No description has been provided for this image
hb = plt.hexbin(tips['total_bill'], tips['tip'], gridsize=20, cmap='viridis')
plt.xlabel('Total Bill ($)')
plt.ylabel('Tip ($)')
plt.title('Hexbin Plot of Total Bill vs. Tip')
cb = plt.colorbar(hb)
cb.set_label('Number of Bills')
plt.show()
No description has been provided for this image

Two Y-Axes on One Chart#

  • Sometimes you want to compare two numbers with very different scales, like a count and a dollar amount, on the very same chart.
  • twinx creates a second y-axis that shares the same x-axis, so both can be plotted together without one squashing the other.
counts_by_day = tips.groupby('day', observed=True).size()
avg_tip_by_day = tips.groupby('day', observed=True)['tip'].mean()
fig, ax1 = plt.subplots()
ax1.bar(counts_by_day.index.astype(str), counts_by_day.values, color='skyblue', label='Number of Bills')
ax1.set_xlabel('Day')
ax1.set_ylabel('Number of Bills', color='skyblue')
ax2 = ax1.twinx()
ax2.plot(avg_tip_by_day.index.astype(str), avg_tip_by_day.values, color='darkred', marker='o')
ax2.set_ylabel('Average Tip ($)', color='darkred')
plt.title('Bill Count and Average Tip by Day')
plt.show()
No description has been provided for this image

A Common Beginner Mistake: Mismatched Data Lengths#

  • plt.scatter and plt.plot expect the x-values and y-values to have exactly the same length, one y for every x.
  • The next cell deliberately triggers this error on purpose, so you recognize it immediately if you ever see it in your own code.
short_list = [1, 2, 3]
long_list = [10, 20, 30, 40]
plt.scatter(short_list, long_list)
---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
Cell In[30], line 3
      1 short_list = [1, 2, 3]
      2 long_list = [10, 20, 30, 40]
----> 3 plt.scatter(short_list, long_list)

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\matplotlib\_api\deprecation.py:453, in make_keyword_only.<locals>.wrapper(*args, **kwargs)
    447 if len(args) > name_idx:
    448     warn_deprecated(
    449         since, message="Passing the %(name)s %(obj_type)s "
    450         "positionally is deprecated since Matplotlib %(since)s; the "
    451         "parameter will become keyword-only in %(removal)s.",
    452         name=name, obj_type=f"parameter of {func.__name__}()")
--> 453 return func(*args, **kwargs)

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\matplotlib\pyplot.py:3948, in scatter(x, y, s, c, marker, cmap, norm, vmin, vmax, alpha, linewidths, edgecolors, colorizer, plotnonfinite, data, **kwargs)
   3928 @_copy_docstring_and_deprecators(Axes.scatter)
   3929 def scatter(
   3930     x: float | ArrayLike,
   (...)
   3946     **kwargs,
   3947 ) -> PathCollection:
-> 3948     __ret = gca().scatter(
   3949         x,
   3950         y,
   3951         s=s,
   3952         c=c,
   3953         marker=marker,
   3954         cmap=cmap,
   3955         norm=norm,
   3956         vmin=vmin,
   3957         vmax=vmax,
   3958         alpha=alpha,
   3959         linewidths=linewidths,
   3960         edgecolors=edgecolors,
   3961         colorizer=colorizer,
   3962         plotnonfinite=plotnonfinite,
   3963         **({"data": data} if data is not None else {}),
   3964         **kwargs,
   3965     )
   3966     sci(__ret)
   3967     return __ret

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\matplotlib\_api\deprecation.py:453, in make_keyword_only.<locals>.wrapper(*args, **kwargs)
    447 if len(args) > name_idx:
    448     warn_deprecated(
    449         since, message="Passing the %(name)s %(obj_type)s "
    450         "positionally is deprecated since Matplotlib %(since)s; the "
    451         "parameter will become keyword-only in %(removal)s.",
    452         name=name, obj_type=f"parameter of {func.__name__}()")
--> 453 return func(*args, **kwargs)

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\matplotlib\__init__.py:1524, in _preprocess_data.<locals>.inner(ax, data, *args, **kwargs)
   1521 @functools.wraps(func)
   1522 def inner(ax, *args, data=None, **kwargs):
   1523     if data is None:
-> 1524         return func(
   1525             ax,
   1526             *map(cbook.sanitize_sequence, args),
   1527             **{k: cbook.sanitize_sequence(v) for k, v in kwargs.items()})
   1529     bound = new_sig.bind(ax, *args, **kwargs)
   1530     auto_label = (bound.arguments.get(label_namer)
   1531                   or bound.kwargs.get(label_namer))

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\matplotlib\axes\_axes.py:4936, in Axes.scatter(self, x, y, s, c, marker, cmap, norm, vmin, vmax, alpha, linewidths, edgecolors, colorizer, plotnonfinite, **kwargs)
   4934 y = np.ma.ravel(y)
   4935 if x.size != y.size:
-> 4936     raise ValueError("x and y must be the same size")
   4938 if s is None:
   4939     s = (20 if mpl.rcParams['_internal.classic_mode'] else
   4940          mpl.rcParams['lines.markersize'] ** 2.0)

ValueError: x and y must be the same size
No description has been provided for this image

Catching Errors Gracefully#

  • Sometimes you want to detect a problem and handle it in your code, rather than letting the whole notebook stop.
  • Python's try and except lets you catch a specific error type and respond to it instead of crashing.
try:
    missing = pd.read_csv('does_not_exist.csv')
except FileNotFoundError as e:
    print('FileNotFoundError:', e)
FileNotFoundError: [Errno 2] No such file or directory: 'does_not_exist.csv'

Best Practice: Wrapping a Chart in a Function#

  • If you find yourself copy-pasting the same labels-and-title pattern for every chart, turn it into a function instead.
def labeled_scatter(x, y, xlabel, ylabel, title):
    plt.scatter(x, y, edgecolors='k')
    plt.xlabel(xlabel)
    plt.ylabel(ylabel)
    plt.title(title)
    plt.show()

labeled_scatter(tips['size'], tips['tip'], 'Party Size', 'Tip ($)', 'Party Size vs. Tip')
No description has been provided for this image

End-to-End Mini Project: A Tips Dashboard#

  • Let's combine several chart types into one figure, so a viewer can see multiple angles on the same dataset at once.
fig, axes = plt.subplots(nrows=2, ncols=2, figsize=(12, 9))

axes[0, 0].scatter(tips['total_bill'], tips['tip'], edgecolors='k')
axes[0, 0].set_xlabel('Total Bill ($)')
axes[0, 0].set_ylabel('Tip ($)')
axes[0, 0].set_title('Total Bill vs. Tip')

axes[0, 1].hist(tips['total_bill'], bins=20, color='green')
axes[0, 1].set_xlabel('Total Bill ($)')
axes[0, 1].set_ylabel('Frequency')
axes[0, 1].set_title('Distribution of Total Bill')

avg_by_day = tips.groupby('day', observed=True)['tip'].mean()
axes[1, 0].bar(avg_by_day.index.astype(str), avg_by_day.values, color='orange')
axes[1, 0].set_xlabel('Day')
axes[1, 0].set_ylabel('Average Tip ($)')
axes[1, 0].set_title('Average Tip by Day')

sex_counts = tips['sex'].value_counts()
axes[1, 1].pie(sex_counts, labels=sex_counts.index, autopct='%.1f%%')
axes[1, 1].set_title('Bills Paid by Sex')

plt.tight_layout()
plt.savefig('tips_dashboard.png')
plt.show()
print('Dashboard saved to tips_dashboard.png')
No description has been provided for this image
Dashboard saved to tips_dashboard.png

Wrap-Up: What You Learned#

  • The core matplotlib pattern: describe a chart, label it, then show it.
  • Line plots for trends, and scatter plots for comparing two numbers, using color and marker shape to encode extra categories.
  • Histograms and overlaid histograms for seeing how one number is distributed, plus subplots for placing multiple charts side by side.
  • Figure size and built-in styles for controlling how a chart looks overall.
  • Axis limits, gridlines, and annotations for guiding a viewer's eye to exactly what matters.
  • Bar charts, horizontal bar charts, error bars, stacked bar charts, and pie charts for comparing and breaking down categories.
  • Continuous color mapping with a colorbar, plus density-style charts, kdeplot and hexbin, for crowded, overlapping data.
  • Two y-axes on one chart, using twinx, for comparing numbers on very different scales.
  • Wrapping repeated chart code in a function, and combining several charts into one saved dashboard image.
  • Practice prompt: pick a dataset of your own, even a CSV export from a spreadsheet you already have, and try building a 2 by 2 dashboard like the one in this lesson.
  • If this lesson helped, consider subscribing for more hands-on data tutorials and drop a comment with which chart type you want to see covered next!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.