Mathew K Analytics

Lesson 13 · Python For Time Series

Mastering Resampling and Aggregation of Time Series Data with Pandas in Python

Welcome to a beginner-friendly journey! In this lesson, you will learn how to visualize time series data with pandas in Python. You are going to load,…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Time Series Plotting with Pandas#

Welcome to a beginner-friendly journey! In this lesson, you will learn how to visualize time series data with pandas in Python.

You are going to load, explore, and plot real-world time series data step by step.

What is a Time Series?#

A time series is a sequence of data points collected over time.

Common time series examples include daily temperatures, stock prices, and sales data.

We analyze and plot time series to discover trends, patterns, and changes.

import warnings; warnings.filterwarnings("ignore")
# Import pandas and matplotlib for data handling and plotting
import pandas as pd
import matplotlib.pyplot as plt

# Set up plots to display in the notebook
%matplotlib inline

Data setup#

You will use real retail sales data. The dataset shows monthly shampoo sales over three years.

Let us download and inspect this time series.

# Load shampoo sales time series from a CSV online
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/shampoo.csv"
df = pd.read_csv(url)

# Show data size and first five rows
print("Data shape:", df.shape)
df.head()
Data shape: (36, 2)
Month Sales
0 1-01 266.0
1 1-02 145.9
2 1-03 183.1
3 1-04 119.3
4 1-05 180.3
# Rename columns for easier access
df.columns = ["Month", "Sales"]
# Convert Month to datetime so pandas understands dates
df["Month"] = pd.date_range(start="1901-01", periods=len(df), freq="M")
df["Month"] = pd.to_datetime(df["Month"])

df.head()
Month Sales
0 1901-01-31 266.0
1 1901-02-28 145.9
2 1901-03-31 183.1
3 1901-04-30 119.3
4 1901-05-31 180.3

Your First Time Series Plot#

Let us make a simple line chart to see how sales changed over time.

Seeing the whole story at once is a powerful step!

# Plot shampoo sales over time
plt.figure(figsize=(10,5))
plt.plot(df["Month"], df["Sales"], marker="o")
plt.xlabel("Month")
plt.ylabel("Sales")
plt.title("Monthly Shampoo Sales")
plt.grid(True)
plt.tight_layout()
plt.show()
No description has been provided for this image
# Check for missing sales values
missing = df["Sales"].isnull().sum()
print("Missing sales values:", missing)
Missing sales values: 0
# Fill missing values with the previous month's sales
df["Sales"] = df["Sales"].fillna(method="ffill")
df.head()
Month Sales
0 1901-01-31 266.0
1 1901-02-28 145.9
2 1901-03-31 183.1
3 1901-04-30 119.3
4 1901-05-31 180.3

Highlighting Peaks and Lows#

Sometimes you want to know when sales were highest or lowest.

Let us find and visualize these key moments.

# Find months with highest and lowest sales
max_idx = df["Sales"].idxmax()
min_idx = df["Sales"].idxmin()
max_month = df.loc[max_idx, "Month"]
min_month = df.loc[min_idx, "Month"]

print("Peak sales in:", max_month.strftime("%Y-%m"))
print("Low sales in:", min_month.strftime("%Y-%m"))
Peak sales in: 1903-09
Low sales in: 1901-04
# Add peak and low points to the chart
plt.figure(figsize=(10,5))
plt.plot(df["Month"], df["Sales"], label="Monthly Sales", marker="o")
plt.scatter(df.loc[max_idx, "Month"], df.loc[max_idx, "Sales"], color="red", label="Peak")
plt.scatter(df.loc[min_idx, "Month"], df.loc[min_idx, "Sales"], color="blue", label="Low")
plt.xlabel("Month")
plt.ylabel("Sales")
plt.title("Shampoo Sales with Peaks and Lows")
plt.legend()
plt.grid(True)
plt.tight_layout()
plt.show()
No description has been provided for this image

Making Plots More Readable#

You can change plot styles and add helpful markers to focus on important points.

Let us try making your chart easier to understand!

# Change style and add markers to plot
plt.style.use("seaborn-v0_8-colorblind")
plt.figure(figsize=(10,5))
plt.plot(df["Month"], df["Sales"], marker="D", linestyle="--", color="purple")
plt.xlabel("Month")
plt.ylabel("Sales (units)")
plt.title("Shampoo Sales: Styled Plot")
plt.grid(True)
plt.tight_layout()
plt.show()
No description has been provided for this image
# Zoom in on just the first year of sales
first_year = df[df["Month"].dt.year == 1901]
plt.figure(figsize=(8,4))
plt.plot(first_year["Month"], first_year["Sales"], marker="o", color="green")
plt.title("Shampoo Sales in 1901")
plt.xlabel("Month")
plt.ylabel("Sales")
plt.tight_layout()
plt.show()
No description has been provided for this image
# Ask the user to choose a year and plot it
year = int(input("Enter a year between 1901 and 1903: "))
year_data = df[df["Month"].dt.year == year]
plt.figure(figsize=(8,4))
plt.plot(year_data["Month"], year_data["Sales"], marker="s", color="orange")
plt.title(f"Shampoo Sales in {year}")
plt.xlabel("Month")
plt.ylabel("Sales")
plt.tight_layout()
plt.show()
No description has been provided for this image
 
# Running statistics to see trends
df["Rolling Mean"] = df["Sales"].rolling(window=3).mean()
plt.figure(figsize=(10,5))
plt.plot(df["Month"], df["Sales"], label="Monthly Sales")
plt.plot(df["Month"], df["Rolling Mean"], label="3-Month Average", color="red")
plt.xlabel("Month")
plt.ylabel("Sales")
plt.title("Sales with 3-Month Moving Average")
plt.legend()
plt.grid(True)
plt.tight_layout()
plt.show()
No description has been provided for this image
# Filter months with sales above 400
high_sales = df[df["Sales"] > 400]
print("Months with sales above 400:")
print(high_sales[["Month", "Sales"]])
Months with sales above 400:
        Month  Sales
21 1902-10-31  421.6
25 1903-02-28  440.4
27 1903-04-30  439.3
28 1903-05-31  401.3
29 1903-06-30  437.4
30 1903-07-31  575.5
31 1903-08-31  407.6
32 1903-09-30  682.0
33 1903-10-31  475.3
34 1903-11-30  581.3
35 1903-12-31  646.9
# Sort dataset by sales from lowest to highest
sorted_df = df.sort_values("Sales")
print(sorted_df[["Month", "Sales"]].head())
        Month  Sales
3  1901-04-30  119.3
9  1901-10-31  122.9
1  1901-02-28  145.9
13 1902-02-28  149.5
5  1901-06-30  168.5
# Calculate monthly sales increase
df["Change"] = df["Sales"].diff()
print(df[["Month", "Sales", "Change"]].head(10))
       Month  Sales  Change
0 1901-01-31  266.0     NaN
1 1901-02-28  145.9  -120.1
2 1901-03-31  183.1    37.2
3 1901-04-30  119.3   -63.8
4 1901-05-31  180.3    61.0
5 1901-06-30  168.5   -11.8
6 1901-07-31  231.8    63.3
7 1901-08-31  224.5    -7.3
8 1901-09-30  192.8   -31.7
9 1901-10-31  122.9   -69.9
# Plot all years as separate lines for easy comparison
plt.figure(figsize=(10,6))
for year, group in df.groupby(df["Month"].dt.year):
    plt.plot(group["Month"], group["Sales"], marker="o", label=str(year))
plt.xlabel("Month")
plt.ylabel("Sales")
plt.title("Shampoo Sales by Year")
plt.legend(title="Year")
plt.grid(True)
plt.tight_layout()
plt.show()
No description has been provided for this image

Mini Project: Market Report#

You are the business analyst. The boss asks for a report on shampoo sales:

  • Plot the trend for all years
  • Mark the biggest jump
  • Find the month with the sharpest drop

Let us put your new skills to the test!

# Plot all years and mark biggest jump and sharpest drop
plt.figure(figsize=(12,6))
for year, group in df.groupby(df["Month"].dt.year):
    plt.plot(group["Month"], group["Sales"], marker="o", label=str(year))
jump_idx = df["Change"].idxmax()
drop_idx = df["Change"].idxmin()
plt.scatter(df.loc[jump_idx, "Month"], df.loc[jump_idx, "Sales"],
            color="green", s=100, label="Biggest Jump")
plt.scatter(df.loc[drop_idx, "Month"], df.loc[drop_idx, "Sales"],
            color="red", s=100, label="Sharpest Drop")
plt.xlabel("Month")
plt.ylabel("Sales")
plt.title("Market Report: Sales Jumps and Drops")
plt.legend()
plt.grid(True)
plt.tight_layout()
plt.show()
No description has been provided for this image
# Save your latest plot to a file
plt.figure(figsize=(12,6))
for year, group in df.groupby(df["Month"].dt.year):
    plt.plot(group["Month"], group["Sales"], marker="o", label=str(year))
plt.xlabel("Month")
plt.ylabel("Sales")
plt.title("Shampoo Sales by Year")
plt.legend()
plt.grid(True)
plt.tight_layout()
plt.savefig("shampoo_sales_by_year.png")
plt.close()
print("Plot saved as shampoo_sales_by_year.png")
Plot saved as shampoo_sales_by_year.png
# Basic troubleshooting: Check for bad data types
print(df.dtypes)
Month           datetime64[ns]
Sales                  float64
Rolling Mean           float64
Change                 float64
dtype: object

Final Tips#

  • Try out different chart styles and color themes.
  • Always check for missing or odd values before you plot.
  • Learn by changing markers, colors, or date ranges to make stories clearer.

The more you experiment, the faster you learn!

Recap: Time Series Plotting with Pandas#

You have loaded real data, made clear line charts, and found important sales moments.

Now you know how to explore and tell stories with time series data step by step.

Great job reaching the end keep practicing and share your results!

Next Steps & YouTube Call-to-Action#

Try your new skills on another time series dataset or your own records.

If you enjoyed this, please subscribe and leave a comment with what you want to see next!

See you in the next Python lesson!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.