Mathew K Analytics

Lesson 15 · Python For Time Series

Foundations of Time Series Analysis: Understanding Trend and Seasonality in Data

Have you ever wondered how we predict weather, sales, or stock prices? In this lesson, we will explore time series data: what it is, why we care, and how to…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Welcome to Time Series in Python!#

Have you ever wondered how we predict weather, sales, or stock prices?

In this lesson, we will explore time series data: what it is, why we care, and how to handle it in Python.

We will work with real-world examples and build your skills step by step.

What is a Time Series?#

A time series is simply a list of data points ordered over time. For example: temperature each day, or sales every month.

Order matters because time flows in one direction!

Why do we care? Time series help us look for trends, patterns, and make forecasts.

import warnings; warnings.filterwarnings("ignore")  # Turn off warnings to keep our notebook clean.

# Let us make sure pandas and matplotlib are available for our data exploration.
import pandas as pd
import matplotlib.pyplot as plt
# Data setup
# We will use a real dataset: Daily minimum temperatures in Melbourne, Australia.
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/daily-min-temperatures.csv"

temps = pd.read_csv(url)
print("Shape of the data:", temps.shape)
temps.head()
Shape of the data: (3650, 2)
Date Temp
0 1981-01-01 20.7
1 1981-01-02 17.9
2 1981-01-03 18.8
3 1981-01-04 14.6
4 1981-01-05 15.8
# Let us look at a simple line plot of the temperature values over time.
plt.figure(figsize=(10,4))
plt.plot(pd.to_datetime(temps['Date']), temps['Temp'])
plt.title('Daily Minimum Temperatures Over Time')
plt.xlabel('Date')
plt.ylabel('Temperature (Celsius)')
plt.show()
No description has been provided for this image

Core Components in a Time Series#

Most time series have patterns such as:

  • Trends (long-term increase or decrease)
  • Seasonality (regular cycles, like monthly or yearly)
  • Noise (random changes, like weather swings)

Spotting these makes it easier to predict what comes next.

# Let us look at basic dataset info: types, missing values, quick summary.
temps.info()

print("\nMissing values per column:")
print(temps.isnull().sum())

print("\nSummary stats:")
print(temps.describe())
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 3650 entries, 0 to 3649
Data columns (total 2 columns):
 #   Column  Non-Null Count  Dtype  
---  ------  --------------  -----  
 0   Date    3650 non-null   object 
 1   Temp    3650 non-null   float64
dtypes: float64(1), object(1)
memory usage: 57.2+ KB

Missing values per column:
Date    0
Temp    0
dtype: int64

Summary stats:
              Temp
count  3650.000000
mean     11.177753
std       4.071837
min       0.000000
25%       8.300000
50%      11.000000
75%      14.000000
max      26.300000
# Convert the 'Date' to a datetime object, so pandas treats it as a timeline.
temps['Date'] = pd.to_datetime(temps['Date'])
temps = temps.set_index('Date')

Safe Access and Common Errors#

When working with time series, it is easy to accidentally try looking up a date that does not exist.

Always double-check your time values before slicing or searching!

Python will show a KeyError if you try to look up a missing value.

# Look up the temperature on January 3, 1990.
date_to_find = "1990-01-03"
if date_to_find in temps.index:
    print("Temperature on", date_to_find, ":", temps.loc[date_to_find, 'Temp'], "Celsius")
else:
    print("Date not found!")
    
Temperature on 1990-01-03 : 15.6 Celsius
# What happens if we look for a date not in the data? Let us try.
try:
    print("Temperature on 2040-01-01:", temps.loc['2040-01-01', 'Temp'])
except KeyError:
    print("Date not found - be careful when picking your date!")
    
Date not found - be careful when picking your date!
# Slice a few days to see a small time window.
sample_days = temps.loc['1990-01-01':'1990-01-05']
print(sample_days)
            Temp
Date            
1990-01-01  14.8
1990-01-02  13.3
1990-01-03  15.6
1990-01-04  14.5
1990-01-05  14.3

Update and Modify Time Series Values#

We can change temperatures, add new days, or update mistakes.

Be careful: changing history changes your results!

Let us try fixing a value.

# Fixing a single bad value. Suppose January 2, 1990's temperature was recorded wrong.
old_value = temps.loc['1990-01-02', 'Temp']
print("Before fix: ", old_value)
temps.loc['1990-01-02', 'Temp'] = 15.0
print("After fix: ", temps.loc['1990-01-02', 'Temp'])
Before fix:  13.3
After fix:  15.0
# Add a new date at the end of the data.
from datetime import timedelta
last_date = temps.index[-1]
next_date = last_date + timedelta(days=1)
temps.loc[next_date] = 10.0
print(f"Added {next_date.date()} with temperature 10.0.")
Added 1991-01-01 with temperature 10.0.
# Remove a day: maybe something was entered by mistake.
temps = temps.drop(next_date)
print(f"Removed {next_date.date()} from the dataset.")
Removed 1991-01-01 from the dataset.
# Calculate average temperature for a month.
monthly_mean = temps['Temp'].resample('M').mean()
print(monthly_mean.head())
Date
1981-01-31    17.712903
1981-02-28    17.678571
1981-03-31    13.500000
1981-04-30    12.356667
1981-05-31     9.490323
Freq: ME, Name: Temp, dtype: float64
# Plot monthly average temperatures for a better visual trend.
plt.figure(figsize=(10,4))
plt.plot(monthly_mean.index, monthly_mean.values, marker='o')
plt.title('Monthly Average Temperatures')
plt.xlabel('Month')
plt.ylabel('Avg Temperature (Celsius)')
plt.show()
No description has been provided for this image
# Filter for hot days: show all days above 20C.
hot_days = temps[temps['Temp'] > 20]
print(hot_days.head())
            Temp
Date            
1981-01-01  20.7
1981-01-09  21.8
1981-01-14  21.5
1981-01-15  25.0
1981-01-16  20.7
# Find coldest temperature and the date it happened.
min_temp = temps['Temp'].min()
coldest_day = temps['Temp'].idxmin()
print("Coldest day:", coldest_day.date(), "with", min_temp, "Celsius")
Coldest day: 1982-06-05 with 0.0 Celsius
# Time series split: training vs. test data for modeling.
total_days = temps.shape[0]
train_size = int(total_days * 0.8)
train = temps.iloc[:train_size]
test = temps.iloc[train_size:]
print("Train shape:", train.shape, ", Test shape:", test.shape)
Train shape: (2920, 1) , Test shape: (730, 1)
# Mini-project part 1: summarize the hottest week in the dataset.
weekly_mean = temps['Temp'].resample('W').mean()
hottest_week = weekly_mean.idxmax()
print("Hottest week starts on:", hottest_week.date())

# Show all days in that week.
hot_week_days = temps.loc[hottest_week : hottest_week + pd.Timedelta(days=6)]
print(hot_week_days)
Hottest week starts on: 1981-01-18
            Temp
Date            
1981-01-18  24.8
1981-01-19  17.7
1981-01-20  15.5
1981-01-21  18.2
1981-01-22  12.1
1981-01-23  14.4
1981-01-24  16.0
# Mini-project part 2: plot the hottest week as a zoomed-in line chart.
plt.figure(figsize=(8,3))
plt.plot(hot_week_days.index, hot_week_days['Temp'], marker='s', color='red')
plt.title('Temperatures During the Hottest Week')
plt.xlabel('Date')
plt.ylabel('Temperature (Celsius)')
plt.grid()
plt.show()
No description has been provided for this image
# Troubleshooting: check for gaps in time (missing dates).
all_days = pd.date_range(start=temps.index.min(), end=temps.index.max(), freq='D')
missing_dates = all_days.difference(temps.index)
if len(missing_dates) > 0:
    print('Missing dates in data:')
    print(missing_dates)
else:
    print('No missing dates  our time series is complete!')
    
Missing dates in data:
DatetimeIndex(['1984-12-31', '1988-12-31'], dtype='datetime64[ns]', freq=None)
# Extra tip: convert to yearly averages for very long-term trends.
yearly_mean = temps['Temp'].resample('A').mean()
plt.figure(figsize=(8,3))
plt.plot(yearly_mean.index.year, yearly_mean.values, '-o', color='green')
plt.title('Yearly Average Temperatures')
plt.xlabel('Year')
plt.ylabel('Avg Temp (Celsius)')
plt.show()
No description has been provided for this image
# Challenge: input a year and print its monthly average temperatures.
year = input("Enter a year (e.g., 1990): ")
year = int(year)
subset = temps.loc[temps.index.year == year]
monthly_avg = subset['Temp'].resample('M').mean()
print(f"Monthly averages for {year}:")
print(monthly_avg)
Monthly averages for 1990:
Date
1990-01-31    15.632258
1990-02-28    15.417857
1990-03-31    14.835484
1990-04-30    13.433333
1990-05-31     9.748387
1990-06-30     7.720000
1990-07-31     8.183871
1990-08-31     7.825806
1990-09-30     9.166667
1990-10-31    11.345161
1990-11-30    12.656667
1990-12-31    14.367742
Freq: ME, Name: Temp, dtype: float64
 

Recap#

You have learned how to load, view, and modify real time series data in Python.

You can now plot, filter, and summarize data across days, months, and years.

Feel free to experiment with different filters and custom charts!

Want More?#

If you found this lesson helpful, please give our video a thumbs up and subscribe for more beginner tips!

Keep exploring new datasets, and remember: practice makes perfect.

See you in the next lesson!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.