Mathew K Analytics

Lesson 44 · Data Science Projects

Anomaly Detection in Time Series Data: Techniques and Practical Applications

Welcome to this practical introduction. Today we will learn about finding unusual patterns in time series. Detecting anomalies helps in preventing failures,…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Anomaly Detection in Time Series Data#

  • Welcome to this practical introduction.
  • Today we will learn about finding unusual patterns in time series.
  • Detecting anomalies helps in preventing failures, spotting fraud, and much more.
  • Time series are just data points measured over time, like temperatures.
  • Anomaly detection means finding things that do not fit normal behavior.
  • Let us get started and see why this is important!

What is a Time Series?#

  • Time series data are measurements collected in a sequence over time.
  • Some examples are daily stock prices, temperature readings, or website visits.
  • Seeing trends and spotting surprises helps us understand and predict events.
  • An anomaly is something that stands outlike a big spike or an unexpected drop.
  • In real life, anomaly detection is used in industry, health, banking, and more.
# Always start by suppressing warnings
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# Data setup: Load a real world time series dataset (ambient temperature with system failure events)
import pandas as pd
url = 'https://raw.githubusercontent.com/numenta/NAB/master/data/realKnownCause/ambient_temperature_system_failure.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(7267, 2)
             timestamp      value
0  2013-07-04 00:00:00  69.880835
1  2013-07-04 01:00:00  71.220227
2  2013-07-04 02:00:00  70.877805
# Convert time column to datetime and sort data
df['timestamp'] = pd.to_datetime(df['timestamp'])
df = df.sort_values('timestamp').reset_index(drop=True)
print(df.dtypes[:2])
print(df.head(3))
timestamp    datetime64[ns]
value               float64
dtype: object
            timestamp      value
0 2013-07-04 00:00:00  69.880835
1 2013-07-04 01:00:00  71.220227
2 2013-07-04 02:00:00  70.877805
# Plot the temperature data over time
import matplotlib.pyplot as plt
plt.figure(figsize=(12,4))
plt.plot(df['timestamp'], df['value'], color='skyblue')
plt.title('Ambient Temperature Over Time')
plt.xlabel('Time')
plt.ylabel('Temperature')
plt.show()
No description has been provided for this image
# Describe the temperature data to find basic statistics
stats = df['value'].describe()
print(stats)
count    7267.000000
mean       71.242433
std         4.247509
min        57.458406
25%        68.369411
50%        71.858493
75%        74.430958
max        86.223213
Name: value, dtype: float64
# Check for missing temperature values
missing = df['value'].isna().sum()
print(f"Missing temperature readings: {missing}")
Missing temperature readings: 0
# Remove any missing values just in case
df = df.dropna(subset=['value']).reset_index(drop=True)
print(df.shape)
(7267, 2)

Simple Approaches to Anomaly Detection#

  • The easiest way to find anomalies is with the mean and standard deviation.
  • If a value is far from average, it could be an anomaly.
  • We can mark values that are more than 3 standard deviations from the mean.
  • In real systems, many other methods are possible too.
  • Let us try this simple method first.
# Mark anomalies using the mean and standard deviation rule
mean = df['value'].mean()
std = df['value'].std()
threshold = 3 * std
df['anomaly'] = abs(df['value'] - mean) > threshold
print(df['anomaly'].value_counts())
anomaly
False    7248
True       19
Name: count, dtype: int64
# Plot anomalies on the time series chart
plt.figure(figsize=(12,4))
plt.plot(df['timestamp'], df['value'], label='Temperature', color='lightgray')
plt.scatter(df.loc[df['anomaly'],'timestamp'], df.loc[df['anomaly'],'value'],
            color='red', label='Anomaly', zorder=5)
plt.legend()
plt.title('Temperature with Anomalies Highlighted')
plt.xlabel('Time')
plt.ylabel('Temperature')
plt.show()
No description has been provided for this image

Why Use More Advanced Methods?#

  • Simple threshold rules work but might miss gradual changes or complex patterns.
  • Real systems often use rolling averages, machine learning, or deep neural networks.
  • More advanced methods can catch subtle or contextual anomalies.
  • Let us try a moving average approach, often used in real time systems.
# Detect anomalies using moving average and moving standard deviation
window = 50  # adjust as needed
df['rolling_mean'] = df['value'].rolling(window).mean()
df['rolling_std'] = df['value'].rolling(window).std()
df['anomaly_ma'] = abs(df['value'] - df['rolling_mean']) > 3 * df['rolling_std']
print(df['anomaly_ma'].value_counts())
anomaly_ma
False    7228
True       39
Name: count, dtype: int64
# Visualize anomalies found by the moving average method
plt.figure(figsize=(12,4))
plt.plot(df['timestamp'], df['value'], color='lightgray', label='Temperature')
plt.plot(df['timestamp'], df['rolling_mean'], color='blue', alpha=0.7, label='Moving Average')
plt.scatter(df.loc[df['anomaly_ma'], 'timestamp'], df.loc[df['anomaly_ma'], 'value'],
            color='orange', label='Anomaly (MA)', zorder=5)
plt.legend()
plt.title('Moving Average Anomaly Detection')
plt.xlabel('Time')
plt.ylabel('Temperature')
plt.show()
No description has been provided for this image
# Mini project: Try changing the rolling window size
window_size = int(input("Choose a new window size (like 25, 50, 100): "))
df['roll_mean2'] = df['value'].rolling(window_size).mean()
df['roll_std2'] = df['value'].rolling(window_size).std()
df['anomaly_custom'] = abs(df['value'] - df['roll_mean2']) > 3 * df['roll_std2']
plt.figure(figsize=(12,4))
plt.plot(df['timestamp'], df['value'], color='gray')
plt.plot(df['timestamp'], df['roll_mean2'], color='green', alpha=0.7)
plt.scatter(df.loc[df['anomaly_custom'],'timestamp'], df.loc[df['anomaly_custom'],'value'],
            c='magenta', label='Custom Anomaly', zorder=5)
plt.legend()
plt.title(f'Moving Average Window: {window_size}')
plt.show()
No description has been provided for this image
# Save detected anomalies to a CSV file for review
anomalies = df.loc[df['anomaly_custom'], ['timestamp', 'value']]
anomalies.to_csv('anomalies_detected.csv', index=False)
print(f"Saved {anomalies.shape[0]} anomalies to anomalies_detected.csv")
Saved 39 anomalies to anomalies_detected.csv

Tips for Real World Anomaly Detection#

  • Always visualize your data firstpictures can show things numbers hide.
  • Start simple with averages and gradually try more advanced tools.
  • Adjust your settings (like window size) based on your use case.
  • Be aware of seasonalityrepeating patterns can look like anomalies.
  • Try machine learning methods for complex data if simple tricks fail.
  • Most importantly, talk to a domain expertthey know what matters most!

Recap and Next Steps#

  • You learned to load, clean, and explore real world time series data.
  • You practiced both basic and moving average based anomaly detection.
  • You visualized results and saved findings for reporting.
  • Next, you could apply these steps to your own sensor, finance, or web data!
  • For advanced anomaly detection, look up algorithms like Isolation Forest or LSTM.
  • If you enjoyed this, do not forget to subscribe for more tutorials and hands-on guides!
  • Thank you for learning and practicing with us!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.