Lesson 44 · Data Science Projects
Anomaly Detection in Time Series Data: Techniques and Practical Applications
Welcome to this practical introduction. Today we will learn about finding unusual patterns in time series. Detecting anomalies helps in preventing failures,…
- CourseData Science Projects
- Lesson44 of 33
- Video21 min
- FormatJupyter notebook · 13 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbAnomaly Detection in Time Series Data#
- Welcome to this practical introduction.
- Today we will learn about finding unusual patterns in time series.
- Detecting anomalies helps in preventing failures, spotting fraud, and much more.
- Time series are just data points measured over time, like temperatures.
- Anomaly detection means finding things that do not fit normal behavior.
- Let us get started and see why this is important!
What is a Time Series?#
- Time series data are measurements collected in a sequence over time.
- Some examples are daily stock prices, temperature readings, or website visits.
- Seeing trends and spotting surprises helps us understand and predict events.
- An anomaly is something that stands outlike a big spike or an unexpected drop.
- In real life, anomaly detection is used in industry, health, banking, and more.
# Always start by suppressing warnings
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# Data setup: Load a real world time series dataset (ambient temperature with system failure events)
import pandas as pd
url = 'https://raw.githubusercontent.com/numenta/NAB/master/data/realKnownCause/ambient_temperature_system_failure.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
# Convert time column to datetime and sort data
df['timestamp'] = pd.to_datetime(df['timestamp'])
df = df.sort_values('timestamp').reset_index(drop=True)
print(df.dtypes[:2])
print(df.head(3))
# Plot the temperature data over time
import matplotlib.pyplot as plt
plt.figure(figsize=(12,4))
plt.plot(df['timestamp'], df['value'], color='skyblue')
plt.title('Ambient Temperature Over Time')
plt.xlabel('Time')
plt.ylabel('Temperature')
plt.show()
# Describe the temperature data to find basic statistics
stats = df['value'].describe()
print(stats)
# Check for missing temperature values
missing = df['value'].isna().sum()
print(f"Missing temperature readings: {missing}")
# Remove any missing values just in case
df = df.dropna(subset=['value']).reset_index(drop=True)
print(df.shape)
Simple Approaches to Anomaly Detection#
- The easiest way to find anomalies is with the mean and standard deviation.
- If a value is far from average, it could be an anomaly.
- We can mark values that are more than 3 standard deviations from the mean.
- In real systems, many other methods are possible too.
- Let us try this simple method first.
# Mark anomalies using the mean and standard deviation rule
mean = df['value'].mean()
std = df['value'].std()
threshold = 3 * std
df['anomaly'] = abs(df['value'] - mean) > threshold
print(df['anomaly'].value_counts())
# Plot anomalies on the time series chart
plt.figure(figsize=(12,4))
plt.plot(df['timestamp'], df['value'], label='Temperature', color='lightgray')
plt.scatter(df.loc[df['anomaly'],'timestamp'], df.loc[df['anomaly'],'value'],
color='red', label='Anomaly', zorder=5)
plt.legend()
plt.title('Temperature with Anomalies Highlighted')
plt.xlabel('Time')
plt.ylabel('Temperature')
plt.show()
Why Use More Advanced Methods?#
- Simple threshold rules work but might miss gradual changes or complex patterns.
- Real systems often use rolling averages, machine learning, or deep neural networks.
- More advanced methods can catch subtle or contextual anomalies.
- Let us try a moving average approach, often used in real time systems.
# Detect anomalies using moving average and moving standard deviation
window = 50 # adjust as needed
df['rolling_mean'] = df['value'].rolling(window).mean()
df['rolling_std'] = df['value'].rolling(window).std()
df['anomaly_ma'] = abs(df['value'] - df['rolling_mean']) > 3 * df['rolling_std']
print(df['anomaly_ma'].value_counts())
# Visualize anomalies found by the moving average method
plt.figure(figsize=(12,4))
plt.plot(df['timestamp'], df['value'], color='lightgray', label='Temperature')
plt.plot(df['timestamp'], df['rolling_mean'], color='blue', alpha=0.7, label='Moving Average')
plt.scatter(df.loc[df['anomaly_ma'], 'timestamp'], df.loc[df['anomaly_ma'], 'value'],
color='orange', label='Anomaly (MA)', zorder=5)
plt.legend()
plt.title('Moving Average Anomaly Detection')
plt.xlabel('Time')
plt.ylabel('Temperature')
plt.show()
# Mini project: Try changing the rolling window size
window_size = int(input("Choose a new window size (like 25, 50, 100): "))
df['roll_mean2'] = df['value'].rolling(window_size).mean()
df['roll_std2'] = df['value'].rolling(window_size).std()
df['anomaly_custom'] = abs(df['value'] - df['roll_mean2']) > 3 * df['roll_std2']
plt.figure(figsize=(12,4))
plt.plot(df['timestamp'], df['value'], color='gray')
plt.plot(df['timestamp'], df['roll_mean2'], color='green', alpha=0.7)
plt.scatter(df.loc[df['anomaly_custom'],'timestamp'], df.loc[df['anomaly_custom'],'value'],
c='magenta', label='Custom Anomaly', zorder=5)
plt.legend()
plt.title(f'Moving Average Window: {window_size}')
plt.show()
# Save detected anomalies to a CSV file for review
anomalies = df.loc[df['anomaly_custom'], ['timestamp', 'value']]
anomalies.to_csv('anomalies_detected.csv', index=False)
print(f"Saved {anomalies.shape[0]} anomalies to anomalies_detected.csv")
Tips for Real World Anomaly Detection#
- Always visualize your data firstpictures can show things numbers hide.
- Start simple with averages and gradually try more advanced tools.
- Adjust your settings (like window size) based on your use case.
- Be aware of seasonalityrepeating patterns can look like anomalies.
- Try machine learning methods for complex data if simple tricks fail.
- Most importantly, talk to a domain expertthey know what matters most!
Recap and Next Steps#
- You learned to load, clean, and explore real world time series data.
- You practiced both basic and moving average based anomaly detection.
- You visualized results and saved findings for reporting.
- Next, you could apply these steps to your own sensor, finance, or web data!
- For advanced anomaly detection, look up algorithms like Isolation Forest or LSTM.
- If you enjoyed this, do not forget to subscribe for more tutorials and hands-on guides!
- Thank you for learning and practicing with us!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



