Lesson 35 · Data Mining
Foundations of Anomaly and Outlier Detection in Data Analysis Explained
Welcome to Week 10! This week, we dive into the world of anomaly and outlier detection in data mining. Anomalies are data points that do not fit the…
- CourseData Mining
- Lesson35 of 31
- Video19 min
- FormatJupyter notebook · 13 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 10: Introduction to Anomaly and Outlier Detection#
Welcome to Week 10! This week, we dive into the world of anomaly and outlier detection in data mining.
Anomalies are data points that do not fit the expected pattern or behavior.
Detecting anomalies can help you find fraud, rare events, or errors in data.
We will use the Credit Card Fraud dataset for hands-on learning.
# Data setup (Credit Card Fraud Dataset)
import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://storage.googleapis.com/download.tensorflow.org/data/creditcard.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
What is an Anomaly?#
An anomaly, or outlier, is any item in your data that is different from most other items.
For example, a credit card purchase of $10,000 when most are below $200 may be an anomaly.
Finding these helps stop fraud or detect data errors.
# Basic summary of the dataset
print(df.info())
print(df.describe().T)
# Checking for missing values
missing = df.isnull().sum().sum()
print(f"Total missing values: {missing}")
# Class balance: Fraud vs. Normal
print(df['Class'].value_counts())
# Visualizing transaction amount by class
import matplotlib.pyplot as plt
df[df['Class'] == 0]['Amount'].hist(bins=50, alpha=0.7, label='Normal')
df[df['Class'] == 1]['Amount'].hist(bins=50, alpha=0.7, label='Fraud')
plt.legend()
plt.xlabel('Transaction Amount')
plt.ylabel('Count')
plt.title('Transaction Amount Distribution by Class')
plt.show()
Why Do We Need Anomaly Detection?#
Sometimes, problems are hard to spot if we only look at a summary or average.
Anomalies can tell us about system errors, fraud, or rare healthy events.
We use detection methods to find these cases automatically.
# Quick scatter plot: Amount vs. Time, highlighting fraud
fraud = df[df['Class'] == 1]
normal = df[df['Class'] == 0].sample(len(fraud) * 2, random_state=42)
plt.scatter(normal['Time'], normal['Amount'], alpha=0.3, label='Normal', s=10)
plt.scatter(fraud['Time'], fraud['Amount'], alpha=0.6, label='Fraud', c='red', s=15)
plt.xlabel('Time (seconds from first transaction)')
plt.ylabel('Amount')
plt.legend()
plt.title('Fraud vs. Normal Transactions Over Time')
plt.show()
Common Approaches to Detect Anomalies#
Some methods include:
- Z-score method (statistical distance)
- Isolation Forest (tree-based model)
- Local Outlier Factor (similarity-based)
We will try a few of these next.
# Z-score outlier detection for 'Amount'
mean = df['Amount'].mean()
std = df['Amount'].std()
df['z_score'] = (df['Amount'] - mean) / std
df['z_outlier'] = df['z_score'].abs() > 3
print(df['z_outlier'].value_counts())
# Using Isolation Forest for anomaly detection
from sklearn.ensemble import IsolationForest
iso = IsolationForest(n_estimators=100, random_state=42)
features = df[['Amount', 'Time']]
preds = iso.fit_predict(features)
df['iso_outlier'] = preds == -1
print(df['iso_outlier'].value_counts())
# Local Outlier Factor (LOF) for anomaly detection
from sklearn.neighbors import LocalOutlierFactor
lof = LocalOutlierFactor(n_neighbors=20)
df['lof_outlier'] = lof.fit_predict(features) == -1
print(df['lof_outlier'].value_counts())
# Comparing overlap with true frauds
print('Z-score hits:', df[df['z_outlier']]['Class'].value_counts())
print('Isolation Forest hits:', df[df['iso_outlier']]['Class'].value_counts())
print('LOF hits:', df[df['lof_outlier']]['Class'].value_counts())
Interpreting Results and Limitations#
No method gets perfect results on its own.
Some anomalies are not fraud, and some fraud might not look odd to a model.
Always check results by exploring and plotting.
# Mini-project: Manual inspection of random flagged outliers
random_outlier = df[df['iso_outlier']].sample(1, random_state=99)
print(random_outlier[['Time', 'Amount', 'Class', 'z_outlier', 'lof_outlier']])
Best Practices for Anomaly Detection#
- Normalize features before using algorithms
- Use several models to spot different types of oddities
- Always combine with expert review when possible
- Watch for class imbalance, as rare events are easily missed
# Bonus: Using StandardScaler to normalize features
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaled_features = scaler.fit_transform(df[['Amount','Time']])
print('Means after scaling:', scaled_features.mean(axis=0))
print('Stds after scaling:', scaled_features.std(axis=0))
# Challenge: Let us try an input - predict if an Amount looks like an outlier
amt = float(input("Enter a transaction amount: "))
z = (amt - mean) / std
if abs(z) > 3:
print("This amount may be an outlier!")
else:
print("This amount looks normal.")
Recap: Week 10 at a Glance#
- You learned what anomalies and outliers are in data mining.
- You practiced with the Credit Card Fraud dataset.
- You tried several detection methods: z-score, Isolation Forest, Local Outlier Factor.
- You saw that no single method is perfectalways check and combine results.
Keep experimenting and discovering hidden signals in data!
Thanks for joining Week 10!#
If you enjoyed this session, please like, subscribe, and comment with your favorite outlier story.
See you next week for more data mining fun!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



