Lesson 36 · Data Mining
Understanding Statistical Outlier Detection: Methods and Practical Applications in Python
Welcome to our beginner's guide to statistical outlier detection in Python! In this lesson, you will see how to spot strange data points using statistics.…
- CourseData Mining
- Lesson36 of 31
- Video19 min
- FormatJupyter notebook · 12 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 10: Statistical Methods for Outlier Detection#
Welcome to our beginner's guide to statistical outlier detection in Python! In this lesson, you will see how to spot strange data points using statistics.
We will start by loading a real credit card fraud dataset, explore what outliers are, and try several detection methods with hands-on code.
Let us get started and remember: outlier detection helps us find fraud, errors, and rare events in all kinds of real data.
What Are Outliers?#
Outliers are data points that look very different from most others. They might be errors or special cases.
Why care? In credit card fraud detection, outliers often point to suspicious transactions.
Today you will learn to find outliers using both simple and advanced statistics!
# Data setup (Credit Card Fraud Dataset)
import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
np.random.seed(42)
url = "https://storage.googleapis.com/download.tensorflow.org/data/creditcard.csv"
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
# Quick Data Overview
print(df.columns)
print(df.describe())
What Should We Check for Outliers?#
Outliers are most interesting in columns with numbers, such as 'Amount', 'Time', or features named 'V1', 'V2', and so on.
We will focus on the 'Amount' column, as it shows the value of each credit card purchase.
# Plot the distribution of the Amount column
import matplotlib.pyplot as plt
plt.figure(figsize=(8,4))
plt.hist(df['Amount'], bins=50, color='teal', edgecolor='black')
plt.title('Transaction Amounts')
plt.xlabel('Amount')
plt.ylabel('Frequency')
plt.show()
Outlier Detection with the Z-Score#
One way to spot outliers is to check how many standard deviations a number is from the average. This is called a Z-score.
If a value's Z-score is above 3, it is often an outlier. Let us code this.
# Z-score calculation for Amount column
amount_mean = df['Amount'].mean()
amount_std = df['Amount'].std()
df['z_score'] = (df['Amount'] - amount_mean) / amount_std
outliers = df[df['z_score'].abs() > 3]
print('Number of outliers (Z-score > 3):', outliers.shape[0])
# Show a few detected outliers
print(outliers[['Amount', 'z_score']].head())
Using the IQR (Interquartile Range) Method#
Another popular approach uses the IQR, which looks at the middle 50 percent of data.
Amounts much higher than this range can be marked as outliers.
# IQR outlier detection for Amount
Q1 = df['Amount'].quantile(0.25)
Q3 = df['Amount'].quantile(0.75)
IQR = Q3 - Q1
lower_bound = Q1 - 1.5 * IQR
upper_bound = Q3 + 1.5 * IQR
iqr_outliers = df[(df['Amount'] < lower_bound) | (df['Amount'] > upper_bound)]
print('Number of outliers (IQR):', iqr_outliers.shape[0])
# Show IQR outliers
print(iqr_outliers[['Amount']].head(5))
Outlier Detection Using Boxplots#
A boxplot is a simple way to visualize outliers. Outliers show up as points far from the 'box' part of the plot.
Let us use one to see Amount outliers quickly.
# Boxplot for Amount
plt.figure(figsize=(8,2))
plt.boxplot(df['Amount'], vert=False, patch_artist=True, boxprops=dict(facecolor='lightblue'))
plt.xlabel('Amount')
plt.title('Boxplot of Transaction Amounts')
plt.show()
Multivariate Outlier Detection: Mahalanobis Distance#
Sometimes outliers do not stand out in a single column, but in combination across several columns. The Mahalanobis distance is a way to measure how far away a row is from the center, taking into account many features at once.
Let us try this on a few columns together.
# Mahalanobis distance calculation on V1, V2, Amount
import numpy as np
from scipy.stats import chi2
cols = ['V1','V2','Amount']
X = df[cols].values
mean_vec = np.mean(X, axis=0)
cov_mat = np.cov(X, rowvar=False)
inv_cov_mat = np.linalg.inv(cov_mat)
diff = X - mean_vec
md = np.sqrt(np.sum(diff.dot(inv_cov_mat) * diff, axis=1))
chi2_threshold = np.sqrt(chi2.ppf(0.997, len(cols)))
multivariate_outliers = df[md > chi2_threshold]
print('Number of Mahalanobis outliers:', multivariate_outliers.shape[0])
# See Mahalanobis outlier examples
print(multivariate_outliers[cols + ['z_score']].head())
Mini-Project: Outlier Removal and Model Training#
For our mini-project, we will build a simple classifier to tell fraud from normal.
First, let us remove the most extreme outliers, then train a logistic regression model.
Let us see if cleaning data first helps our results!
# Remove outliers using Z-score, then split train and test
df_clean = df[df['z_score'].abs() <= 3].copy()
from sklearn.model_selection import train_test_split
X = df_clean.drop(['Class', 'Time', 'z_score'], axis=1)
y = df_clean['Class']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
print('Training rows:', X_train.shape[0], 'Test rows:', X_test.shape[0])
# Train a classifier and score it
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
model = LogisticRegression(max_iter=200, random_state=42)
model.fit(X_train, y_train)
y_pred = model.predict(X_test)
accuracy = accuracy_score(y_test, y_pred)
print(f'Model accuracy after outlier removal: {accuracy:.3f}')
Review and Best Practices for Outlier Detection#
- Outliers can hurt models and hide patterns.
- Always visualize and explore before removing data.
- Try multiple methods: Z-score, IQR, and multivariate distance.
- For big data, consider using automated tools in pandas or sklearn.
Remember, outlier removal is an art, not a strict rule!
Challenge Exercise#
Try finding outliers in a different column or with a different cutoff.
How does changing the threshold affect your results?
Bonus: Try plotting transactions with and without outliers.
Recap#
- You learned what outliers are and why they matter.
- You tried Z-score, IQR, and Mahalanobis distance.
- You saw how outlier removal can make fraud detection models better.
Keep practicing with other datasets and try outlier detection anytime your data looks a little strange!
Thanks for following along! Did you enjoy learning about outlier detection?
Do not forget to subscribe and hit the bell to catch our next lessons on data mining and AI.
What topic should we cover next? Comment below!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



