Mathew K Analytics

Lesson 36 · Data Mining

Understanding Statistical Outlier Detection: Methods and Practical Applications in Python

Welcome to our beginner's guide to statistical outlier detection in Python! In this lesson, you will see how to spot strange data points using statistics.…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Week 10: Statistical Methods for Outlier Detection#

Welcome to our beginner's guide to statistical outlier detection in Python! In this lesson, you will see how to spot strange data points using statistics.

We will start by loading a real credit card fraud dataset, explore what outliers are, and try several detection methods with hands-on code.

Let us get started and remember: outlier detection helps us find fraud, errors, and rare events in all kinds of real data.

What Are Outliers?#

Outliers are data points that look very different from most others. They might be errors or special cases.

Why care? In credit card fraud detection, outliers often point to suspicious transactions.

Today you will learn to find outliers using both simple and advanced statistics!

# Data setup (Credit Card Fraud Dataset)
import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
np.random.seed(42)
url = "https://storage.googleapis.com/download.tensorflow.org/data/creditcard.csv"
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(284807, 31)
   Time        V1        V2        V3        V4        V5        V6        V7  \
0   0.0 -1.359807 -0.072781  2.536347  1.378155 -0.338321  0.462388  0.239599   
1   0.0  1.191857  0.266151  0.166480  0.448154  0.060018 -0.082361 -0.078803   
2   1.0 -1.358354 -1.340163  1.773209  0.379780 -0.503198  1.800499  0.791461   

         V8        V9  ...       V21       V22       V23       V24       V25  \
0  0.098698  0.363787  ... -0.018307  0.277838 -0.110474  0.066928  0.128539   
1  0.085102 -0.255425  ... -0.225775 -0.638672  0.101288 -0.339846  0.167170   
2  0.247676 -1.514654  ...  0.247998  0.771679  0.909412 -0.689281 -0.327642   

        V26       V27       V28  Amount  Class  
0 -0.189115  0.133558 -0.021053  149.62      0  
1  0.125895 -0.008983  0.014724    2.69      0  
2 -0.139097 -0.055353 -0.059752  378.66      0  

[3 rows x 31 columns]
# Quick Data Overview
print(df.columns)
print(df.describe())
Index(['Time', 'V1', 'V2', 'V3', 'V4', 'V5', 'V6', 'V7', 'V8', 'V9', 'V10',
       'V11', 'V12', 'V13', 'V14', 'V15', 'V16', 'V17', 'V18', 'V19', 'V20',
       'V21', 'V22', 'V23', 'V24', 'V25', 'V26', 'V27', 'V28', 'Amount',
       'Class'],
      dtype='object')
                Time            V1            V2            V3            V4  \
count  284807.000000  2.848070e+05  2.848070e+05  2.848070e+05  2.848070e+05   
mean    94813.859575  1.175161e-15  3.384974e-16 -1.379537e-15  2.094852e-15   
std     47488.145955  1.958696e+00  1.651309e+00  1.516255e+00  1.415869e+00   
min         0.000000 -5.640751e+01 -7.271573e+01 -4.832559e+01 -5.683171e+00   
25%     54201.500000 -9.203734e-01 -5.985499e-01 -8.903648e-01 -8.486401e-01   
50%     84692.000000  1.810880e-02  6.548556e-02  1.798463e-01 -1.984653e-02   
75%    139320.500000  1.315642e+00  8.037239e-01  1.027196e+00  7.433413e-01   
max    172792.000000  2.454930e+00  2.205773e+01  9.382558e+00  1.687534e+01   

                 V5            V6            V7            V8            V9  \
count  2.848070e+05  2.848070e+05  2.848070e+05  2.848070e+05  2.848070e+05   
mean   1.021879e-15  1.494498e-15 -5.620335e-16  1.149614e-16 -2.414189e-15   
std    1.380247e+00  1.332271e+00  1.237094e+00  1.194353e+00  1.098632e+00   
min   -1.137433e+02 -2.616051e+01 -4.355724e+01 -7.321672e+01 -1.343407e+01   
25%   -6.915971e-01 -7.682956e-01 -5.540759e-01 -2.086297e-01 -6.430976e-01   
50%   -5.433583e-02 -2.741871e-01  4.010308e-02  2.235804e-02 -5.142873e-02   
75%    6.119264e-01  3.985649e-01  5.704361e-01  3.273459e-01  5.971390e-01   
max    3.480167e+01  7.330163e+01  1.205895e+02  2.000721e+01  1.559499e+01   

       ...           V21           V22           V23           V24  \
count  ...  2.848070e+05  2.848070e+05  2.848070e+05  2.848070e+05   
mean   ...  1.628620e-16 -3.576577e-16  2.618565e-16  4.473914e-15   
std    ...  7.345240e-01  7.257016e-01  6.244603e-01  6.056471e-01   
min    ... -3.483038e+01 -1.093314e+01 -4.480774e+01 -2.836627e+00   
25%    ... -2.283949e-01 -5.423504e-01 -1.618463e-01 -3.545861e-01   
50%    ... -2.945017e-02  6.781943e-03 -1.119293e-02  4.097606e-02   
75%    ...  1.863772e-01  5.285536e-01  1.476421e-01  4.395266e-01   
max    ...  2.720284e+01  1.050309e+01  2.252841e+01  4.584549e+00   

                V25           V26           V27           V28         Amount  \
count  2.848070e+05  2.848070e+05  2.848070e+05  2.848070e+05  284807.000000   
mean   5.109395e-16  1.686100e-15 -3.661401e-16 -1.227452e-16      88.349619   
std    5.212781e-01  4.822270e-01  4.036325e-01  3.300833e-01     250.120109   
min   -1.029540e+01 -2.604551e+00 -2.256568e+01 -1.543008e+01       0.000000   
25%   -3.171451e-01 -3.269839e-01 -7.083953e-02 -5.295979e-02       5.600000   
50%    1.659350e-02 -5.213911e-02  1.342146e-03  1.124383e-02      22.000000   
75%    3.507156e-01  2.409522e-01  9.104512e-02  7.827995e-02      77.165000   
max    7.519589e+00  3.517346e+00  3.161220e+01  3.384781e+01   25691.160000   

               Class  
count  284807.000000  
mean        0.001727  
std         0.041527  
min         0.000000  
25%         0.000000  
50%         0.000000  
75%         0.000000  
max         1.000000  

[8 rows x 31 columns]

What Should We Check for Outliers?#

Outliers are most interesting in columns with numbers, such as 'Amount', 'Time', or features named 'V1', 'V2', and so on.

We will focus on the 'Amount' column, as it shows the value of each credit card purchase.

# Plot the distribution of the Amount column
import matplotlib.pyplot as plt
plt.figure(figsize=(8,4))
plt.hist(df['Amount'], bins=50, color='teal', edgecolor='black')
plt.title('Transaction Amounts')
plt.xlabel('Amount')
plt.ylabel('Frequency')
plt.show()
No description has been provided for this image

Outlier Detection with the Z-Score#

One way to spot outliers is to check how many standard deviations a number is from the average. This is called a Z-score.

If a value's Z-score is above 3, it is often an outlier. Let us code this.

# Z-score calculation for Amount column
amount_mean = df['Amount'].mean()
amount_std = df['Amount'].std()
df['z_score'] = (df['Amount'] - amount_mean) / amount_std
outliers = df[df['z_score'].abs() > 3]
print('Number of outliers (Z-score > 3):', outliers.shape[0])
Number of outliers (Z-score > 3): 4076
# Show a few detected outliers
print(outliers[['Amount', 'z_score']].head())
      Amount    z_score
51   1402.95   5.255876
89   1142.02   4.212658
140   919.60   3.323405
150   937.69   3.395730
164  3828.04  14.951578

Using the IQR (Interquartile Range) Method#

Another popular approach uses the IQR, which looks at the middle 50 percent of data.

Amounts much higher than this range can be marked as outliers.

# IQR outlier detection for Amount
Q1 = df['Amount'].quantile(0.25)
Q3 = df['Amount'].quantile(0.75)
IQR = Q3 - Q1
lower_bound = Q1 - 1.5 * IQR
upper_bound = Q3 + 1.5 * IQR
iqr_outliers = df[(df['Amount'] < lower_bound) | (df['Amount'] > upper_bound)]
print('Number of outliers (IQR):', iqr_outliers.shape[0])
Number of outliers (IQR): 31904
# Show IQR outliers
print(iqr_outliers[['Amount']].head(5))
     Amount
2    378.66
20   231.71
51  1402.95
64   243.66
85   200.01

Outlier Detection Using Boxplots#

A boxplot is a simple way to visualize outliers. Outliers show up as points far from the 'box' part of the plot.

Let us use one to see Amount outliers quickly.

# Boxplot for Amount
plt.figure(figsize=(8,2))
plt.boxplot(df['Amount'], vert=False, patch_artist=True, boxprops=dict(facecolor='lightblue'))
plt.xlabel('Amount')
plt.title('Boxplot of Transaction Amounts')
plt.show()
No description has been provided for this image

Multivariate Outlier Detection: Mahalanobis Distance#

Sometimes outliers do not stand out in a single column, but in combination across several columns. The Mahalanobis distance is a way to measure how far away a row is from the center, taking into account many features at once.

Let us try this on a few columns together.

# Mahalanobis distance calculation on V1, V2, Amount
import numpy as np
from scipy.stats import chi2
cols = ['V1','V2','Amount']
X = df[cols].values
mean_vec = np.mean(X, axis=0)
cov_mat = np.cov(X, rowvar=False)
inv_cov_mat = np.linalg.inv(cov_mat)
diff = X - mean_vec
md = np.sqrt(np.sum(diff.dot(inv_cov_mat) * diff, axis=1))
chi2_threshold = np.sqrt(chi2.ppf(0.997, len(cols)))
multivariate_outliers = df[md > chi2_threshold]
print('Number of Mahalanobis outliers:', multivariate_outliers.shape[0])
Number of Mahalanobis outliers: 8084
# See Mahalanobis outlier examples
print(multivariate_outliers[cols + ['z_score']].head())
           V1        V2   Amount   z_score
18  -5.401258 -5.450148    46.80 -0.166119
51  -1.004929 -0.985978  1402.95  5.255876
85  -4.575093 -4.429184   200.01  0.446427
89  -0.773293 -4.146007  1142.02  4.212658
140 -5.101877  1.897022   919.60  3.323405

Mini-Project: Outlier Removal and Model Training#

For our mini-project, we will build a simple classifier to tell fraud from normal.

First, let us remove the most extreme outliers, then train a logistic regression model.

Let us see if cleaning data first helps our results!

# Remove outliers using Z-score, then split train and test
df_clean = df[df['z_score'].abs() <= 3].copy()
from sklearn.model_selection import train_test_split
X = df_clean.drop(['Class', 'Time', 'z_score'], axis=1)
y = df_clean['Class']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
print('Training rows:', X_train.shape[0], 'Test rows:', X_test.shape[0])
Training rows: 196511 Test rows: 84220
# Train a classifier and score it
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
model = LogisticRegression(max_iter=200, random_state=42)
model.fit(X_train, y_train)
y_pred = model.predict(X_test)
accuracy = accuracy_score(y_test, y_pred)
print(f'Model accuracy after outlier removal: {accuracy:.3f}')
Model accuracy after outlier removal: 0.999

Review and Best Practices for Outlier Detection#

  • Outliers can hurt models and hide patterns.
  • Always visualize and explore before removing data.
  • Try multiple methods: Z-score, IQR, and multivariate distance.
  • For big data, consider using automated tools in pandas or sklearn.

Remember, outlier removal is an art, not a strict rule!

Challenge Exercise#

Try finding outliers in a different column or with a different cutoff.

How does changing the threshold affect your results?

Bonus: Try plotting transactions with and without outliers.

Recap#

  • You learned what outliers are and why they matter.
  • You tried Z-score, IQR, and Mahalanobis distance.
  • You saw how outlier removal can make fraud detection models better.

Keep practicing with other datasets and try outlier detection anytime your data looks a little strange!

Thanks for following along! Did you enjoy learning about outlier detection?

Do not forget to subscribe and hit the bell to catch our next lessons on data mining and AI.

What topic should we cover next? Comment below!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.