Mathew K Analytics

Lesson 35 · Data Mining

Foundations of Anomaly and Outlier Detection in Data Analysis Explained

Welcome to Week 10! This week, we dive into the world of anomaly and outlier detection in data mining. Anomalies are data points that do not fit the…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 10: Introduction to Anomaly and Outlier Detection#

Welcome to Week 10! This week, we dive into the world of anomaly and outlier detection in data mining.

Anomalies are data points that do not fit the expected pattern or behavior.

Detecting anomalies can help you find fraud, rare events, or errors in data.

We will use the Credit Card Fraud dataset for hands-on learning.

# Data setup (Credit Card Fraud Dataset)
import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://storage.googleapis.com/download.tensorflow.org/data/creditcard.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(284807, 31)
   Time        V1        V2        V3        V4        V5        V6        V7  \
0   0.0 -1.359807 -0.072781  2.536347  1.378155 -0.338321  0.462388  0.239599   
1   0.0  1.191857  0.266151  0.166480  0.448154  0.060018 -0.082361 -0.078803   
2   1.0 -1.358354 -1.340163  1.773209  0.379780 -0.503198  1.800499  0.791461   

         V8        V9  ...       V21       V22       V23       V24       V25  \
0  0.098698  0.363787  ... -0.018307  0.277838 -0.110474  0.066928  0.128539   
1  0.085102 -0.255425  ... -0.225775 -0.638672  0.101288 -0.339846  0.167170   
2  0.247676 -1.514654  ...  0.247998  0.771679  0.909412 -0.689281 -0.327642   

        V26       V27       V28  Amount  Class  
0 -0.189115  0.133558 -0.021053  149.62      0  
1  0.125895 -0.008983  0.014724    2.69      0  
2 -0.139097 -0.055353 -0.059752  378.66      0  

[3 rows x 31 columns]

What is an Anomaly?#

An anomaly, or outlier, is any item in your data that is different from most other items.

For example, a credit card purchase of $10,000 when most are below $200 may be an anomaly.

Finding these helps stop fraud or detect data errors.

# Basic summary of the dataset
print(df.info())
print(df.describe().T)
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 284807 entries, 0 to 284806
Data columns (total 31 columns):
 #   Column  Non-Null Count   Dtype  
---  ------  --------------   -----  
 0   Time    284807 non-null  float64
 1   V1      284807 non-null  float64
 2   V2      284807 non-null  float64
 3   V3      284807 non-null  float64
 4   V4      284807 non-null  float64
 5   V5      284807 non-null  float64
 6   V6      284807 non-null  float64
 7   V7      284807 non-null  float64
 8   V8      284807 non-null  float64
 9   V9      284807 non-null  float64
 10  V10     284807 non-null  float64
 11  V11     284807 non-null  float64
 12  V12     284807 non-null  float64
 13  V13     284807 non-null  float64
 14  V14     284807 non-null  float64
 15  V15     284807 non-null  float64
 16  V16     284807 non-null  float64
 17  V17     284807 non-null  float64
 18  V18     284807 non-null  float64
 19  V19     284807 non-null  float64
 20  V20     284807 non-null  float64
 21  V21     284807 non-null  float64
 22  V22     284807 non-null  float64
 23  V23     284807 non-null  float64
 24  V24     284807 non-null  float64
 25  V25     284807 non-null  float64
 26  V26     284807 non-null  float64
 27  V27     284807 non-null  float64
 28  V28     284807 non-null  float64
 29  Amount  284807 non-null  float64
 30  Class   284807 non-null  int64  
dtypes: float64(30), int64(1)
memory usage: 67.4 MB
None
           count          mean           std         min           25%  \
Time    284807.0  9.481386e+04  47488.145955    0.000000  54201.500000   
V1      284807.0  1.175161e-15      1.958696  -56.407510     -0.920373   
V2      284807.0  3.384974e-16      1.651309  -72.715728     -0.598550   
V3      284807.0 -1.379537e-15      1.516255  -48.325589     -0.890365   
V4      284807.0  2.094852e-15      1.415869   -5.683171     -0.848640   
V5      284807.0  1.021879e-15      1.380247 -113.743307     -0.691597   
V6      284807.0  1.494498e-15      1.332271  -26.160506     -0.768296   
V7      284807.0 -5.620335e-16      1.237094  -43.557242     -0.554076   
V8      284807.0  1.149614e-16      1.194353  -73.216718     -0.208630   
V9      284807.0 -2.414189e-15      1.098632  -13.434066     -0.643098   
V10     284807.0  2.238554e-15      1.088850  -24.588262     -0.535426   
V11     284807.0  1.724421e-15      1.020713   -4.797473     -0.762494   
V12     284807.0 -1.245415e-15      0.999201  -18.683715     -0.405571   
V13     284807.0  8.238900e-16      0.995274   -5.791881     -0.648539   
V14     284807.0  1.213481e-15      0.958596  -19.214325     -0.425574   
V15     284807.0  4.866699e-15      0.915316   -4.498945     -0.582884   
V16     284807.0  1.436219e-15      0.876253  -14.129855     -0.468037   
V17     284807.0 -3.768179e-16      0.849337  -25.162799     -0.483748   
V18     284807.0  9.707851e-16      0.838176   -9.498746     -0.498850   
V19     284807.0  1.036249e-15      0.814041   -7.213527     -0.456299   
V20     284807.0  6.418678e-16      0.770925  -54.497720     -0.211721   
V21     284807.0  1.628620e-16      0.734524  -34.830382     -0.228395   
V22     284807.0 -3.576577e-16      0.725702  -10.933144     -0.542350   
V23     284807.0  2.618565e-16      0.624460  -44.807735     -0.161846   
V24     284807.0  4.473914e-15      0.605647   -2.836627     -0.354586   
V25     284807.0  5.109395e-16      0.521278  -10.295397     -0.317145   
V26     284807.0  1.686100e-15      0.482227   -2.604551     -0.326984   
V27     284807.0 -3.661401e-16      0.403632  -22.565679     -0.070840   
V28     284807.0 -1.227452e-16      0.330083  -15.430084     -0.052960   
Amount  284807.0  8.834962e+01    250.120109    0.000000      5.600000   
Class   284807.0  1.727486e-03      0.041527    0.000000      0.000000   

                 50%            75%            max  
Time    84692.000000  139320.500000  172792.000000  
V1          0.018109       1.315642       2.454930  
V2          0.065486       0.803724      22.057729  
V3          0.179846       1.027196       9.382558  
V4         -0.019847       0.743341      16.875344  
V5         -0.054336       0.611926      34.801666  
V6         -0.274187       0.398565      73.301626  
V7          0.040103       0.570436     120.589494  
V8          0.022358       0.327346      20.007208  
V9         -0.051429       0.597139      15.594995  
V10        -0.092917       0.453923      23.745136  
V11        -0.032757       0.739593      12.018913  
V12         0.140033       0.618238       7.848392  
V13        -0.013568       0.662505       7.126883  
V14         0.050601       0.493150      10.526766  
V15         0.048072       0.648821       8.877742  
V16         0.066413       0.523296      17.315112  
V17        -0.065676       0.399675       9.253526  
V18        -0.003636       0.500807       5.041069  
V19         0.003735       0.458949       5.591971  
V20        -0.062481       0.133041      39.420904  
V21        -0.029450       0.186377      27.202839  
V22         0.006782       0.528554      10.503090  
V23        -0.011193       0.147642      22.528412  
V24         0.040976       0.439527       4.584549  
V25         0.016594       0.350716       7.519589  
V26        -0.052139       0.240952       3.517346  
V27         0.001342       0.091045      31.612198  
V28         0.011244       0.078280      33.847808  
Amount     22.000000      77.165000   25691.160000  
Class       0.000000       0.000000       1.000000  
# Checking for missing values
missing = df.isnull().sum().sum()
print(f"Total missing values: {missing}")
Total missing values: 0
# Class balance: Fraud vs. Normal
print(df['Class'].value_counts())
Class
0    284315
1       492
Name: count, dtype: int64
# Visualizing transaction amount by class
import matplotlib.pyplot as plt
df[df['Class'] == 0]['Amount'].hist(bins=50, alpha=0.7, label='Normal')
df[df['Class'] == 1]['Amount'].hist(bins=50, alpha=0.7, label='Fraud')
plt.legend()
plt.xlabel('Transaction Amount')
plt.ylabel('Count')
plt.title('Transaction Amount Distribution by Class')
plt.show()
No description has been provided for this image

Why Do We Need Anomaly Detection?#

Sometimes, problems are hard to spot if we only look at a summary or average.

Anomalies can tell us about system errors, fraud, or rare healthy events.

We use detection methods to find these cases automatically.

# Quick scatter plot: Amount vs. Time, highlighting fraud
fraud = df[df['Class'] == 1]
normal = df[df['Class'] == 0].sample(len(fraud) * 2, random_state=42)
plt.scatter(normal['Time'], normal['Amount'], alpha=0.3, label='Normal', s=10)
plt.scatter(fraud['Time'], fraud['Amount'], alpha=0.6, label='Fraud', c='red', s=15)
plt.xlabel('Time (seconds from first transaction)')
plt.ylabel('Amount')
plt.legend()
plt.title('Fraud vs. Normal Transactions Over Time')
plt.show()
No description has been provided for this image

Common Approaches to Detect Anomalies#

Some methods include:

  • Z-score method (statistical distance)
  • Isolation Forest (tree-based model)
  • Local Outlier Factor (similarity-based)

We will try a few of these next.

# Z-score outlier detection for 'Amount'
mean = df['Amount'].mean()
std = df['Amount'].std()
df['z_score'] = (df['Amount'] - mean) / std
df['z_outlier'] = df['z_score'].abs() > 3
print(df['z_outlier'].value_counts())
z_outlier
False    280731
True       4076
Name: count, dtype: int64
# Using Isolation Forest for anomaly detection
from sklearn.ensemble import IsolationForest
iso = IsolationForest(n_estimators=100, random_state=42)
features = df[['Amount', 'Time']]
preds = iso.fit_predict(features)
df['iso_outlier'] = preds == -1
print(df['iso_outlier'].value_counts())
iso_outlier
False    233131
True      51676
Name: count, dtype: int64
# Local Outlier Factor (LOF) for anomaly detection
from sklearn.neighbors import LocalOutlierFactor
lof = LocalOutlierFactor(n_neighbors=20)
df['lof_outlier'] = lof.fit_predict(features) == -1
print(df['lof_outlier'].value_counts())
lof_outlier
False    280693
True       4114
Name: count, dtype: int64
# Comparing overlap with true frauds
print('Z-score hits:', df[df['z_outlier']]['Class'].value_counts())
print('Isolation Forest hits:', df[df['iso_outlier']]['Class'].value_counts())
print('LOF hits:', df[df['lof_outlier']]['Class'].value_counts())
Z-score hits: Class
0    4065
1      11
Name: count, dtype: int64
Isolation Forest hits: Class
0    51507
1      169
Name: count, dtype: int64
LOF hits: Class
0    4094
1      20
Name: count, dtype: int64

Interpreting Results and Limitations#

No method gets perfect results on its own.

Some anomalies are not fraud, and some fraud might not look odd to a model.

Always check results by exploring and plotting.

# Mini-project: Manual inspection of random flagged outliers
random_outlier = df[df['iso_outlier']].sample(1, random_state=99)
print(random_outlier[['Time', 'Amount', 'Class', 'z_outlier', 'lof_outlier']])
            Time  Amount  Class  z_outlier  lof_outlier
284527  172527.0    6.46      0      False        False

Best Practices for Anomaly Detection#

  • Normalize features before using algorithms
  • Use several models to spot different types of oddities
  • Always combine with expert review when possible
  • Watch for class imbalance, as rare events are easily missed
# Bonus: Using StandardScaler to normalize features
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaled_features = scaler.fit_transform(df[['Amount','Time']])
print('Means after scaling:', scaled_features.mean(axis=0))
print('Stds after scaling:', scaled_features.std(axis=0))
Means after scaling: [-3.67237781e-17 -5.10939521e-17]
Stds after scaling: [1. 1.]
# Challenge: Let us try an input - predict if an Amount looks like an outlier
amt = float(input("Enter a transaction amount: "))
z = (amt - mean) / std
if abs(z) > 3:
    print("This amount may be an outlier!")
else:
    print("This amount looks normal.")
    
This amount may be an outlier!

Recap: Week 10 at a Glance#

  • You learned what anomalies and outliers are in data mining.
  • You practiced with the Credit Card Fraud dataset.
  • You tried several detection methods: z-score, Isolation Forest, Local Outlier Factor.
  • You saw that no single method is perfectalways check and combine results.

Keep experimenting and discovering hidden signals in data!

Thanks for joining Week 10!#

If you enjoyed this session, please like, subscribe, and comment with your favorite outlier story.

See you next week for more data mining fun!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.