Mathew K Analytics

Lesson 36 · Data Science Projects

Online Payment Fraud Detection Using Machine Learning in Python | Data Science Project

This lesson introduces you to detecting online payment fraud with Python. Fraud detection is vital for keeping online transactions safe. You will learn to…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Online Payment Fraud Detection using Machine Learning in Python#

  • This lesson introduces you to detecting online payment fraud with Python.
  • Fraud detection is vital for keeping online transactions safe.
  • You will learn to load real data, clean it, explore it, and build a model.
  • No prior experience is needed! All steps are explained simply.
  • You will use real world tools like pandas and scikit learn.
  • Together, we will uncover how machine learning can spot suspicious transactions.
  • Let us get started.

What you will learn in this lesson#

  • How to load a credit card fraud dataset into Python.
  • How to prepare the data for analysis.
  • How to explore patterns in fraudulent vs. normal transactions.
  • How to train a machine learning model to recognize fraud.
  • How to check model accuracy and important warning signs.
  • How to make predictions using your own sample values.
  • The real world impact and challenges of fraud detection.
  • How to keep learning with free resources and the YouTube channel.
# Always suppress warnings for a beginner friendly experience
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
 
# Data setup
import pandas as pd
url = 'https://storage.googleapis.com/download.tensorflow.org/data/creditcard.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(284807, 31)
   Time        V1        V2        V3        V4        V5        V6        V7  \
0   0.0 -1.359807 -0.072781  2.536347  1.378155 -0.338321  0.462388  0.239599   
1   0.0  1.191857  0.266151  0.166480  0.448154  0.060018 -0.082361 -0.078803   
2   1.0 -1.358354 -1.340163  1.773209  0.379780 -0.503198  1.800499  0.791461   

         V8        V9  ...       V21       V22       V23       V24       V25  \
0  0.098698  0.363787  ... -0.018307  0.277838 -0.110474  0.066928  0.128539   
1  0.085102 -0.255425  ... -0.225775 -0.638672  0.101288 -0.339846  0.167170   
2  0.247676 -1.514654  ...  0.247998  0.771679  0.909412 -0.689281 -0.327642   

        V26       V27       V28  Amount  Class  
0 -0.189115  0.133558 -0.021053  149.62      0  
1  0.125895 -0.008983  0.014724    2.69      0  
2 -0.139097 -0.055353 -0.059752  378.66      0  

[3 rows x 31 columns]

What is in the credit card fraud dataset?#

  • Each row stands for a single transaction.
  • The columns are anonymized numerical values (V1, V2, etc.).
  • The 'Amount' column shows the amount of the transaction.
  • The 'Class' column is the target: 1 means fraud, 0 means normal.
  • This is a real world, imbalanced dataset: frauds are very rare.
  • You will learn special tricks to work with this kind of data.
# Let us check for missing values
print(df.isnull().sum().sum())
0
# Take a quick look at class distribution
print(df['Class'].value_counts())
Class
0    284315
1       492
Name: count, dtype: int64
# Plot class distribution
import matplotlib.pyplot as plt
import seaborn as sns
sns.countplot(x='Class', data=df, palette='Set2')
plt.title('Transaction type counts: 0 = normal, 1 = fraud')
plt.show()
No description has been provided for this image
# Summary stats for fraudulent and normal transactions
print(df.groupby('Class')['Amount'].describe())
          count        mean         std  min   25%    50%     75%       max
Class                                                                      
0      284315.0   88.291022  250.105092  0.0  5.65  22.00   77.05  25691.16
1         492.0  122.211321  256.683288  0.0  1.00   9.25  105.89   2125.87

Feature understanding and selection#

  • All features except for 'Time', 'Amount', and 'Class' are anonymized because of privacy.
  • You can build a strong model using these features.
  • Removing the 'Time' column may help, as it is not always useful.
  • The 'Amount' column sometimes gives away large suspicious transactions.
# Drop the 'Time' column just to keep things simple
df = df.drop(['Time'], axis=1)
# Scale the 'Amount' column for fair modeling
from sklearn.preprocessing import StandardScaler
df['Amount_Scaled'] = StandardScaler().fit_transform(df[['Amount']])
df = df.drop(['Amount'], axis=1)
# Visualize feature distributions for fraud and normal
sns.kdeplot(df[df['Class']==0]['Amount_Scaled'], label='Normal', fill=True)
sns.kdeplot(df[df['Class']==1]['Amount_Scaled'], label='Fraud', fill=True)
plt.legend()
plt.title('Distribution of scaled amounts')
plt.show()
No description has been provided for this image
# Split the dataset into train and test sets (80 percent train, 20 percent test)
from sklearn.model_selection import train_test_split
X = df.drop('Class', axis=1)
y = df['Class']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42, stratify=y)
print(X_train.shape, X_test.shape)
(227845, 29) (56962, 29)

Choosing a machine learning model#

  • For fraud detection, you often use algorithms like random forest, logistic regression, or support vector machine.
  • Here, random forest is a common beginner friendly choice.
  • It works well when you have many features and want to spot tricky rules.
  • Random forest uses many small decision trees and averages their answers. This reduces mistakes.
# Train a Random Forest Classifier on training data
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=100, random_state=42, n_jobs=-1)
rf.fit(X_train, y_train)
RandomForestClassifier(n_jobs=-1, random_state=42)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Predict labels for your test set
y_pred = rf.predict(X_test)
# Evaluate results: accuracy, recall, and confusion matrix
from sklearn.metrics import classification_report, confusion_matrix
print(classification_report(y_test, y_pred, digits=4))
print(confusion_matrix(y_test, y_pred))
              precision    recall  f1-score   support

           0     0.9997    0.9999    0.9998     56864
           1     0.9419    0.8265    0.8804        98

    accuracy                         0.9996     56962
   macro avg     0.9708    0.9132    0.9401     56962
weighted avg     0.9996    0.9996    0.9996     56962

[[56859     5]
 [   17    81]]
# Visualize confusion matrix for clearer understanding
import numpy as np
cm = confusion_matrix(y_test, y_pred)
sns.heatmap(cm, annot=True, fmt='d', cmap='Blues', xticklabels=['Normal','Fraud'], yticklabels=['Normal','Fraud'])
plt.xlabel('Predicted')
plt.ylabel('Actual')
plt.title('Confusion matrix')
plt.show()
No description has been provided for this image
# See the top features that help detect fraud
importances = rf.feature_importances_
indices = np.argsort(importances)[-10:][::-1]
plt.figure(figsize=(8,4))
sns.barplot(x=importances[indices], y=X.columns[indices], orient='h')
plt.title('Top 10 important features in fraud detection')
plt.show()
No description has been provided for this image
# Try your model on a custom made up transaction
sample = X_test.sample(1, random_state=1)
prediction = rf.predict(sample)
print('Fraudulent' if prediction[0] == 1 else 'Normal')
Normal
# Interactive prediction: enter your own transaction details
inputs = []
for col in X.columns:
    val = float(input(f'Enter value for {col}: '))
    inputs.append(val)
custom_pred = rf.predict([inputs])
print('Fraudulent' if custom_pred[0]==1 else 'Normal')
Normal

Real world challenges and model limitations#

  • Fraudsters keep inventing new tricks and patterns.
  • A model trained on past data can miss new, unseen fraud methods.
  • Transaction data may have errors, noise, or missing values.
  • Class imbalance means that even high accuracy can hide missed frauds.
  • Human oversight and constant model review are crucial.

What to try next? Expand your learning#

  • Try other models such as logistic regression or support vector machine.
  • Explore feature engineering like adding rolling averages or time gaps.
  • Use advanced techniques for imbalanced data, such as SMOTE or cost sensitive learning.
  • Visualize individual transaction paths for suspicious cases.
  • Look for new open datasets on fraud or anomaly detection.
  • Search YouTube for video follow ups and step by step guides.
  • Comment below what topic you want next, and subscribe for more practical data mining lessons.

Lesson recap: you learned to#

  • Load and inspect a real payment fraud dataset.
  • Clean and prepare data for safe modeling.
  • Explore patterns and understand class imbalance.
  • Train a random forest fraud detection model.
  • Test and visualize how this model performs.
  • Make manual and automated predictions.
  • Understand real world model limitations.
  • Find resources to grow your skills.
  • Thank you for learning with us! Do not forget to like and subscribe for new Data Mining in Python videos on YouTube.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.