Lesson 36 · Data Science Projects
Online Payment Fraud Detection Using Machine Learning in Python | Data Science Project
This lesson introduces you to detecting online payment fraud with Python. Fraud detection is vital for keeping online transactions safe. You will learn to…
- CourseData Science Projects
- Lesson36 of 33
- Video27 min
- FormatJupyter notebook · 17 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbOnline Payment Fraud Detection using Machine Learning in Python#
- This lesson introduces you to detecting online payment fraud with Python.
- Fraud detection is vital for keeping online transactions safe.
- You will learn to load real data, clean it, explore it, and build a model.
- No prior experience is needed! All steps are explained simply.
- You will use real world tools like pandas and scikit learn.
- Together, we will uncover how machine learning can spot suspicious transactions.
- Let us get started.
What you will learn in this lesson#
- How to load a credit card fraud dataset into Python.
- How to prepare the data for analysis.
- How to explore patterns in fraudulent vs. normal transactions.
- How to train a machine learning model to recognize fraud.
- How to check model accuracy and important warning signs.
- How to make predictions using your own sample values.
- The real world impact and challenges of fraud detection.
- How to keep learning with free resources and the YouTube channel.
# Always suppress warnings for a beginner friendly experience
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
# Data setup
import pandas as pd
url = 'https://storage.googleapis.com/download.tensorflow.org/data/creditcard.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
What is in the credit card fraud dataset?#
- Each row stands for a single transaction.
- The columns are anonymized numerical values (V1, V2, etc.).
- The 'Amount' column shows the amount of the transaction.
- The 'Class' column is the target: 1 means fraud, 0 means normal.
- This is a real world, imbalanced dataset: frauds are very rare.
- You will learn special tricks to work with this kind of data.
# Let us check for missing values
print(df.isnull().sum().sum())
# Take a quick look at class distribution
print(df['Class'].value_counts())
# Plot class distribution
import matplotlib.pyplot as plt
import seaborn as sns
sns.countplot(x='Class', data=df, palette='Set2')
plt.title('Transaction type counts: 0 = normal, 1 = fraud')
plt.show()
# Summary stats for fraudulent and normal transactions
print(df.groupby('Class')['Amount'].describe())
Feature understanding and selection#
- All features except for 'Time', 'Amount', and 'Class' are anonymized because of privacy.
- You can build a strong model using these features.
- Removing the 'Time' column may help, as it is not always useful.
- The 'Amount' column sometimes gives away large suspicious transactions.
# Drop the 'Time' column just to keep things simple
df = df.drop(['Time'], axis=1)
# Scale the 'Amount' column for fair modeling
from sklearn.preprocessing import StandardScaler
df['Amount_Scaled'] = StandardScaler().fit_transform(df[['Amount']])
df = df.drop(['Amount'], axis=1)
# Visualize feature distributions for fraud and normal
sns.kdeplot(df[df['Class']==0]['Amount_Scaled'], label='Normal', fill=True)
sns.kdeplot(df[df['Class']==1]['Amount_Scaled'], label='Fraud', fill=True)
plt.legend()
plt.title('Distribution of scaled amounts')
plt.show()
# Split the dataset into train and test sets (80 percent train, 20 percent test)
from sklearn.model_selection import train_test_split
X = df.drop('Class', axis=1)
y = df['Class']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42, stratify=y)
print(X_train.shape, X_test.shape)
Choosing a machine learning model#
- For fraud detection, you often use algorithms like random forest, logistic regression, or support vector machine.
- Here, random forest is a common beginner friendly choice.
- It works well when you have many features and want to spot tricky rules.
- Random forest uses many small decision trees and averages their answers. This reduces mistakes.
# Train a Random Forest Classifier on training data
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=100, random_state=42, n_jobs=-1)
rf.fit(X_train, y_train)
# Predict labels for your test set
y_pred = rf.predict(X_test)
# Evaluate results: accuracy, recall, and confusion matrix
from sklearn.metrics import classification_report, confusion_matrix
print(classification_report(y_test, y_pred, digits=4))
print(confusion_matrix(y_test, y_pred))
# Visualize confusion matrix for clearer understanding
import numpy as np
cm = confusion_matrix(y_test, y_pred)
sns.heatmap(cm, annot=True, fmt='d', cmap='Blues', xticklabels=['Normal','Fraud'], yticklabels=['Normal','Fraud'])
plt.xlabel('Predicted')
plt.ylabel('Actual')
plt.title('Confusion matrix')
plt.show()
# See the top features that help detect fraud
importances = rf.feature_importances_
indices = np.argsort(importances)[-10:][::-1]
plt.figure(figsize=(8,4))
sns.barplot(x=importances[indices], y=X.columns[indices], orient='h')
plt.title('Top 10 important features in fraud detection')
plt.show()
# Try your model on a custom made up transaction
sample = X_test.sample(1, random_state=1)
prediction = rf.predict(sample)
print('Fraudulent' if prediction[0] == 1 else 'Normal')
# Interactive prediction: enter your own transaction details
inputs = []
for col in X.columns:
val = float(input(f'Enter value for {col}: '))
inputs.append(val)
custom_pred = rf.predict([inputs])
print('Fraudulent' if custom_pred[0]==1 else 'Normal')
Real world challenges and model limitations#
- Fraudsters keep inventing new tricks and patterns.
- A model trained on past data can miss new, unseen fraud methods.
- Transaction data may have errors, noise, or missing values.
- Class imbalance means that even high accuracy can hide missed frauds.
- Human oversight and constant model review are crucial.
What to try next? Expand your learning#
- Try other models such as logistic regression or support vector machine.
- Explore feature engineering like adding rolling averages or time gaps.
- Use advanced techniques for imbalanced data, such as SMOTE or cost sensitive learning.
- Visualize individual transaction paths for suspicious cases.
- Look for new open datasets on fraud or anomaly detection.
- Search YouTube for video follow ups and step by step guides.
- Comment below what topic you want next, and subscribe for more practical data mining lessons.
Lesson recap: you learned to#
- Load and inspect a real payment fraud dataset.
- Clean and prepare data for safe modeling.
- Explore patterns and understand class imbalance.
- Train a random forest fraud detection model.
- Test and visualize how this model performs.
- Make manual and automated predictions.
- Understand real world model limitations.
- Find resources to grow your skills.
- Thank you for learning with us! Do not forget to like and subscribe for new Data Mining in Python videos on YouTube.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



