Lesson 40 · Data Science Projects
Building a Credit Card Fraud Detection System with Machine Learning Techniques
Understand what credit card fraud is Learn why machine learning helps spot fraud See how to use Python for fraud detection Work through a real world data…
- CourseData Science Projects
- Lesson40 of 33
- Video26 min
- FormatJupyter notebook · 16 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbCredit Card Fraud Detection 101#
- Understand what credit card fraud is
- Learn why machine learning helps spot fraud
- See how to use Python for fraud detection
- Work through a real world data science project
- No coding experience needed
Why is fraud detection important?#
- Credit card fraud costs billions each year
- Banks use data mining to catch stolen cards fast
- Early detection protects both customers and banks
- Machine learning can spot hidden patterns in money flows
- You are about to do what real data scientists do
# Always suppress warnings in Jupyter notebooks
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
What you will learn in this notebook#
- How to load a real credit card transaction dataset
- What 'imbalanced data' really means and why it matters
- Simple tricks for cleaning and exploring your data
- Training your first fraud detector with Python
- Measuring accuracy: how do you know if it works?
- Mini challenge: try detecting fraud on new data
- Support each other in the comments! Like and subscribe for more
# Data setup
import pandas as pd
url = 'https://storage.googleapis.com/download.tensorflow.org/data/creditcard.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
# Quickly check basic info about each column
df.info()
# How many fraudulent versus normal transactions do we have?
print(df['Class'].value_counts())
# See some fraudulent cases
df[df['Class'] == 1].head(3)
# Explore column names and the first few more rows
print(df.columns.tolist())
df.head()
# Check for missing values: is anything blank?
print(df.isnull().sum().sum())
Visualizing card transaction amounts#
- Fraud can be easier to spot by visualizing money flows
- Let us plot the amount of each transaction for both fraud and normal
- See if amounts look different for fraud cases
# Simple histogram plot of amounts by class
import seaborn as sns
import matplotlib.pyplot as plt
sns.histplot(df[df['Class'] == 0]['Amount'], bins=50, color='blue', label='Normal', alpha=0.5)
sns.histplot(df[df['Class'] == 1]['Amount'], bins=50, color='red', label='Fraudulent', alpha=0.7)
plt.legend()
plt.title('Transaction Amount By Fraud Status')
plt.xlabel('Amount')
plt.ylabel('Count')
plt.show()
# Preview how data looks over time
df['Hour'] = (df['Time'] // 3600) % 24
sns.countplot(x='Hour', hue='Class', data=df, palette={0:'blue',1:'red'})
plt.title('Fraud Counts By Hour of Day')
plt.ylabel('Transactions')
plt.xlabel('Hour')
plt.legend(['Normal','Fraud'], loc='upper right')
plt.show()
Prepare features for machine learning#
- Time to split data for training and testing our model
- Aim: teach the computer what fraud looks like using labeled data
- Test on data it has never seen before
# Split data into inputs (X) and label (y)
X = df.drop(['Class'], axis=1)
y = df['Class']
# Split into train and test sets
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
print('Training set:', X_train.shape, y_train.shape)
print('Test set:', X_test.shape, y_test.shape)
# Let us build a simple model: Logistic Regression
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(max_iter=1000, random_state=42, class_weight='balanced')
model.fit(X_train, y_train)
# Predict on test set and check accuracy
y_pred = model.predict(X_test)
from sklearn.metrics import accuracy_score, classification_report, confusion_matrix
print('Accuracy:', accuracy_score(y_test, y_pred))
print(classification_report(y_test, y_pred, target_names=['Not Fraud', 'Fraud']))
print(confusion_matrix(y_test, y_pred))
# Visualize confusion matrix as a heatmap
import seaborn as sns
cm = confusion_matrix(y_test, y_pred)
sns.heatmap(cm, annot=True, fmt='d', cmap='Reds', xticklabels=['Not Fraud','Fraud'], yticklabels=['Not Fraud','Fraud'])
plt.xlabel('Predicted')
plt.ylabel('Actual')
plt.title('Confusion Matrix')
plt.show()
# Try using a decision tree for comparison
from sklearn.tree import DecisionTreeClassifier
tree = DecisionTreeClassifier(max_depth=5, class_weight='balanced', random_state=42)
tree.fit(X_train, y_train)
y_tree_pred = tree.predict(X_test)
print('Accuracy (Decision Tree):', accuracy_score(y_test, y_tree_pred))
print(classification_report(y_test, y_tree_pred, target_names=['Not Fraud', 'Fraud']))
# See which features matter most for our decision tree
importances = tree.feature_importances_
features = X.columns
feat_importances = pd.Series(importances, index=features)
feat_importances.nlargest(10).plot(kind='barh')
plt.xlabel('Importance Score')
plt.title('Top 10 Feature Importances for Fraud Detection')
plt.show()
YouTube Mini Project: Your Turn!#
- Try changing the max_depth for the tree above
- What happens to fraud detection accuracy?
- Can you balance between catching more fraud and causing less false alarms?
- Share your results and ideas in the comments
- Like and subscribe for more real-world data science tutorials
# Optional: try your own test amount and get the prediction
test_values = X_test.iloc[0].values.reshape(1, -1)
prediction = model.predict(test_values)[0]
if prediction == 1:
result = "Fraud!"
else:
result = "Not Fraud."
print("Prediction:", result)
What you achieved today!#
- Loaded real world credit card data in Python
- Learned why fraud detection is hard and important
- Explored, visualized, and split the data safely
- Trained both a logistic regression and a decision tree model
- Measured their ability to detect rare frauds
- Used feature importances and heatmaps for deep understanding
- You now have practical data mining experience
- Keep practicing, and you will spot trickier patterns and build better tools!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



