Mathew K Analytics

Lesson 30 · Python For Machine Learning

Building a Credit Card Fraud Detection System with Python and Machine Learning

Welcome! Today we will explore Python by working through a real-world problem: detecting credit card fraud. You will learn the basics of Python as we gently…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Beginner Python: Credit Card Fraud Detection#

Welcome! Today we will explore Python by working through a real-world problem: detecting credit card fraud.

You will learn the basics of Python as we gently build skills toward working with actual data.

# Let us start with a simple message
print("Hello, future fraud detector!")
Hello, future fraud detector!

What is Fraud Detection?#

Fraud detection means finding suspicious activities, like fake credit card purchases.

We use data and Python to spot these problems early.

# Python ignores comments like this line when running your code.
# Now we will assign numbers to variables.
apple_count = 3
orange_count = 4

print("I have", apple_count, "apples and", orange_count, "oranges.")
I have 3 apples and 4 oranges.
# Let us use a string (text) variable.
owner_name = "Eva"
print(owner_name, "owns these fruits.")
Eva owns these fruits.
# We can ask the user to enter their name.
your_name = input("What is your name? ")
print("Welcome,", your_name, "!")
Welcome, Alex !
 

Lists: Storing Many Values#

Lists help us store many items at once. For example, several numbers, or names.

We will use lists a lot in fraud detection.

# Here is a list of sample card transactions.
transactions = [23.5, 99.9, 120.0, 5.30, 500.00]
print("First transaction:", transactions[0])
print("Last transaction:", transactions[-1])
First transaction: 23.5
Last transaction: 500.0
# Let us loop through the list and print each transaction.
for amount in transactions:
    print("Transaction amount:", amount)
    
Transaction amount: 23.5
Transaction amount: 99.9
Transaction amount: 120.0
Transaction amount: 5.3
Transaction amount: 500.0
# Stay safe: import warning library and filter warnings.
import warnings
warnings.filterwarnings('ignore')
# Data setup: download credit card fraud data sample
import pandas as pd
url = "https://raw.githubusercontent.com/DiptarajSinha/credit-card-fraud-detection/main/sample_creditcard.csv"
df = pd.read_csv(url)
print("Data shape:", df.shape)
print("First five rows:")
print(df.head())
Data shape: (1000, 31)
First five rows:
       Time        V1        V2        V3        V4        V5        V6  \
0 -1.980572 -1.054986 -0.587028  0.149669  1.024162  0.694136 -1.179325   
1 -0.025789  0.713820  1.739140 -0.938167 -0.525705 -0.095533 -0.974531   
2  0.319268  0.904688  0.150941  0.620014  0.417378  0.254456  1.107156   
3  0.553876  1.256750  1.033976 -0.528597  0.019071 -0.032556 -2.151053   
4  1.066870 -0.322938  0.401462  1.426497  1.401516  0.428816  0.170106   

         V7        V8        V9  ...       V21       V22       V23       V24  \
0 -0.371712  0.076376 -0.262207  ...  0.743197 -0.308961 -0.421760  1.097233   
1  0.978560 -0.070095  0.121957  ...  0.031745  0.320404 -2.143686  0.689849   
2  0.156505  2.285149  0.960532  ...  0.808576 -0.267068 -0.663552 -0.423959   
3 -1.217260 -0.167287 -0.172896  ...  0.428772 -0.368768  0.261633  1.992982   
4 -1.518633 -0.630195 -0.280756  ... -0.755859 -0.208787 -1.134253 -1.091640   

        V25       V26       V27       V28    Amount  Class  
0  0.232014  1.044697 -2.388974  0.098485 -1.577409      0  
1 -1.169669 -0.022348 -0.659382 -0.551957  1.693958      0  
2 -1.794692 -0.842164 -1.239584 -0.269169  0.664193      0  
3  0.393495  0.212170 -0.869220 -1.699980 -0.072547      0  
4 -1.591054  0.767514  1.205345 -2.161808  1.838306      0  

[5 rows x 31 columns]
# Check the class balance (fraud or not fraud)
print(df['Class'].value_counts())
Class
0    979
1     21
Name: count, dtype: int64
# Let us rename columns to keep things clear.
df = df.rename(columns={'Class': 'Fraud'})
print(df.columns)
Index(['Time', 'V1', 'V2', 'V3', 'V4', 'V5', 'V6', 'V7', 'V8', 'V9', 'V10',
       'V11', 'V12', 'V13', 'V14', 'V15', 'V16', 'V17', 'V18', 'V19', 'V20',
       'V21', 'V22', 'V23', 'V24', 'V25', 'V26', 'V27', 'V28', 'Amount',
       'Fraud'],
      dtype='object')
# Check for missing data.
print(df.isna().sum())
Time      0
V1        0
V2        0
V3        0
V4        0
V5        0
V6        0
V7        0
V8        0
V9        0
V10       0
V11       0
V12       0
V13       0
V14       0
V15       0
V16       0
V17       0
V18       0
V19       0
V20       0
V21       0
V22       0
V23       0
V24       0
V25       0
V26       0
V27       0
V28       0
Amount    0
Fraud     0
dtype: int64
# Select all fraud cases.
frauds = df[df['Fraud'] == 1]
print("Number of frauds:", len(frauds))
print(frauds[['Amount','Fraud']].head())
Number of frauds: 21
       Amount  Fraud
25  -0.998365      1
70   1.427817      1
138 -1.169861      1
158 -0.374499      1
244  0.094256      1
# Mark any big purchases (over 200) as suspicious.
df['Suspicious'] = df['Amount'] > 200
print(df[['Amount','Suspicious']].head(10))
     Amount  Suspicious
0 -1.577409       False
1  1.693958       False
2  0.664193       False
3 -0.072547       False
4  1.838306       False
5 -0.249944       False
6 -0.453654       False
7  0.921517       False
8  0.409870       False
9 -1.319403       False
# Simple rule: flag transactions over 500 as 'Alert'.
alerts = []
for amt in df['Amount']:
    if amt > 500:
        alerts.append('Alert')
    else:
        alerts.append('Safe')
df['AlertFlag'] = alerts
print(df[['Amount','AlertFlag']].head(10))
     Amount AlertFlag
0 -1.577409      Safe
1  1.693958      Safe
2  0.664193      Safe
3 -0.072547      Safe
4  1.838306      Safe
5 -0.249944      Safe
6 -0.453654      Safe
7  0.921517      Safe
8  0.409870      Safe
9 -1.319403      Safe
# Calculate average transaction amount for fraud and not-fraud.
avg_fraud = df[df['Fraud'] == 1]['Amount'].mean()
avg_normal = df[df['Fraud'] == 0]['Amount'].mean()
print("Average fraud amount:", round(avg_fraud,2))
print("Average normal amount:", round(avg_normal,2))
Average fraud amount: -0.08
Average normal amount: -0.04

Mini-Project: Spotting Fake Transactions#

Let us try building a simple rule to find unusual transactions.

We want to see which are likely to be fake.

# Mini-project Part 1: Ask the user for a custom limit.
limit = input("What is your suspicious amount limit? ")
limit = float(limit)
df['UserAlert'] = df['Amount'] > limit
print(df[['Amount', 'UserAlert']].head(10))
     Amount  UserAlert
0 -1.577409      False
1  1.693958      False
2  0.664193      False
3 -0.072547      False
4  1.838306      False
5 -0.249944      False
6 -0.453654      False
7  0.921517      False
8  0.409870      False
9 -1.319403      False
 
# Mini-project Part 2: Review flagged transactions.
flagged = df[df['UserAlert']]
print("Total flagged by your rule:", len(flagged))
print(flagged[['Amount','UserAlert']].head())
Total flagged by your rule: 0
Empty DataFrame
Columns: [Amount, UserAlert]
Index: []
# Best practices: always check your data before making decisions.
if df.empty:
    print("No data to check!")
else:
    print("Ready to analyze fraud data.")
    
Ready to analyze fraud data.
# Troubleshooting: handle errors gently.
try:
    missing_value = df.loc[10000, 'Amount']
except Exception as e:
    print("Handled error:", e)
    
Handled error: 10000

Recap#

We explored lists, variables, data tables, and simple fraud flags.

These are building blocks for deeper data science and fraud prevention.

Thank you for joining this beginner Python lesson!

If you enjoyed this, like and subscribe for more fun projects.

Now try making up your own suspicious transaction rulesand happy coding!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.