Mathew K Analytics

Lesson 26 · Python for Banking and Finance

Understanding Rule-Based Fraud Detection in Banking: A Practical Python Guide

Banks face the constant threat of fraudulent transactions that cost billions every year. Rule-based systems are widely used for detecting suspicious…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Rule-Based Fraud Detection in Banking#

  • Banks face the constant threat of fraudulent transactions that cost billions every year.
  • Rule-based systems are widely used for detecting suspicious patterns in banking transactions.
  • In this lesson, we will use Python to build, test, and understand rule-based fraud detection.
  • You will:
    • Understand transaction and customer data models.
    • Learn to design and implement simple and complex fraud detection rules.
    • Handle errors and edge-cases.
    • Build a mini end-to-end fraud detection engine.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Understanding the data we use for fraud detection#

  • Our main data is a list of banking transactions.
  • Each transaction has unique identifiers, amounts, types, channels, and timestamps.
  • Customers have segments (retail or business) and regions (metro or regional).
  • Common mistakes:
    • Forgetting to match customer IDs across tables.
    • Using unrealistic transaction amounts for rules.
    • Ignoring the channel or time when designing rules.
# Create synthetic banking transactions data
np.random.seed(42)
n_transactions = 1000
n_customers = 200
df = pd.DataFrame({
    'transaction_id': range(1, n_transactions + 1),
    'customer_id': np.random.choice([f'CUST_{i:04d}' for i in range(1, n_customers + 1)], n_transactions),
    'amount': np.round(np.random.normal(150, 60, n_transactions), 2),
    'transaction_type': np.random.choice(['Debit', 'Credit'], n_transactions),
    'channel': np.random.choice(['ATM', 'Online', 'Branch', 'POS'], n_transactions),
    'date': pd.date_range(start='2024-01-01', periods=n_transactions, freq='h')
})

print(df.shape)
print(df.head(3))
(1000, 6)
   transaction_id customer_id  amount transaction_type channel  \
0               1   CUST_0103  238.77           Credit     ATM   
1               2   CUST_0180  269.17           Credit     POS   
2               3   CUST_0093   58.62           Credit  Online   

                 date  
0 2024-01-01 00:00:00  
1 2024-01-01 01:00:00  
2 2024-01-01 02:00:00  
# Create synthetic customers table
customer_ids = [f'CUST_{i:04d}' for i in range(1, 201)]
customers = pd.DataFrame({
    'customer_id': customer_ids,
    'segment': ['Retail'] * 150 + ['Business'] * 50,
    'region': ['Metro'] * 100 + ['Regional'] * 100
})

print(customers.shape)
print(customers.head(3))
(200, 3)
  customer_id segment region
0   CUST_0001  Retail  Metro
1   CUST_0002  Retail  Metro
2   CUST_0003  Retail  Metro
# Basic: Flag all transactions above a simple threshold
rule1_thresh = 500
df['flag_high_amount'] = df['amount'] > rule1_thresh
print(df[['transaction_id', 'amount', 'flag_high_amount']].head(5))
print(f"Total flagged: {df['flag_high_amount'].sum()}")
   transaction_id  amount  flag_high_amount
0               1  238.77             False
1               2  269.17             False
2               3   58.62             False
3               4   81.85             False
4               5  163.56             False
Total flagged: 0
# Basic: Flag transactions done at odd hours (e.g. midnight to 4AM)
df['hour'] = df['date'].dt.hour
df['flag_odd_hour'] = df['hour'].isin([0,1,2,3,4])
print(df[['transaction_id', 'hour', 'flag_odd_hour']].head(5))
print(f"Flagged odd-hour transactions: {df['flag_odd_hour'].sum()}")
   transaction_id  hour  flag_odd_hour
0               1     0           True
1               2     1           True
2               3     2           True
3               4     3           True
4               5     4           True
Flagged odd-hour transactions: 210
# Basic: Flag transactions done on online channel
df['flag_online'] = df['channel'] == 'Online'
print(df[['transaction_id', 'channel', 'flag_online']].head(5))
print(f"Flagged online transactions: {df['flag_online'].sum()}")
   transaction_id channel  flag_online
0               1     ATM        False
1               2     POS        False
2               3  Online         True
3               4  Branch        False
4               5  Online         True
Flagged online transactions: 236
# Intermediate: Flag transactions where both high amount and odd hour rules apply
df['flag_high_amount_odd_hour'] = df['flag_high_amount'] & df['flag_odd_hour']
print(df[['transaction_id', 'amount', 'hour', 'flag_high_amount_odd_hour']].head(5))
print(f"Transactions flagged for both: {df['flag_high_amount_odd_hour'].sum()}")
   transaction_id  amount  hour  flag_high_amount_odd_hour
0               1  238.77     0                      False
1               2  269.17     1                      False
2               3   58.62     2                      False
3               4   81.85     3                      False
4               5  163.56     4                      False
Transactions flagged for both: 0
# Intermediate: Flag for retail customers only if amount is very high
cust_types = customers.set_index('customer_id')['segment']
df['segment'] = df['customer_id'].map(cust_types)
df['flag_retail_very_high'] = (df['segment'] == 'Retail') & (df['amount'] > 800)
print(df[['transaction_id', 'customer_id', 'segment', 'amount', 'flag_retail_very_high']].head(5))
print(f"Number of retail transactions flagged: {df['flag_retail_very_high'].sum()}")
   transaction_id customer_id   segment  amount  flag_retail_very_high
0               1   CUST_0103    Retail  238.77                  False
1               2   CUST_0180  Business  269.17                  False
2               3   CUST_0093    Retail   58.62                  False
3               4   CUST_0015    Retail   81.85                  False
4               5   CUST_0107    Retail  163.56                  False
Number of retail transactions flagged: 0
# Intermediate: Flag if three or more online transactions from the same customer in one hour
df['count_online_same_hour'] = df.groupby(['customer_id', 'date'])['flag_online'].transform('sum')
df['flag_online_burst'] = (df['count_online_same_hour'] >= 3) & df['flag_online']
print(df[df['flag_online_burst']][['customer_id', 'date', 'flag_online_burst']].head(5))
print(f"Number of online bursts flagged: {df['flag_online_burst'].sum()}")
Empty DataFrame
Columns: [customer_id, date, flag_online_burst]
Index: []
Number of online bursts flagged: 0
# Advanced: Flag for channel-region anomaly (ATM in regional area > threshold)
cust_region = customers.set_index('customer_id')['region']
df['region'] = df['customer_id'].map(cust_region)
df['flag_atm_regional_anom'] = (df['channel'] == 'ATM') & (df['region'] == 'Regional') & (df['amount'] > 600)
print(df[df['flag_atm_regional_anom']][['transaction_id', 'customer_id', 'region', 'channel', 'amount']].head())
print(f"Count: {df['flag_atm_regional_anom'].sum()}")
Empty DataFrame
Columns: [transaction_id, customer_id, region, channel, amount]
Index: []
Count: 0
# Advanced: Multiple rules combine - weighted fraud score
df['fraud_score'] = (
    df['flag_high_amount'].astype(int)*2 +
    df['flag_odd_hour'].astype(int) +
    df['flag_online'].astype(int) +
    df['flag_high_amount_odd_hour'].astype(int)*2 +
    df['flag_online_burst'].astype(int)*3 +
    df['flag_atm_regional_anom'].astype(int)*2
)
print(df[['transaction_id', 'fraud_score']].head())
print(df['fraud_score'].value_counts())
   transaction_id  fraud_score
0               1            1
1               2            1
2               3            2
3               4            1
4               5            2
fraud_score
0    606
1    342
2     52
Name: count, dtype: int64
# Advanced: Flag only the highest risk transactions as 'potential fraud' where score is above 5
df['potential_fraud'] = df['fraud_score'] > 5
print(df[df['potential_fraud']][['transaction_id', 'fraud_score']].head())
print(f"Count of potential fraud: {df['potential_fraud'].sum()}")
Empty DataFrame
Columns: [transaction_id, fraud_score]
Index: []
Count of potential fraud: 0
# Error handling: guard against missing values in amount
df_with_missing = df.copy()
df_with_missing.loc[5:8, 'amount'] = np.nan
try:
    df_with_missing['flag_missing_error'] = df_with_missing['amount'] > 500
except Exception as e:
    print(f"Error: {e}")
print(df_with_missing[['amount', 'flag_missing_error']].head(10))
   amount  flag_missing_error
0  238.77               False
1  269.17               False
2   58.62               False
3   81.85               False
4  163.56               False
5     NaN               False
6     NaN               False
7     NaN               False
8     NaN               False
9  138.34               False
# Error handling: fill missing amounts with 0 for flag purposes
df_with_missing['amount_filled'] = df_with_missing['amount'].fillna(0)
df_with_missing['flag_filled'] = df_with_missing['amount_filled'] > 500
print(df_with_missing[['amount', 'amount_filled', 'flag_filled']].head(10))
   amount  amount_filled  flag_filled
0  238.77         238.77        False
1  269.17         269.17        False
2   58.62          58.62        False
3   81.85          81.85        False
4  163.56         163.56        False
5     NaN           0.00        False
6     NaN           0.00        False
7     NaN           0.00        False
8     NaN           0.00        False
9  138.34         138.34        False
# Error handling: record if any critical information is missing
df['missing_info'] = df[['amount', 'channel', 'customer_id']].isnull().any(axis=1)
print(df[['amount', 'channel', 'customer_id', 'missing_info']].head(10))
print(f"Transactions missing info: {df['missing_info'].sum()}")
   amount channel customer_id  missing_info
0  238.77     ATM   CUST_0103         False
1  269.17     POS   CUST_0180         False
2   58.62  Online   CUST_0093         False
3   81.85  Branch   CUST_0015         False
4  163.56  Online   CUST_0107         False
5  200.38  Online   CUST_0072         False
6  149.33  Branch   CUST_0189         False
7   51.72     POS   CUST_0021         False
8  179.79  Branch   CUST_0103         False
9  138.34  Online   CUST_0122         False
Transactions missing info: 0
# Best practice: use clear function for reusable fraud logic
def flag_large_transaction(amount, threshold=500):
    if pd.isnull(amount):
        return False
    return amount > threshold

df['flag_func_large'] = df['amount'].apply(flag_large_transaction)
print(df[['amount', 'flag_func_large']].head(7))
   amount  flag_func_large
0  238.77            False
1  269.17            False
2   58.62            False
3   81.85            False
4  163.56            False
5  200.38            False
6  149.33            False
# Pattern: collect all flagged rules into a summary column
rule_flags = [
    'flag_high_amount', 'flag_odd_hour', 'flag_online',
    'flag_high_amount_odd_hour', 'flag_online_burst',
    'flag_retail_very_high', 'flag_atm_regional_anom',
    'flag_func_large'
]
df['flag_summary'] = df[rule_flags].any(axis=1)
print(df[['transaction_id', 'flag_summary']].head(10))
print(df['flag_summary'].value_counts())
   transaction_id  flag_summary
0               1          True
1               2          True
2               3          True
3               4          True
4               5          True
5               6          True
6               7         False
7               8         False
8               9         False
9              10          True
flag_summary
False    606
True     394
Name: count, dtype: int64
# End-to-end: simulate fraud screening and output flagged cases
flagged = df[df['flag_summary']].copy()
selected_columns = [
    'transaction_id', 'customer_id', 'amount', 'channel', 'hour',
    'region', 'segment', 'fraud_score', 'flag_summary'
]
flagged_summary = flagged[selected_columns]
flagged_summary.to_csv('flagged_transactions.csv', index=False)
print(flagged_summary.head())
print(f"Total flagged cases output: {len(flagged_summary)}")
   transaction_id customer_id  amount channel  hour    region   segment  \
0               1   CUST_0103  238.77     ATM     0  Regional    Retail   
1               2   CUST_0180  269.17     POS     1  Regional  Business   
2               3   CUST_0093   58.62  Online     2     Metro    Retail   
3               4   CUST_0015   81.85  Branch     3     Metro    Retail   
4               5   CUST_0107  163.56  Online     4  Regional    Retail   

   fraud_score  flag_summary  
0            1          True  
1            1          True  
2            2          True  
3            1          True  
4            2          True  
Total flagged cases output: 394
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.