Lesson 26 · Python for Banking and Finance
Understanding Rule-Based Fraud Detection in Banking: A Practical Python Guide
Banks face the constant threat of fraudulent transactions that cost billions every year. Rule-based systems are widely used for detecting suspicious…
- CoursePython for Banking and Finance
- Lesson26 of 24
- Video20 min
- FormatJupyter notebook · 19 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbRule-Based Fraud Detection in Banking#
- Banks face the constant threat of fraudulent transactions that cost billions every year.
- Rule-based systems are widely used for detecting suspicious patterns in banking transactions.
- In this lesson, we will use Python to build, test, and understand rule-based fraud detection.
- You will:
- Understand transaction and customer data models.
- Learn to design and implement simple and complex fraud detection rules.
- Handle errors and edge-cases.
- Build a mini end-to-end fraud detection engine.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
Understanding the data we use for fraud detection#
- Our main data is a list of banking transactions.
- Each transaction has unique identifiers, amounts, types, channels, and timestamps.
- Customers have segments (retail or business) and regions (metro or regional).
- Common mistakes:
- Forgetting to match customer IDs across tables.
- Using unrealistic transaction amounts for rules.
- Ignoring the channel or time when designing rules.
# Create synthetic banking transactions data
np.random.seed(42)
n_transactions = 1000
n_customers = 200
df = pd.DataFrame({
'transaction_id': range(1, n_transactions + 1),
'customer_id': np.random.choice([f'CUST_{i:04d}' for i in range(1, n_customers + 1)], n_transactions),
'amount': np.round(np.random.normal(150, 60, n_transactions), 2),
'transaction_type': np.random.choice(['Debit', 'Credit'], n_transactions),
'channel': np.random.choice(['ATM', 'Online', 'Branch', 'POS'], n_transactions),
'date': pd.date_range(start='2024-01-01', periods=n_transactions, freq='h')
})
print(df.shape)
print(df.head(3))
# Create synthetic customers table
customer_ids = [f'CUST_{i:04d}' for i in range(1, 201)]
customers = pd.DataFrame({
'customer_id': customer_ids,
'segment': ['Retail'] * 150 + ['Business'] * 50,
'region': ['Metro'] * 100 + ['Regional'] * 100
})
print(customers.shape)
print(customers.head(3))
# Basic: Flag all transactions above a simple threshold
rule1_thresh = 500
df['flag_high_amount'] = df['amount'] > rule1_thresh
print(df[['transaction_id', 'amount', 'flag_high_amount']].head(5))
print(f"Total flagged: {df['flag_high_amount'].sum()}")
# Basic: Flag transactions done at odd hours (e.g. midnight to 4AM)
df['hour'] = df['date'].dt.hour
df['flag_odd_hour'] = df['hour'].isin([0,1,2,3,4])
print(df[['transaction_id', 'hour', 'flag_odd_hour']].head(5))
print(f"Flagged odd-hour transactions: {df['flag_odd_hour'].sum()}")
# Basic: Flag transactions done on online channel
df['flag_online'] = df['channel'] == 'Online'
print(df[['transaction_id', 'channel', 'flag_online']].head(5))
print(f"Flagged online transactions: {df['flag_online'].sum()}")
# Intermediate: Flag transactions where both high amount and odd hour rules apply
df['flag_high_amount_odd_hour'] = df['flag_high_amount'] & df['flag_odd_hour']
print(df[['transaction_id', 'amount', 'hour', 'flag_high_amount_odd_hour']].head(5))
print(f"Transactions flagged for both: {df['flag_high_amount_odd_hour'].sum()}")
# Intermediate: Flag for retail customers only if amount is very high
cust_types = customers.set_index('customer_id')['segment']
df['segment'] = df['customer_id'].map(cust_types)
df['flag_retail_very_high'] = (df['segment'] == 'Retail') & (df['amount'] > 800)
print(df[['transaction_id', 'customer_id', 'segment', 'amount', 'flag_retail_very_high']].head(5))
print(f"Number of retail transactions flagged: {df['flag_retail_very_high'].sum()}")
# Intermediate: Flag if three or more online transactions from the same customer in one hour
df['count_online_same_hour'] = df.groupby(['customer_id', 'date'])['flag_online'].transform('sum')
df['flag_online_burst'] = (df['count_online_same_hour'] >= 3) & df['flag_online']
print(df[df['flag_online_burst']][['customer_id', 'date', 'flag_online_burst']].head(5))
print(f"Number of online bursts flagged: {df['flag_online_burst'].sum()}")
# Advanced: Flag for channel-region anomaly (ATM in regional area > threshold)
cust_region = customers.set_index('customer_id')['region']
df['region'] = df['customer_id'].map(cust_region)
df['flag_atm_regional_anom'] = (df['channel'] == 'ATM') & (df['region'] == 'Regional') & (df['amount'] > 600)
print(df[df['flag_atm_regional_anom']][['transaction_id', 'customer_id', 'region', 'channel', 'amount']].head())
print(f"Count: {df['flag_atm_regional_anom'].sum()}")
# Advanced: Multiple rules combine - weighted fraud score
df['fraud_score'] = (
df['flag_high_amount'].astype(int)*2 +
df['flag_odd_hour'].astype(int) +
df['flag_online'].astype(int) +
df['flag_high_amount_odd_hour'].astype(int)*2 +
df['flag_online_burst'].astype(int)*3 +
df['flag_atm_regional_anom'].astype(int)*2
)
print(df[['transaction_id', 'fraud_score']].head())
print(df['fraud_score'].value_counts())
# Advanced: Flag only the highest risk transactions as 'potential fraud' where score is above 5
df['potential_fraud'] = df['fraud_score'] > 5
print(df[df['potential_fraud']][['transaction_id', 'fraud_score']].head())
print(f"Count of potential fraud: {df['potential_fraud'].sum()}")
# Error handling: guard against missing values in amount
df_with_missing = df.copy()
df_with_missing.loc[5:8, 'amount'] = np.nan
try:
df_with_missing['flag_missing_error'] = df_with_missing['amount'] > 500
except Exception as e:
print(f"Error: {e}")
print(df_with_missing[['amount', 'flag_missing_error']].head(10))
# Error handling: fill missing amounts with 0 for flag purposes
df_with_missing['amount_filled'] = df_with_missing['amount'].fillna(0)
df_with_missing['flag_filled'] = df_with_missing['amount_filled'] > 500
print(df_with_missing[['amount', 'amount_filled', 'flag_filled']].head(10))
# Error handling: record if any critical information is missing
df['missing_info'] = df[['amount', 'channel', 'customer_id']].isnull().any(axis=1)
print(df[['amount', 'channel', 'customer_id', 'missing_info']].head(10))
print(f"Transactions missing info: {df['missing_info'].sum()}")
# Best practice: use clear function for reusable fraud logic
def flag_large_transaction(amount, threshold=500):
if pd.isnull(amount):
return False
return amount > threshold
df['flag_func_large'] = df['amount'].apply(flag_large_transaction)
print(df[['amount', 'flag_func_large']].head(7))
# Pattern: collect all flagged rules into a summary column
rule_flags = [
'flag_high_amount', 'flag_odd_hour', 'flag_online',
'flag_high_amount_odd_hour', 'flag_online_burst',
'flag_retail_very_high', 'flag_atm_regional_anom',
'flag_func_large'
]
df['flag_summary'] = df[rule_flags].any(axis=1)
print(df[['transaction_id', 'flag_summary']].head(10))
print(df['flag_summary'].value_counts())
# End-to-end: simulate fraud screening and output flagged cases
flagged = df[df['flag_summary']].copy()
selected_columns = [
'transaction_id', 'customer_id', 'amount', 'channel', 'hour',
'region', 'segment', 'fraud_score', 'flag_summary'
]
flagged_summary = flagged[selected_columns]
flagged_summary.to_csv('flagged_transactions.csv', index=False)
print(flagged_summary.head())
print(f"Total flagged cases output: {len(flagged_summary)}")
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



