Lesson 20 · Python for Banking and Finance
Fundamentals of Credit Risk Data Analysis in Banking Using Python
In this lesson, we learn how banks use data to make decisions about lending money. You will see how real-world credit risk data is structured. We will build…
- CoursePython for Banking and Finance
- Lesson20 of 24
- Video20 min
- FormatJupyter notebook · 19 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIntroduction to Credit Risk Data#
- In this lesson, we learn how banks use data to make decisions about lending money.
- You will see how real-world credit risk data is structured.
- We will build and examine datasets to understand credit risk.
- Credit risk analysis helps banks avoid business losses.
- You will practice with synthetic and public credit datasets.
- By the end, you will understand how basic data analysis supports smarter banking.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
Core Data Concepts: What is Credit Risk Data?#
- Credit risk data helps banks assess customers who may not repay loans.
- Data may include transactions, account details, and loan information.
- Understanding how to join tables and handle missing data is important.
- Beginners often assume data is always clean or correctly labeled.
- Data can be messy, with missing values, duplicate rows, or unclear columns.
# Let us create a small synthetic banking transactions table to practice.
np.random.seed(42) # Set seed for reproducibility
n_transactions = 1000
n_customers = 200
df = pd.DataFrame({
'transaction_id': range(1, n_transactions + 1),
'customer_id': np.random.choice([f'CUST_{i:04d}' for i in range(1, n_customers + 1)], n_transactions),
'amount': np.round(np.random.normal(150, 60, n_transactions), 2),
'transaction_type': np.random.choice(['Debit', 'Credit'], n_transactions),
'channel': np.random.choice(['ATM', 'Online', 'Branch', 'POS'], n_transactions),
'date': pd.date_range(start='2024-01-01', periods=n_transactions, freq='h')
})
print(df.shape)
print(df.head(3))
# We also need a table for customer details.
customer_ids = [f'CUST_{i:04d}' for i in range(1, 201)]
customers = pd.DataFrame({
'customer_id': customer_ids,
'segment': ['Retail'] * 150 + ['Business'] * 50,
'region': ['Metro'] * 100 + ['Regional'] * 100
})
print(customers.shape)
print(customers.head(3))
# Accounts are tied to both customers and their products.
np.random.seed(42)
accounts = pd.DataFrame({
'account_id': [f'ACC_{i:05d}' for i in range(1, 201)],
'customer_id': customer_ids,
'account_type': np.random.choice(['Savings', 'Cheque', 'Credit'], size=200),
'open_date': pd.date_range(start='2015-01-01', periods=200, freq='30D')
})
print(accounts.shape)
print(accounts.head(3))
# Now let us load a well-known public credit risk dataset.
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/statlog/german/german.data'
cols = [f'feature_{i}' for i in range(1, 25)] + ['credit_status']
german = pd.read_csv(url, sep=' ', names=cols)
german['default_flag'] = (german['credit_status'] == 2).astype(int)
print(german.shape)
print(german[['credit_status', 'default_flag']].head(3))
# Beginner Example 1: Count how many transactions are Debits vs Credits.
print(df['transaction_type'].value_counts())
# Beginner Example 2: Find average transaction amount.
print(df['amount'].mean())
# Beginner Example 3: What regions do customers belong to?
print(customers['region'].value_counts())
# Intermediate Example 1: Join transactions to customer detail by customer_id.
merged = pd.merge(df, customers, on='customer_id', how='left')
print(merged[['transaction_id', 'customer_id', 'segment', 'region']].head(3))
# Intermediate Example 2: Calculate total debit amount by customer.
debits = df[df['transaction_type'] == 'Debit']
totals = debits.groupby('customer_id')['amount'].sum().reset_index()
print(totals.head(3))
# Intermediate Example 3: Look for customers who have both debit and credit transactions.
has_debit = df[df['transaction_type']=='Debit']['customer_id'].unique()
has_credit = df[df['transaction_type']=='Credit']['customer_id'].unique()
both = np.intersect1d(has_debit, has_credit)
print('Number of customers with both types:', len(both))
# Advanced Example 1: Flag transactions above a certain amount as 'large'.
threshold = 300
df['is_large'] = df['amount'] > threshold
print(df[['transaction_id', 'amount', 'is_large']].head(3))
# Advanced Example 2: Estimate risk by region using the German Credit dataset.
risk_by_region = german.groupby('credit_status')['default_flag'].mean()
print(risk_by_region)
# Advanced Example 3: Join all three tables for a holistic view.
df_accounts = pd.merge(df, accounts, on='customer_id', how='left')
result = pd.merge(df_accounts, customers, on='customer_id', how='left')
print(result[['transaction_id','customer_id','account_id','segment','region','amount']].head(3))
# Error Handling Example: What happens if we try to join on the wrong column?
try:
wrong_merge = pd.merge(df, accounts, left_on='transaction_id', right_on='account_id', how='left')
print(wrong_merge.head(3))
except Exception as e:
print('Error:', e)
# Debugging Example: Check for missing values after merges.
merged = pd.merge(df, customers, on='customer_id', how='left')
missing = merged['segment'].isnull().sum()
print('Rows with missing segment after merge:', missing)
# Best Practice: Always set random seed for reproducibility.
np.random.seed(42)
# Common Pattern: Use groupby to summarize banking data easily.
summary = df.groupby('transaction_type')['amount'].agg(['count','mean','sum'])
print(summary)
# End-to-End Mini Problem: Find customers whose total transaction amount is above average and are in Metro region.
totals = df.groupby('customer_id')['amount'].sum().reset_index()
avg_total = totals['amount'].mean()
high_spenders = totals[totals['amount'] > avg_total]
final = pd.merge(high_spenders, customers, on='customer_id', how='left')
metro_high = final[final['region'] == 'Metro']
print('High spenders (Metro region):')
print(metro_high[['customer_id', 'amount', 'region']].head())
Want more? Subscribe to our YouTube channel for practical Python banking tutorials!#
- YouTube is a great place to watch more detailed code walkthroughs and banking cases.
- Hit the Subscribe button to keep learning!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



