Lesson 22 · Python for Banking and Finance
Feature Engineering Techniques to Improve Credit Risk Models Using Python
Feature engineering means creating new data features that improve the performance of credit models. This is important in banking because banks use credit…
- CoursePython for Banking and Finance
- Lesson22 of 24
- Video20 min
- FormatJupyter notebook · 15 code cells
What you'll learn
- Core Data Concepts
- Beginner Example 1: Simple Transaction Counts
- Beginner Example 2: Total Amount Spent
- Beginner Example 3: Average Transaction Value
- Intermediate Example 1: Ratio of Debit to Credit Transactions
- Intermediate Example 2: Max Transaction Value Per Channel
- Intermediate Example 3: Days Since Last Transaction
- Advanced Example 1: Rolling Mean of Amounts (3-Month Window)
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbFeature Engineering for Credit Models#
- Feature engineering means creating new data features that improve the performance of credit models.
- This is important in banking because banks use credit models to decide if a customer is likely to default or repay a loan.
- In this lesson, you will learn to build features from real-world style banking datasets for use with credit risk models.
- You will see step-by-step examples from beginner to advanced, using pandas and numpy.
- By the end, you will know how to extract, aggregate, and transform raw banking data into model-ready features.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
np.random.seed(42)
Core Data Concepts#
- In practice, banks store customer, account, and transaction data in different tables.
- Transactions record each time money moves, with amounts, types, and time stamps.
- Customers may have one or more accounts, which they use for transactions.
- Credit models often require joining these tables and aggregating data over time.
- Common mistakes include: misaligning joins, counting transactions incorrectly, or leaking future data when creating features.
customer_ids = [f'CUST_{i:04d}' for i in range(1, 201)]
customers = pd.DataFrame({
'customer_id': customer_ids,
'segment': ['Retail'] * 150 + ['Business'] * 50,
'region': ['Metro'] * 100 + ['Regional'] * 100
})
print(customers.shape)
print(customers.head(3))
accounts = pd.DataFrame({
'account_id': [f'ACC_{i:05d}' for i in range(1, 201)],
'customer_id': customer_ids,
'account_type': np.random.choice(['Savings', 'Cheque', 'Credit'], size=200),
'open_date': pd.date_range(start='2015-01-01', periods=200, freq='30D')
})
print(accounts.shape)
print(accounts.head(3))
n_transactions = 1000
df = pd.DataFrame({
'transaction_id': range(1, n_transactions + 1),
'customer_id': np.random.choice(customer_ids, n_transactions),
'amount': np.round(np.random.normal(150, 60, n_transactions), 2),
'transaction_type': np.random.choice(['Debit', 'Credit'], n_transactions),
'channel': np.random.choice(['ATM', 'Online', 'Branch', 'POS'], n_transactions),
'date': pd.date_range(start='2024-01-01', periods=n_transactions, freq='h')
})
print(df.shape)
print(df.head(3))
Beginner Example 1: Simple Transaction Counts#
- A common first step in feature engineering is to count the number of transactions per customer.
- This captures overall customer activity and is easy to create.
txn_counts = df.groupby('customer_id').transaction_id.count().reset_index(name='txn_count')
print(txn_counts.head(3))
Beginner Example 2: Total Amount Spent#
- Sums of amounts per customer show overall value of their transactions.
- You can later break this down by debit and credit flows.
txn_sum = df.groupby('customer_id').amount.sum().reset_index(name='total_amount')
print(txn_sum.head(3))
Beginner Example 3: Average Transaction Value#
- Average size of customer transactions highlights their typical behavior.
- This helps distinguish between small frequent and large infrequent spend patterns.
txn_avg = df.groupby('customer_id').amount.mean().reset_index(name='avg_amount')
print(txn_avg.head(3))
Intermediate Example 1: Ratio of Debit to Credit Transactions#
- Customers with high debit to credit ratios may have different risk profiles.
- This feature captures transaction type behaviors.
debit_counts = df[df.transaction_type == 'Debit'].groupby('customer_id').transaction_id.count().reset_index(name='debit_count')
credit_counts = df[df.transaction_type == 'Credit'].groupby('customer_id').transaction_id.count().reset_index(name='credit_count')
txn_ratio = pd.merge(debit_counts, credit_counts, on='customer_id', how='outer').fillna(0)
txn_ratio['debit_credit_ratio'] = txn_ratio['debit_count'] / (txn_ratio['credit_count'] + 1)
print(txn_ratio.head(3))
Intermediate Example 2: Max Transaction Value Per Channel#
- The largest transaction per channel can be a signal for credit risk or fraud.
- High-value transactions are often monitored by banks.
max_amt_channel = df.groupby(['customer_id', 'channel']).amount.max().unstack().fillna(0).reset_index()
print(max_amt_channel.head(3))
Intermediate Example 3: Days Since Last Transaction#
- Recency of activity can be important in dynamic risk monitoring.
- Customers who have not transacted recently may have changed their behavior.
latest_txn = df.groupby('customer_id').date.max().reset_index(name='last_txn_date')
latest_txn['days_since_last_txn'] = (pd.Timestamp('2024-02-11') - latest_txn['last_txn_date']).dt.days
print(latest_txn.head(3))
Advanced Example 1: Rolling Mean of Amounts (3-Month Window)#
- Credit models can capture changing customer behavior by using rolling features.
- Here we compute the rolling mean of transaction amounts, grouped by customer.
df['date'] = pd.to_datetime(df['date'])
df_sorted = df.sort_values(['customer_id', 'date'])
df_sorted['rolling_mean_amt'] = df_sorted.groupby('customer_id')['amount'].transform(lambda x: x.rolling(window=12, min_periods=1).mean())
print(df_sorted[['customer_id', 'amount', 'rolling_mean_amt']].head(8))
Advanced Example 2: Transaction Standard Deviation (Variability Feature)#
- Customers who show a sudden increase in transaction variability may be changing risk profile.
- Let us engineer standard deviation of transaction amounts per customer.
amt_std = df.groupby('customer_id').amount.std().reset_index(name='amount_stddev').fillna(0)
print(amt_std.head(3))
Advanced Example 3: One-Hot Encoding for Channel#
- Models need categorical variables converted into numeric features.
- Here we one-hot encode the transaction channel for each customer.
channel_dummies = pd.get_dummies(df[['customer_id', 'channel']], columns=['channel'])
channel_summary = channel_dummies.groupby('customer_id').sum().reset_index()
print(channel_summary.head(3))
Error Handling and Debugging: Handling Missing Values#
- It is common for transaction or account features to have missing data.
- Let us simulate missing total_amount for a few customers and fill these gaps.
txn_sum_missing = txn_sum.copy()
txn_sum_missing.loc[2:5, 'total_amount'] = np.nan
print('Before filling:')
print(txn_sum_missing.head(7))
txn_sum_missing['total_amount'] = txn_sum_missing['total_amount'].fillna(txn_sum_missing['total_amount'].mean())
print('After filling:')
print(txn_sum_missing.head(7))
Best Practices and Common Patterns#
- Always set seeds (e.g., np.random.seed(42)) for reproducibility.
- Check your feature distributions for outlier values and missing data.
- Avoid 'data leakage': never use future information to create present features.
- Break down complex features into smaller steps (like counting then aggregating).
- Test engineered features for meaningful variation, not just technical correctness.
Tiny End-to-End Problem: Predicting Default Risk Features#
- Let us simulate a simple workflow: Bring together features, merge with customer segments, and prepare for modeling.
- We will use a synthetic default flag as the target label.
# Use previously engineered features: txn_counts, txn_sum, amt_std, channel_summary
feature_df = transactions = txn_counts.merge(txn_sum, on='customer_id', how='left').merge(amt_std, on='customer_id', how='left').merge(channel_summary, on='customer_id', how='left')
# Add customer segment information
feature_df = feature_df.merge(customers, on='customer_id', how='left')
# Add a synthetic default flag: 10% chance of default
feature_df['default_flag'] = np.random.binomial(1, 0.10, feature_df.shape[0])
print(feature_df.head(3))
Takeaways and Next Steps#
- Developing strong feature engineering skills is key for banking model success.
- Practice aggregating, encoding, and handling banking data regularly.
- Try exploring public credit datasets for more feature engineering practice.
- For more deep dives, look for feature engineering walkthroughs on our channel.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



