Mathew K Analytics

Lesson 22 · Python for Banking and Finance

Feature Engineering Techniques to Improve Credit Risk Models Using Python

Feature engineering means creating new data features that improve the performance of credit models. This is important in banking because banks use credit…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Feature Engineering for Credit Models#

  • Feature engineering means creating new data features that improve the performance of credit models.
  • This is important in banking because banks use credit models to decide if a customer is likely to default or repay a loan.
  • In this lesson, you will learn to build features from real-world style banking datasets for use with credit risk models.
  • You will see step-by-step examples from beginner to advanced, using pandas and numpy.
  • By the end, you will know how to extract, aggregate, and transform raw banking data into model-ready features.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
np.random.seed(42)

Core Data Concepts#

  • In practice, banks store customer, account, and transaction data in different tables.
  • Transactions record each time money moves, with amounts, types, and time stamps.
  • Customers may have one or more accounts, which they use for transactions.
  • Credit models often require joining these tables and aggregating data over time.
  • Common mistakes include: misaligning joins, counting transactions incorrectly, or leaking future data when creating features.
customer_ids = [f'CUST_{i:04d}' for i in range(1, 201)]
customers = pd.DataFrame({
    'customer_id': customer_ids,
    'segment': ['Retail'] * 150 + ['Business'] * 50,
    'region': ['Metro'] * 100 + ['Regional'] * 100
})
print(customers.shape)
print(customers.head(3))
(200, 3)
  customer_id segment region
0   CUST_0001  Retail  Metro
1   CUST_0002  Retail  Metro
2   CUST_0003  Retail  Metro
accounts = pd.DataFrame({
    'account_id': [f'ACC_{i:05d}' for i in range(1, 201)],
    'customer_id': customer_ids,
    'account_type': np.random.choice(['Savings', 'Cheque', 'Credit'], size=200),
    'open_date': pd.date_range(start='2015-01-01', periods=200, freq='30D')
})
print(accounts.shape)
print(accounts.head(3))
(200, 4)
  account_id customer_id account_type  open_date
0  ACC_00001   CUST_0001       Credit 2015-01-01
1  ACC_00002   CUST_0002      Savings 2015-01-31
2  ACC_00003   CUST_0003       Credit 2015-03-02
n_transactions = 1000
df = pd.DataFrame({
    'transaction_id': range(1, n_transactions + 1),
    'customer_id': np.random.choice(customer_ids, n_transactions),
    'amount': np.round(np.random.normal(150, 60, n_transactions), 2),
    'transaction_type': np.random.choice(['Debit', 'Credit'], n_transactions),
    'channel': np.random.choice(['ATM', 'Online', 'Branch', 'POS'], n_transactions),
    'date': pd.date_range(start='2024-01-01', periods=n_transactions, freq='h')
})
print(df.shape)
print(df.head(3))
(1000, 6)
   transaction_id customer_id  amount transaction_type channel  \
0               1   CUST_0180   88.88            Debit     ATM   
1               2   CUST_0113  328.57           Credit     ATM   
2               3   CUST_0062  234.86           Credit  Branch   

                 date  
0 2024-01-01 00:00:00  
1 2024-01-01 01:00:00  
2 2024-01-01 02:00:00  

Beginner Example 1: Simple Transaction Counts#

  • A common first step in feature engineering is to count the number of transactions per customer.
  • This captures overall customer activity and is easy to create.
txn_counts = df.groupby('customer_id').transaction_id.count().reset_index(name='txn_count')
print(txn_counts.head(3))
  customer_id  txn_count
0   CUST_0001          9
1   CUST_0002          3
2   CUST_0003          6

Beginner Example 2: Total Amount Spent#

  • Sums of amounts per customer show overall value of their transactions.
  • You can later break this down by debit and credit flows.
txn_sum = df.groupby('customer_id').amount.sum().reset_index(name='total_amount')
print(txn_sum.head(3))
  customer_id  total_amount
0   CUST_0001       1340.49
1   CUST_0002        479.36
2   CUST_0003        882.76

Beginner Example 3: Average Transaction Value#

  • Average size of customer transactions highlights their typical behavior.
  • This helps distinguish between small frequent and large infrequent spend patterns.
txn_avg = df.groupby('customer_id').amount.mean().reset_index(name='avg_amount')
print(txn_avg.head(3))
  customer_id  avg_amount
0   CUST_0001  148.943333
1   CUST_0002  159.786667
2   CUST_0003  147.126667

Intermediate Example 1: Ratio of Debit to Credit Transactions#

  • Customers with high debit to credit ratios may have different risk profiles.
  • This feature captures transaction type behaviors.
debit_counts = df[df.transaction_type == 'Debit'].groupby('customer_id').transaction_id.count().reset_index(name='debit_count')
credit_counts = df[df.transaction_type == 'Credit'].groupby('customer_id').transaction_id.count().reset_index(name='credit_count')
txn_ratio = pd.merge(debit_counts, credit_counts, on='customer_id', how='outer').fillna(0)
txn_ratio['debit_credit_ratio'] = txn_ratio['debit_count'] / (txn_ratio['credit_count'] + 1)
print(txn_ratio.head(3))
  customer_id  debit_count  credit_count  debit_credit_ratio
0   CUST_0001          4.0           5.0            0.666667
1   CUST_0002          2.0           1.0            1.000000
2   CUST_0003          1.0           5.0            0.166667

Intermediate Example 2: Max Transaction Value Per Channel#

  • The largest transaction per channel can be a signal for credit risk or fraud.
  • High-value transactions are often monitored by banks.
max_amt_channel = df.groupby(['customer_id', 'channel']).amount.max().unstack().fillna(0).reset_index()
print(max_amt_channel.head(3))
channel customer_id     ATM  Branch  Online     POS
0         CUST_0001  211.05  175.85  178.24  219.87
1         CUST_0002  138.96  206.53  133.87    0.00
2         CUST_0003  108.11  214.68    0.00  214.28

Intermediate Example 3: Days Since Last Transaction#

  • Recency of activity can be important in dynamic risk monitoring.
  • Customers who have not transacted recently may have changed their behavior.
latest_txn = df.groupby('customer_id').date.max().reset_index(name='last_txn_date')
latest_txn['days_since_last_txn'] = (pd.Timestamp('2024-02-11') - latest_txn['last_txn_date']).dt.days
print(latest_txn.head(3))
  customer_id       last_txn_date  days_since_last_txn
0   CUST_0001 2024-02-08 01:00:00                    2
1   CUST_0002 2024-01-15 04:00:00                   26
2   CUST_0003 2024-02-06 11:00:00                    4

Advanced Example 1: Rolling Mean of Amounts (3-Month Window)#

  • Credit models can capture changing customer behavior by using rolling features.
  • Here we compute the rolling mean of transaction amounts, grouped by customer.
df['date'] = pd.to_datetime(df['date'])
df_sorted = df.sort_values(['customer_id', 'date'])
df_sorted['rolling_mean_amt'] = df_sorted.groupby('customer_id')['amount'].transform(lambda x: x.rolling(window=12, min_periods=1).mean())
print(df_sorted[['customer_id', 'amount', 'rolling_mean_amt']].head(8))
    customer_id  amount  rolling_mean_amt
288   CUST_0001   79.87         79.870000
334   CUST_0001  175.85        127.860000
507   CUST_0001  101.40        119.040000
539   CUST_0001  219.87        144.247500
623   CUST_0001   78.60        131.118000
725   CUST_0001  122.46        129.675000
748   CUST_0001  178.24        136.612857
865   CUST_0001  211.05        145.917500

Advanced Example 2: Transaction Standard Deviation (Variability Feature)#

  • Customers who show a sudden increase in transaction variability may be changing risk profile.
  • Let us engineer standard deviation of transaction amounts per customer.
amt_std = df.groupby('customer_id').amount.std().reset_index(name='amount_stddev').fillna(0)
print(amt_std.head(3))
  customer_id  amount_stddev
0   CUST_0001      54.471439
1   CUST_0002      40.560836
2   CUST_0003      76.176081

Advanced Example 3: One-Hot Encoding for Channel#

  • Models need categorical variables converted into numeric features.
  • Here we one-hot encode the transaction channel for each customer.
channel_dummies = pd.get_dummies(df[['customer_id', 'channel']], columns=['channel'])
channel_summary = channel_dummies.groupby('customer_id').sum().reset_index()
print(channel_summary.head(3))
  customer_id  channel_ATM  channel_Branch  channel_Online  channel_POS
0   CUST_0001            2               2               2            3
1   CUST_0002            1               1               1            0
2   CUST_0003            2               3               0            1

Error Handling and Debugging: Handling Missing Values#

  • It is common for transaction or account features to have missing data.
  • Let us simulate missing total_amount for a few customers and fill these gaps.
txn_sum_missing = txn_sum.copy()
txn_sum_missing.loc[2:5, 'total_amount'] = np.nan
print('Before filling:')
print(txn_sum_missing.head(7))
txn_sum_missing['total_amount'] = txn_sum_missing['total_amount'].fillna(txn_sum_missing['total_amount'].mean())
print('After filling:')
print(txn_sum_missing.head(7))
Before filling:
  customer_id  total_amount
0   CUST_0001       1340.49
1   CUST_0002        479.36
2   CUST_0003           NaN
3   CUST_0004           NaN
4   CUST_0005           NaN
5   CUST_0006           NaN
6   CUST_0007        563.23
After filling:
  customer_id  total_amount
0   CUST_0001   1340.490000
1   CUST_0002    479.360000
2   CUST_0003    771.654821
3   CUST_0004    771.654821
4   CUST_0005    771.654821
5   CUST_0006    771.654821
6   CUST_0007    563.230000

Best Practices and Common Patterns#

  • Always set seeds (e.g., np.random.seed(42)) for reproducibility.
  • Check your feature distributions for outlier values and missing data.
  • Avoid 'data leakage': never use future information to create present features.
  • Break down complex features into smaller steps (like counting then aggregating).
  • Test engineered features for meaningful variation, not just technical correctness.

Tiny End-to-End Problem: Predicting Default Risk Features#

  • Let us simulate a simple workflow: Bring together features, merge with customer segments, and prepare for modeling.
  • We will use a synthetic default flag as the target label.
# Use previously engineered features: txn_counts, txn_sum, amt_std, channel_summary
feature_df = transactions = txn_counts.merge(txn_sum, on='customer_id', how='left').merge(amt_std, on='customer_id', how='left').merge(channel_summary, on='customer_id', how='left')
# Add customer segment information
feature_df = feature_df.merge(customers, on='customer_id', how='left')
# Add a synthetic default flag: 10% chance of default
feature_df['default_flag'] = np.random.binomial(1, 0.10, feature_df.shape[0])
print(feature_df.head(3))
  customer_id  txn_count  total_amount  amount_stddev  channel_ATM  \
0   CUST_0001          9       1340.49      54.471439            2   
1   CUST_0002          3        479.36      40.560836            1   
2   CUST_0003          6        882.76      76.176081            2   

   channel_Branch  channel_Online  channel_POS segment region  default_flag  
0               2               2            3  Retail  Metro             0  
1               1               1            0  Retail  Metro             0  
2               3               0            1  Retail  Metro             0  

Takeaways and Next Steps#

  • Developing strong feature engineering skills is key for banking model success.
  • Practice aggregating, encoding, and handling banking data regularly.
  • Try exploring public credit datasets for more feature engineering practice.
  • For more deep dives, look for feature engineering walkthroughs on our channel.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.