Mathew K Analytics

Lesson 76 · Data Science Projects

4 - Loan Default Prediction - Risk Modeling

Welcome to your hands-on journey! In this lesson, you will learn how data mining is used to predict loan default risk. We will use a real dataset and guide…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Loan Default Prediction: Modeling Risk Through Data Mining#

Welcome to your hands-on journey! In this lesson, you will learn how data mining is used to predict loan default risk.

We will use a real dataset and guide you step by step, from basic data handling to building a powerful risk model.

Ready to find out what makes someone more likely to default on a loan?

Let us get started!

import warnings; warnings.filterwarnings("ignore")  # Ignore warnings for a smoother learning experience

# Import must-have tools
import pandas as pd
import numpy as np
np.random.seed(42)

Step 1: Load and Explore the Loan Dataset#

We will start by loading a loan dataset.

Let us see how many records and columns we have, and take a quick peek at the first few entries.

# Data setup (Loan Dataset Placeholder)
import pandas as pd
url = 'https://raw.githubusercontent.com/plotly/datasets/master/finance-charts-apple.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(506, 11)
         Date   AAPL.Open   AAPL.High    AAPL.Low  AAPL.Close  AAPL.Volume  \
0  2015-02-17  127.489998  128.880005  126.919998  127.830002     63152400   
1  2015-02-18  127.629997  128.779999  127.449997  128.720001     44891700   
2  2015-02-19  128.479996  129.029999  128.330002  128.449997     37362400   

   AAPL.Adjusted          dn        mavg          up   direction  
0     122.905254  106.741052  117.927667  129.114281  Increasing  
1     123.760965  107.842423  118.940333  130.038244  Increasing  
2     123.501363  108.894245  119.889167  130.884089  Decreasing  
# Always check for missing data before analysis
missing = df.isnull().sum()
print(missing)
Date             0
AAPL.Open        0
AAPL.High        0
AAPL.Low         0
AAPL.Close       0
AAPL.Volume      0
AAPL.Adjusted    0
dn               0
mavg             0
up               0
direction        0
dtype: int64

Step 2: Clean and Prepare the Loan Data#

Data cleaning helps your model make better, more reliable predictions.

Let us remove rows with missing values for this first exploration.

df_clean = df.dropna()  # Remove any rows where important data is missing
print(f"Rows before cleaning: {df.shape[0]}")
print(f"Rows after cleaning: {df_clean.shape[0]}")
Rows before cleaning: 506
Rows after cleaning: 506
# Rename columns to lower case with underscores for easier coding
df_clean.columns = [c.lower().replace(' ', '_') for c in df_clean.columns]
print(df_clean.columns)
Index(['date', 'aapl.open', 'aapl.high', 'aapl.low', 'aapl.close',
       'aapl.volume', 'aapl.adjusted', 'dn', 'mavg', 'up', 'direction'],
      dtype='object')
# Get a summary of all the columns: type, basic stats
print(df_clean.info())
print(df_clean.describe(include='all').T)
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 506 entries, 0 to 505
Data columns (total 11 columns):
 #   Column         Non-Null Count  Dtype  
---  ------         --------------  -----  
 0   date           506 non-null    object 
 1   aapl.open      506 non-null    float64
 2   aapl.high      506 non-null    float64
 3   aapl.low       506 non-null    float64
 4   aapl.close     506 non-null    float64
 5   aapl.volume    506 non-null    int64  
 6   aapl.adjusted  506 non-null    float64
 7   dn             506 non-null    float64
 8   mavg           506 non-null    float64
 9   up             506 non-null    float64
 10  direction      506 non-null    object 
dtypes: float64(8), int64(1), object(2)
memory usage: 43.6+ KB
None
               count unique         top freq             mean  \
date             506    506  2015-02-17    1              NaN   
aapl.open      506.0    NaN         NaN  NaN          112.935   
aapl.high      506.0    NaN         NaN  NaN       113.919447   
aapl.low       506.0    NaN         NaN  NaN       111.942016   
aapl.close     506.0    NaN         NaN  NaN        112.95834   
aapl.volume    506.0    NaN         NaN  NaN  43178420.948617   
aapl.adjusted  506.0    NaN         NaN  NaN       110.459312   
dn             506.0    NaN         NaN  NaN       107.311385   
mavg           506.0    NaN         NaN  NaN       112.739865   
up             506.0    NaN         NaN  NaN       118.168345   
direction        506      2  Increasing  278              NaN   

                           std         min         25%         50%  \
date                       NaN         NaN         NaN         NaN   
aapl.open             11.28749        90.0    105.4825  112.889999   
aapl.high            11.251892   91.669998  106.349999  114.145001   
aapl.low             11.263687   89.470001  104.657501  111.800003   
aapl.close           11.244744   90.339996  105.672499  113.025002   
aapl.volume    19852531.300948  11475900.0  29742400.0  37474600.0   
aapl.adjusted        10.537529    89.00837  103.484803  110.821123   
dn                   11.095804   85.508858   97.011245  107.351628   
mavg                 10.595315   94.047166  104.954875   112.79975   
up                   10.670752   97.572721  111.052267  118.472542   
direction                  NaN         NaN         NaN         NaN   

                      75%          max  
date                  NaN          NaN  
aapl.open      122.267498   135.669998  
aapl.high        123.4975   136.270004  
aapl.low       121.599998   134.839996  
aapl.close     122.179998   135.509995  
aapl.volume    50763950.0  162206300.0  
aapl.adjusted  119.255457   135.509995  
dn             114.812152   127.289258  
mavg           121.889416      129.845  
up             128.515793   138.805366  
direction             NaN          NaN  

Visualizations help us spot trends quickly.

Let us plot some columns to see if anything looks unusual.

import matplotlib.pyplot as plt
plt.figure(figsize=(10, 5))
plt.plot(df_clean['date'], df_clean['aapl.close'])
plt.title('Loan-like feature trend over time')
plt.xlabel('Date')
plt.ylabel('Close Price (treat as loan risk score)')
plt.show()
No description has been provided for this image
# Make up a target column: Let us say loans with low 'aapl.close' values are at higher risk.
df_clean['loan_default'] = (df_clean['aapl.close'] < df_clean['aapl.close'].median()).astype(int)
print(df_clean[['aapl.close', 'loan_default']].head(5))
   aapl.close  loan_default
0  127.830002             0
1  128.720001             0
2  128.449997             0
3  129.500000             0
4  133.000000             0

Step 4: Prepare Data for Modeling (Features and Target)#

Machine learning models need features and a target.

Let us set up what predicts loan default, and what we want to predict.

features = ['aapl.open', 'aapl.high', 'aapl.low', 'aapl.volume']
X = df_clean[features]
y = df_clean['loan_default']
# Split data for training and testing to measure model accuracy
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print(f'Training samples: {X_train.shape[0]}')
print(f'Testing samples: {X_test.shape[0]}')
Training samples: 404
Testing samples: 102
# Train a simple classifier: Logistic Regression
from sklearn.linear_model import LogisticRegression
clf = LogisticRegression()
clf.fit(X_train, y_train)
LogisticRegression()
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Check how well our model predicts loan default
y_pred = clf.predict(X_test)
from sklearn.metrics import accuracy_score, classification_report
print("Accuracy:", accuracy_score(y_test, y_pred))
print(classification_report(y_test, y_pred))
Accuracy: 0.5784313725490197
              precision    recall  f1-score   support

           0       0.57      0.89      0.70        56
           1       0.60      0.20      0.30        46

    accuracy                           0.58       102
   macro avg       0.59      0.54      0.50       102
weighted avg       0.59      0.58      0.52       102

# Bonus: Show the effect of a single feature (open price) with a simple plot
import seaborn as sns
sns.boxplot(data=df_clean, x='loan_default', y='aapl.open')
plt.title('Open Price by Loan Default Risk Group')
plt.xlabel('Loan Default (0 = No, 1 = Yes)')
plt.ylabel('AAPL Open Price (Feature)')
plt.show()
No description has been provided for this image

Recap: You Just Built a Risk Model for Loan Defaults!#

With basic data mining, you loaded data, cleaned it, explored features, visualized patterns, and built a simple risk model.

Even with a small sample, this matches real steps used by banks and lenders.

Try switching out features and experiment with different classifiers like Decision Tree or Random Forest next.

Your Turn!#

Pick a different feature or a new dataset and try these steps again.

The only way to master data mining is to practice and experiment!

Subscribe to our YouTube channel for more hands-on tutorials.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.