Mathew K Analytics

Lesson 53 · Mastering Pandas

Essential Data Preparation Techniques for Effective scikit-learn Machine Learning Models

In this lesson, we will master how to ready our data with pandas for scikit-learn models. You will work through loading, cleaning, encoding, splitting,…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Preparing Data for scikit-learn Models#

In this lesson, we will master how to ready our data with pandas for scikit-learn models.

You will work through loading, cleaning, encoding, splitting, scaling, and more.

We will use the Titanic Datasetclassic for teaching.

Begin by making sure pandas is installed. Let us set up our workspace!

import warnings
warnings.filterwarnings('ignore')

# Data setup (Titanic Dataset)
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  

Exploring Titanic Data#

Before prepping data, let us explore what columns we have, their types, and a few summaries.

Understanding the data helps you choose what to fix, drop, or transform.

# Get a quick summary of the DataFrame
print(df.info())

# Show summary statistics for numerical columns
print(df.describe())
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 891 entries, 0 to 890
Data columns (total 12 columns):
 #   Column       Non-Null Count  Dtype  
---  ------       --------------  -----  
 0   PassengerId  891 non-null    int64  
 1   Survived     891 non-null    int64  
 2   Pclass       891 non-null    int64  
 3   Name         891 non-null    object 
 4   Sex          891 non-null    object 
 5   Age          714 non-null    float64
 6   SibSp        891 non-null    int64  
 7   Parch        891 non-null    int64  
 8   Ticket       891 non-null    object 
 9   Fare         891 non-null    float64
 10  Cabin        204 non-null    object 
 11  Embarked     889 non-null    object 
dtypes: float64(2), int64(5), object(5)
memory usage: 83.7+ KB
None
       PassengerId    Survived      Pclass         Age       SibSp  \
count   891.000000  891.000000  891.000000  714.000000  891.000000   
mean    446.000000    0.383838    2.308642   29.699118    0.523008   
std     257.353842    0.486592    0.836071   14.526497    1.102743   
min       1.000000    0.000000    1.000000    0.420000    0.000000   
25%     223.500000    0.000000    2.000000   20.125000    0.000000   
50%     446.000000    0.000000    3.000000   28.000000    0.000000   
75%     668.500000    1.000000    3.000000   38.000000    1.000000   
max     891.000000    1.000000    3.000000   80.000000    8.000000   

            Parch        Fare  
count  891.000000  891.000000  
mean     0.381594   32.204208  
std      0.806057   49.693429  
min      0.000000    0.000000  
25%      0.000000    7.910400  
50%      0.000000   14.454200  
75%      0.000000   31.000000  
max      6.000000  512.329200  
# Count missing values per column
print(df.isnull().sum())
PassengerId      0
Survived         0
Pclass           0
Name             0
Sex              0
Age            177
SibSp            0
Parch            0
Ticket           0
Fare             0
Cabin          687
Embarked         2
dtype: int64

Handling Missing Values#

Machine learning models cannot work with missing data.

Let us fill or drop missing values depending on the case.

# Fill missing ages with the median age
df['Age'].fillna(df['Age'].median(), inplace=True)

# Drop rows where 'Embarked' is missing
df.dropna(subset=['Embarked'], inplace=True)

# Print number of missing values after cleaning
print(df.isnull().sum())
PassengerId      0
Survived         0
Pclass           0
Name             0
Sex              0
Age              0
SibSp            0
Parch            0
Ticket           0
Fare             0
Cabin          687
Embarked         0
dtype: int64

Selecting Features for Modeling#

We do not always need every column for machine learning.

It is smart to focus on columns that relate to what we want to predict.

Let us keep only a few useful features.

# Pick features we want to use
features = ['Pclass', 'Sex', 'Age', 'SibSp', 'Parch', 'Fare', 'Embarked']
target = 'Survived'
X = df[features].copy()
y = df[target].copy()
print(X.head(3))
   Pclass     Sex   Age  SibSp  Parch     Fare Embarked
0       3    male  22.0      1      0   7.2500        S
1       1  female  38.0      1      0  71.2833        C
2       3  female  26.0      0      0   7.9250        S

Dealing with Categorical Features#

scikit-learn models need numbers, but 'Sex' and 'Embarked' are words.

We need to turn categories into numbers for our models.

Let us use pandas for easy encoding.

# Turn 'Sex' and 'Embarked' into one-hot encoded dummy variables
X_enc = pd.get_dummies(X, columns=['Sex', 'Embarked'], drop_first=True)
print(X_enc.head(3))
   Pclass   Age  SibSp  Parch     Fare  Sex_male  Embarked_Q  Embarked_S
0       3  22.0      1      0   7.2500      True       False        True
1       1  38.0      1      0  71.2833     False       False       False
2       3  26.0      0      0   7.9250     False       False        True

Feature Scaling Is Important#

Some models, especially those based on distance, need features scaled to similar ranges.

Let us standardize age and fare columns using sklearns StandardScaler.

from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()
X_enc[['Age', 'Fare']] = scaler.fit_transform(X_enc[['Age', 'Fare']])
print(X_enc[['Age', 'Fare']].head())
        Age      Fare
0 -0.563674 -0.500240
1  0.669217  0.788947
2 -0.255451 -0.486650
3  0.438050  0.422861
4  0.438050 -0.484133

Train-Test Split for Real Machine Learning#

Never train and test your model on the same rows.

Let us shuffle and split our data into a training set and a smaller test set using sklearn.

from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    X_enc, y, test_size=0.2, random_state=42, stratify=y)

print(f"Training set shape: {X_train.shape}")
print(f"Test set shape: {X_test.shape}")
Training set shape: (711, 8)
Test set shape: (178, 8)

Checking for Data Leakage#

Be careful: Never let data from the test set influence training steps before splitting.

Did you scale or fill missing values before the split? That is okay only when done with training set stats.

For production, use pipelines to avoid leaks!

# Example of creating an sklearn pipeline with preprocessing
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import OneHotEncoder
from sklearn.compose import ColumnTransformer

num_features = ['Age', 'Fare', 'SibSp', 'Parch', 'Pclass']
cat_features = ['Sex', 'Embarked']

numeric_transformer = Pipeline([
    ('imputer', SimpleImputer(strategy='median')),
    ('scaler', StandardScaler()),
])

categorical_transformer = Pipeline([
    ('imputer', SimpleImputer(strategy='most_frequent')),
    ('encoder', OneHotEncoder(drop='first'))
])

preprocessor = ColumnTransformer([
    ('num', numeric_transformer, num_features),
    ('cat', categorical_transformer, cat_features)
])

Mini-Project: Prepare Titanic Data for Modeling#

Now it is your turn!

Use the steps above to get a fresh Titanic dataset ready for scikit-learn.

Practice: Try different fill strategies. Try creating your own feature columns.

Then split your data and scale it.

# Practice: Prompt the learner for a feature to keep/drop
feature = input("Which feature column would you like to drop? (e.g. 'Fare') ")
if feature in X.columns:
    X2 = X.drop(columns=[feature])
    print(f"Feature '{feature}' dropped!")
    print(X2.head())
else:
    print("Feature not found. No changes made.")
    
Feature 'Fare' dropped!
   Pclass     Sex   Age  SibSp  Parch Embarked
0       3    male  22.0      1      0        S
1       1  female  38.0      1      0        C
2       3  female  26.0      0      0        S
3       1  female  35.0      1      0        S
4       3    male  35.0      0      0        S
# Challenge: Try a model fit
from sklearn.linear_model import LogisticRegression

logreg = LogisticRegression(max_iter=200)
logreg.fit(X_enc, y)

acc = logreg.score(X_enc, y)
print(f"Training accuracy: {acc:.3f}")
Training accuracy: 0.800
# See which features are most important (coefficient size)
import numpy as np
feature_names = X_enc.columns
coefs = logreg.coef_[0]
feat_rank = np.argsort(np.abs(coefs))[::-1]
print("Most influential features:")
for idx in feat_rank[:5]:
    print(f"{feature_names[idx]}: {coefs[idx]:.2f}")
    
Most influential features:
Sex_male: -2.61
Pclass: -1.06
Age: -0.49
Embarked_S: -0.39
SibSp: -0.31

Common Pitfalls & Tips#

  • Always fill or drop missing data before usage.

  • Make sure all features are numeric for scikit-learn.

  • Use pipelines for production models.

  • Try scikit-learn's ColumnTransformer to keep all prep steps together.

# Quick troubleshooting for type errors
try:
    logreg.fit(df[['Sex']], y)
except Exception as e:
    print("Error:", e)
    
Error: could not convert string to float: 'male'

Stretch Challenge#

Ready for more?

  • Try encoding features with OrdinalEncoder or LabelEncoder instead of get_dummies.
  • Make a new feature: FamilySize = SibSp + Parch + 1.
  • Try cross-validation for better accuracy estimates.

Recap#

Today you learned hands-on ways to get data ready for machine learning using pandas.

You saw data exploration, cleaning, selection, encoding, scaling, splitting, and pipeline tricks.

With these tools, you can prep almost any dataset for scikit-learn models!

What Next? Subscribe for More Lessons!#

Keep practicing: Try it with the Iris or Tips dataset next.

Like and subscribe for more pandas and scikit-learn walkthroughs. Have a topic request? Leave a comment on YouTube. Happy coding!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.