Lesson 53 · Mastering Pandas
Essential Data Preparation Techniques for Effective scikit-learn Machine Learning Models
In this lesson, we will master how to ready our data with pandas for scikit-learn models. You will work through loading, cleaning, encoding, splitting,…
- CourseMastering Pandas
- Lesson53 of 44
- Video19 min
- FormatJupyter notebook · 13 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbPreparing Data for scikit-learn Models#
In this lesson, we will master how to ready our data with pandas for scikit-learn models.
You will work through loading, cleaning, encoding, splitting, scaling, and more.
We will use the Titanic Datasetclassic for teaching.
Begin by making sure pandas is installed. Let us set up our workspace!
import warnings
warnings.filterwarnings('ignore')
# Data setup (Titanic Dataset)
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
Exploring Titanic Data#
Before prepping data, let us explore what columns we have, their types, and a few summaries.
Understanding the data helps you choose what to fix, drop, or transform.
# Get a quick summary of the DataFrame
print(df.info())
# Show summary statistics for numerical columns
print(df.describe())
# Count missing values per column
print(df.isnull().sum())
Handling Missing Values#
Machine learning models cannot work with missing data.
Let us fill or drop missing values depending on the case.
# Fill missing ages with the median age
df['Age'].fillna(df['Age'].median(), inplace=True)
# Drop rows where 'Embarked' is missing
df.dropna(subset=['Embarked'], inplace=True)
# Print number of missing values after cleaning
print(df.isnull().sum())
Selecting Features for Modeling#
We do not always need every column for machine learning.
It is smart to focus on columns that relate to what we want to predict.
Let us keep only a few useful features.
# Pick features we want to use
features = ['Pclass', 'Sex', 'Age', 'SibSp', 'Parch', 'Fare', 'Embarked']
target = 'Survived'
X = df[features].copy()
y = df[target].copy()
print(X.head(3))
Dealing with Categorical Features#
scikit-learn models need numbers, but 'Sex' and 'Embarked' are words.
We need to turn categories into numbers for our models.
Let us use pandas for easy encoding.
# Turn 'Sex' and 'Embarked' into one-hot encoded dummy variables
X_enc = pd.get_dummies(X, columns=['Sex', 'Embarked'], drop_first=True)
print(X_enc.head(3))
Feature Scaling Is Important#
Some models, especially those based on distance, need features scaled to similar ranges.
Let us standardize age and fare columns using sklearns StandardScaler.
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_enc[['Age', 'Fare']] = scaler.fit_transform(X_enc[['Age', 'Fare']])
print(X_enc[['Age', 'Fare']].head())
Train-Test Split for Real Machine Learning#
Never train and test your model on the same rows.
Let us shuffle and split our data into a training set and a smaller test set using sklearn.
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X_enc, y, test_size=0.2, random_state=42, stratify=y)
print(f"Training set shape: {X_train.shape}")
print(f"Test set shape: {X_test.shape}")
Checking for Data Leakage#
Be careful: Never let data from the test set influence training steps before splitting.
Did you scale or fill missing values before the split? That is okay only when done with training set stats.
For production, use pipelines to avoid leaks!
# Example of creating an sklearn pipeline with preprocessing
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import OneHotEncoder
from sklearn.compose import ColumnTransformer
num_features = ['Age', 'Fare', 'SibSp', 'Parch', 'Pclass']
cat_features = ['Sex', 'Embarked']
numeric_transformer = Pipeline([
('imputer', SimpleImputer(strategy='median')),
('scaler', StandardScaler()),
])
categorical_transformer = Pipeline([
('imputer', SimpleImputer(strategy='most_frequent')),
('encoder', OneHotEncoder(drop='first'))
])
preprocessor = ColumnTransformer([
('num', numeric_transformer, num_features),
('cat', categorical_transformer, cat_features)
])
Mini-Project: Prepare Titanic Data for Modeling#
Now it is your turn!
Use the steps above to get a fresh Titanic dataset ready for scikit-learn.
Practice: Try different fill strategies. Try creating your own feature columns.
Then split your data and scale it.
# Practice: Prompt the learner for a feature to keep/drop
feature = input("Which feature column would you like to drop? (e.g. 'Fare') ")
if feature in X.columns:
X2 = X.drop(columns=[feature])
print(f"Feature '{feature}' dropped!")
print(X2.head())
else:
print("Feature not found. No changes made.")
# Challenge: Try a model fit
from sklearn.linear_model import LogisticRegression
logreg = LogisticRegression(max_iter=200)
logreg.fit(X_enc, y)
acc = logreg.score(X_enc, y)
print(f"Training accuracy: {acc:.3f}")
# See which features are most important (coefficient size)
import numpy as np
feature_names = X_enc.columns
coefs = logreg.coef_[0]
feat_rank = np.argsort(np.abs(coefs))[::-1]
print("Most influential features:")
for idx in feat_rank[:5]:
print(f"{feature_names[idx]}: {coefs[idx]:.2f}")
Common Pitfalls & Tips#
Always fill or drop missing data before usage.
Make sure all features are numeric for scikit-learn.
Use pipelines for production models.
Try scikit-learn's ColumnTransformer to keep all prep steps together.
# Quick troubleshooting for type errors
try:
logreg.fit(df[['Sex']], y)
except Exception as e:
print("Error:", e)
Stretch Challenge#
Ready for more?
- Try encoding features with OrdinalEncoder or LabelEncoder instead of get_dummies.
- Make a new feature: FamilySize = SibSp + Parch + 1.
- Try cross-validation for better accuracy estimates.
Recap#
Today you learned hands-on ways to get data ready for machine learning using pandas.
You saw data exploration, cleaning, selection, encoding, scaling, splitting, and pipeline tricks.
With these tools, you can prep almost any dataset for scikit-learn models!
What Next? Subscribe for More Lessons!#
Keep practicing: Try it with the Iris or Tips dataset next.
Like and subscribe for more pandas and scikit-learn walkthroughs. Have a topic request? Leave a comment on YouTube. Happy coding!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



