Lesson 5 · Python For Machine Learning
3 Cross-Validation in Python: Machine Learning Model Evaluation Techniques
Welcome! In this lesson, we will learn about cross validation using Python. You will see how splitting data improves models. We will use real data and…
- CoursePython For Machine Learning
- Lesson5 of 16
- Video10 min
- FormatJupyter notebook · 20 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Cross Validation in Python: Beginner Guide#
Welcome! In this lesson, we will learn about cross validation using Python.
You will see how splitting data improves models.
We will use real data and simple code.
Let us get started!
# Import needed libraries
import warnings
warnings.filterwarnings('ignore')
import pandas as pd
import numpy as np
Why Cross Validation?#
Cross validation helps check if a machine learning model is really good.
It splits our data into pieces. Each piece gets a turn for testing.
This helps us avoid mistakes like overfitting.
# Data setup
titanic_url = "https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv"
df = pd.read_csv(titanic_url)
print("Data shape:", df.shape)
df.head()
# Check target variable: Survived
df['Survived'].value_counts().plot(kind="bar", title="Survival Counts")
What is 'train_test_split'?#
Often, we split our data once for training and once for testing.
But, one split may not tell the whole story. That is where cross validation helps.
# Splitting data - one way
from sklearn.model_selection import train_test_split
X = df[['Pclass', 'Age', 'SibSp', 'Parch']]
y = df['Survived']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print("Training set size:", X_train.shape)
print("Testing set size:", X_test.shape)
# Simple fit to show training/testing
from sklearn.linear_model import LogisticRegression
model = LogisticRegression()
model.fit(X_train.fillna(X_train.mean()), y_train)
score = model.score(X_test.fillna(X_train.mean()), y_test)
print(f"Test accuracy: {score:.2f}")
Limitation: Only testing once#
If we test once, maybe we are lucky or unlucky.
Cross validation repeats the split many times for a better check.
# Basic cross validation
from sklearn.model_selection import cross_val_score
scores = cross_val_score(LogisticRegression(), X.fillna(X.mean()), y, cv=5)
print("Cross validation accuracies:", scores)
print(f"Mean accuracy: {scores.mean():.2f}")
# Using a different model (Decision Tree)
from sklearn.tree import DecisionTreeClassifier
dt_scores = cross_val_score(DecisionTreeClassifier(random_state=42), X.fillna(X.mean()), y, cv=5)
print("Decision tree scores:", dt_scores)
print(f"Average: {dt_scores.mean():.2f}")
What is K-Fold?#
K-Fold splits our data into 'k' equal groups. Each group takes a turn as the test set.
This method is strong because all data gets tested.
# Manual K-Fold example
from sklearn.model_selection import KFold
kf = KFold(n_splits=5, shuffle=True, random_state=1)
for i, (train_idx, test_idx) in enumerate(kf.split(X)):
print(f"Fold {i+1}: Train IDs {train_idx[:3]}... Test IDs {test_idx[:3]}...")
# Cross validation with scoring options
scores = cross_val_score(LogisticRegression(), X.fillna(X.mean()), y, cv=5, scoring='recall')
print("Recall scores:", scores)
print(f"Average recall: {np.mean(scores):.2f}")
# Stratified K-Fold keeps ratio
from sklearn.model_selection import StratifiedKFold
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=1)
for i, (train_idx, test_idx) in enumerate(skf.split(X, y)):
print(f"Fold {i+1}: Survive percent in test: {y.iloc[test_idx].mean():.2f}")
Cross Validating Preprocessing#
Sometimes, we need to clean or scale data. Let us use pipelines to combine steps safely.
# Pipeline use for preprocessing
from sklearn.pipeline import make_pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler
pipe = make_pipeline(SimpleImputer(strategy='mean'), StandardScaler(), LogisticRegression())
scores = cross_val_score(pipe, X, y, cv=5)
print(f"With pipeline, cv accuracy: {scores.mean():.2f}")
# Challenge: Let us try your own cv split count
n_folds = input("How many folds for cross validation? (try 3-10): ")
n_folds = int(n_folds)
cv_scores = cross_val_score(LogisticRegression(), X.fillna(X.mean()), y, cv=n_folds)
print(f"You picked {n_folds} folds. CV result: {cv_scores.mean():.2f}")
Mini Project: Cross Validating on Titanic#
Let us do a mini project. Can we predict survival with more columns?
Step 1: Pick columns. Step 2: Build a pipeline. Step 3: Run cross validation.
# Mini project - more features
X2 = df[['Pclass', 'Age', 'SibSp', 'Parch', 'Fare', 'Sex']].copy()
X2['Sex'] = X2['Sex'].map({'male': 0, 'female': 1})
pipe = make_pipeline(SimpleImputer(strategy='mean'), StandardScaler(), LogisticRegression())
scores = cross_val_score(pipe, X2, y, cv=5)
print(f"CV accuracy with extra columns: {scores.mean():.2f}")
# Error check: What if we forget to map 'Sex'?
try:
cross_val_score(pipe, df[['Pclass', 'Age', 'SibSp', 'Parch', 'Fare', 'Sex']], y, cv=5)
except Exception as e:
print("Error:", e)
# Advanced: Cross_val_predict to see predictions
from sklearn.model_selection import cross_val_predict
preds = cross_val_predict(pipe, X2, y, cv=5)
print("Predicted values:", preds[:10])
# Best practice: Shuffle and random_state always
scores1 = cross_val_score(pipe, X2, y, cv=StratifiedKFold(5, shuffle=True, random_state=42))
print(f"With shuffle and set random seed: {scores1.mean():.2f}")
# Troubleshooting: Pipeline errors
# Let us force a mistake to see the message.
bad_pipe = make_pipeline(StandardScaler(), LogisticRegression())
try:
cross_val_score(bad_pipe, X2, y, cv=5)
except Exception as err:
print("Pipeline error:", err)
Extra Tips#
- Use StratifiedKFold for balanced classes.
- Use pipelines to keep things tidy.
- Practice with more datasets for skill.
Remember, practicing small steps grows knowledge!
# Challenge: Try cross validation on a different model
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(random_state=42)
rf_scores = cross_val_score(rf, X2.fillna(X2.mean()), y, cv=5)
print(f"Random Forest cv accuracy: {rf_scores.mean():.2f}")
Summary: Cross Validation Key Points#
- Cross validation splits data into parts, tests many times.
- Safer than a single train-test split.
- Helps judge how models might do on new data.
- Use pipelines for clean, fair steps.
- Real data needs careful cleaning.
What will you try with cross validation next?
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



