Mathew K Analytics

Lesson 5 · Python For Machine Learning

3 Cross-Validation in Python: Machine Learning Model Evaluation Techniques

Welcome! In this lesson, we will learn about cross validation using Python. You will see how splitting data improves models. We will use real data and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Cross Validation in Python: Beginner Guide#

Welcome! In this lesson, we will learn about cross validation using Python.

You will see how splitting data improves models.

We will use real data and simple code.

Let us get started!

# Import needed libraries
import warnings
warnings.filterwarnings('ignore')
import pandas as pd
import numpy as np

Why Cross Validation?#

Cross validation helps check if a machine learning model is really good.

It splits our data into pieces. Each piece gets a turn for testing.

This helps us avoid mistakes like overfitting.

# Data setup
titanic_url = "https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv"
df = pd.read_csv(titanic_url)
print("Data shape:", df.shape)
df.head()
Data shape: (891, 12)
PassengerId Survived Pclass Name Sex Age SibSp Parch Ticket Fare Cabin Embarked
0 1 0 3 Braund, Mr. Owen Harris male 22.0 1 0 A/5 21171 7.2500 NaN S
1 2 1 1 Cumings, Mrs. John Bradley (Florence Briggs Th... female 38.0 1 0 PC 17599 71.2833 C85 C
2 3 1 3 Heikkinen, Miss. Laina female 26.0 0 0 STON/O2. 3101282 7.9250 NaN S
3 4 1 1 Futrelle, Mrs. Jacques Heath (Lily May Peel) female 35.0 1 0 113803 53.1000 C123 S
4 5 0 3 Allen, Mr. William Henry male 35.0 0 0 373450 8.0500 NaN S
# Check target variable: Survived
df['Survived'].value_counts().plot(kind="bar", title="Survival Counts")
<Axes: title={'center': 'Survival Counts'}, xlabel='Survived'>
No description has been provided for this image

What is 'train_test_split'?#

Often, we split our data once for training and once for testing.

But, one split may not tell the whole story. That is where cross validation helps.

# Splitting data - one way
from sklearn.model_selection import train_test_split
X = df[['Pclass', 'Age', 'SibSp', 'Parch']]
y = df['Survived']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print("Training set size:", X_train.shape)
print("Testing set size:", X_test.shape)
Training set size: (712, 4)
Testing set size: (179, 4)
# Simple fit to show training/testing
from sklearn.linear_model import LogisticRegression
model = LogisticRegression()
model.fit(X_train.fillna(X_train.mean()), y_train)
score = model.score(X_test.fillna(X_train.mean()), y_test)
print(f"Test accuracy: {score:.2f}")
Test accuracy: 0.74

Limitation: Only testing once#

If we test once, maybe we are lucky or unlucky.

Cross validation repeats the split many times for a better check.

# Basic cross validation
from sklearn.model_selection import cross_val_score
scores = cross_val_score(LogisticRegression(), X.fillna(X.mean()), y, cv=5)
print("Cross validation accuracies:", scores)
print(f"Mean accuracy: {scores.mean():.2f}")
Cross validation accuracies: [0.62569832 0.67977528 0.71910112 0.71910112 0.70224719]
Mean accuracy: 0.69
# Using a different model (Decision Tree)
from sklearn.tree import DecisionTreeClassifier
dt_scores = cross_val_score(DecisionTreeClassifier(random_state=42), X.fillna(X.mean()), y, cv=5)
print("Decision tree scores:", dt_scores)
print(f"Average: {dt_scores.mean():.2f}")
Decision tree scores: [0.6424581  0.65730337 0.71910112 0.71910112 0.70786517]
Average: 0.69

What is K-Fold?#

K-Fold splits our data into 'k' equal groups. Each group takes a turn as the test set.

This method is strong because all data gets tested.

# Manual K-Fold example
from sklearn.model_selection import KFold
kf = KFold(n_splits=5, shuffle=True, random_state=1)
for i, (train_idx, test_idx) in enumerate(kf.split(X)):
    print(f"Fold {i+1}: Train IDs {train_idx[:3]}... Test IDs {test_idx[:3]}...")
    
Fold 1: Train IDs [0 1 4]... Test IDs [2 3 6]...
Fold 2: Train IDs [1 2 3]... Test IDs [ 0 11 14]...
Fold 3: Train IDs [0 1 2]... Test IDs [4 5 9]...
Fold 4: Train IDs [0 2 3]... Test IDs [ 1 21 24]...
Fold 5: Train IDs [0 1 2]... Test IDs [ 7 10 15]...
# Cross validation with scoring options
scores = cross_val_score(LogisticRegression(), X.fillna(X.mean()), y, cv=5, scoring='recall')
print("Recall scores:", scores)
print(f"Average recall: {np.mean(scores):.2f}")
Recall scores: [0.33333333 0.5        0.47058824 0.47058824 0.49275362]
Average recall: 0.45
# Stratified K-Fold keeps ratio
from sklearn.model_selection import StratifiedKFold
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=1)
for i, (train_idx, test_idx) in enumerate(skf.split(X, y)):
    print(f"Fold {i+1}: Survive percent in test: {y.iloc[test_idx].mean():.2f}")
    
Fold 1: Survive percent in test: 0.39
Fold 2: Survive percent in test: 0.38
Fold 3: Survive percent in test: 0.38
Fold 4: Survive percent in test: 0.38
Fold 5: Survive percent in test: 0.39

Cross Validating Preprocessing#

Sometimes, we need to clean or scale data. Let us use pipelines to combine steps safely.

# Pipeline use for preprocessing
from sklearn.pipeline import make_pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler

pipe = make_pipeline(SimpleImputer(strategy='mean'), StandardScaler(), LogisticRegression())
scores = cross_val_score(pipe, X, y, cv=5)
print(f"With pipeline, cv accuracy: {scores.mean():.2f}")
With pipeline, cv accuracy: 0.69
# Challenge: Let us try your own cv split count
n_folds = input("How many folds for cross validation? (try 3-10): ")
n_folds = int(n_folds)
cv_scores = cross_val_score(LogisticRegression(), X.fillna(X.mean()), y, cv=n_folds)
print(f"You picked {n_folds} folds. CV result: {cv_scores.mean():.2f}")
You picked 7 folds. CV result: 0.70
 

Mini Project: Cross Validating on Titanic#

Let us do a mini project. Can we predict survival with more columns?

Step 1: Pick columns. Step 2: Build a pipeline. Step 3: Run cross validation.

# Mini project - more features
X2 = df[['Pclass', 'Age', 'SibSp', 'Parch', 'Fare', 'Sex']].copy()
X2['Sex'] = X2['Sex'].map({'male': 0, 'female': 1})
pipe = make_pipeline(SimpleImputer(strategy='mean'), StandardScaler(), LogisticRegression())
scores = cross_val_score(pipe, X2, y, cv=5)
print(f"CV accuracy with extra columns: {scores.mean():.2f}")
CV accuracy with extra columns: 0.79
# Error check: What if we forget to map 'Sex'?
try:
    cross_val_score(pipe, df[['Pclass', 'Age', 'SibSp', 'Parch', 'Fare', 'Sex']], y, cv=5)
except Exception as e:
    print("Error:", e)
    
Error: 
All the 5 fits failed.
It is very likely that your model is misconfigured.
You can try to debug the error by setting error_score='raise'.

Below are more details about the failures:
--------------------------------------------------------------------------------
5 fits failed with the following error:
Traceback (most recent call last):
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\model_selection\_validation.py", line 859, in _fit_and_score
    estimator.fit(X_train, y_train, **fit_params)
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\base.py", line 1365, in wrapper
    return fit_method(estimator, *args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\pipeline.py", line 655, in fit
    Xt = self._fit(X, y, routed_params, raw_params=params)
         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\pipeline.py", line 589, in _fit
    X, fitted_transformer = fit_transform_one_cached(
                            ^^^^^^^^^^^^^^^^^^^^^^^^^
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\joblib\memory.py", line 326, in __call__
    return self.func(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\pipeline.py", line 1540, in _fit_transform_one
    res = transformer.fit_transform(X, y, **params.get("fit_transform", {}))
          ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\utils\_set_output.py", line 316, in wrapped
    data_to_wrap = f(self, X, *args, **kwargs)
                   ^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\base.py", line 897, in fit_transform
    return self.fit(X, y, **fit_params).transform(X)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\base.py", line 1365, in wrapper
    return fit_method(estimator, *args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\impute\_base.py", line 436, in fit
    X = self._validate_input(X, in_fit=True)
        ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\impute\_base.py", line 361, in _validate_input
    raise new_ve from None
ValueError: Cannot use mean strategy with non-numeric data:
could not convert string to float: 'male'

# Advanced: Cross_val_predict to see predictions
from sklearn.model_selection import cross_val_predict
preds = cross_val_predict(pipe, X2, y, cv=5)
print("Predicted values:", preds[:10])
Predicted values: [0 1 1 1 0 0 0 0 1 1]
# Best practice: Shuffle and random_state always
scores1 = cross_val_score(pipe, X2, y, cv=StratifiedKFold(5, shuffle=True, random_state=42))
print(f"With shuffle and set random seed: {scores1.mean():.2f}")
With shuffle and set random seed: 0.79
# Troubleshooting: Pipeline errors
# Let us force a mistake to see the message.
bad_pipe = make_pipeline(StandardScaler(), LogisticRegression())
try:
    cross_val_score(bad_pipe, X2, y, cv=5)
except Exception as err:
    print("Pipeline error:", err)
    
Pipeline error: 
All the 5 fits failed.
It is very likely that your model is misconfigured.
You can try to debug the error by setting error_score='raise'.

Below are more details about the failures:
--------------------------------------------------------------------------------
5 fits failed with the following error:
Traceback (most recent call last):
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\model_selection\_validation.py", line 859, in _fit_and_score
    estimator.fit(X_train, y_train, **fit_params)
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\base.py", line 1365, in wrapper
    return fit_method(estimator, *args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\pipeline.py", line 663, in fit
    self._final_estimator.fit(Xt, y, **last_step_params["fit"])
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\base.py", line 1365, in wrapper
    return fit_method(estimator, *args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\linear_model\_logistic.py", line 1247, in fit
    X, y = validate_data(
           ^^^^^^^^^^^^^^
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\utils\validation.py", line 2971, in validate_data
    X, y = check_X_y(X, y, **check_params)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\utils\validation.py", line 1368, in check_X_y
    X = check_array(
        ^^^^^^^^^^^^
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\utils\validation.py", line 1105, in check_array
    _assert_all_finite(
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\utils\validation.py", line 120, in _assert_all_finite
    _assert_all_finite_element_wise(
  File "c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\utils\validation.py", line 169, in _assert_all_finite_element_wise
    raise ValueError(msg_err)
ValueError: Input X contains NaN.
LogisticRegression does not accept missing values encoded as NaN natively. For supervised learning, you might want to consider sklearn.ensemble.HistGradientBoostingClassifier and Regressor which accept missing values encoded as NaNs natively. Alternatively, it is possible to preprocess the data, for instance by using an imputer transformer in a pipeline or drop samples with missing values. See https://scikit-learn.org/stable/modules/impute.html You can find a list of all estimators that handle NaN values at the following page: https://scikit-learn.org/stable/modules/impute.html#estimators-that-handle-nan-values

Extra Tips#

  • Use StratifiedKFold for balanced classes.
  • Use pipelines to keep things tidy.
  • Practice with more datasets for skill.

Remember, practicing small steps grows knowledge!

# Challenge: Try cross validation on a different model
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(random_state=42)
rf_scores = cross_val_score(rf, X2.fillna(X2.mean()), y, cv=5)
print(f"Random Forest cv accuracy: {rf_scores.mean():.2f}")
Random Forest cv accuracy: 0.81

Summary: Cross Validation Key Points#

  • Cross validation splits data into parts, tests many times.
  • Safer than a single train-test split.
  • Helps judge how models might do on new data.
  • Use pipelines for clean, fair steps.
  • Real data needs careful cleaning.

What will you try with cross validation next?

Thank you for learning!#

Want more hands-on Python?

Like, comment, and subscribe for more lessons.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.