Mathew K Analytics

Lesson 61 · Python Fundamentals

Master Hyperparameter Tuning with GridSearchCV in Python for Optimal ML Models

In this lesson, we will explore how to make machine learning models smarter with hyperparameter tuning. We will use GridSearchCV in Python. No experience…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Hyperparameter Tuning with GridSearchCV: A Beginner's Guide#

In this lesson, we will explore how to make machine learning models smarter with hyperparameter tuning.

We will use GridSearchCV in Python.

No experience required!

By the end, you will know how to set up, search, and pick the best model params using real-world data.

# Import libraries and filter warnings
import warnings
warnings.filterwarnings('ignore')
# Data setup
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
titanic = pd.read_csv(url)
print('Shape:', titanic.shape)
titanic.head()
Shape: (891, 12)
PassengerId Survived Pclass Name Sex Age SibSp Parch Ticket Fare Cabin Embarked
0 1 0 3 Braund, Mr. Owen Harris male 22.0 1 0 A/5 21171 7.2500 NaN S
1 2 1 1 Cumings, Mrs. John Bradley (Florence Briggs Th... female 38.0 1 0 PC 17599 71.2833 C85 C
2 3 1 3 Heikkinen, Miss. Laina female 26.0 0 0 STON/O2. 3101282 7.9250 NaN S
3 4 1 1 Futrelle, Mrs. Jacques Heath (Lily May Peel) female 35.0 1 0 113803 53.1000 C123 S
4 5 0 3 Allen, Mr. William Henry male 35.0 0 0 373450 8.0500 NaN S
# See the target distribution
titanic['Survived'].value_counts().plot(kind='bar', title='Survival Counts')
<Axes: title={'center': 'Survival Counts'}, xlabel='Survived'>
No description has been provided for this image

What is a Hyperparameter?#

Machine learning models have settings called hyperparameters.

Hyperparameters do not get learned from the data.

We decide their values before training.

Examples: how many neighbors in KNN, or the max depth of a decision tree.

# Prepare features and target
predictors = ['Pclass', 'Age', 'SibSp', 'Fare']
target = 'Survived'
X = titanic[predictors]
y = titanic[target]
X = X.fillna(X.mean())
X.head()
Pclass Age SibSp Fare
0 3 22.0 1 7.2500
1 1 38.0 1 71.2833
2 3 26.0 0 7.9250
3 1 35.0 1 53.1000
4 3 35.0 0 8.0500
# Split data for training and testing
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)
# Simple decision tree, no tuning yet
from sklearn.tree import DecisionTreeClassifier
tree = DecisionTreeClassifier(random_state=42)
tree.fit(X_train, y_train)
score = tree.score(X_test, y_test)
print('Test accuracy:', score)
Test accuracy: 0.6480446927374302

Why Tune Hyperparameters?#

Each model's settings can change its predictions and results.

Hyperparameter tuning means testing many options to find what works best.

GridSearchCV helps us try all combinations, with no guesswork.

# Import GridSearchCV
from sklearn.model_selection import GridSearchCV
# Set search parameter grid
param_grid = {
    'max_depth': [2, 4, 6, 8],
    'min_samples_split': [2, 5, 10],
    'criterion': ['gini', 'entropy']
}
# Set up the grid search
grid = GridSearchCV(
    DecisionTreeClassifier(random_state=42),
    param_grid,
    cv=5,
    scoring='accuracy'
)
# Fit the grid search
grid.fit(X_train, y_train)
GridSearchCV(cv=5, estimator=DecisionTreeClassifier(random_state=42),
             param_grid={'criterion': ['gini', 'entropy'],
                         'max_depth': [2, 4, 6, 8],
                         'min_samples_split': [2, 5, 10]},
             scoring='accuracy')
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# See the best parameters found
print('Best parameters:', grid.best_params_)
print('Best cross-val accuracy:', grid.best_score_)
Best parameters: {'criterion': 'gini', 'max_depth': 4, 'min_samples_split': 2}
Best cross-val accuracy: 0.7092189500640205
# Test the best model on the test set
best = grid.best_estimator_
test_score = best.score(X_test, y_test)
print('Final test accuracy:', test_score)
Final test accuracy: 0.7150837988826816
# View all tested settings and results
results = pd.DataFrame(grid.cv_results_)
results[['params', 'mean_test_score', 'rank_test_score']].sort_values(by='mean_test_score', ascending=False).head()
params mean_test_score rank_test_score
3 {'criterion': 'gini', 'max_depth': 4, 'min_sam... 0.709219 1
4 {'criterion': 'gini', 'max_depth': 4, 'min_sam... 0.709219 1
5 {'criterion': 'gini', 'max_depth': 4, 'min_sam... 0.709219 1
15 {'criterion': 'entropy', 'max_depth': 4, 'min_... 0.702275 4
16 {'criterion': 'entropy', 'max_depth': 4, 'min_... 0.702275 4
# Use input() for simple user story
depth = input('Pick a tree depth to try (eg. 2, 3, 4): ')
custom_tree = DecisionTreeClassifier(max_depth=int(depth), random_state=42)
custom_tree.fit(X_train, y_train)
print('Test accuracy:', custom_tree.score(X_test, y_test))
Test accuracy: 0.7039106145251397
 
# Tips: Avoiding errors if you make a typo
try:
    depth = int(input('Pick another tree depth: '))
except ValueError:
    print('That is not a valid number! Using depth=2.')
    depth = 2
tree2 = DecisionTreeClassifier(max_depth=depth, random_state=42)
tree2.fit(X_train, y_train)
print('Accuracy:', tree2.score(X_test, y_test))
That is not a valid number! Using depth=2.
Accuracy: 0.6759776536312849
 

Mini Project: Your Own Tuning Grid#

Change the search space above.

For example: Try different numbers for min_samples_split and new options for max_depth.

Run GridSearchCV and see what best parameters it finds.

# Best practices: Use less data or fewer options for speed
## Try a small grid for quick learning
quick_grid = {
    'max_depth': [2, 3],
    'min_samples_split': [2, 5]
}
quick_search = GridSearchCV(DecisionTreeClassifier(random_state=42), quick_grid, cv=3)
quick_search.fit(X_train, y_train)
print('Best found:', quick_search.best_params_)
Best found: {'max_depth': 3, 'min_samples_split': 2}
# Troubleshooting: What if all results look the same?
if results['mean_test_score'].nunique() == 1:
    print('All parameter combos have the same accuracy! Try tweaking grid or preprocessing.')
    
# Extra: Use GridSearchCV with another model
from sklearn.ensemble import RandomForestClassifier
forest_grid = {
    'n_estimators': [10, 20],
    'max_depth': [2, 5]
}
forest_search = GridSearchCV(RandomForestClassifier(random_state=42), forest_grid, cv=3)
forest_search.fit(X_train, y_train)
print('Random forest best:', forest_search.best_params_)
Random forest best: {'max_depth': 2, 'n_estimators': 20}

Recap#

  • We learned what hyperparameters are.

  • We used GridSearchCV to try many model settings.

  • We picked the best model using real-world data.

  • Practice and explore with new parameter values!

Thanks for Learning Hyperparameter Tuning!#

If you enjoyed this lesson, please like the video and subscribe.

You can ask questions in the comments below.

Happy tuning!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.