Lesson 4 · Python For Machine Learning
2 Train-Test Split in Python: Step-by-Step Guide for Machine Learning Beginners
Welcome! Today you will learn how to split data for machine learning. We'll practice using simple examples and a real dataset. By the end, you'll be ready…
- CoursePython For Machine Learning
- Lesson4 of 16
- Video12 min
- FormatJupyter notebook · 19 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Train Test Split in Python: A Beginner's Guide#
Welcome! Today you will learn how to split data for machine learning.
We'll practice using simple examples and a real dataset.
By the end, you'll be ready to split data like a pro!
# Let's start by telling Python to silence warnings.
import warnings
warnings.filterwarnings("ignore")
Why split data?#
Machine learning models need to learn and to be tested.
Splitting data helps you check if your model can really predict new things.
You train on one part and test on another!
# Let's import the tools we need.
import pandas as pd
from sklearn.model_selection import train_test_split
Example: A Simple List Split#
Let's try splitting a tiny list before we use real data.
This will help you see what splitting is all about.
# Make a list of numbers from 0 to 9.
numbers = list(range(10))
print("Numbers:", numbers)
# Split our list into two parts.
train, test = train_test_split(numbers, test_size=0.3, random_state=42)
print("Train:", train)
print("Test:", test)
That's it! Splitting data works for all kinds of data, not just lists.
Next, let's use a real-world dataset: The Titanic passengers.
# Data setup: Load Titanic data from a CSV file.
url = "https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv"
df = pd.read_csv(url)
print("Shape:", df.shape)
df.head()
# Let's see how many people survived and how many did not.
df["Survived"].value_counts()
# Split Titanic data into features and labels.
X = df.drop("Survived", axis=1)
y = df["Survived"]
Now it's time for the real split!
We will keep 20 percent of the data for testing.
Random state makes our split repeatable.
# Split Titanic data into training and testing sets.
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print("Train features shape:", X_train.shape)
print("Test features shape:", X_test.shape)
# Check the training and test target balance.
print("Train set survival:")
print(y_train.value_counts(normalize=True))
print("Test set survival:")
print(y_test.value_counts(normalize=True))
# Try splitting with user choices.
split = float(input("Enter test size as a decimal (try 0.25): "))
X_train2, X_test2, y_train2, y_test2 = train_test_split(X, y, test_size=split, random_state=42)
print("Train shape:", X_train2.shape, "Test shape:", X_test2.shape)
Practice prompt: Change the random_state number#
See if your splits change.
Try different numbers and compare train and test shapes.
# Challenge: What happens if you skip random_state?
split1 = train_test_split(X, y, test_size=0.3)
split2 = train_test_split(X, y, test_size=0.3)
print("Same split?", (split1[0] == split2[0]).all())
# Best practice: Always set random_state for repeats.
split3a = train_test_split(X, y, test_size=0.3, random_state=99)
split3b = train_test_split(X, y, test_size=0.3, random_state=99)
print("Same split?", (split3a[0] == split3b[0]).all())
# What if your label is not balanced?
from sklearn.model_selection import StratifiedShuffleSplit
sss = StratifiedShuffleSplit(n_splits=1, test_size=0.2, random_state=42)
for train_idx, test_idx in sss.split(X, y):
Xs_train, Xs_test = X.iloc[train_idx], X.iloc[test_idx]
ys_train, ys_test = y.iloc[train_idx], y.iloc[test_idx]
print("Stratified train label freq:")
print(ys_train.value_counts(normalize=True))
print("Stratified test label freq:")
print(ys_test.value_counts(normalize=True))
# Quick tip: Shuffle your data for fairness.
df_shuffled = df.sample(frac=1, random_state=24)
df_shuffled.head()
Mini-project: Predicting Titanic survival#
Let's use everything you learned so far.
In this part, we will:
- Prepare the data
- Split into train/test sets
- Print out the shapes
- Show target distribution
You are ready to practice!
# Step 1: Remove rows with missing age values.
df_clean = df.dropna(subset=["Age"])
print("Rows left:", df_clean.shape[0])
# Step 2: Prepare features and labels again.
X_mini = df_clean.drop("Survived", axis=1)
y_mini = df_clean["Survived"]
# Step 3: Split and check label proportions.
X_train_m, X_test_m, y_train_m, y_test_m = train_test_split(
X_mini, y_mini, test_size=0.3, random_state=7
)
print("Train size:", X_train_m.shape[0], "Test size:", X_test_m.shape[0])
print("Train survive %:", y_train_m.mean().round(2), "Test survive %:", y_test_m.mean().round(2))
Challenge: Try it with a different dataset#
Can you split the data from another famous dataset?
Try the Iris dataset from sklearn for practice.
Recap#
You learned why and how to split data.
You tried on lists and real tables.
Using train_test_split is the first step for model testing.
Keep trying more datasets for better skills.
If you found this helpful, please like and subscribe.
Try out a split now for practice!
What project will you split data for next?
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



