Lesson 17 · Mastering Pandas
Effective Techniques for Handling Outliers and Validating Data in Pandas
In this lesson, we will learn how to find and handle outliers, and how to validate data in pandas DataFrames. You will gain tools to clean, explore, and…
- CourseMastering Pandas
- Lesson17 of 44
- Video19 min
- FormatJupyter notebook · 19 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbHandling Outliers and Data Validation in Pandas#
In this lesson, we will learn how to find and handle outliers, and how to validate data in pandas DataFrames.
You will gain tools to clean, explore, and ensure the trustworthiness of your data.
Outliers are unusual values in your data that can skew results and cause misleading analysis.
Data validation helps catch errors, missing information, and impossible values that need fixing.
We will practice on the Titanic dataset, which records information about passengers on the famous ship.
Lets dive in!
import warnings; warnings.filterwarnings('ignore')
# Data setup (Titanic Dataset)
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
# What do the columns look like?
print(df.columns.tolist())
# Describe the numeric columns
df.describe()
# Look for missing values
df.isnull().sum()
# Visualize the distribution of Age
import matplotlib.pyplot as plt
plt.figure(figsize=(7,4))
df['Age'].hist(bins=30, edgecolor='k')
plt.xlabel('Age')
plt.ylabel('Count')
plt.title('Passenger Age Distribution')
plt.show()
# Detect outliers using the IQR (interquartile range) for Fare
Q1 = df['Fare'].quantile(0.25)
Q3 = df['Fare'].quantile(0.75)
IQR = Q3 - Q1
lower = Q1 - 1.5 * IQR
upper = Q3 + 1.5 * IQR
outliers = df[(df['Fare'] < lower) | (df['Fare'] > upper)]
print('Number of Fare outliers:', outliers.shape[0])
# Show a few of the Fare outliers
outliers[['PassengerId', 'Fare', 'Pclass', 'Embarked']].head()
# Remove Fare outliers for further analysis (optional)
df_no_outliers = df[(df['Fare'] >= lower) & (df['Fare'] <= upper)]
print('Rows after removing Fare outliers:', df_no_outliers.shape[0])
# Visualize Fare with outliers removed
plt.figure(figsize=(7,4))
df_no_outliers['Fare'].hist(bins=30, edgecolor='k')
plt.xlabel('Fare')
plt.ylabel('Count')
plt.title('Fare Distribution (No Outliers)')
plt.show()
# Spot impossible or suspicious values: Any negative ages?
neg_ages = df[df['Age'] < 0]
print(neg_ages.shape[0])
# Validate Embarked: List unexpected port codes
valid_ports = ['C','Q','S']
unexpected = df.loc[~df['Embarked'].isin(valid_ports), 'Embarked'].unique()
print('Unexpected values in Embarked:', unexpected)
# Check for duplicate rows
duplicates = df.duplicated().sum()
print('Number of duplicate rows:', duplicates)
# Fill missing Ages with the median age
median_age = df['Age'].median()
df['Age_filled'] = df['Age'].fillna(median_age)
print(df['Age_filled'].isnull().sum())
# Create a boolean column for outlier detection in Fare
df['Fare_outlier'] = (df['Fare'] < lower) | (df['Fare'] > upper)
print(df['Fare_outlier'].value_counts())
# Advanced: Apply a custom function to validate cabin format
def is_valid_cabin(x):
if pd.isnull(x):
return True
return str(x)[0].isalpha()
df['Cabin_valid'] = df['Cabin'].apply(is_valid_cabin)
print(df['Cabin_valid'].value_counts())
# Practice: Type the number of Fare outliers you found earlier
user_input = input('How many Fare outliers did we find before removal? ')
answer = str(outliers.shape[0])
if user_input == answer:
print('Great job!')
else:
print('Please check the earlier output and try again.')
Mini-Project: Outlier Handling and Data Validation EDA#
Let us do a quick End-to-End Analysis:
- Find passengers with very high fares and list their details.
- Fill all missing ages, and check if outliers exist in Parch (parents/children).
- List any passengers with negative or zero values in numeric fields.
Try this as a small personal project after the video.
# Troubleshooting tip: What if you see a KeyError?
# This often happens if you mistype a column name.
# You can double check column names using df.columns.
# Best practice: Use .copy() when subsetting
subset = df[['Age', 'Fare']].copy()
print(subset.head(2))
# Speed tip: For very large datasets, use .astype to reduce memory use
df['Pclass'] = df['Pclass'].astype('int8')
print(df['Pclass'].dtype)
Lesson Recap#
We explored outlier detection and data validation in pandas using the Titanic dataset.
You learned to spot, visualize, and handle extreme or invalid values, fill in missing entries, troubleshoot errors, and keep your analysis solid.
Remember, clean data is the foundation for meaningful results.
Keep practicing and stay curious!
If you learned something new today, please like and subscribe for more lessons!
Comment below with your favorite pandas tip or any problems you would like to see solved.
Happy coding!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



