Mathew K Analytics

Lesson 17 · Mastering Pandas

Effective Techniques for Handling Outliers and Validating Data in Pandas

In this lesson, we will learn how to find and handle outliers, and how to validate data in pandas DataFrames. You will gain tools to clean, explore, and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Handling Outliers and Data Validation in Pandas#

In this lesson, we will learn how to find and handle outliers, and how to validate data in pandas DataFrames.

You will gain tools to clean, explore, and ensure the trustworthiness of your data.

Outliers are unusual values in your data that can skew results and cause misleading analysis.

Data validation helps catch errors, missing information, and impossible values that need fixing.

We will practice on the Titanic dataset, which records information about passengers on the famous ship.

Lets dive in!

import warnings; warnings.filterwarnings('ignore')
# Data setup (Titanic Dataset)
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  
# What do the columns look like?
print(df.columns.tolist())
['PassengerId', 'Survived', 'Pclass', 'Name', 'Sex', 'Age', 'SibSp', 'Parch', 'Ticket', 'Fare', 'Cabin', 'Embarked']
# Describe the numeric columns
df.describe()
PassengerId Survived Pclass Age SibSp Parch Fare
count 891.000000 891.000000 891.000000 714.000000 891.000000 891.000000 891.000000
mean 446.000000 0.383838 2.308642 29.699118 0.523008 0.381594 32.204208
std 257.353842 0.486592 0.836071 14.526497 1.102743 0.806057 49.693429
min 1.000000 0.000000 1.000000 0.420000 0.000000 0.000000 0.000000
25% 223.500000 0.000000 2.000000 20.125000 0.000000 0.000000 7.910400
50% 446.000000 0.000000 3.000000 28.000000 0.000000 0.000000 14.454200
75% 668.500000 1.000000 3.000000 38.000000 1.000000 0.000000 31.000000
max 891.000000 1.000000 3.000000 80.000000 8.000000 6.000000 512.329200
# Look for missing values
df.isnull().sum()
PassengerId      0
Survived         0
Pclass           0
Name             0
Sex              0
Age            177
SibSp            0
Parch            0
Ticket           0
Fare             0
Cabin          687
Embarked         2
dtype: int64
# Visualize the distribution of Age
import matplotlib.pyplot as plt
plt.figure(figsize=(7,4))
df['Age'].hist(bins=30, edgecolor='k')
plt.xlabel('Age')
plt.ylabel('Count')
plt.title('Passenger Age Distribution')
plt.show()
No description has been provided for this image
# Detect outliers using the IQR (interquartile range) for Fare
Q1 = df['Fare'].quantile(0.25)
Q3 = df['Fare'].quantile(0.75)
IQR = Q3 - Q1
lower = Q1 - 1.5 * IQR
upper = Q3 + 1.5 * IQR
outliers = df[(df['Fare'] < lower) | (df['Fare'] > upper)]
print('Number of Fare outliers:', outliers.shape[0])
Number of Fare outliers: 116
# Show a few of the Fare outliers
outliers[['PassengerId', 'Fare', 'Pclass', 'Embarked']].head()
PassengerId Fare Pclass Embarked
1 2 71.2833 1 C
27 28 263.0000 1 S
31 32 146.5208 1 C
34 35 82.1708 1 C
52 53 76.7292 1 C
# Remove Fare outliers for further analysis (optional)
df_no_outliers = df[(df['Fare'] >= lower) & (df['Fare'] <= upper)]
print('Rows after removing Fare outliers:', df_no_outliers.shape[0])
Rows after removing Fare outliers: 775
# Visualize Fare with outliers removed
plt.figure(figsize=(7,4))
df_no_outliers['Fare'].hist(bins=30, edgecolor='k')
plt.xlabel('Fare')
plt.ylabel('Count')
plt.title('Fare Distribution (No Outliers)')
plt.show()
No description has been provided for this image
# Spot impossible or suspicious values: Any negative ages?
neg_ages = df[df['Age'] < 0]
print(neg_ages.shape[0])
0
# Validate Embarked: List unexpected port codes
valid_ports = ['C','Q','S']
unexpected = df.loc[~df['Embarked'].isin(valid_ports), 'Embarked'].unique()
print('Unexpected values in Embarked:', unexpected)
Unexpected values in Embarked: [nan]
# Check for duplicate rows
duplicates = df.duplicated().sum()
print('Number of duplicate rows:', duplicates)
Number of duplicate rows: 0
# Fill missing Ages with the median age
median_age = df['Age'].median()
df['Age_filled'] = df['Age'].fillna(median_age)
print(df['Age_filled'].isnull().sum())
0
# Create a boolean column for outlier detection in Fare
df['Fare_outlier'] = (df['Fare'] < lower) | (df['Fare'] > upper)
print(df['Fare_outlier'].value_counts())
Fare_outlier
False    775
True     116
Name: count, dtype: int64
# Advanced: Apply a custom function to validate cabin format
def is_valid_cabin(x):
    if pd.isnull(x):
        return True
    return str(x)[0].isalpha()
df['Cabin_valid'] = df['Cabin'].apply(is_valid_cabin)
print(df['Cabin_valid'].value_counts())
Cabin_valid
True    891
Name: count, dtype: int64
# Practice: Type the number of Fare outliers you found earlier
user_input = input('How many Fare outliers did we find before removal? ')
answer = str(outliers.shape[0])
if user_input == answer:
    print('Great job!')
else:
    print('Please check the earlier output and try again.')
    
Please check the earlier output and try again.

Mini-Project: Outlier Handling and Data Validation EDA#

Let us do a quick End-to-End Analysis:

  1. Find passengers with very high fares and list their details.
  2. Fill all missing ages, and check if outliers exist in Parch (parents/children).
  3. List any passengers with negative or zero values in numeric fields.

Try this as a small personal project after the video.

# Troubleshooting tip: What if you see a KeyError?
# This often happens if you mistype a column name.
# You can double check column names using df.columns.
# Best practice: Use .copy() when subsetting
subset = df[['Age', 'Fare']].copy()
print(subset.head(2))
    Age     Fare
0  22.0   7.2500
1  38.0  71.2833
# Speed tip: For very large datasets, use .astype to reduce memory use
df['Pclass'] = df['Pclass'].astype('int8')
print(df['Pclass'].dtype)
int8

Lesson Recap#

We explored outlier detection and data validation in pandas using the Titanic dataset.

You learned to spot, visualize, and handle extreme or invalid values, fill in missing entries, troubleshoot errors, and keep your analysis solid.

Remember, clean data is the foundation for meaningful results.

Keep practicing and stay curious!

If you learned something new today, please like and subscribe for more lessons!

Comment below with your favorite pandas tip or any problems you would like to see solved.

Happy coding!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.