Lesson 15 · Mastering Pandas
Comprehensive Guide to String Operations and Text Cleaning with Pandas in Python
Let us explore how to clean and manipulate text data in pandas. You will learn to tidy messy strings, extract features, and handle missing or inconsistent…
- CourseMastering Pandas
- Lesson15 of 44
- Video14 min
- FormatJupyter notebook · 15 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbString Operations and Text Cleaning in Pandas#
Let us explore how to clean and manipulate text data in pandas.
You will learn to tidy messy strings, extract features, and handle missing or inconsistent text.
These skills are essential for data analysis or machine learning on real-world data.
We will use the Titanic dataset for hands-on practice.
import warnings; warnings.filterwarnings('ignore')
# Data setup (Titanic Dataset)
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
Why clean text data?#
Real-world data is rarely tidy.
Names may be inconsistent. Some fields are empty or filled with symbols or typos.
We must fix these to get meaningful results.
# Inspect text columns
print(df[['Name', 'Sex', 'Cabin']].head())
# Counting missing values in text columns
print(df[['Name', 'Sex', 'Cabin']].isnull().sum())
# Fill missing Cabin values with 'Unknown'
df['Cabin'] = df['Cabin'].fillna('Unknown')
print(df['Cabin'].unique()[:5])
# Convert names to all lowercase
df['Name_lower'] = df['Name'].str.lower()
print(df['Name_lower'].head())
# Remove leading and trailing spaces from 'Cabin'
df['Cabin_clean'] = df['Cabin'].str.strip()
print(df[['Cabin', 'Cabin_clean']].head())
# Replace all '/' with '-' in 'Cabin_clean'
df['Cabin_dash'] = df['Cabin_clean'].str.replace('/', '-', regex=False)
print(df[['Cabin_clean', 'Cabin_dash']].head())
# Extract the first letter from Cabin codes
df['Cabin_letter'] = df['Cabin_clean'].str[0]
print(df[['Cabin_clean', 'Cabin_letter']].head())
# Does Name contain 'Miss'? Mark as True or False
df['Is_Miss'] = df['Name'].str.contains('Miss')
print(df[['Name', 'Is_Miss']].head())
# Count how many unique titles appear in Name
import re
df['Title'] = df['Name'].str.extract(r',\s*([^\.]+)\.')
print(df['Title'].unique())
# Standardize rare titles as 'Other' in Title column
common_titles = ['Mr', 'Mrs', 'Miss', 'Master']
df['Title_clean'] = df['Title'].where(df['Title'].isin(common_titles), 'Other')
print(df['Title_clean'].value_counts())
# Replace any digits in names with blank ('')
df['Name_no_num'] = df['Name'].str.replace(r'\d+', '', regex=True)
print(df[['Name', 'Name_no_num']].head())
# Combine Last Name and Title as a new feature
df['Last_Title'] = df['Name'].str.split(',').str[0] + '_' + df['Title_clean']
print(df[['Last_Title']].head())
# Split 'Name' into 'Last' and 'Rest' columns
df[['Last', 'Rest']] = df['Name'].str.split(',', n=1, expand=True)
print(df[['Last', 'Rest']].head())
# Make every title uppercase
df['Title_upper'] = df['Title_clean'].str.upper()
print(df[['Title_clean', 'Title_upper']].drop_duplicates().head())
Recap: What you learned#
You tried common string cleaning and manipulation tools in pandas:
- Handling missing values
- Changing case
- Trimming whitespace
- Replacing and cleaning unwanted characters
- Extracting substrings and using regular expressions
- Grouping and engineering new features
These skills make your data trustworthy and ready for analysis!
Ready for more?#
Explore the pandas string methods documentation for even more powerful text tricks.
See you in the next video!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



