Lesson 6 · Data Mining
Data Cleaning Fundamentals: Managing Missing Values, Noise, and Outliers in Datasets
Welcome! This week we learn about making our data trustworthy. Dirty data can lead to bad predictions. Today we practice cleaning real-world datasets and…
- CourseData Mining
- Lesson6 of 31
- Video17 min
- FormatJupyter notebook · 12 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 3: Data Cleaning - Handling Missing Values, Noise, and Outliers#
Welcome! This week we learn about making our data trustworthy.
Dirty data can lead to bad predictions.
Today we practice cleaning real-world datasets and prepare them for data mining.
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore") # Suppress warnings for a cleaner workspace
Data setup: Titanic Dataset#
Let us use the Titanic dataset. It is famous and full of interesting details.
We will work on cleaning this data step by step.
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape) # rows and columns
print(df.head(3)) # a quick look at the data
Step 1: Spotting Missing Values#
Missing values are blanks or NAs where information is lost.
Finding them is the first step to cleaning data.
# Check total and percentage of missing values per column
missing_counts = df.isnull().sum()
missing_percent = (missing_counts / len(df)) * 100
missing_summary = pd.DataFrame({'MissingCount': missing_counts, 'Percent': missing_percent.round(2)})
print(missing_summary[missing_summary.MissingCount > 0])
# Visualize missing values using a heatmap
import seaborn as sns
import matplotlib.pyplot as plt
plt.figure(figsize=(8, 4))
sns.heatmap(df.isnull(), cbar=False, cmap='viridis')
plt.title('Where is data missing?')
plt.show()
Step 2: Handling Missing Data#
Now let us fix missing values.
Different columns might need different solutions.
We will try several common strategies.
# Fill missing Age with median (safer for skewed data)
age_median = df['Age'].median()
df['Age'].fillna(age_median, inplace=True)
print(f"Filled missing Age with: {age_median}")
# Fill missing Embarked values with the most common (mode)
mode_embarked = df['Embarked'].mode()[0]
df['Embarked'].fillna(mode_embarked, inplace=True)
print(f"Filled missing Embarked with: {mode_embarked}")
# Drop the Cabin column (too many missing values)
df.drop('Cabin', axis=1, inplace=True)
print('Cabin column dropped. Too much missing data.')
Step 3: Handling Outliers#
Outliers are unusual values that can skew our analyses.
We will use charts and simple math to spot them.
# Boxplot to visualize outliers for Age and Fare
plt.figure(figsize=(10, 4))
plt.subplot(1, 2, 1)
sns.boxplot(x=df['Age'])
plt.title('Age Boxplot')
plt.subplot(1, 2, 2)
sns.boxplot(x=df['Fare'])
plt.title('Fare Boxplot')
plt.tight_layout()
plt.show()
# Find outliers in Fare using the interquartile range (IQR) rule
Q1 = df['Fare'].quantile(0.25)
Q3 = df['Fare'].quantile(0.75)
IQR = Q3 - Q1
fare_lower = Q1 - 1.5 * IQR
fare_upper = Q3 + 1.5 * IQR
outliers = df[(df['Fare'] < fare_lower) | (df['Fare'] > fare_upper)]
print(f"Number of Fare outliers: {len(outliers)}")
Step 4: Smoothing Noise#
Noise means random errors or small fluctuations.
We can use rounding or binning to smooth it out.
# Bin Age into categories (child, teen, adult, senior)
bins = [0, 12, 18, 59, 120]
labels = ['child', 'teen', 'adult', 'senior']
df['AgeGroup'] = pd.cut(df['Age'], bins=bins, labels=labels)
print(df[['Age', 'AgeGroup']].head(8))
# Preview cleaned data
print(df.head(5))
Mini-Practice: Fill a Missing Value by Hand#
Suppose we spot the very last missing value. Let us try filling it manually and see the effect.
# Practice: Enter a custom Age for a missing spot
df.loc[df.index[1], 'Age'] = np.nan
missing_index = df['Age'].isnull().idxmax() if df['Age'].isnull().sum() > 0 else None
if missing_index is not None:
new_age = input("Enter a value to fill in the missing Age: ")
df.loc[missing_index, 'Age'] = float(new_age)
print(f"Age at index {missing_index} set to {new_age}")
else:
print("No missing Age value left. Well done!")
Recap: You Cleaned Titanic Data!#
Great work today. You loaded a classic dataset, identified and fixed missing data, found outliers, and smoothed noisy details.
With these skills you are ready for any beginner data mining project.
Practice Challenge and Next Steps#
- Try counting missing values in a different column.
- Make a plot that shows outliers for SibSp or Parch.
- Look for a public dataset and check for missing values using methods today.
Subscribe if you enjoyed this lesson! New data mining tips every week.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



