Mathew K Analytics

Lesson 6 · Data Mining

Data Cleaning Fundamentals: Managing Missing Values, Noise, and Outliers in Datasets

Welcome! This week we learn about making our data trustworthy. Dirty data can lead to bad predictions. Today we practice cleaning real-world datasets and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 3: Data Cleaning - Handling Missing Values, Noise, and Outliers#

Welcome! This week we learn about making our data trustworthy.

Dirty data can lead to bad predictions.

Today we practice cleaning real-world datasets and prepare them for data mining.

import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")  # Suppress warnings for a cleaner workspace

Data setup: Titanic Dataset#

Let us use the Titanic dataset. It is famous and full of interesting details.

We will work on cleaning this data step by step.

# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)  # rows and columns
print(df.head(3))  # a quick look at the data
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  

Step 1: Spotting Missing Values#

Missing values are blanks or NAs where information is lost.

Finding them is the first step to cleaning data.

# Check total and percentage of missing values per column
missing_counts = df.isnull().sum()
missing_percent = (missing_counts / len(df)) * 100
missing_summary = pd.DataFrame({'MissingCount': missing_counts, 'Percent': missing_percent.round(2)})
print(missing_summary[missing_summary.MissingCount > 0])
          MissingCount  Percent
Age                177    19.87
Cabin              687    77.10
Embarked             2     0.22
# Visualize missing values using a heatmap
import seaborn as sns
import matplotlib.pyplot as plt
plt.figure(figsize=(8, 4))
sns.heatmap(df.isnull(), cbar=False, cmap='viridis')
plt.title('Where is data missing?')
plt.show()
No description has been provided for this image

Step 2: Handling Missing Data#

Now let us fix missing values.

Different columns might need different solutions.

We will try several common strategies.

# Fill missing Age with median (safer for skewed data)
age_median = df['Age'].median()
df['Age'].fillna(age_median, inplace=True)
print(f"Filled missing Age with: {age_median}")
Filled missing Age with: 28.0
# Fill missing Embarked values with the most common (mode)
mode_embarked = df['Embarked'].mode()[0]
df['Embarked'].fillna(mode_embarked, inplace=True)
print(f"Filled missing Embarked with: {mode_embarked}")
Filled missing Embarked with: S
# Drop the Cabin column (too many missing values)
df.drop('Cabin', axis=1, inplace=True)
print('Cabin column dropped. Too much missing data.')
Cabin column dropped. Too much missing data.

Step 3: Handling Outliers#

Outliers are unusual values that can skew our analyses.

We will use charts and simple math to spot them.

# Boxplot to visualize outliers for Age and Fare
plt.figure(figsize=(10, 4))
plt.subplot(1, 2, 1)
sns.boxplot(x=df['Age'])
plt.title('Age Boxplot')
plt.subplot(1, 2, 2)
sns.boxplot(x=df['Fare'])
plt.title('Fare Boxplot')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Find outliers in Fare using the interquartile range (IQR) rule
Q1 = df['Fare'].quantile(0.25)
Q3 = df['Fare'].quantile(0.75)
IQR = Q3 - Q1
fare_lower = Q1 - 1.5 * IQR
fare_upper = Q3 + 1.5 * IQR
outliers = df[(df['Fare'] < fare_lower) | (df['Fare'] > fare_upper)]
print(f"Number of Fare outliers: {len(outliers)}")
Number of Fare outliers: 116

Step 4: Smoothing Noise#

Noise means random errors or small fluctuations.

We can use rounding or binning to smooth it out.

# Bin Age into categories (child, teen, adult, senior)
bins = [0, 12, 18, 59, 120]
labels = ['child', 'teen', 'adult', 'senior']
df['AgeGroup'] = pd.cut(df['Age'], bins=bins, labels=labels)
print(df[['Age', 'AgeGroup']].head(8))
    Age AgeGroup
0  22.0    adult
1  38.0    adult
2  26.0    adult
3  35.0    adult
4  35.0    adult
5  28.0    adult
6  54.0    adult
7   2.0    child
# Preview cleaned data
print(df.head(5))
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   
3            4         1       1   
4            5         0       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   
3       Futrelle, Mrs. Jacques Heath (Lily May Peel)  female  35.0      1   
4                           Allen, Mr. William Henry    male  35.0      0   

   Parch            Ticket     Fare Embarked AgeGroup  
0      0         A/5 21171   7.2500        S    adult  
1      0          PC 17599  71.2833        C    adult  
2      0  STON/O2. 3101282   7.9250        S    adult  
3      0            113803  53.1000        S    adult  
4      0            373450   8.0500        S    adult  

Mini-Practice: Fill a Missing Value by Hand#

Suppose we spot the very last missing value. Let us try filling it manually and see the effect.

# Practice: Enter a custom Age for a missing spot
df.loc[df.index[1], 'Age'] = np.nan
missing_index = df['Age'].isnull().idxmax() if df['Age'].isnull().sum() > 0 else None
if missing_index is not None:
    new_age = input("Enter a value to fill in the missing Age: ")
    df.loc[missing_index, 'Age'] = float(new_age)
    print(f"Age at index {missing_index} set to {new_age}")
else:
    print("No missing Age value left. Well done!")
    
Age at index 1 set to 27

Recap: You Cleaned Titanic Data!#

Great work today. You loaded a classic dataset, identified and fixed missing data, found outliers, and smoothed noisy details.

With these skills you are ready for any beginner data mining project.

Practice Challenge and Next Steps#

  • Try counting missing values in a different column.
  • Make a plot that shows outliers for SibSp or Parch.
  • Look for a public dataset and check for missing values using methods today.

Subscribe if you enjoyed this lesson! New data mining tips every week.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.