Mathew K Analytics

Lesson 5 · Data Mining

Fundamentals of Exploratory Data Analysis and Visualization for Biblical Data Insights

In this lesson, we will explore how to load data, clean it, and visualize it using Python. We will use the Iris, Titanic, and Groceries datasets. Let's get…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Welcome to Beginner Data Mining in Python!#

In this lesson, we will explore how to load data, clean it, and visualize it using Python.

We will use the Iris, Titanic, and Groceries datasets.

Let's get started on our data adventure together!

import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# This ignores warning messages for a cleaner notebook

Iris Dataset: Data Setup#

The Iris dataset contains measurements for different types of flowers.

We will use it to practice data loading and exploration.

# Data setup (Iris Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(150, 5)
   sepal_length  sepal_width  petal_length  petal_width species
0           5.1          3.5           1.4          0.2  setosa
1           4.9          3.0           1.4          0.2  setosa
2           4.7          3.2           1.3          0.2  setosa

What are the columns?#

Lets print out the list of columns to see what information we have.

print(df.columns.tolist())
['sepal_length', 'sepal_width', 'petal_length', 'petal_width', 'species']
# Checking for missing values
print(df.isnull().sum())
sepal_length    0
sepal_width     0
petal_length    0
petal_width     0
species         0
dtype: int64

Simple Descriptive Statistics#

Exploratory Data Analysis (EDA) helps us learn about the data.

Lets start with some summary statistics.

print(df.describe())
       sepal_length  sepal_width  petal_length  petal_width
count    150.000000   150.000000    150.000000   150.000000
mean       5.843333     3.054000      3.758667     1.198667
std        0.828066     0.433594      1.764420     0.763161
min        4.300000     2.000000      1.000000     0.100000
25%        5.100000     2.800000      1.600000     0.300000
50%        5.800000     3.000000      4.350000     1.300000
75%        6.400000     3.300000      5.100000     1.800000
max        7.900000     4.400000      6.900000     2.500000
# Visualizing petal length distribution
import matplotlib.pyplot as plt
plt.hist(df['petal_length'], bins=20, color='skyblue', edgecolor='black')
plt.title('Distribution of Petal Lengths')
plt.xlabel('Petal Length (cm)')
plt.ylabel('Count')
plt.show()
No description has been provided for this image
# Scatter plot: Sepal length vs petal length
plt.scatter(df['sepal_length'], df['petal_length'], c='green')
plt.title('Sepal Length vs Petal Length')
plt.xlabel('Sepal Length (cm)')
plt.ylabel('Petal Length (cm)')
plt.show()
No description has been provided for this image

Next Dataset: Titanic Survival#

The Titanic dataset helps us practice working with mixed types of data.

Lets load it and get a quick first look.

# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  
# Look at gender counts
print(df['Sex'].value_counts())
Sex
male      577
female    314
Name: count, dtype: int64
# Survival rate
print(df['Survived'].value_counts(normalize=True))
Survived
0    0.616162
1    0.383838
Name: proportion, dtype: float64
# Visualizing survival by class
import seaborn as sns
sns.countplot(x='Pclass', hue='Survived', data=df)
plt.title('Survival Count by Passenger Class')
plt.xlabel('Passenger Class')
plt.ylabel('Count')
plt.legend(['Did not survive', 'Survived'])
plt.show()
No description has been provided for this image

Association Rule Mining: Groceries Dataset#

This dataset tracks what shoppers buy together.

Lets load it and create baskets of items for each shopper.

# Data setup (Groceries Dataset)
import pandas as pd
import kagglehub, os
path = kagglehub.dataset_download('heeraldedhia/groceries-dataset')
files = os.listdir(path)
csv_file = [f for f in files if f.endswith('.csv')][0]
df = pd.read_csv(os.path.join(path, csv_file))
print(df.shape)
print(df.head(3))
(38765, 3)
   Member_number        Date itemDescription
0           1808  21-07-2015  tropical fruit
1           2552  05-01-2015      whole milk
2           2300  19-09-2015       pip fruit
# Group items into baskets for each customer
transactions = df.groupby('Member_number')['itemDescription'].apply(list).values.tolist()
print('Number of baskets:', len(transactions))
print('Example basket:', transactions[0][:5])
Number of baskets: 3898
Example basket: ['soda', 'canned beer', 'sausage', 'sausage', 'whole milk']

Recap: You Did It!#

You have learned how to load, clean, and explore three popular datasets.

We covered summary stats, visualizations, and set the stage for more complex analysis.

Keep practicing these basicsthey are your foundation in data mining.

Keep Learning and Subscribe#

If you found this lesson helpful, please like and subscribe for more beginner-friendly data science tutorials!

Try rerunning parts of the code with your own changes to deepen your understanding.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.