Lesson 5 · Data Mining
Fundamentals of Exploratory Data Analysis and Visualization for Biblical Data Insights
In this lesson, we will explore how to load data, clean it, and visualize it using Python. We will use the Iris, Titanic, and Groceries datasets. Let's get…
- CourseData Mining
- Lesson5 of 31
- Video13 min
- FormatJupyter notebook · 13 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWelcome to Beginner Data Mining in Python!#
In this lesson, we will explore how to load data, clean it, and visualize it using Python.
We will use the Iris, Titanic, and Groceries datasets.
Let's get started on our data adventure together!
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# This ignores warning messages for a cleaner notebook
Iris Dataset: Data Setup#
The Iris dataset contains measurements for different types of flowers.
We will use it to practice data loading and exploration.
# Data setup (Iris Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
What are the columns?#
Lets print out the list of columns to see what information we have.
print(df.columns.tolist())
# Checking for missing values
print(df.isnull().sum())
Simple Descriptive Statistics#
Exploratory Data Analysis (EDA) helps us learn about the data.
Lets start with some summary statistics.
print(df.describe())
# Visualizing petal length distribution
import matplotlib.pyplot as plt
plt.hist(df['petal_length'], bins=20, color='skyblue', edgecolor='black')
plt.title('Distribution of Petal Lengths')
plt.xlabel('Petal Length (cm)')
plt.ylabel('Count')
plt.show()
# Scatter plot: Sepal length vs petal length
plt.scatter(df['sepal_length'], df['petal_length'], c='green')
plt.title('Sepal Length vs Petal Length')
plt.xlabel('Sepal Length (cm)')
plt.ylabel('Petal Length (cm)')
plt.show()
Next Dataset: Titanic Survival#
The Titanic dataset helps us practice working with mixed types of data.
Lets load it and get a quick first look.
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
# Look at gender counts
print(df['Sex'].value_counts())
# Survival rate
print(df['Survived'].value_counts(normalize=True))
# Visualizing survival by class
import seaborn as sns
sns.countplot(x='Pclass', hue='Survived', data=df)
plt.title('Survival Count by Passenger Class')
plt.xlabel('Passenger Class')
plt.ylabel('Count')
plt.legend(['Did not survive', 'Survived'])
plt.show()
Association Rule Mining: Groceries Dataset#
This dataset tracks what shoppers buy together.
Lets load it and create baskets of items for each shopper.
# Data setup (Groceries Dataset)
import pandas as pd
import kagglehub, os
path = kagglehub.dataset_download('heeraldedhia/groceries-dataset')
files = os.listdir(path)
csv_file = [f for f in files if f.endswith('.csv')][0]
df = pd.read_csv(os.path.join(path, csv_file))
print(df.shape)
print(df.head(3))
# Group items into baskets for each customer
transactions = df.groupby('Member_number')['itemDescription'].apply(list).values.tolist()
print('Number of baskets:', len(transactions))
print('Example basket:', transactions[0][:5])
Recap: You Did It!#
You have learned how to load, clean, and explore three popular datasets.
We covered summary stats, visualizations, and set the stage for more complex analysis.
Keep practicing these basicsthey are your foundation in data mining.
Keep Learning and Subscribe#
If you found this lesson helpful, please like and subscribe for more beginner-friendly data science tutorials!
Try rerunning parts of the code with your own changes to deepen your understanding.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



