Lesson 4 · Data Mining
Introduction to Python, Pandas, and scikit-learn for Data Analysis and Machine Learning
This lesson is for absolute beginners. You will learn to load, clean, and explore data. We will cover association rules, classification, clustering, and…
- CourseData Mining
- Lesson4 of 31
- Video26 min
- FormatJupyter notebook · 23 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWelcome to Data Mining with Python!#
This lesson is for absolute beginners.
You will learn to load, clean, and explore data.
We will cover association rules, classification, clustering, and more.
We will use Python libraries like Pandas and scikit-learn.
Do not worry if something is new.
Let us start your journey into data mining!
# Suppress warnings to keep outputs clean
import warnings; warnings.filterwarnings('ignore')
import numpy as np
np.random.seed(42)
# Data setup (Iris Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
What is Data Cleaning?#
Data is not always ready to use. It can have missing values or errors. Cleaning means fixing or removing bad parts of data.
# Check for missing values
print(df.isnull().sum())
# Fill any missing values just in case
df.fillna(method='ffill', inplace=True)
print('Missing values after fill:', df.isnull().sum().sum())
# Remove duplicate rows if any
print('Shape before dropping duplicates:', df.shape)
df.drop_duplicates(inplace=True)
print('Shape after:', df.shape)
Exploratory Data Analysis#
Exploratory Data Analysis (EDA) finds patterns, trends, and surprises in data.
We use simple stats and plots to understand what data shows before applying models.
# Show some descriptive statistics
print(df.describe())
# Simple visualization: scatter plot
import matplotlib.pyplot as plt
plt.scatter(df['sepal_length'], df['petal_length'], c='green', alpha=0.5)
plt.xlabel('Sepal Length')
plt.ylabel('Petal Length')
plt.title('Scatter plot of Sepal and Petal Length')
plt.show()
Association Rule Mining#
Association rules uncover items that appear together in data. Stores use these rules to learn which items are bought together. Let us try it using a small version of a groceries dataset.
# Data setup (Groceries Dataset)
import pandas as pd
import kagglehub, os
path = kagglehub.dataset_download('heeraldedhia/groceries-dataset')
files = os.listdir(path)
csv_file = [f for f in files if f.endswith('.csv')][0]
df_groc = pd.read_csv(os.path.join(path, csv_file))
print(df_groc.shape)
print(df_groc.head(3))
# Create 'baskets' of items bought by each shopper
basket = df_groc.groupby(['Member_number'])['itemDescription'].apply(list).reset_index()
print(basket.head(3))
# Convert baskets from lists to a one-hot format
from mlxtend.preprocessing import TransactionEncoder
te = TransactionEncoder()
te_ary = te.fit_transform(basket['itemDescription'])
df_te = pd.DataFrame(te_ary, columns=te.columns_)
print(df_te.shape)
# Apply the Apriori algorithm for frequent itemsets
from mlxtend.frequent_patterns import apriori, association_rules
freq_items = apriori(df_te, min_support=0.01, use_colnames=True)
print(freq_items.head())
# Mine association rules from itemsets
rules = association_rules(freq_items, metric='lift', min_threshold=1.1)
rules = rules[['antecedents', 'consequents', 'support', 'confidence', 'lift']].sort_values('lift', ascending=False)
print(rules.head(5))
Classification Basics#
Classification puts data into groups or classes. For example: Predicting if a flower is setosa, versicolor, or virginica.
Let us train a simple decision tree on the Iris data.
# Prepare the data for classification
from sklearn.model_selection import train_test_split
X = df.drop('species', axis=1)
y = df['species']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Train a simple decision tree classifier
from sklearn.tree import DecisionTreeClassifier
tree = DecisionTreeClassifier(random_state=42)
tree.fit(X_train, y_train)
score = tree.score(X_test, y_test)
print('Test Accuracy:', score)
Clustering#
Clustering groups data by similarity. Unlike classification, classes are not given in advance. A real-world example is grouping customers by buying habits.
# Use KMeans clustering on two Iris columns
from sklearn.cluster import KMeans
X = df[['sepal_length', 'petal_length']]
kmeans = KMeans(n_clusters=3, random_state=42)
clusters = kmeans.fit_predict(X)
plt.scatter(X['sepal_length'], X['petal_length'], c=clusters, cmap='plasma')
plt.xlabel('Sepal Length')
plt.ylabel('Petal Length')
plt.title('KMeans Clusters on Iris Data')
plt.show()
Ensemble Learning#
Ensembles combine several models to improve predictions. Voting and Random Forests are examples.
Ensemble methods are strong in real-world tasks.
# Try a RandomForest ensemble classifier
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(X_train, y_train)
print('Random Forest Test Accuracy:', rf.score(X_test, y_test))
Anomaly Detection#
Anomaly detection spots data points that are rare or different. This is used to find fraud, errors, or unique behaviors. Let us try a simple one.
# Use IsolationForest for anomaly detection
from sklearn.ensemble import IsolationForest
iso = IsolationForest(random_state=42)
outliers = iso.fit_predict(X)
print('Anomalies found:', list(outliers).count(-1))
Introduction to Time-Series Data#
A time series is data measured over time.
Examples include stock prices, temperature changes, and website visits.
# Simulate a small time series
import numpy as np
dates = pd.date_range('2024-01-01', periods=10)
values = np.random.randint(10, 100, size=10)
ts = pd.Series(values, index=dates)
import matplotlib.pyplot as plt
ts.plot(marker='o')
plt.title('Example Time Series')
plt.ylabel('Value')
plt.show()
Mini-Project: Market Basket Analysis#
You are a store manager. You want to find out which items shoppers buy together.
Use association rules to help marketing or optimize layout!
# Find pairs of items that are frequently bought together
pair_rules = association_rules(freq_items, metric='confidence', min_threshold=0.3)
pair_rules = pair_rules[pair_rules['antecedents'].apply(lambda x: len(x) == 1)]
pair_rules = pair_rules.sort_values('confidence', ascending=False).head(5)
print(pair_rules[['antecedents', 'consequents', 'support', 'confidence']])
Mini-Project: Classification With Titanic Data#
Now, let us use the Titanic dataset to predict who survived.
Try a simple logistic regression model.
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df_titanic = pd.read_csv(url)
print(df_titanic.shape)
print(df_titanic.head(3))
# Prepare Titanic features for modeling
df_titanic['Sex'] = df_titanic['Sex'].map({'male': 0, 'female': 1})
df_titanic['Age'].fillna(df_titanic['Age'].median(), inplace=True)
X = df_titanic[['Pclass', 'Sex', 'Age']]
y = df_titanic['Survived']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Train logistic regression on Titanic
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(max_iter=1000, random_state=42)
model.fit(X_train, y_train)
print('Titanic Test Accuracy:', model.score(X_test, y_test))
Best Practices and Troubleshooting#
- Always preview and check missing values.
- Plot your data before applying models.
- Try more than one model or method.
- Tune parameters and try again if results look odd.
These habits save time and improve accuracy.
# Ask the user a question to reflect
response = input('What was the most surprising thing you learned about your data? ')
print('Thanks for sharing! Each discovery helps build data intuition.')
Extra Tips#
- Use random_state so results can be repeated.
- Keep code simple and build up step by step.
- Check what every line prints; output often helps spot issues.
Coding is like learning a new language. It gets easier with practice!
Challenge Exercise#
Pick a column in the Titanic or Iris dataset. Plot its distribution using matplotlib.
Can you spot any patterns?
Try one new classifier (for example, k-nearest neighbors) on Iris.
Recap#
- You learned to load and clean data.
- You practiced exploring, visualizing, and modeling.
- You tried association rules, classification, and clustering.
- You saw how time series and anomalies are handled.
- You completed two mini-projects.
Practice is key come back and build new projects soon!
Thanks for joining!#
If you enjoyed this lesson, leave a like or comment and subscribe for more beginner-friendly Python data mining videos.
Happy coding!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



