Lesson 1 · Data Mining
Understanding Data Mining and the Knowledge Discovery in Databases (KDD) Process
In this lesson, we will explore the basics of data mining and the KDD process. Get ready to work hands-on with real-world datasets such as Iris, Titanic,…
- CourseData Mining
- Lesson1 of 31
- Video22 min
- FormatJupyter notebook · 19 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWelcome to Data Mining with Python!#
In this lesson, we will explore the basics of data mining and the KDD process.
Get ready to work hands-on with real-world datasets such as Iris, Titanic, and Groceries.
By learning step by step, you will build confidence and skills for your own projects!
What is Data Mining?#
Data mining is discovering useful patterns and knowledge from large amounts of data.
It is a key step in the KDD (Knowledge Discovery in Databases) process.
Common tasks in data mining:
- Cleaning and preparing data
- Finding patterns (like trends or associations)
- Predicting future outcomes
- Detecting unusual data points (anomalies)
The Datasets We'll Use#
Today, we will use these datasets:
- Iris (classifying flower types)
- Titanic (predicting survival)
- Groceries (finding shopping patterns)
# Suppress warnings for a cleaner notebook
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# Data setup (Iris Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
# Basic data exploration
print(df.info())
print(df.describe())
# Checking for missing values
print(df.isnull().sum())
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
titanic = pd.read_csv(url)
print(titanic.shape)
print(titanic.head(3))
# Dropping rows with missing values (Titanic)
before = titanic.shape[0]
titanic = titanic.dropna()
after = titanic.shape[0]
print(f"Rows before: {before}, Rows after cleaning: {after}")
# Data setup (Groceries Dataset)
import pandas as pd
import kagglehub, os
path = kagglehub.dataset_download('heeraldedhia/groceries-dataset')
files = os.listdir(path)
csv_file = [f for f in files if f.endswith('.csv')][0]
df_groc = pd.read_csv(os.path.join(path, csv_file))
print(df_groc.shape)
print(df_groc.head(3))
# Grouping grocery transactions into baskets
baskets = df_groc.groupby('Member_number')['itemDescription'].apply(list).values.tolist()
print(baskets[:2])
# Exploratory data analysis: Iris - simple scatterplot
import seaborn as sns
import matplotlib.pyplot as plt
sns.scatterplot(data=df, x='sepal_length', y='sepal_width', hue='species')
plt.title('Sepal Size by Species')
plt.show()
# Association rule mining: find frequent itemsets in groceries
from mlxtend.preprocessing import TransactionEncoder
from mlxtend.frequent_patterns import apriori, association_rules
te = TransactionEncoder()
te_ary = te.fit(baskets).transform(baskets)
df_basket = pd.DataFrame(te_ary, columns=te.columns_)
frequent = apriori(df_basket, min_support=0.01, use_colnames=True)
rules = association_rules(frequent, metric='lift', min_threshold=1.0)
print(rules[['antecedents','consequents','support','confidence','lift']].head())
# Simple classification: Iris with DecisionTreeClassifier
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
X = df.drop('species', axis=1)
y = df['species']
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42, test_size=0.3)
clf = DecisionTreeClassifier(random_state=42)
clf.fit(X_train, y_train)
print(f"Test accuracy: {clf.score(X_test, y_test):.2f}")
# Clustering: KMeans on Iris features
from sklearn.cluster import KMeans
kmeans = KMeans(n_clusters=3, random_state=42)
labels = kmeans.fit_predict(X)
print(labels[:10])
# Ensemble learning: Random Forest on Titanic survival
from sklearn.ensemble import RandomForestClassifier
features = ['Pclass','Sex','Age','Fare']
X = titanic[features]
y = titanic['Survived']
X = pd.get_dummies(X)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42, test_size=0.3)
rf = RandomForestClassifier(random_state=42)
rf.fit(X_train, y_train)
print(f"Titanic Test accuracy: {rf.score(X_test, y_test):.2f}")
# Anomaly detection: Find outliers in Titanic fares
import numpy as np
fare = titanic['Fare']
q1, q3 = np.percentile(fare, [25, 75])
iqr = q3 - q1
lower = q1 - 1.5 * iqr
upper = q3 + 1.5 * iqr
outliers = titanic[(fare < lower) | (fare > upper)]
print(outliers[['Fare','Survived']].head())
# Time series basics: Simple line plot of purchase counts over days
df_groc['Date'] = pd.to_datetime(df_groc['Date'])
daily_counts = df_groc.groupby('Date').size()
daily_counts.plot(title='Transactions per day')
plt.ylabel('Number of Transactions')
plt.show()
# Mini-project part 1: Find strong item pairs in groceries
rules_pairs = rules[(rules['antecedents'].apply(lambda x: len(x)==1)) & (rules['consequents'].apply(lambda x: len(x)==1))]
print(rules_pairs[['antecedents','consequents','support','confidence','lift']].sort_values('lift', ascending=False).head())
# Mini-project part 2: Titanic - predict survival from user input
print("Predict if you would survive the Titanic. Enter your details:")
pclass = int(input("Passenger class (1, 2, 3): "))
sex = input("Sex (male/female): ")
age = float(input("Age: "))
fare = float(input("Fare paid: "))
row = pd.DataFrame({'Pclass':[pclass],'Sex':[sex],'Age':[age],'Fare':[fare]})
row = pd.get_dummies(row).reindex(columns=X.columns, fill_value=0)
prediction = rf.predict(row)[0]
print("You would have {}survived.".format("" if prediction==1 else "not "))
Best Practices and Troubleshooting#
- Always check for missing data before modeling.
- Visualize data early to spot patterns and odd values.
- Try different models, not just one.
- Use train/test split to avoid overfitting.
- Document your steps for reproducibility.
If you get an error, read the message. Most issues are due to data format mismatches, missing columns, or typos!
# Extra tip: View feature importance
importances = rf.feature_importances_
for name, value in zip(X.columns, importances):
print(f"{name}: {value:.2f}")
# Challenge: Try clustering Titanic passengers by fare and age
from sklearn.preprocessing import StandardScaler
X_titanic = titanic[['Fare','Age']]
X_scaled = StandardScaler().fit_transform(X_titanic)
kmeans = KMeans(n_clusters=3, random_state=42)
titanic_labels = kmeans.fit_predict(X_scaled)
titanic_sample = titanic[['Fare','Age']].copy()
titanic_sample['Cluster'] = titanic_labels
print(titanic_sample.head())
Lesson Recap: What did you learn?#
- How to load and explore data from real sources
- Basic cleaning and visualization
- Association rules and frequent itemsets
- Simple models for prediction and clustering
- How to try interactive predictions
You have built a complete mini data mining workflow in Python!
Keep Practicing and Subscribe!#
For more hands-on guides, subscribe to our YouTube channel or revisit this notebook anytime.
Happy mining!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



