Mathew K Analytics

Lesson 1 · Data Mining

Understanding Data Mining and the Knowledge Discovery in Databases (KDD) Process

In this lesson, we will explore the basics of data mining and the KDD process. Get ready to work hands-on with real-world datasets such as Iris, Titanic,…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Welcome to Data Mining with Python!#

In this lesson, we will explore the basics of data mining and the KDD process.

Get ready to work hands-on with real-world datasets such as Iris, Titanic, and Groceries.

By learning step by step, you will build confidence and skills for your own projects!

What is Data Mining?#

Data mining is discovering useful patterns and knowledge from large amounts of data.

It is a key step in the KDD (Knowledge Discovery in Databases) process.

Common tasks in data mining:

  • Cleaning and preparing data
  • Finding patterns (like trends or associations)
  • Predicting future outcomes
  • Detecting unusual data points (anomalies)

The Datasets We'll Use#

Today, we will use these datasets:

  • Iris (classifying flower types)
  • Titanic (predicting survival)
  • Groceries (finding shopping patterns)
# Suppress warnings for a cleaner notebook
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# Data setup (Iris Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(150, 5)
   sepal_length  sepal_width  petal_length  petal_width species
0           5.1          3.5           1.4          0.2  setosa
1           4.9          3.0           1.4          0.2  setosa
2           4.7          3.2           1.3          0.2  setosa
# Basic data exploration
print(df.info())
print(df.describe())
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 150 entries, 0 to 149
Data columns (total 5 columns):
 #   Column        Non-Null Count  Dtype  
---  ------        --------------  -----  
 0   sepal_length  150 non-null    float64
 1   sepal_width   150 non-null    float64
 2   petal_length  150 non-null    float64
 3   petal_width   150 non-null    float64
 4   species       150 non-null    object 
dtypes: float64(4), object(1)
memory usage: 6.0+ KB
None
       sepal_length  sepal_width  petal_length  petal_width
count    150.000000   150.000000    150.000000   150.000000
mean       5.843333     3.054000      3.758667     1.198667
std        0.828066     0.433594      1.764420     0.763161
min        4.300000     2.000000      1.000000     0.100000
25%        5.100000     2.800000      1.600000     0.300000
50%        5.800000     3.000000      4.350000     1.300000
75%        6.400000     3.300000      5.100000     1.800000
max        7.900000     4.400000      6.900000     2.500000
# Checking for missing values
print(df.isnull().sum())
sepal_length    0
sepal_width     0
petal_length    0
petal_width     0
species         0
dtype: int64
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
titanic = pd.read_csv(url)
print(titanic.shape)
print(titanic.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  
# Dropping rows with missing values (Titanic)
before = titanic.shape[0]
titanic = titanic.dropna()
after = titanic.shape[0]
print(f"Rows before: {before}, Rows after cleaning: {after}")
Rows before: 891, Rows after cleaning: 183
# Data setup (Groceries Dataset)
import pandas as pd
import kagglehub, os
path = kagglehub.dataset_download('heeraldedhia/groceries-dataset')
files = os.listdir(path)
csv_file = [f for f in files if f.endswith('.csv')][0]
df_groc = pd.read_csv(os.path.join(path, csv_file))
print(df_groc.shape)
print(df_groc.head(3))
(38765, 3)
   Member_number        Date itemDescription
0           1808  21-07-2015  tropical fruit
1           2552  05-01-2015      whole milk
2           2300  19-09-2015       pip fruit
# Grouping grocery transactions into baskets
baskets = df_groc.groupby('Member_number')['itemDescription'].apply(list).values.tolist()
print(baskets[:2])
[['soda', 'canned beer', 'sausage', 'sausage', 'whole milk', 'whole milk', 'pickled vegetables', 'misc. beverages', 'semi-finished bread', 'hygiene articles', 'yogurt', 'pastry', 'salty snack'], ['frankfurter', 'frankfurter', 'beef', 'sausage', 'whole milk', 'soda', 'curd', 'white bread', 'whole milk', 'soda', 'whipped/sour cream', 'rolls/buns']]
# Exploratory data analysis: Iris - simple scatterplot
import seaborn as sns
import matplotlib.pyplot as plt
sns.scatterplot(data=df, x='sepal_length', y='sepal_width', hue='species')
plt.title('Sepal Size by Species')
plt.show()
No description has been provided for this image
# Association rule mining: find frequent itemsets in groceries
from mlxtend.preprocessing import TransactionEncoder
from mlxtend.frequent_patterns import apriori, association_rules
te = TransactionEncoder()
te_ary = te.fit(baskets).transform(baskets)
df_basket = pd.DataFrame(te_ary, columns=te.columns_)
frequent = apriori(df_basket, min_support=0.01, use_colnames=True)
rules = association_rules(frequent, metric='lift', min_threshold=1.0)
print(rules[['antecedents','consequents','support','confidence','lift']].head())
      antecedents      consequents   support  confidence      lift
0          (beef)       (UHT-milk)  0.010518    0.087983  1.120775
1      (UHT-milk)           (beef)  0.010518    0.133987  1.120775
2      (UHT-milk)   (bottled beer)  0.014879    0.189542  1.193597
3  (bottled beer)       (UHT-milk)  0.014879    0.093700  1.193597
4      (UHT-milk)  (bottled water)  0.021293    0.271242  1.269268
# Simple classification: Iris with DecisionTreeClassifier
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
X = df.drop('species', axis=1)
y = df['species']
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42, test_size=0.3)
clf = DecisionTreeClassifier(random_state=42)
clf.fit(X_train, y_train)
print(f"Test accuracy: {clf.score(X_test, y_test):.2f}")
Test accuracy: 1.00
# Clustering: KMeans on Iris features
from sklearn.cluster import KMeans
kmeans = KMeans(n_clusters=3, random_state=42)
labels = kmeans.fit_predict(X)
print(labels[:10])
[1 1 1 1 1 1 1 1 1 1]
# Ensemble learning: Random Forest on Titanic survival
from sklearn.ensemble import RandomForestClassifier
features = ['Pclass','Sex','Age','Fare']
X = titanic[features]
y = titanic['Survived']
X = pd.get_dummies(X)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42, test_size=0.3)
rf = RandomForestClassifier(random_state=42)
rf.fit(X_train, y_train)
print(f"Titanic Test accuracy: {rf.score(X_test, y_test):.2f}")
Titanic Test accuracy: 0.73
# Anomaly detection: Find outliers in Titanic fares
import numpy as np
fare = titanic['Fare']
q1, q3 = np.percentile(fare, [25, 75])
iqr = q3 - q1
lower = q1 - 1.5 * iqr
upper = q3 + 1.5 * iqr
outliers = titanic[(fare < lower) | (fare > upper)]
print(outliers[['Fare','Survived']].head())
         Fare  Survived
27   263.0000         0
88   263.0000         1
118  247.5208         0
299  247.5208         1
311  262.3750         1
# Time series basics: Simple line plot of purchase counts over days
df_groc['Date'] = pd.to_datetime(df_groc['Date'])
daily_counts = df_groc.groupby('Date').size()
daily_counts.plot(title='Transactions per day')
plt.ylabel('Number of Transactions')
plt.show()
No description has been provided for this image
# Mini-project part 1: Find strong item pairs in groceries
rules_pairs = rules[(rules['antecedents'].apply(lambda x: len(x)==1)) & (rules['consequents'].apply(lambda x: len(x)==1))]
print(rules_pairs[['antecedents','consequents','support','confidence','lift']].sort_values('lift', ascending=False).head())
          antecedents    consequents   support  confidence      lift
213     (white bread)    (beverages)  0.010518    0.118497  1.908685
212       (beverages)  (white bread)  0.010518    0.169421  1.908685
777         (chicken)      (waffles)  0.011801    0.117347  1.700440
776         (waffles)      (chicken)  0.011801    0.171004  1.700440
1642  (specialty bar)   (newspapers)  0.012314    0.234146  1.674683
# Mini-project part 2: Titanic - predict survival from user input
print("Predict if you would survive the Titanic. Enter your details:")
pclass = int(input("Passenger class (1, 2, 3): "))
sex = input("Sex (male/female): ")
age = float(input("Age: "))
fare = float(input("Fare paid: "))
row = pd.DataFrame({'Pclass':[pclass],'Sex':[sex],'Age':[age],'Fare':[fare]})
row = pd.get_dummies(row).reindex(columns=X.columns, fill_value=0)
prediction = rf.predict(row)[0]
print("You would have {}survived.".format("" if prediction==1 else "not "))
Predict if you would survive the Titanic. Enter your details:
You would have survived.

Best Practices and Troubleshooting#

  1. Always check for missing data before modeling.
  2. Visualize data early to spot patterns and odd values.
  3. Try different models, not just one.
  4. Use train/test split to avoid overfitting.
  5. Document your steps for reproducibility.

If you get an error, read the message. Most issues are due to data format mismatches, missing columns, or typos!

# Extra tip: View feature importance
importances = rf.feature_importances_
for name, value in zip(X.columns, importances):
    print(f"{name}: {value:.2f}")
    
Pclass: 0.02
Age: 0.38
Fare: 0.29
Sex_female: 0.15
Sex_male: 0.16
# Challenge: Try clustering Titanic passengers by fare and age
from sklearn.preprocessing import StandardScaler
X_titanic = titanic[['Fare','Age']]
X_scaled = StandardScaler().fit_transform(X_titanic)
kmeans = KMeans(n_clusters=3, random_state=42)
titanic_labels = kmeans.fit_predict(X_scaled)
titanic_sample = titanic[['Fare','Age']].copy()
titanic_sample['Cluster'] = titanic_labels
print(titanic_sample.head())
       Fare   Age  Cluster
1   71.2833  38.0        2
3   53.1000  35.0        1
6   51.8625  54.0        2
10  16.7000   4.0        1
11  26.5500  58.0        2

Lesson Recap: What did you learn?#

  • How to load and explore data from real sources
  • Basic cleaning and visualization
  • Association rules and frequent itemsets
  • Simple models for prediction and clustering
  • How to try interactive predictions

You have built a complete mini data mining workflow in Python!

Keep Practicing and Subscribe!#

For more hands-on guides, subscribe to our YouTube channel or revisit this notebook anytime.

Happy mining!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.