Mathew K Analytics

Lesson 2 · Data Mining

Data Mining vs Data Science vs Machine Learning: Key Differences Explained

Welcome to your Python Data Mining journey! In this lesson for absolute beginners, we will explore what data mining is, how it relates to data science and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 12: Data Mining vs Data Science vs Machine Learning#

Welcome to your Python Data Mining journey! In this lesson for absolute beginners, we will explore what data mining is, how it relates to data science and machine learning, and start hands-on practical work using Python. No experience neededlet's get started!

What is Data Mining?#

  • Data mining means finding patterns and insights from large data sets.
  • It is a key part of data science, which also includes cleaning, storing, and communicating data.
  • Data mining uses machine learning as well as statistics and basic programming.

Examples#

  • Predicting which passengers survived on the Titanic by analyzing ticket and age data.
  • Grouping grocery purchases to find items often bought together.
# Let us import useful modules and suppress warnings.
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
 
 
# Data setup (Iris Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(150, 5)
   sepal_length  sepal_width  petal_length  petal_width species
0           5.1          3.5           1.4          0.2  setosa
1           4.9          3.0           1.4          0.2  setosa
2           4.7          3.2           1.3          0.2  setosa

The Iris Dataset#

The Iris dataset contains flower measurements from three iris species.

  • Each row is a sample of one flower.
  • Columns include sepal length, sepal width, petal length, and petal width.
  • The last column tells us the species name.

You will see this dataset used often to explain data mining basics.

# Checking for missing data
print(df.isnull().sum())
sepal_length    0
sepal_width     0
petal_length    0
petal_width     0
species         0
dtype: int64
# Quick data cleaning - remove duplicates
df_clean = df.drop_duplicates()
print('Rows after removing duplicates:', df_clean.shape[0])
Rows after removing duplicates: 147
# Exploratory data analysis: summary statistics
print(df_clean.describe())
       sepal_length  sepal_width  petal_length  petal_width
count    147.000000   147.000000    147.000000   147.000000
mean       5.856463     3.055782      3.780272     1.208844
std        0.829100     0.437009      1.759111     0.757874
min        4.300000     2.000000      1.000000     0.100000
25%        5.100000     2.800000      1.600000     0.300000
50%        5.800000     3.000000      4.400000     1.300000
75%        6.400000     3.300000      5.100000     1.800000
max        7.900000     4.400000      6.900000     2.500000
# Visualizing flower features: pairplot
import seaborn as sns
sns.pairplot(df_clean, hue='species')
<seaborn.axisgrid.PairGrid at 0x1897fe3f4d0>
No description has been provided for this image
# Data setup (Groceries Dataset)
import kagglehub, os
path = kagglehub.dataset_download('heeraldedhia/groceries-dataset')
files = os.listdir(path)
csv_file = [f for f in files if f.endswith('.csv')][0]
df_groc = pd.read_csv(os.path.join(path, csv_file))
print(df_groc.shape)
print(df_groc.head(3))
(38765, 3)
   Member_number        Date itemDescription
0           1808  21-07-2015  tropical fruit
1           2552  05-01-2015      whole milk
2           2300  19-09-2015       pip fruit
# Convert transactions into baskets for association rule mining
transactions = df_groc.groupby('Member_number')['itemDescription'].apply(list).values.tolist()
print('First basket:', transactions[0])
First basket: ['soda', 'canned beer', 'sausage', 'sausage', 'whole milk', 'whole milk', 'pickled vegetables', 'misc. beverages', 'semi-finished bread', 'hygiene articles', 'yogurt', 'pastry', 'salty snack']
# Find frequent itemsets using apriori (mlxtend)
from mlxtend.preprocessing import TransactionEncoder
from mlxtend.frequent_patterns import apriori
te = TransactionEncoder()
te_ary = te.fit(transactions).transform(transactions)
df_trans = pd.DataFrame(te_ary, columns=te.columns_)
freq_items = apriori(df_trans, min_support=0.01, use_colnames=True)
print(freq_items.head(3))
    support                 itemsets
0  0.015393  (Instant food products)
1  0.078502               (UHT-milk)
2  0.031042          (baking powder)
# Classification demo: Predict species with logistic regression
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
X = df_clean.drop('species', axis=1)
y = df_clean['species']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
model = LogisticRegression(max_iter=200)
model.fit(X_train, y_train)
score = model.score(X_test, y_test)
print('Classification accuracy:', round(score, 2))
Classification accuracy: 0.93
# Clustering demo: Find natural groups in Iris data (KMeans)
from sklearn.cluster import KMeans
X_nospecies = df_clean.drop('species', axis=1)
kmeans = KMeans(n_clusters=3, random_state=42)
clusters = kmeans.fit_predict(X_nospecies)
df_clean['cluster'] = clusters
print(df_clean[['species', 'cluster']].head(6))
  species  cluster
0  setosa        2
1  setosa        2
2  setosa        2
3  setosa        2
4  setosa        2
5  setosa        2
# Ensemble learning: combine multiple classifiers (Random Forest)
from sklearn.ensemble import RandomForestClassifier
forest = RandomForestClassifier(n_estimators=50, random_state=42)
forest.fit(X_train, y_train)
rf_score = forest.score(X_test, y_test)
print('Random Forest accuracy:', round(rf_score, 2))
Random Forest accuracy: 0.96
# Anomaly detection: find outlier flower samples (IsolationForest)
from sklearn.ensemble import IsolationForest
iso = IsolationForest(contamination=0.05, random_state=42)
anoms = iso.fit_predict(X_nospecies)
outliers = df_clean[anoms == -1]
print('Number of anomalies found:', len(outliers))
print(outliers.head(2))
Number of anomalies found: 8
    sepal_length  sepal_width  petal_length  petal_width species  cluster
13           4.3          3.0           1.1          0.1  setosa        2
14           5.8          4.0           1.2          0.2  setosa        2
# Time series basics: plot number of each species
counts = df_clean['species'].value_counts().sort_index()
counts.plot(kind='bar', title='Number of Samples per Iris Species');
No description has been provided for this image
# Mini-project 1: Market basket analysis - find association rules
from mlxtend.frequent_patterns import association_rules
rules = association_rules(freq_items, metric='confidence', min_threshold=0.3)
print(rules[['antecedents','consequents','support','confidence']].head(3))
  antecedents         consequents   support  confidence
0  (UHT-milk)  (other vegetables)  0.038994    0.496732
1  (UHT-milk)        (rolls/buns)  0.031042    0.395425
2  (UHT-milk)              (soda)  0.027450    0.349673
# Mini-project 2: Ask the user for a sample and predict the species
sample = input("Enter four measurements separated by commas (sepal_length,sepal_width,petal_length,petal_width): ")
parts = [float(x.strip()) for x in sample.split(",")]
pred_species = model.predict([parts])[0]
print("Predicted iris species:", pred_species)
Predicted iris species: setosa
# Troubleshooting checklist for data mining in Python
print("Check these steps if something seems off:")
print("1. Is your dataset loaded and in the right shape?")
print("2. Are there missing or duplicate values?")
print("3. Are your input columns all numbers for modeling?")
print("4. Is your target variable clear and correct?")
print("5. Do you get the same results if you set the random state?")
Check these steps if something seems off:
1. Is your dataset loaded and in the right shape?
2. Are there missing or duplicate values?
3. Are your input columns all numbers for modeling?
4. Is your target variable clear and correct?
5. Do you get the same results if you set the random state?
# Extra: Show top 5 most frequent items in the groceries data
top_items = df_groc['itemDescription'].value_counts().head(5)
print(top_items)
itemDescription
whole milk          2502
other vegetables    1898
rolls/buns          1716
soda                1514
yogurt              1334
Name: count, dtype: int64
# Challenge: Explore your own idea!
print("Try analyzing other columns, find rare items, or test another classifier.")
Try analyzing other columns, find rare items, or test another classifier.

Recap and Next Steps#

  • You just did real-world data mining in Python, even if you are a beginner!
  • You learned to clean data, find patterns, make models, and check results.
  • Next, try more datasets and challenges.

Happy mining!

Thanks for following along!#

If you enjoyed this beginner mini-course, please like, subscribe, and tell us what you want to learn next on YouTube. See you in the next lesson!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.