Lesson 2 · Data Mining
Data Mining vs Data Science vs Machine Learning: Key Differences Explained
Welcome to your Python Data Mining journey! In this lesson for absolute beginners, we will explore what data mining is, how it relates to data science and…
- CourseData Mining
- Lesson2 of 31
- Video22 min
- FormatJupyter notebook · 19 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 12: Data Mining vs Data Science vs Machine Learning#
Welcome to your Python Data Mining journey! In this lesson for absolute beginners, we will explore what data mining is, how it relates to data science and machine learning, and start hands-on practical work using Python. No experience neededlet's get started!
What is Data Mining?#
- Data mining means finding patterns and insights from large data sets.
- It is a key part of data science, which also includes cleaning, storing, and communicating data.
- Data mining uses machine learning as well as statistics and basic programming.
Examples#
- Predicting which passengers survived on the Titanic by analyzing ticket and age data.
- Grouping grocery purchases to find items often bought together.
# Let us import useful modules and suppress warnings.
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# Data setup (Iris Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
The Iris Dataset#
The Iris dataset contains flower measurements from three iris species.
- Each row is a sample of one flower.
- Columns include sepal length, sepal width, petal length, and petal width.
- The last column tells us the species name.
You will see this dataset used often to explain data mining basics.
# Checking for missing data
print(df.isnull().sum())
# Quick data cleaning - remove duplicates
df_clean = df.drop_duplicates()
print('Rows after removing duplicates:', df_clean.shape[0])
# Exploratory data analysis: summary statistics
print(df_clean.describe())
# Visualizing flower features: pairplot
import seaborn as sns
sns.pairplot(df_clean, hue='species')
# Data setup (Groceries Dataset)
import kagglehub, os
path = kagglehub.dataset_download('heeraldedhia/groceries-dataset')
files = os.listdir(path)
csv_file = [f for f in files if f.endswith('.csv')][0]
df_groc = pd.read_csv(os.path.join(path, csv_file))
print(df_groc.shape)
print(df_groc.head(3))
# Convert transactions into baskets for association rule mining
transactions = df_groc.groupby('Member_number')['itemDescription'].apply(list).values.tolist()
print('First basket:', transactions[0])
# Find frequent itemsets using apriori (mlxtend)
from mlxtend.preprocessing import TransactionEncoder
from mlxtend.frequent_patterns import apriori
te = TransactionEncoder()
te_ary = te.fit(transactions).transform(transactions)
df_trans = pd.DataFrame(te_ary, columns=te.columns_)
freq_items = apriori(df_trans, min_support=0.01, use_colnames=True)
print(freq_items.head(3))
# Classification demo: Predict species with logistic regression
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
X = df_clean.drop('species', axis=1)
y = df_clean['species']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
model = LogisticRegression(max_iter=200)
model.fit(X_train, y_train)
score = model.score(X_test, y_test)
print('Classification accuracy:', round(score, 2))
# Clustering demo: Find natural groups in Iris data (KMeans)
from sklearn.cluster import KMeans
X_nospecies = df_clean.drop('species', axis=1)
kmeans = KMeans(n_clusters=3, random_state=42)
clusters = kmeans.fit_predict(X_nospecies)
df_clean['cluster'] = clusters
print(df_clean[['species', 'cluster']].head(6))
# Ensemble learning: combine multiple classifiers (Random Forest)
from sklearn.ensemble import RandomForestClassifier
forest = RandomForestClassifier(n_estimators=50, random_state=42)
forest.fit(X_train, y_train)
rf_score = forest.score(X_test, y_test)
print('Random Forest accuracy:', round(rf_score, 2))
# Anomaly detection: find outlier flower samples (IsolationForest)
from sklearn.ensemble import IsolationForest
iso = IsolationForest(contamination=0.05, random_state=42)
anoms = iso.fit_predict(X_nospecies)
outliers = df_clean[anoms == -1]
print('Number of anomalies found:', len(outliers))
print(outliers.head(2))
# Time series basics: plot number of each species
counts = df_clean['species'].value_counts().sort_index()
counts.plot(kind='bar', title='Number of Samples per Iris Species');
# Mini-project 1: Market basket analysis - find association rules
from mlxtend.frequent_patterns import association_rules
rules = association_rules(freq_items, metric='confidence', min_threshold=0.3)
print(rules[['antecedents','consequents','support','confidence']].head(3))
# Mini-project 2: Ask the user for a sample and predict the species
sample = input("Enter four measurements separated by commas (sepal_length,sepal_width,petal_length,petal_width): ")
parts = [float(x.strip()) for x in sample.split(",")]
pred_species = model.predict([parts])[0]
print("Predicted iris species:", pred_species)
# Troubleshooting checklist for data mining in Python
print("Check these steps if something seems off:")
print("1. Is your dataset loaded and in the right shape?")
print("2. Are there missing or duplicate values?")
print("3. Are your input columns all numbers for modeling?")
print("4. Is your target variable clear and correct?")
print("5. Do you get the same results if you set the random state?")
# Extra: Show top 5 most frequent items in the groceries data
top_items = df_groc['itemDescription'].value_counts().head(5)
print(top_items)
# Challenge: Explore your own idea!
print("Try analyzing other columns, find rare items, or test another classifier.")
Recap and Next Steps#
- You just did real-world data mining in Python, even if you are a beginner!
- You learned to clean data, find patterns, make models, and check results.
- Next, try more datasets and challenges.
Happy mining!
Thanks for following along!#
If you enjoyed this beginner mini-course, please like, subscribe, and tell us what you want to learn next on YouTube. See you in the next lesson!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



