Mathew K Analytics

Lesson 4 · Data Mining

Introduction to Python, Pandas, and scikit-learn for Data Analysis and Machine Learning

This lesson is for absolute beginners. You will learn to load, clean, and explore data. We will cover association rules, classification, clustering, and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Welcome to Data Mining with Python!#

This lesson is for absolute beginners.

You will learn to load, clean, and explore data.

We will cover association rules, classification, clustering, and more.

We will use Python libraries like Pandas and scikit-learn.

Do not worry if something is new.

Let us start your journey into data mining!

# Suppress warnings to keep outputs clean
import warnings; warnings.filterwarnings('ignore')
import numpy as np
np.random.seed(42)
# Data setup (Iris Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(150, 5)
   sepal_length  sepal_width  petal_length  petal_width species
0           5.1          3.5           1.4          0.2  setosa
1           4.9          3.0           1.4          0.2  setosa
2           4.7          3.2           1.3          0.2  setosa

What is Data Cleaning?#

Data is not always ready to use. It can have missing values or errors. Cleaning means fixing or removing bad parts of data.

# Check for missing values
print(df.isnull().sum())
sepal_length    0
sepal_width     0
petal_length    0
petal_width     0
species         0
dtype: int64
# Fill any missing values just in case
df.fillna(method='ffill', inplace=True)
print('Missing values after fill:', df.isnull().sum().sum())
Missing values after fill: 0
# Remove duplicate rows if any
print('Shape before dropping duplicates:', df.shape)
df.drop_duplicates(inplace=True)
print('Shape after:', df.shape)
Shape before dropping duplicates: (150, 5)
Shape after: (147, 5)

Exploratory Data Analysis#

Exploratory Data Analysis (EDA) finds patterns, trends, and surprises in data.

We use simple stats and plots to understand what data shows before applying models.

# Show some descriptive statistics
print(df.describe())
       sepal_length  sepal_width  petal_length  petal_width
count    147.000000   147.000000    147.000000   147.000000
mean       5.856463     3.055782      3.780272     1.208844
std        0.829100     0.437009      1.759111     0.757874
min        4.300000     2.000000      1.000000     0.100000
25%        5.100000     2.800000      1.600000     0.300000
50%        5.800000     3.000000      4.400000     1.300000
75%        6.400000     3.300000      5.100000     1.800000
max        7.900000     4.400000      6.900000     2.500000
# Simple visualization: scatter plot
import matplotlib.pyplot as plt
plt.scatter(df['sepal_length'], df['petal_length'], c='green', alpha=0.5)
plt.xlabel('Sepal Length')
plt.ylabel('Petal Length')
plt.title('Scatter plot of Sepal and Petal Length')
plt.show()
No description has been provided for this image

Association Rule Mining#

Association rules uncover items that appear together in data. Stores use these rules to learn which items are bought together. Let us try it using a small version of a groceries dataset.

# Data setup (Groceries Dataset)
import pandas as pd
import kagglehub, os
path = kagglehub.dataset_download('heeraldedhia/groceries-dataset')
files = os.listdir(path)
csv_file = [f for f in files if f.endswith('.csv')][0]
df_groc = pd.read_csv(os.path.join(path, csv_file))
print(df_groc.shape)
print(df_groc.head(3))
(38765, 3)
   Member_number        Date itemDescription
0           1808  21-07-2015  tropical fruit
1           2552  05-01-2015      whole milk
2           2300  19-09-2015       pip fruit
# Create 'baskets' of items bought by each shopper
basket = df_groc.groupby(['Member_number'])['itemDescription'].apply(list).reset_index()
print(basket.head(3))
   Member_number                                    itemDescription
0           1000  [soda, canned beer, sausage, sausage, whole mi...
1           1001  [frankfurter, frankfurter, beef, sausage, whol...
2           1002  [tropical fruit, butter milk, butter, frozen v...
# Convert baskets from lists to a one-hot format
from mlxtend.preprocessing import TransactionEncoder
te = TransactionEncoder()
te_ary = te.fit_transform(basket['itemDescription'])
df_te = pd.DataFrame(te_ary, columns=te.columns_)
print(df_te.shape)
(3898, 167)
# Apply the Apriori algorithm for frequent itemsets
from mlxtend.frequent_patterns import apriori, association_rules
freq_items = apriori(df_te, min_support=0.01, use_colnames=True)
print(freq_items.head())
    support                 itemsets
0  0.015393  (Instant food products)
1  0.078502               (UHT-milk)
2  0.031042          (baking powder)
3  0.119548                   (beef)
4  0.079785                (berries)
# Mine association rules from itemsets
rules = association_rules(freq_items, metric='lift', min_threshold=1.1)
rules = rules[['antecedents', 'consequents', 'support', 'confidence', 'lift']].sort_values('lift', ascending=False)
print(rules.head(5))
                                   antecedents  \
14108                     (yogurt, rolls/buns)   
14089  (whole milk, other vegetables, sausage)   
14097   (other vegetables, yogurt, rolls/buns)   
14100                    (whole milk, sausage)   
11754                    (whole milk, sausage)   

                                   consequents   support  confidence      lift  
14108  (whole milk, other vegetables, sausage)  0.013597    0.122120  2.428689  
14089                     (yogurt, rolls/buns)  0.013597    0.270408  2.428689  
14097                    (whole milk, sausage)  0.013597    0.259804  2.428575  
14100   (other vegetables, yogurt, rolls/buns)  0.013597    0.127098  2.428575  
11754                           (curd, yogurt)  0.010005    0.093525  2.322046  

Classification Basics#

Classification puts data into groups or classes. For example: Predicting if a flower is setosa, versicolor, or virginica.

Let us train a simple decision tree on the Iris data.

# Prepare the data for classification
from sklearn.model_selection import train_test_split
X = df.drop('species', axis=1)
y = df['species']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Train a simple decision tree classifier
from sklearn.tree import DecisionTreeClassifier
tree = DecisionTreeClassifier(random_state=42)
tree.fit(X_train, y_train)
score = tree.score(X_test, y_test)
print('Test Accuracy:', score)
Test Accuracy: 0.9333333333333333

Clustering#

Clustering groups data by similarity. Unlike classification, classes are not given in advance. A real-world example is grouping customers by buying habits.

# Use KMeans clustering on two Iris columns
from sklearn.cluster import KMeans
X = df[['sepal_length', 'petal_length']]
kmeans = KMeans(n_clusters=3, random_state=42)
clusters = kmeans.fit_predict(X)
plt.scatter(X['sepal_length'], X['petal_length'], c=clusters, cmap='plasma')
plt.xlabel('Sepal Length')
plt.ylabel('Petal Length')
plt.title('KMeans Clusters on Iris Data')
plt.show()
No description has been provided for this image

Ensemble Learning#

Ensembles combine several models to improve predictions. Voting and Random Forests are examples.

Ensemble methods are strong in real-world tasks.

# Try a RandomForest ensemble classifier
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(X_train, y_train)
print('Random Forest Test Accuracy:', rf.score(X_test, y_test))
Random Forest Test Accuracy: 0.9333333333333333

Anomaly Detection#

Anomaly detection spots data points that are rare or different. This is used to find fraud, errors, or unique behaviors. Let us try a simple one.

# Use IsolationForest for anomaly detection
from sklearn.ensemble import IsolationForest
iso = IsolationForest(random_state=42)
outliers = iso.fit_predict(X)
print('Anomalies found:', list(outliers).count(-1))
Anomalies found: 44

Introduction to Time-Series Data#

A time series is data measured over time.

Examples include stock prices, temperature changes, and website visits.

# Simulate a small time series
import numpy as np
dates = pd.date_range('2024-01-01', periods=10)
values = np.random.randint(10, 100, size=10)
ts = pd.Series(values, index=dates)
import matplotlib.pyplot as plt
ts.plot(marker='o')
plt.title('Example Time Series')
plt.ylabel('Value')
plt.show()
No description has been provided for this image

Mini-Project: Market Basket Analysis#

You are a store manager. You want to find out which items shoppers buy together.

Use association rules to help marketing or optimize layout!

# Find pairs of items that are frequently bought together
pair_rules = association_rules(freq_items, metric='confidence', min_threshold=0.3)
pair_rules = pair_rules[pair_rules['antecedents'].apply(lambda x: len(x) == 1)]
pair_rules = pair_rules.sort_values('confidence', ascending=False).head(5)
print(pair_rules[['antecedents', 'consequents', 'support', 'confidence']])
          antecedents   consequents   support  confidence
202          (liquor)  (whole milk)  0.016675    0.631068
223         (mustard)  (whole milk)  0.014110    0.604396
173             (ham)  (whole milk)  0.036942    0.582996
128        (dog food)  (whole milk)  0.010005    0.582090
101  (condensed milk)  (whole milk)  0.013853    0.580645

Mini-Project: Classification With Titanic Data#

Now, let us use the Titanic dataset to predict who survived.

Try a simple logistic regression model.

# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df_titanic = pd.read_csv(url)
print(df_titanic.shape)
print(df_titanic.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  
# Prepare Titanic features for modeling
df_titanic['Sex'] = df_titanic['Sex'].map({'male': 0, 'female': 1})
df_titanic['Age'].fillna(df_titanic['Age'].median(), inplace=True)
X = df_titanic[['Pclass', 'Sex', 'Age']]
y = df_titanic['Survived']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Train logistic regression on Titanic
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(max_iter=1000, random_state=42)
model.fit(X_train, y_train)
print('Titanic Test Accuracy:', model.score(X_test, y_test))
Titanic Test Accuracy: 0.8100558659217877

Best Practices and Troubleshooting#

  • Always preview and check missing values.
  • Plot your data before applying models.
  • Try more than one model or method.
  • Tune parameters and try again if results look odd.

These habits save time and improve accuracy.

# Ask the user a question to reflect
response = input('What was the most surprising thing you learned about your data? ')
print('Thanks for sharing! Each discovery helps build data intuition.')
Thanks for sharing! Each discovery helps build data intuition.

Extra Tips#

  • Use random_state so results can be repeated.
  • Keep code simple and build up step by step.
  • Check what every line prints; output often helps spot issues.

Coding is like learning a new language. It gets easier with practice!

Challenge Exercise#

Pick a column in the Titanic or Iris dataset. Plot its distribution using matplotlib.

Can you spot any patterns?

Try one new classifier (for example, k-nearest neighbors) on Iris.

Recap#

  • You learned to load and clean data.
  • You practiced exploring, visualizing, and modeling.
  • You tried association rules, classification, and clustering.
  • You saw how time series and anomalies are handled.
  • You completed two mini-projects.

Practice is key come back and build new projects soon!

Thanks for joining!#

If you enjoyed this lesson, leave a like or comment and subscribe for more beginner-friendly Python data mining videos.

Happy coding!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.