Mathew K Analytics

Lesson 25 · Data analytics zero to hero

Machine Learning for Beginners in Python | Data Analytics #25

Video twenty-five of the 30-part series, and the start of a machine learning block: your first real predictive model with scikit-learn. We're using a real,…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Data Analytics Zero to Hero, Video 25: Introduction to Machine Learning#

  • Video twenty-five of the 30-part series, and the start of a machine learning block: your first real predictive model with scikit-learn.
  • We're using a real, classic dataset: mushroom specimens labeled edible or poisonous, based on real physical characteristics.
  • Let's jump straight in.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • Install scikit-learn if you haven't already: pip install scikit-learn.
  • Place mushroom_predictions.csv in the same folder as this notebook.

Part 1: Features, Target, and Supervised Learning#

import pandas as pd

df = pd.read_csv('mushroom_predictions.csv')
print(df.shape)
print(df['actual'].value_counts())
(1625, 24)
actual
e    842
p    783
Name: count, dtype: int64
X = df.drop(columns=['actual', 'predicted'])
y = (df['actual'] == 'p').astype(int)
print(X.shape, y.shape)
print(y.value_counts())
(1625, 22) (1625,)
actual
0    842
1    783
Name: count, dtype: int64

Part 2: Splitting into Train and Test Sets#

from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42, stratify=y)
print(f'Train: {X_train.shape[0]} real specimens')
print(f'Test: {X_test.shape[0]} real specimens')
Train: 1300 real specimens
Test: 325 real specimens

Part 3: Training a Decision Tree#

from sklearn.tree import DecisionTreeClassifier

model = DecisionTreeClassifier(max_depth=4, random_state=42)
model.fit(X_train, y_train)
print('Real model trained')
Real model trained
predictions = model.predict(X_test)
print(predictions[:10])
print(y_test.values[:10])
[0 0 0 0 1 0 1 0 0 1]
[0 0 0 0 1 1 1 0 0 1]

Part 4: Evaluating the Model#

from sklearn.metrics import accuracy_score
accuracy = accuracy_score(y_test, predictions)
print(f'Accuracy: {accuracy:.2%}')
Accuracy: 97.85%
from sklearn.metrics import confusion_matrix, classification_report
print(confusion_matrix(y_test, predictions))
print(classification_report(y_test, predictions, target_names=['edible', 'poisonous']))
[[162   6]
 [  1 156]]
              precision    recall  f1-score   support

      edible       0.99      0.96      0.98       168
   poisonous       0.96      0.99      0.98       157

    accuracy                           0.98       325
   macro avg       0.98      0.98      0.98       325
weighted avg       0.98      0.98      0.98       325

Part 5: Which Features Mattered Most#

importances = pd.Series(model.feature_importances_, index=X.columns).sort_values(ascending=False)
print(importances.head(6))
gill-color           0.398922
population           0.174077
spore-print-color    0.163794
gill-size            0.141015
odor                 0.041537
stalk-shape          0.031978
dtype: float64

Wrap-Up: What You Learned#

  • Features and target, and the core idea of supervised learning: training on real labeled examples.
  • train_test_split, for a fair, honest evaluation on data the model never trained on.
  • The standard scikit-learn pattern: create, fit, predict, applied to a DecisionTreeClassifier.
  • Evaluating a real model with accuracy_score, confusion_matrix, and classification_report.
  • Reading feature_importances_ to see which real inputs actually drove the model's decisions.
  • All of it on a real, classic mushroom edibility dataset. Video twenty-six covers feature engineering and model evaluation in more depth, including how to compare multiple real models fairly. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.