Mathew K Analytics

Lesson 10 · Python for Data Analysts

Intro to Machine Learning with scikit-learn

Everything you need to train, evaluate, and understand your first real models: regression, classification, metrics, and cross-validation. No prior machine…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Intro to Machine Learning with scikit-learn#

  • Everything you need to train, evaluate, and understand your first real models: regression, classification, metrics, and cross-validation.
  • No prior machine learning experience needed. Let's get straight into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • If scikit-learn isn't installed yet, open a terminal in VS Code and run: pip install scikit-learn

Part 1: What Is Machine Learning?#

The Core Idea#

  • Instead of writing explicit rules, you show a model examples, features paired with known answers, and it learns the pattern connecting them.
  • Regression predicts a number, like a price. Classification predicts a category, like whether a customer churns.
  • scikit-learn gives every model the exact same interface: fit to train it, predict to use it.

Part 2: Train/Test Split#

import numpy as np
import pandas as pd

rng = np.random.default_rng(seed=42)
n_houses = 200
sqft = rng.integers(600, 3500, size=n_houses)
bedrooms = rng.integers(1, 6, size=n_houses)
age = rng.integers(0, 60, size=n_houses)
noise = rng.normal(0, 15000, size=n_houses)
price = 80 * sqft + 12000 * bedrooms - 800 * age + 20000 + noise

houses = pd.DataFrame({'sqft': sqft, 'bedrooms': bedrooms, 'age': age, 'price': price})
print(houses.head())
   sqft  bedrooms  age          price
0   858         2   59   61057.432213
1  2844         5   46  269167.675979
2  2498         3   19  236860.339327
3  1872         4   58  173648.437682
4  1855         3   29  203272.379595
from sklearn.model_selection import train_test_split

X_reg = houses[['sqft', 'bedrooms', 'age']]
y_reg = houses['price']

X_train_reg, X_test_reg, y_train_reg, y_test_reg = train_test_split(X_reg, y_reg, test_size=0.2, random_state=42)
print(X_train_reg.shape, X_test_reg.shape)
(160, 3) (40, 3)

Part 3: Linear Regression#

from sklearn.linear_model import LinearRegression

reg_model = LinearRegression()
reg_model.fit(X_train_reg, y_train_reg)
print(reg_model.coef_)
print(reg_model.intercept_)
[   80.22391439 12895.74322071  -793.68473862]
17050.152852429805
reg_predictions = reg_model.predict(X_test_reg)
print(reg_predictions[:5])
print(y_test_reg.values[:5])
[151890.48985275 242058.89927988 204675.73965004 158224.8726605
 168874.3101971 ]
[132905.29178841 237392.12165316 219323.75901892 175528.33624858
 186418.26245066]

Part 4: Classification#

rng = np.random.default_rng(seed=7)
n_customers = 300
monthly_usage = rng.uniform(0, 100, size=n_customers)
support_tickets = rng.integers(0, 10, size=n_customers)
tenure_months = rng.integers(1, 60, size=n_customers)

churn_score = -0.05 * monthly_usage + 0.6 * support_tickets - 0.03 * tenure_months + rng.normal(0, 1, size=n_customers)
churned = (churn_score > 0.5).astype(int)

customers = pd.DataFrame({
    'monthly_usage': monthly_usage,
    'support_tickets': support_tickets,
    'tenure_months': tenure_months,
    'churned': churned,
})
print(customers['churned'].value_counts())
churned
0    206
1     94
Name: count, dtype: int64
from sklearn.linear_model import LogisticRegression

X_clf = customers[['monthly_usage', 'support_tickets', 'tenure_months']]
y_clf = customers['churned']
X_train_clf, X_test_clf, y_train_clf, y_test_clf = train_test_split(X_clf, y_clf, test_size=0.2, random_state=7)

log_model = LogisticRegression()
log_model.fit(X_train_clf, y_train_clf)
log_predictions = log_model.predict(X_test_clf)
print(log_predictions[:10])
[0 0 0 1 1 0 0 0 0 0]
from sklearn.tree import DecisionTreeClassifier

tree_model = DecisionTreeClassifier(max_depth=4, random_state=7)
tree_model.fit(X_train_clf, y_train_clf)
tree_predictions = tree_model.predict(X_test_clf)
print(tree_predictions[:10])
[0 0 1 0 1 0 0 0 0 0]

Part 5: Evaluation Metrics#

from sklearn.metrics import mean_squared_error, r2_score

mse = mean_squared_error(y_test_reg, reg_predictions)
r2 = r2_score(y_test_reg, reg_predictions)
print(f'MSE: {mse:,.0f}')
print(f'R-squared: {r2:.3f}')
MSE: 200,165,991
R-squared: 0.957
from sklearn.metrics import accuracy_score, confusion_matrix, classification_report

accuracy = accuracy_score(y_test_clf, log_predictions)
print(f'Accuracy: {accuracy:.3f}')
print(confusion_matrix(y_test_clf, log_predictions))
Accuracy: 0.917
[[40  3]
 [ 2 15]]
print(classification_report(y_test_clf, log_predictions))
              precision    recall  f1-score   support

           0       0.95      0.93      0.94        43
           1       0.83      0.88      0.86        17

    accuracy                           0.92        60
   macro avg       0.89      0.91      0.90        60
weighted avg       0.92      0.92      0.92        60

Part 6: Overfitting and Underfitting#

for depth in [1, 3, 6, None]:
    depth_tree = DecisionTreeClassifier(max_depth=depth, random_state=7)
    depth_tree.fit(X_train_clf, y_train_clf)
    train_acc = depth_tree.score(X_train_clf, y_train_clf)
    test_acc = depth_tree.score(X_test_clf, y_test_clf)
    print(f'max_depth={depth}: train={train_acc:.3f}, test={test_acc:.3f}')
max_depth=1: train=0.717, test=0.650
max_depth=3: train=0.892, test=0.800
max_depth=6: train=0.954, test=0.883
max_depth=None: train=1.000, test=0.850

Part 7: Feature Scaling#

from sklearn.neighbors import KNeighborsClassifier

unscaled_knn = KNeighborsClassifier(n_neighbors=5)
unscaled_knn.fit(X_train_clf, y_train_clf)
unscaled_accuracy = unscaled_knn.score(X_test_clf, y_test_clf)
print(f'Unscaled KNN accuracy: {unscaled_accuracy:.3f}')
Unscaled KNN accuracy: 0.750
from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train_clf)
X_test_scaled = scaler.transform(X_test_clf)
print(X_train_clf.iloc[0].values)
print(X_train_scaled[0])
[61.25396043  1.          5.        ]
[ 0.37624548 -1.21490336 -1.46766695]
scaled_knn = KNeighborsClassifier(n_neighbors=5)
scaled_knn.fit(X_train_scaled, y_train_clf)
scaled_accuracy = scaled_knn.score(X_test_scaled, y_test_clf)
print(f'Scaled KNN accuracy: {scaled_accuracy:.3f}')
Scaled KNN accuracy: 0.783

Part 8: Cross-Validation#

from sklearn.model_selection import cross_val_score

cv_scores = cross_val_score(LogisticRegression(), X_clf, y_clf, cv=5)
print(cv_scores)
print(f'Mean accuracy: {cv_scores.mean():.3f} (+/- {cv_scores.std():.3f})')
[0.91666667 0.86666667 0.83333333 0.9        0.93333333]
Mean accuracy: 0.890 (+/- 0.036)

Capstone Project: An End-to-End Model Pipeline#

from sklearn.pipeline import Pipeline

def evaluate_model(X, y, model, model_name, cv=5, random_state=7):
    pipeline = Pipeline([
        ('scaler', StandardScaler()),
        ('model', model),
    ])

    cv_scores = cross_val_score(pipeline, X, y, cv=cv)

    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=random_state)
    pipeline.fit(X_train, y_train)
    test_predictions = pipeline.predict(X_test)
    test_accuracy = accuracy_score(y_test, test_predictions)

    return {
        'model_name': model_name,
        'cv_mean_accuracy': cv_scores.mean(),
        'cv_std_accuracy': cv_scores.std(),
        'held_out_test_accuracy': test_accuracy,
    }
results = []
results.append(evaluate_model(X_clf, y_clf, LogisticRegression(), 'Logistic Regression'))
results.append(evaluate_model(X_clf, y_clf, DecisionTreeClassifier(max_depth=4, random_state=7), 'Decision Tree'))
results.append(evaluate_model(X_clf, y_clf, KNeighborsClassifier(n_neighbors=5), 'K-Nearest Neighbors'))

results_df = pd.DataFrame(results)
print(results_df)
            model_name  cv_mean_accuracy  cv_std_accuracy  \
0  Logistic Regression          0.896667         0.046428   
1        Decision Tree          0.846667         0.064464   
2  K-Nearest Neighbors          0.843333         0.047842   

   held_out_test_accuracy  
0                0.916667  
1                0.833333  
2                0.783333  

Wrap-Up: What You Learned#

  • The core machine learning idea: learning patterns from examples instead of writing explicit rules.
  • Train/test splitting, and why evaluating on training data alone is misleading.
  • Linear regression for predicting numbers, and classification for predicting categories.
  • Regression metrics like MSE and R-squared, and classification metrics like accuracy, confusion matrices, precision, and recall.
  • Overfitting versus underfitting, and spotting it by comparing train and test scores.
  • Feature scaling with StandardScaler, and why some models need it and others don't.
  • Cross-validation for a more stable, reliable performance estimate.
  • A capstone pipeline comparing multiple models with one reusable, leakage-free evaluation function.
  • You went from raw data to a properly evaluated, multi-model comparison in one sitting. If you want the next build to land in your feed automatically, subscribing is the move see you in the next one.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.