Lesson 10 · Python for Data Analysts
Intro to Machine Learning with scikit-learn
Everything you need to train, evaluate, and understand your first real models: regression, classification, metrics, and cross-validation. No prior machine…
- CoursePython for Data Analysts
- Lesson10 of 12
- Video23 min
- FormatJupyter notebook · 17 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIntro to Machine Learning with scikit-learn#
- Everything you need to train, evaluate, and understand your first real models: regression, classification, metrics, and cross-validation.
- No prior machine learning experience needed. Let's get straight into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- If scikit-learn isn't installed yet, open a terminal in VS Code and run: pip install scikit-learn
Part 1: What Is Machine Learning?#
The Core Idea#
- Instead of writing explicit rules, you show a model examples, features paired with known answers, and it learns the pattern connecting them.
- Regression predicts a number, like a price. Classification predicts a category, like whether a customer churns.
- scikit-learn gives every model the exact same interface: fit to train it, predict to use it.
Part 2: Train/Test Split#
import numpy as np
import pandas as pd
rng = np.random.default_rng(seed=42)
n_houses = 200
sqft = rng.integers(600, 3500, size=n_houses)
bedrooms = rng.integers(1, 6, size=n_houses)
age = rng.integers(0, 60, size=n_houses)
noise = rng.normal(0, 15000, size=n_houses)
price = 80 * sqft + 12000 * bedrooms - 800 * age + 20000 + noise
houses = pd.DataFrame({'sqft': sqft, 'bedrooms': bedrooms, 'age': age, 'price': price})
print(houses.head())
from sklearn.model_selection import train_test_split
X_reg = houses[['sqft', 'bedrooms', 'age']]
y_reg = houses['price']
X_train_reg, X_test_reg, y_train_reg, y_test_reg = train_test_split(X_reg, y_reg, test_size=0.2, random_state=42)
print(X_train_reg.shape, X_test_reg.shape)
Part 3: Linear Regression#
from sklearn.linear_model import LinearRegression
reg_model = LinearRegression()
reg_model.fit(X_train_reg, y_train_reg)
print(reg_model.coef_)
print(reg_model.intercept_)
reg_predictions = reg_model.predict(X_test_reg)
print(reg_predictions[:5])
print(y_test_reg.values[:5])
Part 4: Classification#
rng = np.random.default_rng(seed=7)
n_customers = 300
monthly_usage = rng.uniform(0, 100, size=n_customers)
support_tickets = rng.integers(0, 10, size=n_customers)
tenure_months = rng.integers(1, 60, size=n_customers)
churn_score = -0.05 * monthly_usage + 0.6 * support_tickets - 0.03 * tenure_months + rng.normal(0, 1, size=n_customers)
churned = (churn_score > 0.5).astype(int)
customers = pd.DataFrame({
'monthly_usage': monthly_usage,
'support_tickets': support_tickets,
'tenure_months': tenure_months,
'churned': churned,
})
print(customers['churned'].value_counts())
from sklearn.linear_model import LogisticRegression
X_clf = customers[['monthly_usage', 'support_tickets', 'tenure_months']]
y_clf = customers['churned']
X_train_clf, X_test_clf, y_train_clf, y_test_clf = train_test_split(X_clf, y_clf, test_size=0.2, random_state=7)
log_model = LogisticRegression()
log_model.fit(X_train_clf, y_train_clf)
log_predictions = log_model.predict(X_test_clf)
print(log_predictions[:10])
from sklearn.tree import DecisionTreeClassifier
tree_model = DecisionTreeClassifier(max_depth=4, random_state=7)
tree_model.fit(X_train_clf, y_train_clf)
tree_predictions = tree_model.predict(X_test_clf)
print(tree_predictions[:10])
Part 5: Evaluation Metrics#
from sklearn.metrics import mean_squared_error, r2_score
mse = mean_squared_error(y_test_reg, reg_predictions)
r2 = r2_score(y_test_reg, reg_predictions)
print(f'MSE: {mse:,.0f}')
print(f'R-squared: {r2:.3f}')
from sklearn.metrics import accuracy_score, confusion_matrix, classification_report
accuracy = accuracy_score(y_test_clf, log_predictions)
print(f'Accuracy: {accuracy:.3f}')
print(confusion_matrix(y_test_clf, log_predictions))
print(classification_report(y_test_clf, log_predictions))
Part 6: Overfitting and Underfitting#
for depth in [1, 3, 6, None]:
depth_tree = DecisionTreeClassifier(max_depth=depth, random_state=7)
depth_tree.fit(X_train_clf, y_train_clf)
train_acc = depth_tree.score(X_train_clf, y_train_clf)
test_acc = depth_tree.score(X_test_clf, y_test_clf)
print(f'max_depth={depth}: train={train_acc:.3f}, test={test_acc:.3f}')
Part 7: Feature Scaling#
from sklearn.neighbors import KNeighborsClassifier
unscaled_knn = KNeighborsClassifier(n_neighbors=5)
unscaled_knn.fit(X_train_clf, y_train_clf)
unscaled_accuracy = unscaled_knn.score(X_test_clf, y_test_clf)
print(f'Unscaled KNN accuracy: {unscaled_accuracy:.3f}')
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train_clf)
X_test_scaled = scaler.transform(X_test_clf)
print(X_train_clf.iloc[0].values)
print(X_train_scaled[0])
scaled_knn = KNeighborsClassifier(n_neighbors=5)
scaled_knn.fit(X_train_scaled, y_train_clf)
scaled_accuracy = scaled_knn.score(X_test_scaled, y_test_clf)
print(f'Scaled KNN accuracy: {scaled_accuracy:.3f}')
Part 8: Cross-Validation#
from sklearn.model_selection import cross_val_score
cv_scores = cross_val_score(LogisticRegression(), X_clf, y_clf, cv=5)
print(cv_scores)
print(f'Mean accuracy: {cv_scores.mean():.3f} (+/- {cv_scores.std():.3f})')
Capstone Project: An End-to-End Model Pipeline#
from sklearn.pipeline import Pipeline
def evaluate_model(X, y, model, model_name, cv=5, random_state=7):
pipeline = Pipeline([
('scaler', StandardScaler()),
('model', model),
])
cv_scores = cross_val_score(pipeline, X, y, cv=cv)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=random_state)
pipeline.fit(X_train, y_train)
test_predictions = pipeline.predict(X_test)
test_accuracy = accuracy_score(y_test, test_predictions)
return {
'model_name': model_name,
'cv_mean_accuracy': cv_scores.mean(),
'cv_std_accuracy': cv_scores.std(),
'held_out_test_accuracy': test_accuracy,
}
results = []
results.append(evaluate_model(X_clf, y_clf, LogisticRegression(), 'Logistic Regression'))
results.append(evaluate_model(X_clf, y_clf, DecisionTreeClassifier(max_depth=4, random_state=7), 'Decision Tree'))
results.append(evaluate_model(X_clf, y_clf, KNeighborsClassifier(n_neighbors=5), 'K-Nearest Neighbors'))
results_df = pd.DataFrame(results)
print(results_df)
Wrap-Up: What You Learned#
- The core machine learning idea: learning patterns from examples instead of writing explicit rules.
- Train/test splitting, and why evaluating on training data alone is misleading.
- Linear regression for predicting numbers, and classification for predicting categories.
- Regression metrics like MSE and R-squared, and classification metrics like accuracy, confusion matrices, precision, and recall.
- Overfitting versus underfitting, and spotting it by comparing train and test scores.
- Feature scaling with StandardScaler, and why some models need it and others don't.
- Cross-validation for a more stable, reliable performance estimate.
- A capstone pipeline comparing multiple models with one reusable, leakage-free evaluation function.
- You went from raw data to a properly evaluated, multi-model comparison in one sitting. If you want the next build to land in your feed automatically, subscribing is the move see you in the next one.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



