Lesson 26 · Data analytics zero to hero
Feature Engineering & Model Evaluation | Data Analytics #26
Video twenty-six of the 30-part series, wrapping up the machine learning block: building better real features, cross-validation, and comparing models…
- CourseData analytics zero to hero
- Lesson26 of 30
- Video11 min
- FormatJupyter notebook · 7 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- mtcars.csv1.7 KB
📓 Full notebook
Download .ipynbData Analytics Zero to Hero, Video 26: Feature Engineering and Model Evaluation#
- Video twenty-six of the 30-part series, wrapping up the machine learning block: building better real features, cross-validation, and comparing models fairly.
- Switching to a regression problem this time: predicting real mpg from mtcars' other real specs.
- Let's jump straight in.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- Place mtcars.csv in the same folder as this notebook.
import pandas as pd
df = pd.read_csv('mtcars.csv')
print(df.shape)
Part 1: Engineering New Features#
df['power_to_weight'] = df['hp'] / df['wt']
df['displacement_per_cyl'] = df['disp'] / df['cyl']
print(df[['hp', 'wt', 'power_to_weight', 'disp', 'cyl', 'displacement_per_cyl']].head(3))
dummies = pd.get_dummies(df['gear'], prefix='gear')
print(dummies.head(3))
df = pd.concat([df, dummies], axis=1)
Part 2: Scaling Numeric Features#
from sklearn.preprocessing import StandardScaler
features = ['wt', 'hp', 'qsec', 'power_to_weight']
scaler = StandardScaler()
scaled = scaler.fit_transform(df[features])
scaled_df = pd.DataFrame(scaled, columns=features)
print(scaled_df.describe().round(2))
Part 3: Cross-Validation#
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import cross_val_score
X = df[['wt', 'hp', 'qsec', 'power_to_weight']]
y = df['mpg']
model = LinearRegression()
scores = cross_val_score(model, X, y, cv=5, scoring='r2')
print(f'R2 scores across 5 real folds: {scores.round(3)}')
print(f'Mean R2: {scores.mean():.3f}')
Part 4: Comparing Models Fairly#
from sklearn.ensemble import RandomForestRegressor
from sklearn.tree import DecisionTreeRegressor
models = {
'Linear Regression': LinearRegression(),
'Decision Tree': DecisionTreeRegressor(max_depth=3, random_state=42),
'Random Forest': RandomForestRegressor(n_estimators=100, random_state=42)
}
for name, m in models.items():
cv_scores = cross_val_score(m, X, y, cv=5, scoring='r2')
print(f'{name}: mean R2 = {cv_scores.mean():.3f}')
Part 5: Regression Metrics on a Held-Out Set#
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)
final_model = RandomForestRegressor(n_estimators=100, random_state=42)
final_model.fit(X_train, y_train)
preds = final_model.predict(X_test)
mae = mean_absolute_error(y_test, preds)
rmse = mean_squared_error(y_test, preds) ** 0.5
r2 = r2_score(y_test, preds)
print(f'MAE: {mae:.2f} mpg')
print(f'RMSE: {rmse:.2f} mpg')
print(f'R2: {r2:.3f}')
Wrap-Up: What You Learned#
- Engineering new features, like a power-to-weight ratio, that expose a relationship more directly than raw columns.
- One-hot encoding a categorical column with get_dummies.
- Scaling numeric features with StandardScaler, and why it matters for some real model types.
- Cross-validation with cross_val_score, for a more trustworthy estimate than a single split, especially on small real datasets.
- Comparing multiple real model types fairly, under the identical evaluation setup.
- Regression metrics: MAE, RMSE, and R-squared, each telling a slightly different real story.
- This wraps up the machine learning block. Video twenty-seven covers storytelling with data: building effective real reports that communicate findings, not just charts. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



