Mathew K Analytics

Lesson 26 · Data analytics zero to hero

Feature Engineering & Model Evaluation | Data Analytics #26

Video twenty-six of the 30-part series, wrapping up the machine learning block: building better real features, cross-validation, and comparing models…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Data Analytics Zero to Hero, Video 26: Feature Engineering and Model Evaluation#

  • Video twenty-six of the 30-part series, wrapping up the machine learning block: building better real features, cross-validation, and comparing models fairly.
  • Switching to a regression problem this time: predicting real mpg from mtcars' other real specs.
  • Let's jump straight in.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • Place mtcars.csv in the same folder as this notebook.
import pandas as pd
df = pd.read_csv('mtcars.csv')
print(df.shape)
(32, 12)

Part 1: Engineering New Features#

df['power_to_weight'] = df['hp'] / df['wt']
df['displacement_per_cyl'] = df['disp'] / df['cyl']
print(df[['hp', 'wt', 'power_to_weight', 'disp', 'cyl', 'displacement_per_cyl']].head(3))
    hp     wt  power_to_weight   disp  cyl  displacement_per_cyl
0  110  2.620        41.984733  160.0    6             26.666667
1  110  2.875        38.260870  160.0    6             26.666667
2   93  2.320        40.086207  108.0    4             27.000000
dummies = pd.get_dummies(df['gear'], prefix='gear')
print(dummies.head(3))
df = pd.concat([df, dummies], axis=1)
   gear_3  gear_4  gear_5
0   False    True   False
1   False    True   False
2   False    True   False

Part 2: Scaling Numeric Features#

from sklearn.preprocessing import StandardScaler

features = ['wt', 'hp', 'qsec', 'power_to_weight']
scaler = StandardScaler()
scaled = scaler.fit_transform(df[features])
scaled_df = pd.DataFrame(scaled, columns=features)
print(scaled_df.describe().round(2))
          wt     hp   qsec  power_to_weight
count  32.00  32.00  32.00            32.00
mean   -0.00   0.00  -0.00            -0.00
std     1.02   1.02   1.02             1.02
min    -1.77  -1.40  -1.90            -1.62
25%    -0.66  -0.74  -0.54            -0.60
50%     0.11  -0.35  -0.08            -0.27
75%     0.41   0.49   0.60             0.15
max     2.29   2.79   2.87             3.03

Part 3: Cross-Validation#

from sklearn.linear_model import LinearRegression
from sklearn.model_selection import cross_val_score

X = df[['wt', 'hp', 'qsec', 'power_to_weight']]
y = df['mpg']
model = LinearRegression()
scores = cross_val_score(model, X, y, cv=5, scoring='r2')
print(f'R2 scores across 5 real folds: {scores.round(3)}')
print(f'Mean R2: {scores.mean():.3f}')
R2 scores across 5 real folds: [ 0.394  0.68   0.754  0.611 -8.988]
Mean R2: -1.310

Part 4: Comparing Models Fairly#

from sklearn.ensemble import RandomForestRegressor
from sklearn.tree import DecisionTreeRegressor

models = {
    'Linear Regression': LinearRegression(),
    'Decision Tree': DecisionTreeRegressor(max_depth=3, random_state=42),
    'Random Forest': RandomForestRegressor(n_estimators=100, random_state=42)
}

for name, m in models.items():
    cv_scores = cross_val_score(m, X, y, cv=5, scoring='r2')
    print(f'{name}: mean R2 = {cv_scores.mean():.3f}')
Linear Regression: mean R2 = -1.310
Decision Tree: mean R2 = 0.326
Random Forest: mean R2 = 0.697

Part 5: Regression Metrics on a Held-Out Set#

from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)
final_model = RandomForestRegressor(n_estimators=100, random_state=42)
final_model.fit(X_train, y_train)
preds = final_model.predict(X_test)

mae = mean_absolute_error(y_test, preds)
rmse = mean_squared_error(y_test, preds) ** 0.5
r2 = r2_score(y_test, preds)
print(f'MAE: {mae:.2f} mpg')
print(f'RMSE: {rmse:.2f} mpg')
print(f'R2: {r2:.3f}')
MAE: 2.34 mpg
RMSE: 2.94 mpg
R2: 0.790

Wrap-Up: What You Learned#

  • Engineering new features, like a power-to-weight ratio, that expose a relationship more directly than raw columns.
  • One-hot encoding a categorical column with get_dummies.
  • Scaling numeric features with StandardScaler, and why it matters for some real model types.
  • Cross-validation with cross_val_score, for a more trustworthy estimate than a single split, especially on small real datasets.
  • Comparing multiple real model types fairly, under the identical evaluation setup.
  • Regression metrics: MAE, RMSE, and R-squared, each telling a slightly different real story.
  • This wraps up the machine learning block. Video twenty-seven covers storytelling with data: building effective real reports that communicate findings, not just charts. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.