Mathew K Analytics

Lesson 8 · Python For Machine Learning

Polynomial Regression in Python: Complete Guide for Machine Learning Beginners

In this lesson, you will learn how to use Python to fit curvesnot just straight linesto real-world data. We will explore concepts step-by-step, using small…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Welcome to Polynomial Regression in Python!#

In this lesson, you will learn how to use Python to fit curvesnot just straight linesto real-world data.

We will explore concepts step-by-step, using small examples and a real dataset.

No experience needed. Let us get started!

# First things first: Set up our workspace.
import warnings
warnings.filterwarnings('ignore')

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt

What is Polynomial Regression?#

Polynomial regression is a way to model curved relationships between two things, like 'as one gets bigger, the other goes up and then down'.

You can use it to fit lines, curves, or even more wiggly shapes to your data.

This is handy for things like predicting prices, tracking science results, or following trends over time.

# Here is a basic example of a curved pattern.
x = np.arange(-6, 7)
y = x**2 - 4 * x + 3
plt.plot(x, y, 'bo-')
plt.title('Sample quadratic curve: y = x**2 - 4x + 3')
plt.xlabel('x')
plt.ylabel('y')
plt.show()
No description has been provided for this image

Data setup#

Now, let us get a real dataset to practice with.

We will use California housing data, which contains house prices and features such as location, rooms, and income.

This dataset is popular for learning regression, which means predicting numbers.

# Load sample housing data from scikit-learn
from sklearn.datasets import fetch_california_housing
housing = fetch_california_housing(as_frame=True)
df = housing.frame
print('Data shape:', df.shape)
df.head()
Data shape: (20640, 9)
MedInc HouseAge AveRooms AveBedrms Population AveOccup Latitude Longitude MedHouseVal
0 8.3252 41.0 6.984127 1.023810 322.0 2.555556 37.88 -122.23 4.526
1 8.3014 21.0 6.238137 0.971880 2401.0 2.109842 37.86 -122.22 3.585
2 7.2574 52.0 8.288136 1.073446 496.0 2.802260 37.85 -122.24 3.521
3 5.6431 52.0 5.817352 1.073059 558.0 2.547945 37.85 -122.25 3.413
4 3.8462 52.0 6.281853 1.081081 565.0 2.181467 37.85 -122.25 3.422
# What columns do we have?
print('Column names:', df.columns.tolist())
Column names: ['MedInc', 'HouseAge', 'AveRooms', 'AveBedrms', 'Population', 'AveOccup', 'Latitude', 'Longitude', 'MedHouseVal']
# Let us focus on just one feature to keep things simple.
# We will predict house value from the average income in an area.
X = df[['MedInc']]
y = df['MedHouseVal']
# Plot income vs house value to see their relationship
plt.scatter(X, y, alpha=0.2)
plt.xlabel('Average Income (MedInc)')
plt.ylabel('House Value (MedHouseVal)')
plt.title('Income vs House Value')
plt.show()
No description has been provided for this image
# Split the data into training and test sets
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
# Let us try simple linear regression first
from sklearn.linear_model import LinearRegression
lin_reg = LinearRegression()
lin_reg.fit(X_train, y_train)
y_pred_linear = lin_reg.predict(X_test)
# Plot the linear fit against the real data
plt.scatter(X_test, y_test, alpha=0.2, label='Real data')
plt.plot(X_test, y_pred_linear, 'r-', label='Linear fit')
plt.xlabel('Average Income')
plt.ylabel('House Value')
plt.title('Straight line vs Real Data')
plt.legend()
plt.show()
No description has been provided for this image
# Now let us try polynomial regression for curves
from sklearn.preprocessing import PolynomialFeatures
poly = PolynomialFeatures(degree=2, include_bias=False)
X_poly_train = poly.fit_transform(X_train)
X_poly_test = poly.transform(X_test)
poly_reg = LinearRegression()
poly_reg.fit(X_poly_train, y_train)
y_pred_poly = poly_reg.predict(X_poly_test)
# Plot the curved fit versus the real data
plt.scatter(X_test, y_test, alpha=0.2, label='Real data')
plt.scatter(X_test, y_pred_poly, color='orange', s=8, label='Polynomial fit')
plt.xlabel('Average Income')
plt.ylabel('House Value')
plt.title('Curve fit vs Real Data')
plt.legend()
plt.show()
No description has been provided for this image
# Compare the two approaches using mean squared error (MSE)
from sklearn.metrics import mean_squared_error
mse_linear = mean_squared_error(y_test, y_pred_linear)
mse_poly = mean_squared_error(y_test, y_pred_poly)
print('MSE linear:', mse_linear)
print('MSE polynomial:', mse_poly)
MSE linear: 0.7001962368292408
MSE polynomial: 0.6939366057817791
# Let users experiment: input a degree for the curve
degree = int(input('What polynomial degree would you like to try? (25): '))
poly_user = PolynomialFeatures(degree=degree, include_bias=False)
X_poly_user_train = poly_user.fit_transform(X_train)
X_poly_user_test = poly_user.transform(X_test)
poly_user_reg = LinearRegression()
poly_user_reg.fit(X_poly_user_train, y_train)
y_pred_user = poly_user_reg.predict(X_poly_user_test)
user_mse = mean_squared_error(y_test, y_pred_user)
print('Your model's mean squared error:', user_mse)'
'')
  Cell In[13], line 10
    print('Your model's mean squared error:', user_mse)'
          ^
SyntaxError: invalid syntax. Perhaps you forgot a comma?
bm
### What if we use too high a degree?

If you pick a degree that is too high, the curve follows every tiny wiggleeven noise that does not matter.

This is called overfitting. It could work well on the data you have, but guesses poorly for new data.
    
We want balance: enough curve to match the real pattern, but not so much it memorizes every bump.
  Cell In[14], line 4
    If you pick a degree that is too high, the curve follows every tiny wiggleeven noise that does not matter.
       ^
SyntaxError: invalid syntax
# Visualize predictions with higher degree (overfitting example)
# Try degree 10 for demonstration
poly_too_high = PolynomialFeatures(degree=10, include_bias=False)
X_poly_high_train = poly_too_high.fit_transform(X_train)
X_poly_high_test = poly_too_high.transform(X_test)
poly_high_reg = LinearRegression()
poly_high_reg.fit(X_poly_high_train, y_train)
y_pred_high = poly_high_reg.predict(X_poly_high_test)
plt.scatter(X_test, y_test, alpha=0.2, label='Real data')
plt.scatter(X_test, y_pred_high, color='red', s=8, label='Degree 10 fit')
plt.xlabel('Average Income')
plt.ylabel('House Value')
plt.title('Overfitting: Too much curve')
plt.legend()
plt.show()
No description has been provided for this image
# Tips: Always check the test set, not just training scores
# A good model does well on both seen and unseen data
# This helps spot overfitting early
# Troubleshooting: What if you get strange values or errors?
# Check for missing data or weird numbers with .isnull() or .describe()
print('Any missing values?', df.isnull().any().any())
print(df['MedInc'].describe())
Any missing values? False
count    20640.000000
mean         3.870671
std          1.899822
min          0.499900
25%          2.563400
50%          3.534800
75%          4.743250
max         15.000100
Name: MedInc, dtype: float64
# Extra tip: try plotting the true vs predicted values scatter
plt.scatter(y_test, y_pred_poly, alpha=0.2)
plt.xlabel('True Values')
plt.ylabel('Predicted Values')
plt.title('True vs Predicted Values (Polynomial Regression)')
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'r--')
plt.show()
No description has been provided for this image
# Challenge: Ask the user to try another feature from the dataset
print('Available columns:', df.columns.tolist())
chosen_col = input('Pick another feature to try (type column name): ')
X_new = df[[chosen_col]]
X_new_train, X_new_test, y_new_train, y_new_test = train_test_split(X_new, y, random_state=42)
poly_challenge = PolynomialFeatures(degree=2, include_bias=False)
X_poly_new_train = poly_challenge.fit_transform(X_new_train)
X_poly_new_test = poly_challenge.transform(X_new_test)
challenge_reg = LinearRegression()
challenge_reg.fit(X_poly_new_train, y_new_train)
y_pred_challenge = challenge_reg.predict(X_poly_new_test)
print('New feature mean squared error:', mean_squared_error(y_new_test, y_pred_challenge))
Available columns: ['MedInc', 'HouseAge', 'AveRooms', 'AveBedrms', 'Population', 'AveOccup', 'Latitude', 'Longitude', 'MedHouseVal']
New feature mean squared error: 1.2611116166914895
 

Recap#

Nice work! Today you learned how to:

  • Plot and explore data
  • Fit straight lines using linear regression
  • Fit curves using polynomial regression
  • Compare results with mean squared error
  • Avoid overfitting by checking the test set
  • Experiment with new features

With these basics, you can predict and model all sorts of number-based problems!

Thank you for learning with us!#

Try more experiments, and share your results below!

For more beginner-friendly tutorials, be sure to like and follow our channel.

Happy coding!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.