Lesson 8 · Python For Machine Learning
Polynomial Regression in Python: Complete Guide for Machine Learning Beginners
In this lesson, you will learn how to use Python to fit curvesnot just straight linesto real-world data. We will explore concepts step-by-step, using small…
- CoursePython For Machine Learning
- Lesson8 of 16
- Video11 min
- FormatJupyter notebook · 21 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Welcome to Polynomial Regression in Python!#
In this lesson, you will learn how to use Python to fit curvesnot just straight linesto real-world data.
We will explore concepts step-by-step, using small examples and a real dataset.
No experience needed. Let us get started!
# First things first: Set up our workspace.
import warnings
warnings.filterwarnings('ignore')
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
What is Polynomial Regression?#
Polynomial regression is a way to model curved relationships between two things, like 'as one gets bigger, the other goes up and then down'.
You can use it to fit lines, curves, or even more wiggly shapes to your data.
This is handy for things like predicting prices, tracking science results, or following trends over time.
# Here is a basic example of a curved pattern.
x = np.arange(-6, 7)
y = x**2 - 4 * x + 3
plt.plot(x, y, 'bo-')
plt.title('Sample quadratic curve: y = x**2 - 4x + 3')
plt.xlabel('x')
plt.ylabel('y')
plt.show()
Data setup#
Now, let us get a real dataset to practice with.
We will use California housing data, which contains house prices and features such as location, rooms, and income.
This dataset is popular for learning regression, which means predicting numbers.
# Load sample housing data from scikit-learn
from sklearn.datasets import fetch_california_housing
housing = fetch_california_housing(as_frame=True)
df = housing.frame
print('Data shape:', df.shape)
df.head()
# What columns do we have?
print('Column names:', df.columns.tolist())
# Let us focus on just one feature to keep things simple.
# We will predict house value from the average income in an area.
X = df[['MedInc']]
y = df['MedHouseVal']
# Plot income vs house value to see their relationship
plt.scatter(X, y, alpha=0.2)
plt.xlabel('Average Income (MedInc)')
plt.ylabel('House Value (MedHouseVal)')
plt.title('Income vs House Value')
plt.show()
# Split the data into training and test sets
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
# Let us try simple linear regression first
from sklearn.linear_model import LinearRegression
lin_reg = LinearRegression()
lin_reg.fit(X_train, y_train)
y_pred_linear = lin_reg.predict(X_test)
# Plot the linear fit against the real data
plt.scatter(X_test, y_test, alpha=0.2, label='Real data')
plt.plot(X_test, y_pred_linear, 'r-', label='Linear fit')
plt.xlabel('Average Income')
plt.ylabel('House Value')
plt.title('Straight line vs Real Data')
plt.legend()
plt.show()
# Now let us try polynomial regression for curves
from sklearn.preprocessing import PolynomialFeatures
poly = PolynomialFeatures(degree=2, include_bias=False)
X_poly_train = poly.fit_transform(X_train)
X_poly_test = poly.transform(X_test)
poly_reg = LinearRegression()
poly_reg.fit(X_poly_train, y_train)
y_pred_poly = poly_reg.predict(X_poly_test)
# Plot the curved fit versus the real data
plt.scatter(X_test, y_test, alpha=0.2, label='Real data')
plt.scatter(X_test, y_pred_poly, color='orange', s=8, label='Polynomial fit')
plt.xlabel('Average Income')
plt.ylabel('House Value')
plt.title('Curve fit vs Real Data')
plt.legend()
plt.show()
# Compare the two approaches using mean squared error (MSE)
from sklearn.metrics import mean_squared_error
mse_linear = mean_squared_error(y_test, y_pred_linear)
mse_poly = mean_squared_error(y_test, y_pred_poly)
print('MSE linear:', mse_linear)
print('MSE polynomial:', mse_poly)
# Let users experiment: input a degree for the curve
degree = int(input('What polynomial degree would you like to try? (25): '))
poly_user = PolynomialFeatures(degree=degree, include_bias=False)
X_poly_user_train = poly_user.fit_transform(X_train)
X_poly_user_test = poly_user.transform(X_test)
poly_user_reg = LinearRegression()
poly_user_reg.fit(X_poly_user_train, y_train)
y_pred_user = poly_user_reg.predict(X_poly_user_test)
user_mse = mean_squared_error(y_test, y_pred_user)
print('Your model's mean squared error:', user_mse)'
'')
bm
### What if we use too high a degree?
If you pick a degree that is too high, the curve follows every tiny wiggleeven noise that does not matter.
This is called overfitting. It could work well on the data you have, but guesses poorly for new data.
We want balance: enough curve to match the real pattern, but not so much it memorizes every bump.
# Visualize predictions with higher degree (overfitting example)
# Try degree 10 for demonstration
poly_too_high = PolynomialFeatures(degree=10, include_bias=False)
X_poly_high_train = poly_too_high.fit_transform(X_train)
X_poly_high_test = poly_too_high.transform(X_test)
poly_high_reg = LinearRegression()
poly_high_reg.fit(X_poly_high_train, y_train)
y_pred_high = poly_high_reg.predict(X_poly_high_test)
plt.scatter(X_test, y_test, alpha=0.2, label='Real data')
plt.scatter(X_test, y_pred_high, color='red', s=8, label='Degree 10 fit')
plt.xlabel('Average Income')
plt.ylabel('House Value')
plt.title('Overfitting: Too much curve')
plt.legend()
plt.show()
# Tips: Always check the test set, not just training scores
# A good model does well on both seen and unseen data
# This helps spot overfitting early
# Troubleshooting: What if you get strange values or errors?
# Check for missing data or weird numbers with .isnull() or .describe()
print('Any missing values?', df.isnull().any().any())
print(df['MedInc'].describe())
# Extra tip: try plotting the true vs predicted values scatter
plt.scatter(y_test, y_pred_poly, alpha=0.2)
plt.xlabel('True Values')
plt.ylabel('Predicted Values')
plt.title('True vs Predicted Values (Polynomial Regression)')
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'r--')
plt.show()
# Challenge: Ask the user to try another feature from the dataset
print('Available columns:', df.columns.tolist())
chosen_col = input('Pick another feature to try (type column name): ')
X_new = df[[chosen_col]]
X_new_train, X_new_test, y_new_train, y_new_test = train_test_split(X_new, y, random_state=42)
poly_challenge = PolynomialFeatures(degree=2, include_bias=False)
X_poly_new_train = poly_challenge.fit_transform(X_new_train)
X_poly_new_test = poly_challenge.transform(X_new_test)
challenge_reg = LinearRegression()
challenge_reg.fit(X_poly_new_train, y_new_train)
y_pred_challenge = challenge_reg.predict(X_poly_new_test)
print('New feature mean squared error:', mean_squared_error(y_new_test, y_pred_challenge))
Recap#
Nice work! Today you learned how to:
- Plot and explore data
- Fit straight lines using linear regression
- Fit curves using polynomial regression
- Compare results with mean squared error
- Avoid overfitting by checking the test set
- Experiment with new features
With these basics, you can predict and model all sorts of number-based problems!
Thank you for learning with us!#
Try more experiments, and share your results below!
For more beginner-friendly tutorials, be sure to like and follow our channel.
Happy coding!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



