Lesson 24 · Probability and Statistics in python
Understanding Multiple Linear Regression with Python’s Statsmodels Library
Welcome! In this lesson, you will learn the basics of probability, descriptive statistics, and multiple linear regression using Python. By the end, you will…
- CourseProbability and Statistics in python
- Lesson24 of 35
- Video11 min
- FormatJupyter notebook · 14 code cells
What you'll learn
- Probability: Flipping a Coin
- Descriptive Statistics: Summarize Your Data
- Probability Distributions
- Exploring Relationships: Correlation
- Linear Regression: Predicting MPG From Just Horsepower
- Multiple Linear Regression: Predicting MPG from Several Features
- Best Practices and Troubleshooting
- Challenge Exercise
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbMultiple Linear Regression in Python: Absolute Beginner Lesson#
Welcome! In this lesson, you will learn the basics of probability, descriptive statistics, and multiple linear regression using Python.
By the end, you will understand how to apply regression to real-world data and explore relationships between variables.
Let us get started!
import warnings; warnings.filterwarnings("ignore")
# All warnings will be suppressed for a clean output
Probability: Flipping a Coin#
Let's start with a simple probability example using a coin flip.
Probability tells us how likely something is to happen.
import numpy as np
# Simulate 10 coin flips (0=heads, 1=tails)
coin_flips = np.random.choice([0,1], size=10)
print("Coin flips:", coin_flips)
# Calculate the probability of getting heads
prob_heads = np.mean(coin_flips == 0)
print("Probability of heads:", prob_heads)
Descriptive Statistics: Summarize Your Data#
Descriptive statistics help us quickly describe data using numbers like mean, median, and standard deviation.
Let us work with some fun real data next!
# Data setup: Load the MPG (Miles per Gallon) dataset
import seaborn as sns
import pandas as pd
mpg = sns.load_dataset("mpg")
print("MPG data shape:", mpg.shape)
mpg.head()
# Quickly look at summary statistics for mpg
mpg_describe = mpg['mpg'].describe()
print(mpg_describe)
# Find the median explicitly
median_mpg = mpg['mpg'].median()
print("Median MPG:", median_mpg)
# Check for missing data in all columns
missing = mpg.isnull().sum()
print("Missing data by column:\n", missing)
# Remove any rows with missing mpg or horsepower values
mpg_clean = mpg.dropna(subset=['mpg', 'horsepower'])
print("Data shape after dropna:", mpg_clean.shape)
Probability Distributions#
Probability distributions help us describe how values are spread out.
Let us simulate rolling a die using a discrete distribution!
# Simulate 1000 dice rolls (1 through 6)
rolls = np.random.choice([1,2,3,4,5,6], size=1000)
# Probability of getting a 6
prob_6 = np.mean(rolls == 6)
print("Probability of 6:", prob_6)
# Visualize the distribution of MPG values
import matplotlib.pyplot as plt
plt.figure(figsize=(6, 3))
plt.hist(mpg_clean['mpg'], bins=20, alpha=0.7, color='skyblue')
plt.title('Distribution of Miles Per Gallon')
plt.xlabel('Miles Per Gallon')
plt.ylabel('Count')
plt.show()
Exploring Relationships: Correlation#
Correlation measures how two variables move together.
Values close to 1 or -1 mean strong relationships. Values close to 0 mean weak or no relationship.
Let us check how mpg relates to horsepower.
# Calculate correlation between mpg and horsepower
corr = mpg_clean['mpg'].corr(mpg_clean['horsepower'])
print("Correlation (mpg vs horsepower):", corr)
Linear Regression: Predicting MPG From Just Horsepower#
Regression predicts the value of one variable from one or more other variables.
Let us start by predicting mpg from horsepower using only one variable.
# Fit a simple linear regression (mpg ~ horsepower) with statsmodels
import statsmodels.api as sm
X_simple = mpg_clean[['horsepower']]
X_simple = sm.add_constant(X_simple)
y = mpg_clean['mpg']
model_simple = sm.OLS(y, X_simple).fit()
print(model_simple.summary())
Multiple Linear Regression: Predicting MPG from Several Features#
Multiple linear regression lets us use several predictors at once.
We can use horsepower, weight, and cylinders to better predict mpg.
Let us do it!
# Build the model with more predictors
X_multi = mpg_clean[['horsepower', 'weight', 'cylinders']]
X_multi = sm.add_constant(X_multi)
model_multi = sm.OLS(y, X_multi).fit()
print(model_multi.summary())
# Visualize predictions vs. actual MPG
pred_mpg = model_multi.predict(X_multi)
plt.figure(figsize=(5, 5))
plt.scatter(y, pred_mpg, alpha=0.7, color='teal')
plt.xlabel('Actual MPG')
plt.ylabel('Predicted MPG')
plt.title('Actual vs Predicted MPG (Multiple Regression)')
plt.plot([y.min(), y.max()], [y.min(), y.max()], 'r--')
plt.show()
# Check for multicollinearity (predictors that are too related to each other)
from statsmodels.stats.outliers_influence import variance_inflation_factor
vif_data = pd.DataFrame()
vif_data['feature'] = X_multi.columns
vif_data['VIF'] = [variance_inflation_factor(X_multi.values, i) for i in range(X_multi.shape[1])]
print(vif_data)
# Use the model to predict mpg for a new car
print("Enter new horsepower:")
hp = float(input())
print("Enter new weight:")
wt = float(input())
print("Enter number of cylinders:")
cy = float(input())
new_X = pd.DataFrame({'const': [1], 'horsepower': [hp], 'weight': [wt], 'cylinders': [cy]})
predicted_mpg = model_multi.predict(new_X)[0]
print(f"Predicted MPG for your car: {predicted_mpg:.2f}")
Best Practices and Troubleshooting#
- Always check for missing data and outliers.
- Keep an eye on multicollinearity using VIF.
- Try different models and be sure to visualize your results.
If your predictions do not make sense, double-check your data and calculations.
# Try making your model with only two predictors
X_two = mpg_clean[['weight', 'cylinders']]
X_two = sm.add_constant(X_two)
model_two = sm.OLS(y, X_two).fit()
print(model_two.summary())
Challenge Exercise#
Build a regression model to predict mpg using 'horsepower' and 'acceleration'.
How does the R-squared value compare to your model with three predictors?
Try visualizing the predictions too!
Recap: What You Learned#
- How to load and clean real datasets in Python
- How to calculate descriptive statistics, visualize distributions, and check relationships
- How to build and interpret multiple linear regression models
- How to predict new values using regression
Keep practicing and you will master these skills!
Thanks for Learning Multiple Regression in Python!#
If you found this helpful, like and subscribe to the channel for more beginner data science lessons!
See you next time.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



