Mathew K Analytics

Lesson 17 · Data Science Projects

Building a Car Price Prediction Model Using Python and Data Science Techniques

Learn how to analyze and predict car prices using a real world dataset. Data driven insights help buyers, sellers, and businesses make informed decisions.…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Car Price Prediction Analysis with Python#

  • Learn how to analyze and predict car prices using a real world dataset.
  • Data driven insights help buyers, sellers, and businesses make informed decisions.
  • This lesson guides you step by step no prior experience required.
  • We will clean, explore, visualize, and model the car data together.
  • Ready to discover patterns and build your first price predictor?
# Always suppress warnings for smoother runs
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)

Dataset and Problem Overview#

  • We will use a public car price dataset.
  • Each row is a car with features like brand, model, type, price, and specs.
  • Our goal is to understand what affects car price.
  • Later, we will build a simple model to predict price from features.
# Data setup
import pandas as pd
url = 'https://raw.githubusercontent.com/selva86/datasets/master/Cars93_miss.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(93, 27)
  Manufacturer    Model     Type  Min.Price  Price  Max.Price  MPG.city  \
0        Acura  Integra    Small       12.9   15.9       18.8      25.0   
1          NaN   Legend  Midsize       29.2   33.9       38.7      18.0   
2         Audi       90  Compact       25.9   29.1       32.3      20.0   

   MPG.highway             AirBags DriveTrain  ... Passengers  Length  \
0         31.0                 NaN      Front  ...        5.0   177.0   
1         25.0  Driver & Passenger      Front  ...        5.0   195.0   
2         26.0         Driver only      Front  ...        5.0   180.0   

   Wheelbase  Width  Turn.circle Rear.seat.room  Luggage.room  Weight  \
0      102.0   68.0         37.0           26.5           NaN  2705.0   
1      115.0   71.0         38.0           30.0          15.0  3560.0   
2      102.0   67.0         37.0           28.0          14.0  3375.0   

    Origin           Make  
0  non-USA  Acura Integra  
1  non-USA   Acura Legend  
2  non-USA        Audi 90  

[3 rows x 27 columns]
# Quick data info: see types and missing values
df.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 93 entries, 0 to 92
Data columns (total 27 columns):
 #   Column              Non-Null Count  Dtype  
---  ------              --------------  -----  
 0   Manufacturer        89 non-null     object 
 1   Model               92 non-null     object 
 2   Type                90 non-null     object 
 3   Min.Price           86 non-null     float64
 4   Price               91 non-null     float64
 5   Max.Price           88 non-null     float64
 6   MPG.city            84 non-null     float64
 7   MPG.highway         91 non-null     float64
 8   AirBags             55 non-null     object 
 9   DriveTrain          86 non-null     object 
 10  Cylinders           88 non-null     object 
 11  EngineSize          91 non-null     float64
 12  Horsepower          86 non-null     float64
 13  RPM                 90 non-null     float64
 14  Rev.per.mile        87 non-null     float64
 15  Man.trans.avail     88 non-null     object 
 16  Fuel.tank.capacity  85 non-null     float64
 17  Passengers          91 non-null     float64
 18  Length              89 non-null     float64
 19  Wheelbase           92 non-null     float64
 20  Width               87 non-null     float64
 21  Turn.circle         88 non-null     float64
 22  Rear.seat.room      89 non-null     float64
 23  Luggage.room        74 non-null     float64
 24  Weight              86 non-null     float64
 25  Origin              88 non-null     object 
 26  Make                90 non-null     object 
dtypes: float64(18), object(9)
memory usage: 19.7+ KB
# Count missing values in each column
df.isnull().sum().sort_values(ascending=False)
AirBags               38
Luggage.room          19
MPG.city               9
Fuel.tank.capacity     8
DriveTrain             7
Horsepower             7
Min.Price              7
Weight                 7
Rev.per.mile           6
Width                  6
Cylinders              5
Turn.circle            5
Max.Price              5
Man.trans.avail        5
Origin                 5
Rear.seat.room         4
Length                 4
Manufacturer           4
RPM                    3
Type                   3
Make                   3
Passengers             2
EngineSize             2
MPG.highway            2
Price                  2
Wheelbase              1
Model                  1
dtype: int64
# Remove rows where Price or Manufacturer is missing
df_clean = df.dropna(subset=['Price', 'Manufacturer'])
print('Rows after drop:', df_clean.shape[0])
Rows after drop: 87
# Fill remaining missing numeric values with their column median
num_cols = df_clean.select_dtypes(include='number').columns
df_clean[num_cols] = df_clean[num_cols].fillna(df_clean[num_cols].median())

Exploring Key Features#

  • Price is our target the thing we want to analyze and predict.
  • Manufacturer and Type help explain patterns.
  • Numerical features like horsepower, weight, and MPG affect value.
  • Let us start by visualizing price and its relationships.
# Basic statistics for price
df_clean['Price'].describe()
count    87.000000
mean     19.205747
std       9.643257
min       7.400000
25%      12.150000
50%      16.600000
75%      22.700000
max      61.900000
Name: Price, dtype: float64
# Plot the price distribution
import seaborn as sns
import matplotlib.pyplot as plt
sns.histplot(df_clean['Price'], bins=20, kde=True)
plt.title('Distribution of Car Prices')
plt.xlabel('Price in $1000s')
plt.show()
No description has been provided for this image
# Average price by manufacturer (top 10)
mean_price = df_clean.groupby('Manufacturer')['Price'].mean().sort_values(ascending=False).head(10)
mean_price.plot(kind='bar', figsize=(8,4), color='skyblue')
plt.ylabel('Average Price ($1000s)')
plt.title('Top 10 Manufacturers by Average Car Price')
plt.show()
No description has been provided for this image
# Scatter plot: price vs. horsepower
sns.scatterplot(data=df_clean, x='Horsepower', y='Price', alpha=0.7)
plt.title('Car Price vs Horsepower')
plt.xlabel('Horsepower')
plt.ylabel('Price ($1000s)')
plt.show()
No description has been provided for this image
# Boxplot: price by car type
sns.boxplot(data=df_clean, x='Type', y='Price')
plt.title('Price by Car Type')
plt.xlabel('Type')
plt.ylabel('Price ($1000s)')
plt.xticks(rotation=30)
plt.show()
No description has been provided for this image
# Correlation matrix: how features relate to price and each other
corr = df_clean.select_dtypes(include='number').corr()
sns.heatmap(corr, annot=True, fmt='.1f', cmap='Blues', vmin=-1, vmax=1)
plt.title('Correlation Heatmap (numeric features)')
plt.show()
No description has been provided for this image

What Patterns Did You See?#

  • Take a moment to reflect on which features matter most.
  • Does price rise with horsepower or weight?
  • Is there overlap among car types and brands?
  • Jot down one surprise from the plots.
# Prepare data for modeling: select features and convert categories
features = ['Horsepower', 'RPM', 'EngineSize', 'MPG.city', 'MPG.highway', 'Weight', 'Length', 'Type']
df_model = df_clean[features + ['Price']].dropna()
df_model = pd.get_dummies(df_model, columns=['Type'], drop_first=True)
print(df_model.columns)
Index(['Horsepower', 'RPM', 'EngineSize', 'MPG.city', 'MPG.highway', 'Weight',
       'Length', 'Price', 'Type_Large', 'Type_Midsize', 'Type_Small',
       'Type_Sporty', 'Type_Van'],
      dtype='object')
# Split into training and test sets
from sklearn.model_selection import train_test_split
X = df_model.drop('Price', axis=1)
y = df_model['Price']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print('Training samples:', X_train.shape[0])
Training samples: 67
# Linear regression model: predict price from features
from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(X_train, y_train)
print('Model coefficients:', model.coef_)
print('Intercept:', model.intercept_)
Model coefficients: [ 1.18885414e-01 -5.96844933e-04 -3.72082840e-01  1.30478439e-02
 -2.49091995e-01 -1.23263227e-03  1.44148700e-02  7.69104630e-01
  2.24607040e+00 -3.88410769e+00 -2.41140576e+00 -2.17429120e+00]
Intercept: 15.548841926775339
# Predict and compare to actual prices on test data
y_pred = model.predict(X_test)
results = pd.DataFrame({'Actual': y_test, 'Predicted': y_pred})
print(results.head())
    Actual  Predicted
79     8.4   7.352552
0     15.9  15.700326
62    26.1  29.510208
24    13.3  16.152063
13    15.1  20.210209
# Evaluate accuracy with Root Mean Squared Error
from sklearn.metrics import mean_squared_error
import numpy as np
rmse = np.sqrt(mean_squared_error(y_test, y_pred))
print('Root Mean Squared Error:', round(rmse, 2))
Root Mean Squared Error: 8.59
# Plot actual vs. predicted prices
plt.scatter(y_test, y_pred, alpha=0.7)
plt.xlabel('Actual Price ($1000s)')
plt.ylabel('Predicted Price ($1000s)')
plt.title('Actual vs Predicted Car Prices')
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'r--')
plt.show()
No description has been provided for this image

Practice: Try It Yourself!#

  • Change a feature like using only 'Horsepower' and 'Weight' to predict price.
  • Or swap car types by editing the data setup.
  • What does that do to accuracy?
  • Share your results or questions in the video comments below!

Congratulations and Next Steps#

  • You learned to load, clean, explore, visualize, and model real car data.
  • Price prediction helps us understand data and make smarter choices.
  • Try more datasets or test other regression models.
  • If you enjoyed this, hit Like and Subscribe on our YouTube channel for more practical data mining lessons!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.