Lesson 17 · Data Science Projects
Building a Car Price Prediction Model Using Python and Data Science Techniques
Learn how to analyze and predict car prices using a real world dataset. Data driven insights help buyers, sellers, and businesses make informed decisions.…
- CourseData Science Projects
- Lesson17 of 33
- Video23 min
- FormatJupyter notebook · 18 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbCar Price Prediction Analysis with Python#
- Learn how to analyze and predict car prices using a real world dataset.
- Data driven insights help buyers, sellers, and businesses make informed decisions.
- This lesson guides you step by step no prior experience required.
- We will clean, explore, visualize, and model the car data together.
- Ready to discover patterns and build your first price predictor?
# Always suppress warnings for smoother runs
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
Dataset and Problem Overview#
- We will use a public car price dataset.
- Each row is a car with features like brand, model, type, price, and specs.
- Our goal is to understand what affects car price.
- Later, we will build a simple model to predict price from features.
# Data setup
import pandas as pd
url = 'https://raw.githubusercontent.com/selva86/datasets/master/Cars93_miss.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
# Quick data info: see types and missing values
df.info()
# Count missing values in each column
df.isnull().sum().sort_values(ascending=False)
# Remove rows where Price or Manufacturer is missing
df_clean = df.dropna(subset=['Price', 'Manufacturer'])
print('Rows after drop:', df_clean.shape[0])
# Fill remaining missing numeric values with their column median
num_cols = df_clean.select_dtypes(include='number').columns
df_clean[num_cols] = df_clean[num_cols].fillna(df_clean[num_cols].median())
Exploring Key Features#
- Price is our target the thing we want to analyze and predict.
- Manufacturer and Type help explain patterns.
- Numerical features like horsepower, weight, and MPG affect value.
- Let us start by visualizing price and its relationships.
# Basic statistics for price
df_clean['Price'].describe()
# Plot the price distribution
import seaborn as sns
import matplotlib.pyplot as plt
sns.histplot(df_clean['Price'], bins=20, kde=True)
plt.title('Distribution of Car Prices')
plt.xlabel('Price in $1000s')
plt.show()
# Average price by manufacturer (top 10)
mean_price = df_clean.groupby('Manufacturer')['Price'].mean().sort_values(ascending=False).head(10)
mean_price.plot(kind='bar', figsize=(8,4), color='skyblue')
plt.ylabel('Average Price ($1000s)')
plt.title('Top 10 Manufacturers by Average Car Price')
plt.show()
# Scatter plot: price vs. horsepower
sns.scatterplot(data=df_clean, x='Horsepower', y='Price', alpha=0.7)
plt.title('Car Price vs Horsepower')
plt.xlabel('Horsepower')
plt.ylabel('Price ($1000s)')
plt.show()
# Boxplot: price by car type
sns.boxplot(data=df_clean, x='Type', y='Price')
plt.title('Price by Car Type')
plt.xlabel('Type')
plt.ylabel('Price ($1000s)')
plt.xticks(rotation=30)
plt.show()
# Correlation matrix: how features relate to price and each other
corr = df_clean.select_dtypes(include='number').corr()
sns.heatmap(corr, annot=True, fmt='.1f', cmap='Blues', vmin=-1, vmax=1)
plt.title('Correlation Heatmap (numeric features)')
plt.show()
What Patterns Did You See?#
- Take a moment to reflect on which features matter most.
- Does price rise with horsepower or weight?
- Is there overlap among car types and brands?
- Jot down one surprise from the plots.
# Prepare data for modeling: select features and convert categories
features = ['Horsepower', 'RPM', 'EngineSize', 'MPG.city', 'MPG.highway', 'Weight', 'Length', 'Type']
df_model = df_clean[features + ['Price']].dropna()
df_model = pd.get_dummies(df_model, columns=['Type'], drop_first=True)
print(df_model.columns)
# Split into training and test sets
from sklearn.model_selection import train_test_split
X = df_model.drop('Price', axis=1)
y = df_model['Price']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print('Training samples:', X_train.shape[0])
# Linear regression model: predict price from features
from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(X_train, y_train)
print('Model coefficients:', model.coef_)
print('Intercept:', model.intercept_)
# Predict and compare to actual prices on test data
y_pred = model.predict(X_test)
results = pd.DataFrame({'Actual': y_test, 'Predicted': y_pred})
print(results.head())
# Evaluate accuracy with Root Mean Squared Error
from sklearn.metrics import mean_squared_error
import numpy as np
rmse = np.sqrt(mean_squared_error(y_test, y_pred))
print('Root Mean Squared Error:', round(rmse, 2))
# Plot actual vs. predicted prices
plt.scatter(y_test, y_pred, alpha=0.7)
plt.xlabel('Actual Price ($1000s)')
plt.ylabel('Predicted Price ($1000s)')
plt.title('Actual vs Predicted Car Prices')
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'r--')
plt.show()
Practice: Try It Yourself!#
- Change a feature like using only 'Horsepower' and 'Weight' to predict price.
- Or swap car types by editing the data setup.
- What does that do to accuracy?
- Share your results or questions in the video comments below!
Congratulations and Next Steps#
- You learned to load, clean, explore, visualize, and model real car data.
- Price prediction helps us understand data and make smarter choices.
- Try more datasets or test other regression models.
- If you enjoyed this, hit Like and Subscribe on our YouTube channel for more practical data mining lessons!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



