Lesson 45 · Python for Retail E-commerce Analytics
Demand Prediction Using Machine Learning for Retail E-Commerce Analytics
We are solving the problem of predicting future customer demand for retail products. Accurate demand prediction helps retailers optimize inventory, reduce…
- CoursePython for Retail E-commerce Analytics
- Lesson45 of 43
- Video23 min
- FormatJupyter notebook · 24 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbDemand Prediction Using Machine Learning#
- We are solving the problem of predicting future customer demand for retail products.
- Accurate demand prediction helps retailers optimize inventory, reduce stockouts, and boost sales.
- You will learn to analyze retail transaction data and build machine learning models to estimate product demand.
- Insights produced here can support better sales forecasts and smarter inventory decisions.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import mean_squared_error, r2_score
import warnings
warnings.filterwarnings('ignore')
Core Retail Analytics Concepts#
- Retail datasets often contain transactions, orders, product, and customer details.
- Sales metrics like revenue, quantity, and price are key for performance evaluation.
- Common beginner mistakes include: ignoring missing or duplicate rows, misinterpreting product codes, or confusing revenue with unit sales.
- Clean, structured data enables accurate demand prediction and actionable business insights.
# Beginner Example 1: Load retail orders dataset
np.random.seed(42)
n_orders = 1000
order_ids = list(range(1, n_orders + 1))
customer_ids = np.random.randint(1000, 1500, n_orders)
product_ids = np.random.randint(1001, 1100, n_orders)
quantities = np.random.randint(1, 5, n_orders)
order_dates = pd.date_range('2023-01-01', periods=n_orders, freq='h')
orders_df = pd.DataFrame({'OrderID': order_ids,'CustomerID': customer_ids,'ProductID': product_ids,'Quantity': quantities,'OrderDate': order_dates})
print(orders_df.shape)
print(orders_df.head(3))
# Beginner Example 2: Quickly inspect for missing data
print('Missing values per column:')
print(orders_df.isnull().sum())
# Beginner Example 3: Check unique number of products and customers
n_products = orders_df['ProductID'].nunique()
n_customers = orders_df['CustomerID'].nunique()
print(f'Unique products: {n_products}')
print(f'Unique customers: {n_customers}')
# Intermediate Example 1: Aggregate demand by ProductID
product_demand = orders_df.groupby('ProductID')['Quantity'].sum().sort_values(ascending=False)
print(product_demand.head(5))
# Intermediate Example 2: Visualize top 10 product demand
top10 = product_demand.head(10)
top10.plot(kind='bar', figsize=(8,4))
plt.xlabel('ProductID')
plt.ylabel('Total Quantity Ordered')
plt.title('Top 10 Most Demanded Products')
plt.tight_layout()
plt.show()
# Intermediate Example 3: Join orders to product catalog for price information
categories = ['Electronics','Clothing','Home','Sports','Beauty']
product_ids_cat = list(range(1001,1101))
product_categories = np.random.choice(categories,100)
product_prices = np.round(np.random.uniform(5,500,100),2)
catalog_df = pd.DataFrame({'ProductID':product_ids_cat,'Category':product_categories,'Price':product_prices})
orders_full = pd.merge(orders_df, catalog_df, on='ProductID', how='left')
print(orders_full.head(3))
# Intermediate Example 4: Compute revenue per order
orders_full['Revenue'] = orders_full['Quantity'] * orders_full['Price']
print(orders_full[['OrderID','ProductID','Quantity','Price','Revenue']].head(3))
# Intermediate Example 5: Category-level revenue analysis
category_revenue = orders_full.groupby('Category')['Revenue'].sum().sort_values(ascending=False)
print(category_revenue)
# Intermediate Example 6: Time-based demand aggregation
orders_full['OrderDate'] = pd.to_datetime(orders_full['OrderDate'])
orders_full['Month'] = orders_full['OrderDate'].dt.to_period('M')
monthly_demand = orders_full.groupby('Month')['Quantity'].sum()
monthly_demand.plot(marker='o', figsize=(8,4))
plt.xlabel('Month')
plt.ylabel('Total Quantity Ordered')
plt.title('Monthly Retail Demand')
plt.tight_layout()
plt.show()
# Advanced Example 1: Prepare demand features for prediction
features = orders_full[['Quantity','Price']]
target = orders_full['Revenue']
X_train, X_test, y_train, y_test = train_test_split(features, target, test_size=0.25, random_state=42)
print('Training records:', X_train.shape[0])
print('Test records:', X_test.shape[0])
# Advanced Example 2: Train a linear regression model for revenue prediction
lr = LinearRegression()
lr.fit(X_train, y_train)
y_pred_lr = lr.predict(X_test)
print('Linear Regression RMSE:', np.sqrt(mean_squared_error(y_test, y_pred_lr)))
print('R2 score:', r2_score(y_test, y_pred_lr))
# Advanced Example 3: Random forest for potentially better prediction
rf = RandomForestRegressor(n_estimators=100, random_state=42)
rf.fit(X_train, y_train)
y_pred_rf = rf.predict(X_test)
print('Random Forest RMSE:', np.sqrt(mean_squared_error(y_test, y_pred_rf)))
print('R2 score:', r2_score(y_test, y_pred_rf))
# Advanced Example 4: Feature importance in Random Forest
importances = rf.feature_importances_
for feat, score in zip(['Quantity','Price'], importances):
print(f'{feat}: {score:.2f}')
# Advanced Example 5: Predict next month's demand using model
future_price = [np.mean(orders_full['Price'])] * 10
future_quantity = np.linspace(1, 10, 10)
future_features = pd.DataFrame({'Quantity': future_quantity, 'Price': future_price})
future_prediction = rf.predict(future_features)
for qty, pred in zip(future_quantity, future_prediction):
print(f'Predicted revenue for quantity={qty:.1f}: ${pred:.2f}')
# Error Handling Example 1: Test model with missing input values
test_data = pd.DataFrame({'Quantity': [3, np.nan], 'Price': [40, 25]})
try:
rf.predict(test_data)
except ValueError as e:
print('Error:', e)
# Error Handling Example 2: Incorrect aggregation by product/category
wrong_group = orders_full.groupby('Revenue')['Quantity'].sum()
print(wrong_group.head())
# Error Handling Example 3: Grouping by customer but taking product mean
customer_means = orders_full.groupby('CustomerID')['ProductID'].mean()
print(customer_means.head())
# Best Practice Example 1: Customer segmentation by total spend
customer_revenue = orders_full.groupby('CustomerID')['Revenue'].sum().sort_values(ascending=False)
top_customers = customer_revenue.head(5)
print(top_customers)
# Best Practice Example 2: Product performance analysis at category level
cat_perf = orders_full.groupby('Category')['Quantity'].sum().sort_values(ascending=False)
print(cat_perf)
# Best Practice Example 3: Demand forecasting using previous months' averages
monthly_revenue = orders_full.groupby('Month')['Revenue'].sum()
avg_monthly_revenue = monthly_revenue.mean()
print('Average monthly revenue:', avg_monthly_revenue)
print('Recent months:')
print(monthly_revenue.tail(3))
# Best Practice Example 4: Detecting seasonality in retail demand
orders_full['Weekday'] = orders_full['OrderDate'].dt.day_name()
weekday_demand = orders_full.groupby('Weekday')['Quantity'].sum()
weekday_demand = weekday_demand.reindex(['Monday', 'Tuesday', 'Wednesday', 'Thursday', 'Friday', 'Saturday', 'Sunday'])
weekday_demand.plot(kind='bar', figsize=(7,4))
plt.title('Demand by Day of Week')
plt.ylabel('Total Quantity Ordered')
plt.tight_layout()
plt.show()
# End-to-End Example: Identify best-selling products and make a recommendation
best_sellers = product_demand.head(3).index.tolist()
print('Top 3 products with highest demand:', best_sellers)
rec = f'It is recommended to ensure sufficient inventory and promote these top-selling products: {best_sellers}'
print(rec)
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



