Mathew K Analytics

Lesson 45 · Python for Retail E-commerce Analytics

Demand Prediction Using Machine Learning for Retail E-Commerce Analytics

We are solving the problem of predicting future customer demand for retail products. Accurate demand prediction helps retailers optimize inventory, reduce…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Demand Prediction Using Machine Learning#

  • We are solving the problem of predicting future customer demand for retail products.
  • Accurate demand prediction helps retailers optimize inventory, reduce stockouts, and boost sales.
  • You will learn to analyze retail transaction data and build machine learning models to estimate product demand.
  • Insights produced here can support better sales forecasts and smarter inventory decisions.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import mean_squared_error, r2_score
import warnings
warnings.filterwarnings('ignore')

Core Retail Analytics Concepts#

  • Retail datasets often contain transactions, orders, product, and customer details.
  • Sales metrics like revenue, quantity, and price are key for performance evaluation.
  • Common beginner mistakes include: ignoring missing or duplicate rows, misinterpreting product codes, or confusing revenue with unit sales.
  • Clean, structured data enables accurate demand prediction and actionable business insights.
# Beginner Example 1: Load retail orders dataset
np.random.seed(42)
n_orders = 1000
order_ids = list(range(1, n_orders + 1))
customer_ids = np.random.randint(1000, 1500, n_orders)
product_ids = np.random.randint(1001, 1100, n_orders)
quantities = np.random.randint(1, 5, n_orders)
order_dates = pd.date_range('2023-01-01', periods=n_orders, freq='h')
orders_df = pd.DataFrame({'OrderID': order_ids,'CustomerID': customer_ids,'ProductID': product_ids,'Quantity': quantities,'OrderDate': order_dates})
print(orders_df.shape)
print(orders_df.head(3))
(1000, 5)
   OrderID  CustomerID  ProductID  Quantity           OrderDate
0        1        1102       1049         3 2023-01-01 00:00:00
1        2        1435       1011         1 2023-01-01 01:00:00
2        3        1348       1085         3 2023-01-01 02:00:00
# Beginner Example 2: Quickly inspect for missing data
print('Missing values per column:')
print(orders_df.isnull().sum())
Missing values per column:
OrderID       0
CustomerID    0
ProductID     0
Quantity      0
OrderDate     0
dtype: int64
# Beginner Example 3: Check unique number of products and customers
n_products = orders_df['ProductID'].nunique()
n_customers = orders_df['CustomerID'].nunique()
print(f'Unique products: {n_products}')
print(f'Unique customers: {n_customers}')
Unique products: 99
Unique customers: 423
# Intermediate Example 1: Aggregate demand by ProductID
product_demand = orders_df.groupby('ProductID')['Quantity'].sum().sort_values(ascending=False)
print(product_demand.head(5))
ProductID
1098    54
1026    51
1017    50
1059    48
1040    45
Name: Quantity, dtype: int32
# Intermediate Example 2: Visualize top 10 product demand
top10 = product_demand.head(10)
top10.plot(kind='bar', figsize=(8,4))
plt.xlabel('ProductID')
plt.ylabel('Total Quantity Ordered')
plt.title('Top 10 Most Demanded Products')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Intermediate Example 3: Join orders to product catalog for price information
categories = ['Electronics','Clothing','Home','Sports','Beauty']
product_ids_cat = list(range(1001,1101))
product_categories = np.random.choice(categories,100)
product_prices = np.round(np.random.uniform(5,500,100),2)
catalog_df = pd.DataFrame({'ProductID':product_ids_cat,'Category':product_categories,'Price':product_prices})
orders_full = pd.merge(orders_df, catalog_df, on='ProductID', how='left')
print(orders_full.head(3))
   OrderID  CustomerID  ProductID  Quantity           OrderDate     Category  \
0        1        1102       1049         3 2023-01-01 00:00:00       Beauty   
1        2        1435       1011         1 2023-01-01 01:00:00       Sports   
2        3        1348       1085         3 2023-01-01 02:00:00  Electronics   

    Price  
0  233.43  
1  211.00  
2  314.21  
# Intermediate Example 4: Compute revenue per order
orders_full['Revenue'] = orders_full['Quantity'] * orders_full['Price']
print(orders_full[['OrderID','ProductID','Quantity','Price','Revenue']].head(3))
   OrderID  ProductID  Quantity   Price  Revenue
0        1       1049         3  233.43   700.29
1        2       1011         1  211.00   211.00
2        3       1085         3  314.21   942.63
# Intermediate Example 5: Category-level revenue analysis
category_revenue = orders_full.groupby('Category')['Revenue'].sum().sort_values(ascending=False)
print(category_revenue)
Category
Sports         191575.92
Clothing       122116.40
Beauty         103516.93
Home            90647.86
Electronics     90631.93
Name: Revenue, dtype: float64
# Intermediate Example 6: Time-based demand aggregation
orders_full['OrderDate'] = pd.to_datetime(orders_full['OrderDate'])
orders_full['Month'] = orders_full['OrderDate'].dt.to_period('M')
monthly_demand = orders_full.groupby('Month')['Quantity'].sum()
monthly_demand.plot(marker='o', figsize=(8,4))
plt.xlabel('Month')
plt.ylabel('Total Quantity Ordered')
plt.title('Monthly Retail Demand')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Advanced Example 1: Prepare demand features for prediction
features = orders_full[['Quantity','Price']]
target = orders_full['Revenue']
X_train, X_test, y_train, y_test = train_test_split(features, target, test_size=0.25, random_state=42)
print('Training records:', X_train.shape[0])
print('Test records:', X_test.shape[0])
Training records: 750
Test records: 250
# Advanced Example 2: Train a linear regression model for revenue prediction
lr = LinearRegression()
lr.fit(X_train, y_train)
y_pred_lr = lr.predict(X_test)
print('Linear Regression RMSE:', np.sqrt(mean_squared_error(y_test, y_pred_lr)))
print('R2 score:', r2_score(y_test, y_pred_lr))
Linear Regression RMSE: 163.2153445429137
R2 score: 0.8822716627343812
# Advanced Example 3: Random forest for potentially better prediction
rf = RandomForestRegressor(n_estimators=100, random_state=42)
rf.fit(X_train, y_train)
y_pred_rf = rf.predict(X_test)
print('Random Forest RMSE:', np.sqrt(mean_squared_error(y_test, y_pred_rf)))
print('R2 score:', r2_score(y_test, y_pred_rf))
Random Forest RMSE: 2.8743027977859135
R2 score: 0.999963488970974
# Advanced Example 4: Feature importance in Random Forest
importances = rf.feature_importances_
for feat, score in zip(['Quantity','Price'], importances):
    print(f'{feat}: {score:.2f}')
Quantity: 0.38
Price: 0.62
# Advanced Example 5: Predict next month's demand using model
future_price = [np.mean(orders_full['Price'])] * 10
future_quantity = np.linspace(1, 10, 10)
future_features = pd.DataFrame({'Quantity': future_quantity, 'Price': future_price})
future_prediction = rf.predict(future_features)
for qty, pred in zip(future_quantity, future_prediction):
    print(f'Predicted revenue for quantity={qty:.1f}: ${pred:.2f}')
Predicted revenue for quantity=1.0: $240.92
Predicted revenue for quantity=2.0: $486.63
Predicted revenue for quantity=3.0: $729.41
Predicted revenue for quantity=4.0: $968.09
Predicted revenue for quantity=5.0: $968.09
Predicted revenue for quantity=6.0: $968.09
Predicted revenue for quantity=7.0: $968.09
Predicted revenue for quantity=8.0: $968.09
Predicted revenue for quantity=9.0: $968.09
Predicted revenue for quantity=10.0: $968.09
# Error Handling Example 1: Test model with missing input values
test_data = pd.DataFrame({'Quantity': [3, np.nan], 'Price': [40, 25]})
try:
    rf.predict(test_data)
except ValueError as e:
    print('Error:', e)
# Error Handling Example 2: Incorrect aggregation by product/category
wrong_group = orders_full.groupby('Revenue')['Quantity'].sum()
print(wrong_group.head())
Revenue
10.46    2
10.99    2
11.00    1
14.50    4
15.37    2
Name: Quantity, dtype: int32
# Error Handling Example 3: Grouping by customer but taking product mean
customer_means = orders_full.groupby('CustomerID')['ProductID'].mean()
print(customer_means.head())
CustomerID
1000    1014.666667
1001    1044.666667
1003    1026.500000
1004    1048.000000
1005    1057.000000
Name: ProductID, dtype: float64
# Best Practice Example 1: Customer segmentation by total spend
customer_revenue = orders_full.groupby('CustomerID')['Revenue'].sum().sort_values(ascending=False)
top_customers = customer_revenue.head(5)
print(top_customers)
CustomerID
1053    6249.65
1098    6108.68
1372    5173.33
1416    5021.12
1251    4696.86
Name: Revenue, dtype: float64
# Best Practice Example 2: Product performance analysis at category level
cat_perf = orders_full.groupby('Category')['Quantity'].sum().sort_values(ascending=False)
print(cat_perf)
Category
Sports         707
Clothing       524
Home           452
Beauty         438
Electronics    369
Name: Quantity, dtype: int32
# Best Practice Example 3: Demand forecasting using previous months' averages
monthly_revenue = orders_full.groupby('Month')['Revenue'].sum()
avg_monthly_revenue = monthly_revenue.mean()
print('Average monthly revenue:', avg_monthly_revenue)
print('Recent months:')
print(monthly_revenue.tail(3))
Average monthly revenue: 299244.52
Recent months:
Month
2023-01    450184.27
2023-02    148304.77
Freq: M, Name: Revenue, dtype: float64
# Best Practice Example 4: Detecting seasonality in retail demand
orders_full['Weekday'] = orders_full['OrderDate'].dt.day_name()
weekday_demand = orders_full.groupby('Weekday')['Quantity'].sum()
weekday_demand = weekday_demand.reindex(['Monday', 'Tuesday', 'Wednesday', 'Thursday', 'Friday', 'Saturday', 'Sunday'])
weekday_demand.plot(kind='bar', figsize=(7,4))
plt.title('Demand by Day of Week')
plt.ylabel('Total Quantity Ordered')
plt.tight_layout()
plt.show()
No description has been provided for this image
# End-to-End Example: Identify best-selling products and make a recommendation
best_sellers = product_demand.head(3).index.tolist()
print('Top 3 products with highest demand:', best_sellers)
rec = f'It is recommended to ensure sufficient inventory and promote these top-selling products: {best_sellers}'
print(rec)
Top 3 products with highest demand: [1098, 1026, 1017]
It is recommended to ensure sufficient inventory and promote these top-selling products: [1098, 1026, 1017]
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.