Mathew K Analytics

Lesson 50 · Market Research Analytics in Python

Introductory Market Trend Forecasting

In this lesson, we will learn how to use real-world datasets to forecast market trends using Python. Market trend forecasting helps businesses understand…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Introductory Market Trend Forecasting#

  • In this lesson, we will learn how to use real-world datasets to forecast market trends using Python.
  • Market trend forecasting helps businesses understand future customer behaviors and industry movements.
  • You will analyze datasets to identify patterns, segment customers, and make actionable business recommendations.
  • We focus on practical analytics scenarios such as customer churn, campaign responses, and purchase trends.
import pandas as pd
import numpy as np
import openml
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')

Understanding core market research concepts#

  • Market research data often includes customer demographics, behaviors, opinions, and time-based transactions.
  • Survey responses may be categorical (like preferred brands) or numeric (such as satisfaction ratings and NPS).
  • Customer analytics datasets may also capture purchase and campaign history over specific time periods.
  • Beginners sometimes forget to handle missing values, or misinterpret categorical scales.
  • Always check what each column means and how time and responses are structured before analyzing market trends.
# BEGINNER EXAMPLE 1
# Load online retail transaction data to spot simple sales trends.
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/00502/online_retail_II.xlsx'
df_retail = pd.read_excel(url, sheet_name='Year 2010-2011')
df_retail['InvoiceDate'] = pd.to_datetime(df_retail['InvoiceDate'])
print(df_retail.shape)
print(df_retail.head(3))
(541910, 8)
  Invoice StockCode                         Description  Quantity  \
0  536365    85123A  WHITE HANGING HEART T-LIGHT HOLDER         6   
1  536365     71053                 WHITE METAL LANTERN         6   
2  536365    84406B      CREAM CUPID HEARTS COAT HANGER         8   

          InvoiceDate  Price  Customer ID         Country  
0 2010-12-01 08:26:00   2.55      17850.0  United Kingdom  
1 2010-12-01 08:26:00   3.39      17850.0  United Kingdom  
2 2010-12-01 08:26:00   2.75      17850.0  United Kingdom  
# BEGINNER EXAMPLE 2
# Plot the number of transactions by month to see overall sales trend.
df_retail['Month'] = df_retail['InvoiceDate'].dt.to_period('M')
monthly_sales = df_retail.groupby('Month').size()
monthly_sales.plot(kind='line', marker='o', figsize=(10,5))
plt.title('Number of Transactions Per Month')
plt.ylabel('Transactions')
plt.xlabel('Month')
plt.show()
No description has been provided for this image
# BEGINNER EXAMPLE 3
# Load the customer satisfaction survey dataset for churn trend analysis.
dataset = openml.datasets.get_dataset(42178)
df_survey, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_survey.shape)
print(df_survey[['tenure','MonthlyCharges','Churn']].head(3))
(7043, 20)
   tenure  MonthlyCharges Churn
0       1           29.85    No
1      34           56.95    No
2       2           53.85   Yes
# BEGINNER EXAMPLE 4
# Calculate churn rate over customer tenure groups.
df_survey['tenure_group'] = pd.cut(df_survey['tenure'], bins=[0,12,24,36,48,60,72], labels=['0-12','13-24','25-36','37-48','49-60','61-72'])
churn_trend = df_survey.groupby('tenure_group')['Churn'].value_counts(normalize=True).unstack().fillna(0)
print(churn_trend)
Churn               No       Yes
tenure_group                    
0-12          0.523218  0.476782
13-24         0.712891  0.287109
25-36         0.783654  0.216346
37-48         0.809711  0.190289
49-60         0.855769  0.144231
61-72         0.933902  0.066098
# BEGINNER EXAMPLE 5
# Visualize churn rate trend for different tenure groups.
churn_trend['Yes'].plot(kind='bar', color='salmon', figsize=(8,5))
plt.title('Churn Rate by Customer Tenure Group')
plt.ylabel('Churn Rate')
plt.xlabel('Tenure Group (months)')
plt.ylim(0,1)
plt.show()
No description has been provided for this image
# BEGINNER EXAMPLE 6
# Load marketing campaign response dataset.
dataset = openml.datasets.get_dataset(1461)
df_campaign, _, _, _ = dataset.get_data(dataset_format='dataframe')
df_campaign.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
print(df_campaign.shape)
print(df_campaign[['month','response']].head(3))
(45211, 17)
  month response
0   may        1
1   may        1
2   may        1
# BEGINNER EXAMPLE 7
# Trend: Campaign response rates by month.
monthly_response = df_campaign.groupby('month')['response'].value_counts(normalize=True).unstack().fillna(0)
print(monthly_response)
response         1         2
month                       
apr       0.803206  0.196794
aug       0.889867  0.110133
dec       0.532710  0.467290
feb       0.833522  0.166478
jan       0.898788  0.101212
jul       0.909065  0.090935
jun       0.897772  0.102228
mar       0.480084  0.519916
may       0.932805  0.067195
nov       0.898489  0.101511
oct       0.562331  0.437669
sep       0.535406  0.464594
# BEGINNER EXAMPLE 8
# Visualize trend: Yes response rate by campaign month.
monthly_response['2'].plot(kind='bar', color='green', figsize=(8,5))
plt.title('Campaign Yes Response Rate by Month')
plt.ylabel('Response Rate')
plt.xlabel('Month')
plt.ylim(0,1)
plt.show()
No description has been provided for this image
# INTERMEDIATE EXAMPLE 1
# Calculate and plot average monthly revenue (retail).
df_retail['Revenue'] = df_retail['Quantity'] * df_retail['Price']
monthly_rev = df_retail.groupby('Month')['Revenue'].sum()
monthly_rev.plot(kind='line', marker='o', color='royalblue', figsize=(10,5))
plt.title('Monthly Revenue Trend')
plt.ylabel('Total Revenue')
plt.xlabel('Month')
plt.show()
No description has been provided for this image
# INTERMEDIATE EXAMPLE 2
# Segment revenue trends by country for cross-market insights.
country_rev = df_retail.groupby(['Month','Country'])['Revenue'].sum().unstack().fillna(0)
top_countries = country_rev.sum().sort_values(ascending=False).head(5).index
country_rev[top_countries].plot(figsize=(12,6), marker='o')
plt.title('Monthly Revenue Trend for Top 5 Countries')
plt.ylabel('Total Revenue')
plt.xlabel('Month')
plt.legend(title='Country')
plt.show()
No description has been provided for this image
# INTERMEDIATE EXAMPLE 3
# Trend analysis: Customer churn by monthly charges level.
df_survey['charge_group'] = pd.cut(df_survey['MonthlyCharges'], bins=[0,30,60,90,120], labels=['Low','Medium','High','Very High'])
charge_churn = df_survey.groupby('charge_group')['Churn'].value_counts(normalize=True).unstack().fillna(0)
charge_churn['Yes'].plot(kind='bar', color='tomato', figsize=(8,5))
plt.title('Churn Rate Trend by Monthly Charges Group')
plt.ylabel('Churn Rate')
plt.xlabel('Monthly Charges Group')
plt.ylim(0,1)
plt.show()
No description has been provided for this image
# INTERMEDIATE EXAMPLE 4
# Trend analysis: NPS trends by region (synthetic NPS survey data).
np.random.seed(42)
df_nps = pd.DataFrame({'CustomerID': range(1,501), 'Age': np.random.randint(18,70,500), 'Region': np.random.choice(['North','South','East','West'],500), 'NPS_Score': np.random.randint(0,11,500)})
nps_region = df_nps.groupby('Region')['NPS_Score'].mean()
nps_region.plot(kind='bar', color='teal', figsize=(6,4))
plt.title('Average NPS Score by Region')
plt.ylabel('Average NPS')
plt.xlabel('Region')
plt.ylim(0,10)
plt.show()
No description has been provided for this image
# INTERMEDIATE EXAMPLE 5
# Analyze market response trend by customer segment (age band, campaign data).
df_campaign['age_band'] = pd.cut(df_campaign['age'], bins=[15,30,45,60,95], labels=['<30','30-44','45-59','60+'])
age_response = df_campaign.groupby('age_band')['response'].value_counts(normalize=True).unstack().fillna(0)
age_response['2'].plot(kind='bar', color='violet', figsize=(6,4))
plt.title('Campaign Yes Response Rate by Age Group')
plt.ylabel('Yes Response Rate')
plt.xlabel('Age Group')
plt.ylim(0,1)
plt.show()
No description has been provided for this image
# ADVANCED EXAMPLE 1
# Advanced: Detect recent sales peaks or declines using moving average (retail data).
monthly_rev_ma = monthly_rev.rolling(window=3, center=False).mean()
plt.figure(figsize=(10,5))
plt.plot(monthly_rev.index.to_timestamp(), monthly_rev, label='Monthly Revenue')
plt.plot(monthly_rev.index.to_timestamp(), monthly_rev_ma, label='3-Month Moving Average', linestyle='--')
plt.title('Monthly Revenue and Moving Average Trend')
plt.ylabel('Revenue')
plt.xlabel('Month')
plt.legend()
plt.show()
No description has been provided for this image
# ADVANCED EXAMPLE 2
# Advanced: Forecast next months revenue using simple linear model (retail data).
from sklearn.linear_model import LinearRegression
months = np.arange(len(monthly_rev)).reshape(-1,1)
lr = LinearRegression()
lr.fit(months, monthly_rev.values)
next_month = np.array([[len(months)]])
predicted_revenue = lr.predict(next_month)[0]
print(f'Predicted revenue for next month: {predicted_revenue:.2f}')
Predicted revenue for next month: 990362.65
# ADVANCED EXAMPLE 3
# Advanced: Trend in customer cohort retention (synthetic cohort data).
np.random.seed(0)
dates = pd.date_range('2021-01-01', periods=24, freq='M')
df_cohort = pd.DataFrame({'CustomerID': np.random.randint(1000,2000,len(dates)), 'Signup_Month': dates, 'Active_Users': np.random.randint(50,300,len(dates))})
df_cohort.set_index('Signup_Month', inplace=True)
df_cohort['Active_Users'].plot(figsize=(10,4), marker='o')
plt.title('Active Customer Cohort Trend Over Time')
plt.ylabel('Active Users')
plt.xlabel('Signup Month')
plt.show()
No description has been provided for this image
# ERROR HANDLING 1
# Simulate missing survey responses in churn dataset and fill missing values.
df_survey_missing = df_survey.copy()
df_survey_missing.loc[df_survey_missing.sample(frac=0.05, random_state=42).index,'Churn'] = np.nan
print(f'Missing Churn entries: {df_survey_missing.Churn.isnull().sum()}')
df_survey_missing['Churn_filled'] = df_survey_missing['Churn'].fillna('No')
print(df_survey_missing[['Churn', 'Churn_filled']].head(10))
Missing Churn entries: 352
  Churn Churn_filled
0    No           No
1    No           No
2   Yes          Yes
3    No           No
4   Yes          Yes
5   Yes          Yes
6    No           No
7    No           No
8   Yes          Yes
9    No           No
# ERROR HANDLING 2
# Demonstrate incorrect aggregation: grouping NPS scores by a non-category.
try:
    print(df_nps.groupby('CustomerID')['NPS_Score'].mean().head())
except Exception as e:
    print('Error:', e)
correct = df_nps.groupby('Region')['NPS_Score'].mean()
print('\nCorrect: Mean NPS by region:')
print(correct)
CustomerID
1    2.0
2    0.0
3    4.0
4    3.0
5    9.0
Name: NPS_Score, dtype: float64

Correct: Mean NPS by region:
Region
East     4.504425
North    5.214765
South    4.719008
West     5.025641
Name: NPS_Score, dtype: float64
# ERROR HANDLING 3
# Common mistake: Misinterpreting NPS scale (what is good vs bad).
df_nps['Promoter'] = df_nps['NPS_Score'] >= 9
df_nps['Detractor'] = df_nps['NPS_Score'] <= 6
nps_summary = df_nps.groupby('Region')[['Promoter','Detractor']].mean()
print(nps_summary)
        Promoter  Detractor
Region                     
East    0.150442   0.672566
North   0.167785   0.577181
South   0.148760   0.694215
West    0.213675   0.658120

Market Research Analytics Best Practices#

  • Segment your data by meaningful customer attributes, like tenure, charges, or region.
  • Cross-tabulate outcomes to compare segments and trends over time.
  • Construct scores (like NPS or churn rate) using standardized business rules.
  • Always visualize trends to spot real business shifts, not just random noise.
  • Document your assumptions and investigate possible causes behind trend changes.
# PATTERN 1: Segmentation
# Segment marketing campaign responses by education level.
edu_response = df_campaign.groupby('education')['response'].value_counts(normalize=True).unstack().fillna(0)
print(edu_response)
response          1         2
education                    
primary    0.913735  0.086265
secondary  0.894406  0.105594
tertiary   0.849936  0.150064
unknown    0.864297  0.135703
# PATTERN 2: Trend cross-tabulation
# Compare churn trend by both tenure group and churn status.
pd.crosstab(df_survey['tenure_group'], df_survey['Churn'], normalize='index')
Churn No Yes
tenure_group
0-12 0.523218 0.476782
13-24 0.712891 0.287109
25-36 0.783654 0.216346
37-48 0.809711 0.190289
49-60 0.855769 0.144231
61-72 0.933902 0.066098
# PATTERN 3: Constructing a trend index
# Build a simple Sales Trend Index based on moving average (retail).
sales_index = monthly_rev / monthly_rev_ma
print(sales_index.tail())
Month
2011-08    0.996564
2011-09    1.283343
2011-10    1.158323
2011-11    1.234540
2011-12    0.438636
Freq: M, Name: Revenue, dtype: float64
# END-TO-END MINI PROJECT
# Goal: Forecast next months customer churn for high-charge customers.
high_charge = df_survey[df_survey['MonthlyCharges'] > 80]
churn_counts = high_charge.groupby('tenure_group')['Churn'].value_counts().unstack().fillna(0)
tenure_by_month = churn_counts.index.astype(str)
plt.figure(figsize=(8,4))
plt.bar(tenure_by_month, churn_counts['Yes'], color='coral')
plt.title('Churned High-Charge Customers by Tenure Group')
plt.xlabel('Tenure Group (months)')
plt.ylabel('Churned Customers')
plt.show()
recent_churn = churn_counts['Yes'].iloc[-1]
print(f'Latest period churned customers: {recent_churn}')
print('Recommendation: Target retention offers at customers in this tenure group to reduce forecasted churn.')
No description has been provided for this image
Latest period churned customers: 80
Recommendation: Target retention offers at customers in this tenure group to reduce forecasted churn.
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.