Lesson 50 · Market Research Analytics in Python
Introductory Market Trend Forecasting
In this lesson, we will learn how to use real-world datasets to forecast market trends using Python. Market trend forecasting helps businesses understand…
- CourseMarket Research Analytics in Python
- Lesson50 of 56
- Video27 min
- FormatJupyter notebook · 25 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIntroductory Market Trend Forecasting#
- In this lesson, we will learn how to use real-world datasets to forecast market trends using Python.
- Market trend forecasting helps businesses understand future customer behaviors and industry movements.
- You will analyze datasets to identify patterns, segment customers, and make actionable business recommendations.
- We focus on practical analytics scenarios such as customer churn, campaign responses, and purchase trends.
import pandas as pd
import numpy as np
import openml
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')
Understanding core market research concepts#
- Market research data often includes customer demographics, behaviors, opinions, and time-based transactions.
- Survey responses may be categorical (like preferred brands) or numeric (such as satisfaction ratings and NPS).
- Customer analytics datasets may also capture purchase and campaign history over specific time periods.
- Beginners sometimes forget to handle missing values, or misinterpret categorical scales.
- Always check what each column means and how time and responses are structured before analyzing market trends.
# BEGINNER EXAMPLE 1
# Load online retail transaction data to spot simple sales trends.
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/00502/online_retail_II.xlsx'
df_retail = pd.read_excel(url, sheet_name='Year 2010-2011')
df_retail['InvoiceDate'] = pd.to_datetime(df_retail['InvoiceDate'])
print(df_retail.shape)
print(df_retail.head(3))
# BEGINNER EXAMPLE 2
# Plot the number of transactions by month to see overall sales trend.
df_retail['Month'] = df_retail['InvoiceDate'].dt.to_period('M')
monthly_sales = df_retail.groupby('Month').size()
monthly_sales.plot(kind='line', marker='o', figsize=(10,5))
plt.title('Number of Transactions Per Month')
plt.ylabel('Transactions')
plt.xlabel('Month')
plt.show()
# BEGINNER EXAMPLE 3
# Load the customer satisfaction survey dataset for churn trend analysis.
dataset = openml.datasets.get_dataset(42178)
df_survey, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_survey.shape)
print(df_survey[['tenure','MonthlyCharges','Churn']].head(3))
# BEGINNER EXAMPLE 4
# Calculate churn rate over customer tenure groups.
df_survey['tenure_group'] = pd.cut(df_survey['tenure'], bins=[0,12,24,36,48,60,72], labels=['0-12','13-24','25-36','37-48','49-60','61-72'])
churn_trend = df_survey.groupby('tenure_group')['Churn'].value_counts(normalize=True).unstack().fillna(0)
print(churn_trend)
# BEGINNER EXAMPLE 5
# Visualize churn rate trend for different tenure groups.
churn_trend['Yes'].plot(kind='bar', color='salmon', figsize=(8,5))
plt.title('Churn Rate by Customer Tenure Group')
plt.ylabel('Churn Rate')
plt.xlabel('Tenure Group (months)')
plt.ylim(0,1)
plt.show()
# BEGINNER EXAMPLE 6
# Load marketing campaign response dataset.
dataset = openml.datasets.get_dataset(1461)
df_campaign, _, _, _ = dataset.get_data(dataset_format='dataframe')
df_campaign.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
print(df_campaign.shape)
print(df_campaign[['month','response']].head(3))
# BEGINNER EXAMPLE 7
# Trend: Campaign response rates by month.
monthly_response = df_campaign.groupby('month')['response'].value_counts(normalize=True).unstack().fillna(0)
print(monthly_response)
# BEGINNER EXAMPLE 8
# Visualize trend: Yes response rate by campaign month.
monthly_response['2'].plot(kind='bar', color='green', figsize=(8,5))
plt.title('Campaign Yes Response Rate by Month')
plt.ylabel('Response Rate')
plt.xlabel('Month')
plt.ylim(0,1)
plt.show()
# INTERMEDIATE EXAMPLE 1
# Calculate and plot average monthly revenue (retail).
df_retail['Revenue'] = df_retail['Quantity'] * df_retail['Price']
monthly_rev = df_retail.groupby('Month')['Revenue'].sum()
monthly_rev.plot(kind='line', marker='o', color='royalblue', figsize=(10,5))
plt.title('Monthly Revenue Trend')
plt.ylabel('Total Revenue')
plt.xlabel('Month')
plt.show()
# INTERMEDIATE EXAMPLE 2
# Segment revenue trends by country for cross-market insights.
country_rev = df_retail.groupby(['Month','Country'])['Revenue'].sum().unstack().fillna(0)
top_countries = country_rev.sum().sort_values(ascending=False).head(5).index
country_rev[top_countries].plot(figsize=(12,6), marker='o')
plt.title('Monthly Revenue Trend for Top 5 Countries')
plt.ylabel('Total Revenue')
plt.xlabel('Month')
plt.legend(title='Country')
plt.show()
# INTERMEDIATE EXAMPLE 3
# Trend analysis: Customer churn by monthly charges level.
df_survey['charge_group'] = pd.cut(df_survey['MonthlyCharges'], bins=[0,30,60,90,120], labels=['Low','Medium','High','Very High'])
charge_churn = df_survey.groupby('charge_group')['Churn'].value_counts(normalize=True).unstack().fillna(0)
charge_churn['Yes'].plot(kind='bar', color='tomato', figsize=(8,5))
plt.title('Churn Rate Trend by Monthly Charges Group')
plt.ylabel('Churn Rate')
plt.xlabel('Monthly Charges Group')
plt.ylim(0,1)
plt.show()
# INTERMEDIATE EXAMPLE 4
# Trend analysis: NPS trends by region (synthetic NPS survey data).
np.random.seed(42)
df_nps = pd.DataFrame({'CustomerID': range(1,501), 'Age': np.random.randint(18,70,500), 'Region': np.random.choice(['North','South','East','West'],500), 'NPS_Score': np.random.randint(0,11,500)})
nps_region = df_nps.groupby('Region')['NPS_Score'].mean()
nps_region.plot(kind='bar', color='teal', figsize=(6,4))
plt.title('Average NPS Score by Region')
plt.ylabel('Average NPS')
plt.xlabel('Region')
plt.ylim(0,10)
plt.show()
# INTERMEDIATE EXAMPLE 5
# Analyze market response trend by customer segment (age band, campaign data).
df_campaign['age_band'] = pd.cut(df_campaign['age'], bins=[15,30,45,60,95], labels=['<30','30-44','45-59','60+'])
age_response = df_campaign.groupby('age_band')['response'].value_counts(normalize=True).unstack().fillna(0)
age_response['2'].plot(kind='bar', color='violet', figsize=(6,4))
plt.title('Campaign Yes Response Rate by Age Group')
plt.ylabel('Yes Response Rate')
plt.xlabel('Age Group')
plt.ylim(0,1)
plt.show()
# ADVANCED EXAMPLE 1
# Advanced: Detect recent sales peaks or declines using moving average (retail data).
monthly_rev_ma = monthly_rev.rolling(window=3, center=False).mean()
plt.figure(figsize=(10,5))
plt.plot(monthly_rev.index.to_timestamp(), monthly_rev, label='Monthly Revenue')
plt.plot(monthly_rev.index.to_timestamp(), monthly_rev_ma, label='3-Month Moving Average', linestyle='--')
plt.title('Monthly Revenue and Moving Average Trend')
plt.ylabel('Revenue')
plt.xlabel('Month')
plt.legend()
plt.show()
# ADVANCED EXAMPLE 2
# Advanced: Forecast next months revenue using simple linear model (retail data).
from sklearn.linear_model import LinearRegression
months = np.arange(len(monthly_rev)).reshape(-1,1)
lr = LinearRegression()
lr.fit(months, monthly_rev.values)
next_month = np.array([[len(months)]])
predicted_revenue = lr.predict(next_month)[0]
print(f'Predicted revenue for next month: {predicted_revenue:.2f}')
# ADVANCED EXAMPLE 3
# Advanced: Trend in customer cohort retention (synthetic cohort data).
np.random.seed(0)
dates = pd.date_range('2021-01-01', periods=24, freq='M')
df_cohort = pd.DataFrame({'CustomerID': np.random.randint(1000,2000,len(dates)), 'Signup_Month': dates, 'Active_Users': np.random.randint(50,300,len(dates))})
df_cohort.set_index('Signup_Month', inplace=True)
df_cohort['Active_Users'].plot(figsize=(10,4), marker='o')
plt.title('Active Customer Cohort Trend Over Time')
plt.ylabel('Active Users')
plt.xlabel('Signup Month')
plt.show()
# ERROR HANDLING 1
# Simulate missing survey responses in churn dataset and fill missing values.
df_survey_missing = df_survey.copy()
df_survey_missing.loc[df_survey_missing.sample(frac=0.05, random_state=42).index,'Churn'] = np.nan
print(f'Missing Churn entries: {df_survey_missing.Churn.isnull().sum()}')
df_survey_missing['Churn_filled'] = df_survey_missing['Churn'].fillna('No')
print(df_survey_missing[['Churn', 'Churn_filled']].head(10))
# ERROR HANDLING 2
# Demonstrate incorrect aggregation: grouping NPS scores by a non-category.
try:
print(df_nps.groupby('CustomerID')['NPS_Score'].mean().head())
except Exception as e:
print('Error:', e)
correct = df_nps.groupby('Region')['NPS_Score'].mean()
print('\nCorrect: Mean NPS by region:')
print(correct)
# ERROR HANDLING 3
# Common mistake: Misinterpreting NPS scale (what is good vs bad).
df_nps['Promoter'] = df_nps['NPS_Score'] >= 9
df_nps['Detractor'] = df_nps['NPS_Score'] <= 6
nps_summary = df_nps.groupby('Region')[['Promoter','Detractor']].mean()
print(nps_summary)
Market Research Analytics Best Practices#
- Segment your data by meaningful customer attributes, like tenure, charges, or region.
- Cross-tabulate outcomes to compare segments and trends over time.
- Construct scores (like NPS or churn rate) using standardized business rules.
- Always visualize trends to spot real business shifts, not just random noise.
- Document your assumptions and investigate possible causes behind trend changes.
# PATTERN 1: Segmentation
# Segment marketing campaign responses by education level.
edu_response = df_campaign.groupby('education')['response'].value_counts(normalize=True).unstack().fillna(0)
print(edu_response)
# PATTERN 2: Trend cross-tabulation
# Compare churn trend by both tenure group and churn status.
pd.crosstab(df_survey['tenure_group'], df_survey['Churn'], normalize='index')
# PATTERN 3: Constructing a trend index
# Build a simple Sales Trend Index based on moving average (retail).
sales_index = monthly_rev / monthly_rev_ma
print(sales_index.tail())
# END-TO-END MINI PROJECT
# Goal: Forecast next months customer churn for high-charge customers.
high_charge = df_survey[df_survey['MonthlyCharges'] > 80]
churn_counts = high_charge.groupby('tenure_group')['Churn'].value_counts().unstack().fillna(0)
tenure_by_month = churn_counts.index.astype(str)
plt.figure(figsize=(8,4))
plt.bar(tenure_by_month, churn_counts['Yes'], color='coral')
plt.title('Churned High-Charge Customers by Tenure Group')
plt.xlabel('Tenure Group (months)')
plt.ylabel('Churned Customers')
plt.show()
recent_churn = churn_counts['Yes'].iloc[-1]
print(f'Latest period churned customers: {recent_churn}')
print('Recommendation: Target retention offers at customers in this tenure group to reduce forecasted churn.')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



