Lesson 48 · Market Research Analytics in Python
Customer Cohort Analysis Training in Python for Market Research
In this lesson, we will learn how to analyze cohorts of customers based on when they joined and their retention behavior over time. Understanding cohort…
- CourseMarket Research Analytics in Python
- Lesson48 of 56
- Video21 min
- FormatJupyter notebook · 19 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbCustomer Cohort Analysis: Unlocking Retention Insights#
In this lesson, we will learn how to analyze cohorts of customers based on when they joined and their retention behavior over time.
Understanding cohort retention helps businesses measure customer loyalty, identify critical churn points, and design better engagement strategies.
You will gain the skills to segment customers by signup period, track their activity, and visualize cohort retention trends for actionable market research insights.
Let us get started!
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import warnings
warnings.filterwarnings('ignore')
Core Concepts: What Is Customer Cohort Analysis?#
- A cohort is a group of customers who share a common experiencesuch as signing up in the same month.
- Cohort analysis tracks how these groups behave over time (such as repeated purchases, active usage, or churn).
- Data for cohort analysis often includes customer IDs, signup dates, and activity metrics like purchases or logins.
- Beginners sometimes confuse signup date with individual event dates, or group by calendar periods instead of cohort-based time.
- Cohort tables usually align each customer's measurements relative to their joining date, not the overall calendar.
np.random.seed(0)
dates = pd.date_range('2021-01-01', periods=24, freq='ME')
df = pd.DataFrame({'CustomerID': np.random.randint(1000,2000,len(dates)), 'Signup_Month': dates, 'Active_Users': np.random.randint(50,300,len(dates))})
print(df.shape)
print(df.head(3))
Understanding Our Synthetic Cohort Dataset#
- Each record is a cohort from a particular signup month.
- Columns include:
- CustomerID: Random unique identifier for the group
- Signup_Month: The month the cohort joined
- Active_Users: Number of active users in that cohort for this month
- This setup lets us simulate cohort-based retention calculations.
- Signup_Month: The month the cohort joined
- CustomerID: Random unique identifier for the group
# BEGINNER EXAMPLE 1: Group by Signup Month
grouped = df.groupby('Signup_Month').sum()
print(grouped[['Active_Users']].head())
# BEGINNER EXAMPLE 2: Visualizing New User Cohorts
plt.figure(figsize=(10,5))
plt.plot(grouped.index, grouped['Active_Users'], marker='o')
plt.title('New Users per Cohort Signup Month')
plt.xlabel('Signup Month')
plt.ylabel('Active Users')
plt.grid(True)
plt.show()
# BEGINNER EXAMPLE 3: Calculating Average Activity
avg_activity = df['Active_Users'].mean()
print(f'Average active users per cohort: {avg_activity:.2f}')
# INTERMEDIATE EXAMPLE 1: Simulate Monthly Retention for Each Cohort
np.random.seed(42)
months = pd.date_range('2021-01-01', periods=12, freq='ME')
cohorts = []
for i, month in enumerate(months):
initial_users = np.random.randint(100, 200)
retained = initial_users
row = [month]
for m in range(12):
if m == 0:
row.append(initial_users)
else:
retained = int(retained * np.random.uniform(0.7, 0.95))
row.append(retained)
cohorts.append(row)
cohort_cols = ['Cohort_Month'] + [f'Month_{i}' for i in range(12)]
cohort_df = pd.DataFrame(cohorts, columns=cohort_cols)
cohort_df.head()
# INTERMEDIATE EXAMPLE 2: Cohort Retention Table as Percentages
base_counts = cohort_df['Month_0']
retention_pct = cohort_df.iloc[:,1:].div(base_counts, axis=0) * 100
retention_pct['Cohort_Month'] = cohort_df['Cohort_Month']
retention_pct = retention_pct.set_index('Cohort_Month')
print(retention_pct.head())
# INTERMEDIATE EXAMPLE 3: Visualize Cohort Retention Heatmap
plt.figure(figsize=(12,6))
import seaborn as sns
sns.heatmap(retention_pct.iloc[:,:12], annot=True, fmt='.1f', cmap='Blues')
plt.title('Cohort Retention Rates (%) by Month')
plt.xlabel('Months After Signup')
plt.ylabel('Cohort Signup Month')
plt.show()
# ADVANCED EXAMPLE 1: Identify Critical Churn Months
avg_retention = retention_pct.iloc[:,1:].mean()
churn_point = avg_retention.idxmin()
print(f'Lowest retention rate on average occurs in: {churn_point}')
print(avg_retention)
# ADVANCED EXAMPLE 2: Comparing Cohorts Across Different Signup Seasons
season_labels = ['Q1','Q2','Q3','Q4']*3
retention_pct['Signup_Quarter'] = season_labels[:len(retention_pct)]
quarter_retention = retention_pct.groupby('Signup_Quarter').mean()
print(quarter_retention.iloc[:,:3])
# ADVANCED EXAMPLE 3: Flagging High-Performing Cohorts
top_cohorts = retention_pct[retention_pct['Month_6'] > 70]
print('Cohorts with >70% retention at 6 months:')
print(top_cohorts.index)
# ERROR HANDLING EXAMPLE 1: Missing Active User Data
cohort_df_missing = cohort_df.copy()
cohort_df_missing.loc[2,'Month_3'] = np.nan
print(cohort_df_missing.iloc[2,:])
# ERROR HANDLING EXAMPLE 2: Filling Missing Values Before Retention Calculation
cohort_df_filled = cohort_df_missing.fillna(method='ffill', axis=1)
print(cohort_df_filled.iloc[2,:])
# ERROR HANDLING EXAMPLE 3: Handling Impossible Groupings
try:
wrong_group = df.groupby('Active_Users').count()
print(wrong_group.head())
except Exception as e:
print(f'Error: {e}')
# ERROR HANDLING EXAMPLE 4: Misinterpreting a Retention Metric
try:
# Incorrectly using the sum instead of mean for retention
sum_retention = retention_pct.iloc[:,1:].sum().iloc[0]
print(f'Sum of Month 1 retention rates: {sum_retention}')
except Exception as e:
print(f'Error: {e}')
Best Practices in Cohort and Retention Analytics#
- Always align measurements to cohort-relative time, not calendar dates.
- Segment by key customer attributes like signup month, region, or channel for more actionable insights.
- Use cross-tabulation to compare cohort retention by demographic or product segment.
- Build retention indices and scorecards to track progress over time.
- Visualize both absolute numbers and percentages to communicate trends clearly.
- Watch for batch effectsvery large or small initial cohorts can skew average percentages.
# SEGMENTATION EXAMPLE: Comparing Retention by Randomly Assigned Region
regions = ['North', 'East', 'South', 'West']
cohort_df['Region'] = np.random.choice(regions, len(cohort_df), replace=True)
cohort_ret_by_region = cohort_df.groupby('Region')['Month_6'].mean()
print(cohort_ret_by_region)
# INDEXING EXAMPLE: Build a Retention Health Score
cohort_df['Health_Score'] = cohort_df['Month_6']/cohort_df['Month_0']*100
print('Cohorts with Retention Health Score:')
print(cohort_df[['Cohort_Month','Health_Score']].head())
# TREND EXAMPLE: Plotting Retention Health Over Time
plt.figure(figsize=(10,5))
plt.plot(cohort_df['Cohort_Month'], cohort_df['Health_Score'], marker='o', color='green')
plt.title('Retention Health Score by Cohort Signup Month')
plt.xlabel('Cohort Signup Month')
plt.ylabel('Health Score (%)')
plt.grid(True)
plt.show()
# END-TO-END MINI PROJECT: From Raw Cohort Data to Actionable Insight
raw_df = cohort_df[['Cohort_Month','Month_0','Month_3','Month_6']]
raw_df['Loss_3M'] = raw_df['Month_0'] - raw_df['Month_3']
raw_df['Loss_6M'] = raw_df['Month_3'] - raw_df['Month_6']
fastest_churn = raw_df.sort_values('Loss_3M', ascending=False).iloc[0]
print(f'Greatest initial churn: {fastest_churn.Cohort_Month} with {fastest_churn.Loss_3M} users lost by Month 3.')
Recap and Next Steps#
- You have learned:
- How to build and analyze customer cohorts.
- How to visualize and interpret retention data.
- Best practices for segmenting, error handling, and communicating cohort insights.
- Try cohort analysis on your own customer datasets or survey periods.
- For further study and advanced visualization, watch our in-depth cohort analytics videos.
- How to visualize and interpret retention data.
- How to build and analyze customer cohorts.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



