Mathew K Analytics

Lesson 48 · Market Research Analytics in Python

Customer Cohort Analysis Training in Python for Market Research

In this lesson, we will learn how to analyze cohorts of customers based on when they joined and their retention behavior over time. Understanding cohort…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Customer Cohort Analysis: Unlocking Retention Insights#

  • In this lesson, we will learn how to analyze cohorts of customers based on when they joined and their retention behavior over time.

  • Understanding cohort retention helps businesses measure customer loyalty, identify critical churn points, and design better engagement strategies.

  • You will gain the skills to segment customers by signup period, track their activity, and visualize cohort retention trends for actionable market research insights.

  • Let us get started!

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import warnings
warnings.filterwarnings('ignore')

Core Concepts: What Is Customer Cohort Analysis?#

  • A cohort is a group of customers who share a common experiencesuch as signing up in the same month.
  • Cohort analysis tracks how these groups behave over time (such as repeated purchases, active usage, or churn).
  • Data for cohort analysis often includes customer IDs, signup dates, and activity metrics like purchases or logins.
  • Beginners sometimes confuse signup date with individual event dates, or group by calendar periods instead of cohort-based time.
  • Cohort tables usually align each customer's measurements relative to their joining date, not the overall calendar.
np.random.seed(0)
dates = pd.date_range('2021-01-01', periods=24, freq='ME')
df = pd.DataFrame({'CustomerID': np.random.randint(1000,2000,len(dates)), 'Signup_Month': dates, 'Active_Users': np.random.randint(50,300,len(dates))})
print(df.shape)
print(df.head(3))
(24, 3)
   CustomerID Signup_Month  Active_Users
0        1684   2021-01-31           138
1        1559   2021-02-28           131
2        1629   2021-03-31           215

Understanding Our Synthetic Cohort Dataset#

  • Each record is a cohort from a particular signup month.
  • Columns include:
    • CustomerID: Random unique identifier for the group
      • Signup_Month: The month the cohort joined
        • Active_Users: Number of active users in that cohort for this month
        • This setup lets us simulate cohort-based retention calculations.
# BEGINNER EXAMPLE 1: Group by Signup Month
grouped = df.groupby('Signup_Month').sum()
print(grouped[['Active_Users']].head())
              Active_Users
Signup_Month              
2021-01-31             138
2021-02-28             131
2021-03-31             215
2021-04-30              75
2021-05-31             127
# BEGINNER EXAMPLE 2: Visualizing New User Cohorts
plt.figure(figsize=(10,5))
plt.plot(grouped.index, grouped['Active_Users'], marker='o')
plt.title('New Users per Cohort Signup Month')
plt.xlabel('Signup Month')
plt.ylabel('Active Users')
plt.grid(True)
plt.show()
No description has been provided for this image
# BEGINNER EXAMPLE 3: Calculating Average Activity
avg_activity = df['Active_Users'].mean()
print(f'Average active users per cohort: {avg_activity:.2f}')
Average active users per cohort: 181.50
# INTERMEDIATE EXAMPLE 1: Simulate Monthly Retention for Each Cohort
np.random.seed(42)
months = pd.date_range('2021-01-01', periods=12, freq='ME')
cohorts = []
for i, month in enumerate(months):
    initial_users = np.random.randint(100, 200)
    retained = initial_users
    row = [month]
    for m in range(12):
        if m == 0:
            row.append(initial_users)
        else:
            retained = int(retained * np.random.uniform(0.7, 0.95))
            row.append(retained)
    cohorts.append(row)
cohort_cols = ['Cohort_Month'] + [f'Month_{i}' for i in range(12)]
cohort_df = pd.DataFrame(cohorts, columns=cohort_cols)
cohort_df.head()
Cohort_Month Month_0 Month_1 Month_2 Month_3 Month_4 Month_5 Month_6 Month_7 Month_8 Month_9 Month_10 Month_11
0 2021-01-31 151 141 124 105 77 56 40 36 30 26 18 16
1 2021-02-28 129 97 72 53 41 34 27 20 17 12 9 7
2 2021-03-31 161 116 99 78 73 59 53 46 37 26 24 20
3 2021-04-30 108 76 57 43 37 31 28 20 15 11 9 7
4 2021-05-31 153 128 95 89 79 73 67 56 52 37 27 19
# INTERMEDIATE EXAMPLE 2: Cohort Retention Table as Percentages
base_counts = cohort_df['Month_0']
retention_pct = cohort_df.iloc[:,1:].div(base_counts, axis=0) * 100
retention_pct['Cohort_Month'] = cohort_df['Cohort_Month']
retention_pct = retention_pct.set_index('Cohort_Month')
print(retention_pct.head())
              Month_0    Month_1    Month_2    Month_3    Month_4    Month_5  \
Cohort_Month                                                                   
2021-01-31      100.0  93.377483  82.119205  69.536424  50.993377  37.086093   
2021-02-28      100.0  75.193798  55.813953  41.085271  31.782946  26.356589   
2021-03-31      100.0  72.049689  61.490683  48.447205  45.341615  36.645963   
2021-04-30      100.0  70.370370  52.777778  39.814815  34.259259  28.703704   
2021-05-31      100.0  83.660131  62.091503  58.169935  51.633987  47.712418   

                Month_6    Month_7    Month_8    Month_9   Month_10   Month_11  
Cohort_Month                                                                    
2021-01-31    26.490066  23.841060  19.867550  17.218543  11.920530  10.596026  
2021-02-28    20.930233  15.503876  13.178295   9.302326   6.976744   5.426357  
2021-03-31    32.919255  28.571429  22.981366  16.149068  14.906832  12.422360  
2021-04-30    25.925926  18.518519  13.888889  10.185185   8.333333   6.481481  
2021-05-31    43.790850  36.601307  33.986928  24.183007  17.647059  12.418301  
# INTERMEDIATE EXAMPLE 3: Visualize Cohort Retention Heatmap
plt.figure(figsize=(12,6))
import seaborn as sns
sns.heatmap(retention_pct.iloc[:,:12], annot=True, fmt='.1f', cmap='Blues')
plt.title('Cohort Retention Rates (%) by Month')
plt.xlabel('Months After Signup')
plt.ylabel('Cohort Signup Month')
plt.show()
No description has been provided for this image
# ADVANCED EXAMPLE 1: Identify Critical Churn Months
avg_retention = retention_pct.iloc[:,1:].mean()
churn_point = avg_retention.idxmin()
print(f'Lowest retention rate on average occurs in: {churn_point}')
print(avg_retention)
Lowest retention rate on average occurs in: Month_11
Month_1     81.167365
Month_2     64.061304
Month_3     51.328111
Month_4     41.663934
Month_5     35.112202
Month_6     28.898944
Month_7     23.461685
Month_8     18.854591
Month_9     14.453226
Month_10    11.699347
Month_11     9.156966
dtype: float64
# ADVANCED EXAMPLE 2: Comparing Cohorts Across Different Signup Seasons
season_labels = ['Q1','Q2','Q3','Q4']*3
retention_pct['Signup_Quarter'] = season_labels[:len(retention_pct)]
quarter_retention = retention_pct.groupby('Signup_Quarter').mean()
print(quarter_retention.iloc[:,:3])
                Month_0    Month_1    Month_2
Signup_Quarter                               
Q1                100.0  86.446166  70.784100
Q2                100.0  84.812591  67.445578
Q3                100.0  79.174186  61.269504
Q4                100.0  74.236517  56.746032
# ADVANCED EXAMPLE 3: Flagging High-Performing Cohorts
top_cohorts = retention_pct[retention_pct['Month_6'] > 70]
print('Cohorts with >70% retention at 6 months:')
print(top_cohorts.index)
Cohorts with >70% retention at 6 months:
DatetimeIndex([], dtype='datetime64[ns]', name='Cohort_Month', freq=None)
# ERROR HANDLING EXAMPLE 1: Missing Active User Data
cohort_df_missing = cohort_df.copy()
cohort_df_missing.loc[2,'Month_3'] = np.nan
print(cohort_df_missing.iloc[2,:])
Cohort_Month    2021-03-31 00:00:00
Month_0                         161
Month_1                         116
Month_2                          99
Month_3                         NaN
Month_4                          73
Month_5                          59
Month_6                          53
Month_7                          46
Month_8                          37
Month_9                          26
Month_10                         24
Month_11                         20
Name: 2, dtype: object
# ERROR HANDLING EXAMPLE 2: Filling Missing Values Before Retention Calculation
cohort_df_filled = cohort_df_missing.fillna(method='ffill', axis=1)
print(cohort_df_filled.iloc[2,:])
Cohort_Month    2021-03-31 00:00:00
Month_0                         161
Month_1                         116
Month_2                          99
Month_3                          99
Month_4                          73
Month_5                          59
Month_6                          53
Month_7                          46
Month_8                          37
Month_9                          26
Month_10                         24
Month_11                         20
Name: 2, dtype: object
# ERROR HANDLING EXAMPLE 3: Handling Impossible Groupings
try:
    wrong_group = df.groupby('Active_Users').count()
    print(wrong_group.head())
except Exception as e:
    print(f'Error: {e}')
              CustomerID  Signup_Month
Active_Users                          
59                     1             1
75                     1             1
79                     1             1
122                    1             1
127                    1             1
# ERROR HANDLING EXAMPLE 4: Misinterpreting a Retention Metric
try:
    # Incorrectly using the sum instead of mean for retention
    sum_retention = retention_pct.iloc[:,1:].sum().iloc[0]
    print(f'Sum of Month 1 retention rates: {sum_retention}')
except Exception as e:
    print(f'Error: {e}')
Sum of Month 1 retention rates: 974.0083801254541

Best Practices in Cohort and Retention Analytics#

  • Always align measurements to cohort-relative time, not calendar dates.
  • Segment by key customer attributes like signup month, region, or channel for more actionable insights.
  • Use cross-tabulation to compare cohort retention by demographic or product segment.
  • Build retention indices and scorecards to track progress over time.
  • Visualize both absolute numbers and percentages to communicate trends clearly.
  • Watch for batch effectsvery large or small initial cohorts can skew average percentages.
# SEGMENTATION EXAMPLE: Comparing Retention by Randomly Assigned Region
regions = ['North', 'East', 'South', 'West']
cohort_df['Region'] = np.random.choice(regions, len(cohort_df), replace=True)
cohort_ret_by_region = cohort_df.groupby('Region')['Month_6'].mean()
print(cohort_ret_by_region)
Region
East     41.333333
North    43.200000
South    27.333333
West     36.000000
Name: Month_6, dtype: float64
# INDEXING EXAMPLE: Build a Retention Health Score
cohort_df['Health_Score'] = cohort_df['Month_6']/cohort_df['Month_0']*100
print('Cohorts with Retention Health Score:')
print(cohort_df[['Cohort_Month','Health_Score']].head())
Cohorts with Retention Health Score:
  Cohort_Month  Health_Score
0   2021-01-31     26.490066
1   2021-02-28     20.930233
2   2021-03-31     32.919255
3   2021-04-30     25.925926
4   2021-05-31     43.790850
# TREND EXAMPLE: Plotting Retention Health Over Time
plt.figure(figsize=(10,5))
plt.plot(cohort_df['Cohort_Month'], cohort_df['Health_Score'], marker='o', color='green')
plt.title('Retention Health Score by Cohort Signup Month')
plt.xlabel('Cohort Signup Month')
plt.ylabel('Health Score (%)')
plt.grid(True)
plt.show()
No description has been provided for this image
# END-TO-END MINI PROJECT: From Raw Cohort Data to Actionable Insight
raw_df = cohort_df[['Cohort_Month','Month_0','Month_3','Month_6']]
raw_df['Loss_3M'] = raw_df['Month_0'] - raw_df['Month_3']
raw_df['Loss_6M'] = raw_df['Month_3'] - raw_df['Month_6']
fastest_churn = raw_df.sort_values('Loss_3M', ascending=False).iloc[0]
print(f'Greatest initial churn: {fastest_churn.Cohort_Month} with {fastest_churn.Loss_3M} users lost by Month 3.')
Greatest initial churn: 2021-03-31 00:00:00 with 83 users lost by Month 3.

Recap and Next Steps#

  • You have learned:
    • How to build and analyze customer cohorts.
      • How to visualize and interpret retention data.
        • Best practices for segmenting, error handling, and communicating cohort insights.
        • Try cohort analysis on your own customer datasets or survey periods.
        • For further study and advanced visualization, watch our in-depth cohort analytics videos.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.