Lesson 29 · Market Research Analytics in Python
Rule Based Customer Segmentation in Python: Step-by-Step Training
We are learning to segment customers using rules based on real business data. Segmentation helps companies target the right customers and improve products…
- CourseMarket Research Analytics in Python
- Lesson29 of 56
- Video19 min
- FormatJupyter notebook · 23 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbRule-Based Customer Segmentation in Market Research#
- We are learning to segment customers using rules based on real business data.
- Segmentation helps companies target the right customers and improve products or marketing.
- You will use real survey and behavioral data to design and analyze practical customer segments.
- By the end, you will have learned to define rules, apply them to data, and interpret actionable insights.
import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')
Understanding the Data Used for Rule-Based Segmentation#
- Market research data includes customer demographics, surveys, and behavior.
- Each row represents a person or customer with features like age, gender, and service usage.
- Beginners often forget to handle missing values or misinterpret text responses.
- Columns may be categorical, numerical, or text, each requiring different analysis techniques.
- Getting to know your data structure is essential before creating any rules.
dataset = openml.datasets.get_dataset(42178)
df_satisfaction, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_satisfaction.shape)
print(df_satisfaction.head(3))
churn_counts = df_satisfaction['Churn'].value_counts()
print('Churned vs Not Churned:')
print(churn_counts)
contract_counts = df_satisfaction['Contract'].value_counts()
print('Customers by Contract Type:')
print(contract_counts)
df_satisfaction['Segment'] = np.where(df_satisfaction['tenure'] < 12, 'New', 'Established')
print(df_satisfaction[['tenure','Segment']].head(5))
churn_by_segment = pd.crosstab(df_satisfaction['Segment'], df_satisfaction['Churn'], normalize='index')
print(churn_by_segment)
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/00502/online_retail_II.xlsx'
df_retail = pd.read_excel(url, sheet_name='Year 2010-2011')
df_retail['InvoiceDate'] = pd.to_datetime(df_retail['InvoiceDate'])
print(df_retail.shape)
print(df_retail[['Customer ID','InvoiceDate','Country']].head(3))
customers_by_country = df_retail.groupby('Country')['Customer ID'].nunique().sort_values(ascending=False)
print(customers_by_country.head(5))
customer_value = df_retail.groupby('Customer ID')['Quantity'].sum()
value_threshold = customer_value.quantile(0.9)
high_value_ids = customer_value[customer_value > value_threshold].index
df_retail['ValueSegment'] = np.where(df_retail['Customer ID'].isin(high_value_ids), 'High Value', 'Regular')
print(df_retail[['Customer ID','Quantity','ValueSegment']].drop_duplicates().head(7))
df_satisfaction['Segment2'] = np.select(
[
(df_satisfaction['tenure'] < 12) & (df_satisfaction['Contract']=='Month-to-month'),
(df_satisfaction['tenure'] >= 12) & (df_satisfaction['Contract']=='Two year')
],
['At-Risk New', 'Loyal Long-Term'],
default='Other'
)
print(df_satisfaction[['tenure','Contract','Segment2']].head(7))
multi_seg_churn = pd.crosstab(df_satisfaction['Segment2'], df_satisfaction['Churn'], normalize='index')
print(multi_seg_churn)
np.random.seed(42)
df_nps = pd.DataFrame({
'CustomerID': range(1,501),
'Age': np.random.randint(18,70,500),
'Region': np.random.choice(['North','South','East','West'],500),
'NPS_Score': np.random.randint(0,11,500)
})
print(df_nps.shape)
print(df_nps.head(3))
def nps_segment(score):
if score <= 6:
return 'Detractor'
elif score <= 8:
return 'Passive'
else:
return 'Promoter'
df_nps['NPS_Segment'] = df_nps['NPS_Score'].apply(nps_segment)
print(df_nps.groupby('NPS_Segment').size())
nps_by_region = pd.crosstab(df_nps['Region'], df_nps['NPS_Segment'], normalize='index')
print(nps_by_region.round(2))
summary = df_nps.groupby(['Region','NPS_Segment']).size().unstack().fillna(0)
summary.to_csv('nps_segment_summary.csv')
print('Segmentation summary file saved as nps_segment_summary.csv')
missing = df_retail.isnull().sum()
print('Missing values per column:')
print(missing[missing > 0])
try:
result = df_retail.groupby('ValuSegment')['Customer ID'].nunique()
except KeyError as e:
print('Error:', e)
print('Check for correct segment column name: should be ValueSegment')
def safe_nps_segment(score):
if pd.isnull(score) or not (0 <= score <= 10):
return 'Unknown'
elif score <= 6:
return 'Detractor'
elif score <= 8:
return 'Passive'
else:
return 'Promoter'
df_nps['Safe_NPS_Segment'] = df_nps['NPS_Score'].apply(safe_nps_segment)
print(df_nps['Safe_NPS_Segment'].value_counts())
tenure_contract_ct = pd.crosstab(df_satisfaction['Segment'], df_satisfaction['Contract'])
print(tenure_contract_ct)
df_satisfaction['LoyaltyIndex'] = df_satisfaction['tenure'] * (df_satisfaction['Contract']=='Two year').astype(int)
print(df_satisfaction[['tenure','Contract','LoyaltyIndex']].head(6))
df_retail['Month'] = df_retail['InvoiceDate'].dt.to_period('M')
monthly_counts = df_retail.groupby('Month')['Customer ID'].nunique()
print(monthly_counts.tail(12))
segment_counts = df_satisfaction['Segment2'].value_counts()
at_risk_churn = multi_seg_churn.loc['At-Risk New', 'Yes'] if 'At-Risk New' in multi_seg_churn.index else None
if at_risk_churn is not None and at_risk_churn > 0.3:
recommendation = 'Launch onboarding and specials for At-Risk New to reduce churn.'
else:
recommendation = 'Focus on other segments or maintain current actions.'
print('Segment sizes:')
print(segment_counts)
print('At-Risk New churn rate:', at_risk_churn)
print('Recommendation:', recommendation)
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



