Lesson 28 · Market Research Analytics in Python
Behavioral and Attitudinal Segmentation in Market Research
In this lesson, we will solve the business problem of identifying customer segments based on both behaviors (e.g., purchases, service usage) and attitudes…
- CourseMarket Research Analytics in Python
- Lesson28 of 56
- Video21 min
- FormatJupyter notebook · 21 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
📓 Full notebook
Download .ipynbBehavioral and Attitudinal Segmentation in Market Research#
- In this lesson, we will solve the business problem of identifying customer segments based on both behaviors (e.g., purchases, service usage) and attitudes (e.g., satisfaction, loyalty).
- Behavioral and attitudinal segmentation helps companies target different customer needs, improve experiences, and personalize marketing strategies.
- By analyzing customer data from various sources, we will extract actionable segments, interpret them, and generate business insights for decision-making.
- You will use Python to load, explore, and segment real customer datasets using survey responses, transactional behaviors, and customer attitudes.
import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')
What Are Behavioral and Attitudinal Segmentation?#
- Behavioral segmentation divides customers based on actions like purchasing, website visits, or service usage.
- Attitudinal segmentation groups customers based on subjective data such as opinions, satisfaction surveys, or brand perception.
- Survey and transactional datasets may include both types but structure them in different columns or even different files.
- Beginners often analyze just one dimension, or mislabel attitudes as true behaviors, leading to weak segment definitions.
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/00502/online_retail_II.xlsx'
df_retail = pd.read_excel(url, sheet_name='Year 2010-2011')
df_retail['InvoiceDate'] = pd.to_datetime(df_retail['InvoiceDate'])
print(df_retail.shape)
print(df_retail.head(3))
dataset = openml.datasets.get_dataset(42178)
df_sat, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_sat.shape)
print(df_sat[['gender', 'SeniorCitizen', 'MonthlyCharges', 'Churn']].head(3))
np.random.seed(42)
df_nps = pd.DataFrame({
'CustomerID': range(1, 501),
'Age': np.random.randint(18,70,500),
'Region': np.random.choice(['North','South','East','West'],500),
'NPS_Score': np.random.randint(0,11,500)
})
print(df_nps.shape)
print(df_nps.head(3))
recent_date = df_retail['InvoiceDate'].max()
df_retail['Recency'] = (recent_date - df_retail['InvoiceDate']).dt.days
customer_recency = df_retail.groupby('Customer ID')['Recency'].min().reset_index()
customer_recency['Segment'] = pd.cut(customer_recency['Recency'], bins=[-1,30,90,180,10000], labels=['Active','Warm','Cool','Dormant'])
print(customer_recency.head(5))
def nps_category(score):
if score >= 9:
return 'Promoter'
elif score >= 7:
return 'Passive'
else:
return 'Detractor'
df_nps['NPS_Category'] = df_nps['NPS_Score'].apply(nps_category)
print(df_nps[['CustomerID','NPS_Score','NPS_Category']].head(6))
customer_recency['CustomerID'] = customer_recency['Customer ID']
merged = pd.merge(customer_recency[['CustomerID','Segment']], df_nps[['CustomerID','NPS_Category']], on='CustomerID', how='inner')
crosstab = pd.crosstab(merged['Segment'], merged['NPS_Category'])
print(crosstab)
order_counts = df_retail.groupby('Customer ID').size().reset_index(name='Frequency')
order_counts['FreqSegment'] = pd.qcut(order_counts['Frequency'], q=4, labels=['Low','Medium','High','Very High'])
print(order_counts[['Customer ID','Frequency','FreqSegment']].head(6))
churn_rates = df_sat.groupby(['SeniorCitizen','Partner'])['Churn'].value_counts(normalize=True).unstack().fillna(0)
churn_rates = churn_rates[['Yes','No']]
print(churn_rates)
customer_profile = pd.merge(order_counts[['Customer ID','FreqSegment']], df_nps[['CustomerID','NPS_Category']],
left_on='Customer ID', right_on='CustomerID', how='inner')
summary = customer_profile.groupby(['FreqSegment','NPS_Category']).size().unstack().fillna(0).astype(int)
print(summary)
df_retail['Amount'] = df_retail['Quantity'] * df_retail['Price']
rfm = df_retail.groupby('Customer ID').agg({
'Recency': 'min',
'Invoice': 'nunique',
'Amount': 'sum'
}).rename(columns={'Recency':'Recency','Invoice':'Frequency','Amount':'Monetary'})
rfm['R_Score'] = pd.qcut(rfm['Recency'],q=4,labels=[4,3,2,1]).astype(int)
rfm['F_Score'] = pd.qcut(rfm['Frequency'].rank(method='first'),q=4,labels=[1,2,3,4]).astype(int)
rfm['M_Score'] = pd.qcut(rfm['Monetary'].rank(method='first'),q=4,labels=[1,2,3,4]).astype(int)
rfm['RFM_Segment'] = rfm['R_Score'].astype(str) + rfm['F_Score'].astype(str) + rfm['M_Score'].astype(str)
print(rfm.head(6))
import pandas as pd
df_feedback = pd.DataFrame({
'CustomerID':[1,2,3,4,5],
'Feedback':['Great service and friendly staff',
'Delivery was slow and packaging was poor',
'Excellent quality, will buy again',
'Customer support needs improvement',
'Good value for money']
})
# Simple sentiment assignment using keywords
def simple_sentiment(text):
if any(w in text.lower() for w in ['poor','slow','needs improvement']):
return 'Negative'
if any(w in text.lower() for w in ['great','excellent','good']):
return 'Positive'
return 'Neutral'
df_feedback['Sentiment'] = df_feedback['Feedback'].apply(simple_sentiment)
print(df_feedback)
df_nps_missing = df_nps.copy()
df_nps_missing.loc[[3,8,15], 'NPS_Score'] = np.nan
missing_count = df_nps_missing['NPS_Score'].isna().sum()
print(f'Missing NPS scores: {missing_count}')
mean_score = df_nps_missing['NPS_Score'].mean()
print(f'Mean NPS (with missing): {mean_score:.2f}')
mean_score_filled = df_nps_missing['NPS_Score'].fillna(df_nps_missing['NPS_Score'].mean()).mean()
print(f'Mean NPS (after filling): {mean_score_filled:.2f}')
try:
bad_group = df_retail.groupby('Country')['Recency'].mean()
print('Mean recency by country:', bad_group.head())
except Exception as e:
print('Error:', e)
sample_scores = [5,6,7,8,9,10]
cats = [nps_category(x) for x in sample_scores]
for s,c in zip(sample_scores, cats):
print(f'Score {s}: Segment {c}')
freq_att = pd.crosstab(order_counts['FreqSegment'], df_nps['NPS_Category'])
print(freq_att)
region_nps_avg = df_nps.groupby('Region')['NPS_Score'].mean().sort_values()
print(region_nps_avg)
np.random.seed(42)
monthly_nps = pd.DataFrame({
'Month': pd.date_range('2022-01-01', periods=12, freq='MS'),
'NPS_Score': np.random.randint(5,11,12)
})
monthly_nps['NPS_Category'] = monthly_nps['NPS_Score'].apply(nps_category)
print(monthly_nps)
# End-to-end: Identify at-risk high-value customers with negative attitudes
risky_customers = rfm[(rfm['R_Score'] <= 2) & (rfm['M_Score'] >= 3)].reset_index()
risky_merged = pd.merge(risky_customers, df_nps[['CustomerID','NPS_Category']],
left_on='Customer ID', right_on='CustomerID', how='inner')
risky_segment = risky_merged[risky_merged['NPS_Category'] == 'Detractor']
print('Number of at-risk high-value Detractors:', risky_segment.shape[0])
if not risky_segment.empty:
print(risky_segment[['Customer ID','Recency','Monetary','NPS_Category']].head(10))
output_file = 'at_risk_high_value_detractors.csv'
risky_segment[['Customer ID', 'Recency', 'Monetary', 'NPS_Category']].to_csv(output_file, index=False)
print(f'Saved: {output_file}')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



