Lesson 2 · Real-World Data Analytics
Python Data Analytics #02: Customer Segmentation with RFM Analysis in Python
Video two of the hundred-video real-world data analytics series. Real Recency, Frequency, and Monetary scoring, on the real cleaned retail data from last…
- CourseReal-World Data Analytics
- Lesson2 of 26
- Video28 min
- FormatJupyter notebook · 30 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- online_retail_clean.csv36.2 MB
📓 Full notebook
Download .ipynbData Analytics 100, Video 2: Customer Segmentation with RFM Analysis#
- Video two of the hundred-video real-world data analytics series.
- Real Recency, Frequency, and Monetary scoring, on the real cleaned retail data from last video.
- Let's get into it.
Part 1: What RFM Actually Measures#
import pandas as pd
clean = pd.read_csv('online_retail_clean.csv', parse_dates=['InvoiceDate'])
clean.shape
Part 2: Choosing a Real Snapshot Date#
snapshot_date = clean['InvoiceDate'].max() + pd.Timedelta(days=1)
snapshot_date
Part 3: Real Recency per Customer#
recency = clean.groupby('CustomerID')['InvoiceDate'].max()
recency = (snapshot_date - recency).dt.days
recency.describe()
Part 4: Real Frequency per Customer#
frequency = clean.groupby('CustomerID')['InvoiceNo'].nunique()
frequency.describe()
Part 5: Real Monetary Value per Customer#
monetary = clean.groupby('CustomerID')['Revenue'].sum()
monetary.describe()
Part 6: Combining Into One Real RFM Table#
rfm = pd.DataFrame({'Recency': recency, 'Frequency': frequency, 'Monetary': monetary})
rfm = rfm.reset_index()
rfm.head()
rfm.shape[0]
Part 7: Confirming the Real Customer Count Lines Up#
rfm.shape[0] == clean['CustomerID'].nunique()
rfm.isna().sum().sum()
Part 8: Scoring Recency Into Real Quartiles#
rfm['R_score'] = pd.qcut(rfm['Recency'], 4, labels=[4, 3, 2, 1]).astype(int)
rfm['R_score'].value_counts().sort_index()
Part 9: Scoring Frequency Into Real Quartiles#
rfm['F_score'] = pd.qcut(rfm['Frequency'].rank(method='first'), 4, labels=[1, 2, 3, 4]).astype(int)
rfm['F_score'].value_counts().sort_index()
Part 10: Scoring Monetary Into Real Quartiles#
rfm['M_score'] = pd.qcut(rfm['Monetary'], 4, labels=[1, 2, 3, 4]).astype(int)
rfm['M_score'].value_counts().sort_index()
Part 11: A Real Combined RFM Score#
rfm['RFM_Sum'] = rfm['R_score'] + rfm['F_score'] + rfm['M_score']
rfm['RFM_Sum'].describe()
Part 12: A Real Rule-Based Segment Function#
def segment_customer(row):
if row['R_score'] >= 3 and row['F_score'] >= 3 and row['M_score'] >= 3:
return 'Champions'
elif row['R_score'] >= 3 and row['F_score'] >= 2:
return 'Loyal Customers'
elif row['R_score'] >= 3:
return 'New Customers'
elif row['R_score'] == 2:
return 'At Risk'
else:
return 'Lost'
Part 13: Applying the Real Segmentation#
rfm['Segment'] = rfm.apply(segment_customer, axis=1)
rfm['Segment'].value_counts()
Part 14: Real Revenue Contribution by Segment#
segment_revenue = rfm.groupby('Segment')['Monetary'].sum().sort_values(ascending=False)
segment_revenue
(segment_revenue / segment_revenue.sum() * 100).round(1)
Part 15: Real Average Behavior by Segment#
rfm.groupby('Segment')[['Recency', 'Frequency', 'Monetary']].mean().round(1)
Part 16: Real Customers vs Real Revenue Share#
segment_customers = rfm['Segment'].value_counts()
customer_share = (segment_customers / segment_customers.sum() * 100).round(1)
revenue_share = (segment_revenue / segment_revenue.sum() * 100).round(1)
pd.DataFrame({'Customer %': customer_share, 'Revenue %': revenue_share})
Part 17: Visualizing Real Segment Sizes#
import matplotlib.pyplot as plt
order = rfm['Segment'].value_counts().index
plt.figure(figsize=(8, 5))
plt.bar(order, rfm['Segment'].value_counts()[order], color='seagreen')
plt.title('Real Customer Count by RFM Segment')
plt.ylabel('Number of Real Customers')
plt.tight_layout()
plt.savefig('rfm_segment_sizes.png', dpi=120)
plt.close()
Part 18: Visualizing Real Recency Against Monetary Value#
colors = {'Champions': 'gold', 'Loyal Customers': 'seagreen', 'New Customers': 'steelblue', 'At Risk': 'orange', 'Lost': 'gray'}
plt.figure(figsize=(9, 6))
for seg, group in rfm.groupby('Segment'):
plt.scatter(group['Recency'], group['Monetary'], label=seg, alpha=0.6, color=colors[seg])
plt.yscale('log')
plt.xlabel('Real Recency (days since last order)')
plt.ylabel('Real Monetary Value (log scale)')
plt.title('Real Recency vs Monetary Value by Segment')
plt.legend()
plt.tight_layout()
plt.savefig('rfm_recency_monetary_scatter.png', dpi=120)
plt.close()
Part 19: Real Champions Worth Naming#
rfm[rfm['Segment'] == 'Champions'].sort_values('Monetary', ascending=False).head(10)
Part 20: Real At-Risk Customers Worth a Win-Back#
rfm[rfm['Segment'] == 'At Risk'].sort_values('Monetary', ascending=False).head(10)
Part 21: Saving the Real RFM Table#
rfm.to_csv('online_retail_rfm.csv', index=False)
reloaded_rfm = pd.read_csv('online_retail_rfm.csv')
reloaded_rfm.shape == rfm.shape
Part 22: One Last Real Sanity Check#
rfm['Segment'].isna().sum()
rfm.groupby('Segment').size().sum() == rfm.shape[0]
Part 23: Real Correlation Between Recency, Frequency, and Monetary#
rfm[['Recency', 'Frequency', 'Monetary']].corr().round(2)
Part 24: Does a Higher Real RFM Score Actually Mean More Spend?#
rfm.groupby('RFM_Sum')['Monetary'].mean().round(1)
Part 25: Real Distribution of the Combined Score#
plt.figure(figsize=(8, 5))
plt.hist(rfm['RFM_Sum'], bins=range(3, 14), color='slateblue', edgecolor='white')
plt.xlabel('Real Combined RFM Score')
plt.ylabel('Number of Real Customers')
plt.title('Real Distribution of Combined RFM Scores')
plt.tight_layout()
plt.savefig('rfm_score_distribution.png', dpi=120)
plt.close()
Part 26: Where the Real Champions Actually Are#
customer_country = clean.groupby('CustomerID')['Country'].first()
rfm_with_country = rfm.merge(customer_country, on='CustomerID')
rfm_with_country[rfm_with_country['Segment'] == 'Champions']['Country'].value_counts().head(5)
Part 27: The Real Quartile Boundaries Behind Each Score#
_, r_bins = pd.qcut(rfm['Recency'], 4, labels=[4, 3, 2, 1], retbins=True)
r_bins.round(1)
Part 28: Real Median vs Real Mean Spend by Segment#
rfm.groupby('Segment')['Monetary'].median().round(1)
rfm.groupby('Segment')['Monetary'].mean().round(1)
Part 29: A Real Edge Case Worth Knowing#
one_order_champions = rfm[(rfm['Segment'] == 'Champions') & (rfm['Frequency'] == 1)]
one_order_champions.shape[0]
Part 30: Reconciling Row Counts One Last Time#
rfm_with_country.shape[0] == rfm.shape[0]
Wrap-Up: What You Learned#
- RFM genuinely scores every real customer on Recency, Frequency, and Monetary value, each split into real quartiles.
- Combining the three real individual scores into rule-based segments beats using just the real combined total alone.
- In this real dataset, a small real Champions segment genuinely drives a hugely disproportionate share of real revenue.
- The real At Risk segment is the real highest-value group still worth an active win-back effort.
- Next video: real market basket analysis, finding which real products actually get bought together.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



