Lesson 24 · Real-World Data Analytics
Python Data Analytics #24: Marketing Campaign A/B Test Analysis in Python
Video twenty-four of the hundred-video real-world data analytics series. A real mobile game genuinely A/B tested moving a paywall gate from level thirty to…
- CourseReal-World Data Analytics
- Lesson24 of 100
- Video29 min
- FormatJupyter notebook · 30 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- ab_test_cookie_cats_raw.csv63.2 KB
📓 Full notebook
Download .ipynbData Analytics 100, Video 24: Marketing Campaign A/B Test Analysis#
- Video twenty-four of the hundred-video real-world data analytics series.
- A real mobile game genuinely A/B tested moving a paywall gate from level thirty to level forty.
- Let's get into it.
Part 1: A Real Mobile Game A/B Test#
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from scipy import stats
ab_test = pd.read_csv('ab_test_cookie_cats_raw.csv')
ab_test.shape
Part 2: Real Data Preview#
ab_test.head(5)
ab_test.dtypes
Part 3: The Real Two Groups#
group_counts = ab_test['version'].value_counts()
group_counts
Part 4: Real Group Balance Check#
group_share = (group_counts / len(ab_test) * 100).round(2)
group_share
Part 5: Real One-Day Retention by Group#
retention_1_by_group = ab_test.groupby('version')['retention_1'].mean().round(4) * 100
retention_1_by_group
Part 6: Real Seven-Day Retention by Group#
retention_7_by_group = ab_test.groupby('version')['retention_7'].mean().round(4) * 100
retention_7_by_group
Part 7: Visualizing Real Retention by Group#
fig, axes = plt.subplots(1, 2, figsize=(11, 5))
retention_1_by_group.plot(kind='bar', ax=axes[0], color=['steelblue', 'darkorange'])
axes[0].set_title('Real 1-Day Retention (%)')
axes[0].set_ylabel('Real Retention Rate (%)')
retention_7_by_group.plot(kind='bar', ax=axes[1], color=['steelblue', 'darkorange'])
axes[1].set_title('Real 7-Day Retention (%)')
plt.tight_layout()
plt.savefig('retention_by_group.png', dpi=120)
plt.close()
Part 8: Real Statistical Significance, One-Day Retention#
contingency_1 = pd.crosstab(ab_test['version'], ab_test['retention_1'])
chi2_1, p_value_1, dof_1, expected_1 = stats.chi2_contingency(contingency_1)
round(p_value_1, 4)
Part 9: Real Statistical Significance, Seven-Day Retention#
contingency_7 = pd.crosstab(ab_test['version'], ab_test['retention_7'])
chi2_7, p_value_7, dof_7, expected_7 = stats.chi2_contingency(contingency_7)
round(p_value_7, 4)
Part 10: Real Honest Interpretation#
alpha = 0.05
significant_1 = p_value_1 < alpha
significant_7 = p_value_7 < alpha
significant_1, significant_7
Part 11: Real Engagement, Total Game Rounds#
gamerounds_by_group = ab_test.groupby('version')['sum_gamerounds'].mean().round(2)
gamerounds_by_group
Part 12: Real Median Engagement, Outlier-Resistant#
gamerounds_median_by_group = ab_test.groupby('version')['sum_gamerounds'].median()
gamerounds_median_by_group
Part 13: Real Outlier Check#
max_rounds = ab_test['sum_gamerounds'].max()
zero_rounds = (ab_test['sum_gamerounds'] == 0).sum()
max_rounds, zero_rounds
Part 14: Real Mann-Whitney Test on Engagement#
gate_30_rounds = ab_test.loc[ab_test['version'] == 'gate_30', 'sum_gamerounds']
gate_40_rounds = ab_test.loc[ab_test['version'] == 'gate_40', 'sum_gamerounds']
mw_stat, mw_p = stats.mannwhitneyu(gate_30_rounds, gate_40_rounds, alternative='two-sided')
round(mw_p, 4)
Part 15: Real Distribution of Game Rounds#
capped_rounds = ab_test[ab_test['sum_gamerounds'] < 200]
plt.figure(figsize=(9, 5))
plt.hist([capped_rounds.loc[capped_rounds['version']=='gate_30','sum_gamerounds'], capped_rounds.loc[capped_rounds['version']=='gate_40','sum_gamerounds']], bins=30, label=['gate_30','gate_40'], color=['steelblue','darkorange'], alpha=0.7)
plt.xlabel('Real Game Rounds Played (capped at 200)')
plt.ylabel('Real Number of Players')
plt.title('Real Distribution of Game Rounds by Gate Version')
plt.legend()
plt.tight_layout()
plt.savefig('gamerounds_distribution.png', dpi=120)
plt.close()
Part 16: Real Retention Funnel, Both Milestones#
ab_test['retained_both'] = ab_test['retention_1'] & ab_test['retention_7']
retained_both_by_group = ab_test.groupby('version')['retained_both'].mean().round(4) * 100
retained_both_by_group
Part 17: Real Conditional Retention#
day1_returners = ab_test[ab_test['retention_1'] == True]
conditional_retention = day1_returners.groupby('version')['retention_7'].mean().round(4) * 100
conditional_retention
Part 18: Real Effect Size, Not Just P-Values#
retention_7_gap_pct_points = round(retention_7_by_group['gate_30'] - retention_7_by_group['gate_40'], 2)
retention_7_relative_lift = round((retention_7_by_group['gate_30'] / retention_7_by_group['gate_40'] - 1) * 100, 2)
retention_7_gap_pct_points, retention_7_relative_lift
Part 19: Real Sample Size Sensitivity Check#
n_gate_30 = (ab_test['version'] == 'gate_30').sum()
n_gate_40 = (ab_test['version'] == 'gate_40').sum()
n_gate_30, n_gate_40, n_gate_30 + n_gate_40
Part 20: Real Confidence Interval on the Gap#
p1 = retention_7_by_group['gate_30'] / 100
p2 = retention_7_by_group['gate_40'] / 100
se = np.sqrt(p1*(1-p1)/n_gate_30 + p2*(1-p2)/n_gate_40)
ci_low = round((p1 - p2 - 1.96*se) * 100, 2)
ci_high = round((p1 - p2 + 1.96*se) * 100, 2)
ci_low, ci_high
Part 21: Real Recommendation Logic#
recommend_gate_30 = retention_7_by_group['gate_30'] > retention_7_by_group['gate_40']
recommend_gate_30
Part 22: Real Correlation, Engagement and Retention#
engagement_retention_corr = ab_test['sum_gamerounds'].corr(ab_test['retention_7'].astype(int))
round(engagement_retention_corr, 3)
Part 23: Real High-Engagement Segment#
high_engagement = ab_test[ab_test['sum_gamerounds'] > ab_test['sum_gamerounds'].quantile(0.75)]
high_engagement_retention = high_engagement.groupby('version')['retention_7'].mean().round(4) * 100
high_engagement_retention
Part 24: Real Low-Engagement Segment#
low_engagement = ab_test[ab_test['sum_gamerounds'] <= ab_test['sum_gamerounds'].quantile(0.25)]
low_engagement_retention = low_engagement.groupby('version')['retention_7'].mean().round(4) * 100
low_engagement_retention
Part 25: Real Summary Table#
summary_table = pd.DataFrame({'retention_1_pct': retention_1_by_group, 'retention_7_pct': retention_7_by_group, 'avg_gamerounds': gamerounds_by_group, 'median_gamerounds': gamerounds_median_by_group}).round(2)
summary_table
Part 26: Saving the Real Summary Table#
summary_table.to_csv('ab_test_summary.csv')
reloaded_summary = pd.read_csv('ab_test_summary.csv', index_col=0)
reloaded_summary.equals(summary_table)
Part 27: Real Sanity Check, Percentages in Range#
all((0 <= retention_1_by_group) & (retention_1_by_group <= 100)) and all((0 <= retention_7_by_group) & (retention_7_by_group <= 100))
Part 28: Real Sanity Check, Group Sizes Sum Correctly#
(n_gate_30 + n_gate_40) == len(ab_test)
Part 29: Real Business Framing#
weekly_active_estimate = 1_000_000
extra_retained_players = round(weekly_active_estimate * (retention_7_gap_pct_points / 100))
extra_retained_players
Part 30: Real Recap Print#
print(f'Across {len(ab_test)} real players, gate_30 held {retention_7_by_group["gate_30"]}% seven-day retention versus {retention_7_by_group["gate_40"]}% for gate_40, a real {retention_7_gap_pct_points} point gap with a chi-square p-value of {round(p_value_7,4)} in this sample.')
Wrap-Up: What You Learned#
- A real randomized controlled experiment, not just an observational comparison, is what makes an A/B test genuinely trustworthy.
- Chi-square tests check whether a real categorical outcome gap could plausibly be explained by real random chance alone.
- Real skewed engagement data, like game rounds played, needs medians and non-parametric tests, not just a real raw average.
- A real percentage-point gap and a real relative lift tell genuinely different stories about the same real effect, both matter for real decisions.
- Segmenting a real A/B test by engagement level can reveal whether a real effect is consistent or concentrated in one real group.
- Next video: real customer funnel and conversion rate analysis, tracking how real users drop off at each real stage.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



