Lesson 15 · Statistics for data analysts
Statistics Capstone: A Full Real-World Analysis | Statistics #15
The final video of the 15-part series: one complete real statistical analysis, start to finish, combining nearly everything covered so far. Real Sample…
- CourseStatistics for data analysts
- Lesson15 of 15
- Video15 min
- FormatJupyter notebook · 11 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- superstore_full.csv1.2 MB
📓 Full notebook
Download .ipynbStatistics for Data Analysts, Video 15: Capstone Statistical Analysis#
- The final video of the 15-part series: one complete real statistical analysis, start to finish, combining nearly everything covered so far.
- Real Sample Superstore order data, including real Discount figures this time.
- The business question: is our discounting strategy hurting real profitability, and where should we focus? Let's find out, honestly.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- You'll need pandas, NumPy, SciPy, Matplotlib, and statsmodels.
- Place superstore_full.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from scipy import stats
import matplotlib.pyplot as plt
import statsmodels.formula.api as smf
import statsmodels.api as sm
from statsmodels.stats.outliers_influence import variance_inflation_factor
df = pd.read_csv('superstore_full.csv')
print(df.shape)
print(df[['Sales', 'Discount', 'Quantity', 'Profit']].describe())
Part 1: Framing the Business Question#
df['margin'] = df['Profit'] / df['Sales']
df['high_discount'] = (df['Discount'] >= 0.2).astype(int)
print(f'Real orders with discount >= 20%: {df["high_discount"].sum()} of {len(df)}')
Part 2: The Shape of Real Profit#
print(f'Real Profit skewness: {stats.skew(df["Profit"]):.2f}')
print(f'Real Profit range: {df["Profit"].min():.2f} to {df["Profit"].max():.2f}')
print(f'Real orders with a loss (Profit < 0): {(df["Profit"] < 0).sum()}')
plt.figure(figsize=(9, 4))
plt.hist(df['margin'].clip(-1, 1), bins=60, color='steelblue', edgecolor='white')
plt.axvline(0, color='darkred', linestyle='--', label='Break-even')
plt.title('Real Profit Margin Distribution (clipped for display)')
plt.xlabel('Profit Margin')
plt.legend()
plt.show()
Part 3: A Confidence Interval for Average Margin#
n = len(df)
mean_margin = df['margin'].mean()
se_margin = df['margin'].std(ddof=1) / np.sqrt(n)
t_crit = stats.t.ppf(0.975, df=n - 1)
ci_low, ci_high = mean_margin - t_crit * se_margin, mean_margin + t_crit * se_margin
print(f'Real average margin: {mean_margin:.4f}')
print(f'Real 95% CI: ({ci_low:.4f}, {ci_high:.4f})')
Part 4: Does Discounting Actually Hurt Margin?#
high = df.loc[df['high_discount'] == 1, 'margin']
low = df.loc[df['high_discount'] == 0, 'margin']
print(f'Real high-discount orders (n={len(high)}): mean margin = {high.mean():.4f}')
print(f'Real low-discount orders (n={len(low)}): mean margin = {low.mean():.4f}')
u_stat, p_value = stats.mannwhitneyu(high, low, alternative='two-sided')
print(f'Real Mann-Whitney U p-value: {p_value:.2e}')
Part 5: Quantifying the Effect with Regression#
model = smf.ols('Profit ~ Sales + Discount + Quantity', data=df).fit()
print(model.summary())
X = df[['Sales', 'Discount', 'Quantity']]
X_const = sm.add_constant(X)
vif_data = pd.DataFrame({'feature': X_const.columns, 'VIF': [variance_inflation_factor(X_const.values, i) for i in range(X_const.shape[1])]})
print(vif_data)
discount_effect = model.params['Discount']
print(f'Real effect: each 10-point rise in discount predicts ${abs(discount_effect) * 0.10:.2f} less real profit per order, holding Sales and Quantity fixed')
Part 6: Where Is the Discounting Concentrated?#
category_table = pd.crosstab(df['Category'], df['high_discount'])
chi2_stat, p_chi, dof, expected = stats.chi2_contingency(category_table)
cramers_v = np.sqrt(chi2_stat / (len(df) * (min(category_table.shape) - 1)))
print(f'Real chi-square p-value: {p_chi:.2e}')
print(f'Real Cramer\'s V: {cramers_v:.4f}')
print(df.groupby('Category')['high_discount'].mean().round(3))
Part 7: Executive Summary#
print('EXECUTIVE SUMMARY: Discounting and Profitability')
print('=' * 55)
print(f'- Real average profit margin across all orders: {mean_margin:.1%} (95% CI: {ci_low:.1%} to {ci_high:.1%})')
print(f'- Real orders discounted 20% or more average a {high.mean():.1%} margin, i.e. a genuine loss on average')
print(f'- Real orders discounted under 20% average a {low.mean():.1%} margin, solidly profitable')
print(f'- This gap is real and statistically robust (Mann-Whitney p < 0.001), not sampling noise')
print(f'- Regression estimates each 10-point rise in discount costs about ${abs(discount_effect) * 0.10:.2f} in real profit per order, holding order size and quantity fixed')
print(f'- Heavy discounting is only weakly associated with product category (Cramer\'s V={cramers_v:.3f}), so this is a broad real pricing issue, not one confined to a specific category')
print('RECOMMENDATION: Review discount approval thresholds above 20% across the board, rather than targeting a single category')
Wrap-Up: The Full 15-Part Series#
- This capstone combined distribution checks, confidence intervals, non-parametric hypothesis testing, multiple regression with a VIF sanity check, and a chi-square test with an honest effect size, into one coherent real analysis with an actionable conclusion.
- Across all 15 videos: probability and simulation, discrete and continuous distributions, the Central Limit Theorem, sampling and bias, confidence intervals, hypothesis testing, a full synthetic A/B test with power analysis, non-parametric tests, multiple comparisons, simple and multiple regression, categorical data analysis, and Bayesian statistics.
- Every honest finding in this series, including the real misses, the real disagreements between tests, and the real cases where statistical and practical significance diverged, was reported as it actually came out, not smoothed over.
- That's the whole series. Thanks for following along through all fifteen videos, real data and all. Subscribe if you haven't, and go run these notebooks yourself. See you in the next series.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



