Lesson 23 · Data analytics zero to hero
Statistics Fundamentals Every Analyst Needs | Data Analytics #23
Video twenty-three of the 30-part series, and the start of a statistics block: the descriptive math underneath everything you've plotted and summarized so…
- CourseData analytics zero to hero
- Lesson23 of 30
- Video10 min
- FormatJupyter notebook · 7 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- superstore_sales.csv715.5 KB
📓 Full notebook
Download .ipynbData Analytics Zero to Hero, Video 23: Statistics Fundamentals#
- Video twenty-three of the 30-part series, and the start of a statistics block: the descriptive math underneath everything you've plotted and summarized so far.
- Real Sample Superstore sales and profit data throughout.
- Let's jump straight in.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- Install SciPy if you haven't already:
pip install scipy. - Place superstore_sales.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from scipy import stats
df = pd.read_csv('superstore_sales.csv')
sales = df['Sales']
print(sales.describe())
Part 1: Measures of Central Tendency#
mean_sales = sales.mean()
median_sales = sales.median()
mode_sales = sales.mode()[0]
print(f'Mean: {mean_sales:.2f}, Median: {median_sales:.2f}, Mode: {mode_sales:.2f}')
Part 2: Measures of Spread#
variance = sales.var()
std_dev = sales.std()
data_range = sales.max() - sales.min()
print(f'Variance: {variance:.2f}')
print(f'Std Dev: {std_dev:.2f}')
print(f'Range: {data_range:.2f}')
q1 = sales.quantile(0.25)
q3 = sales.quantile(0.75)
iqr = q3 - q1
print(f'Q1: {q1:.2f}, Q3: {q3:.2f}, IQR: {iqr:.2f}')
Part 3: Skewness and Kurtosis#
skewness = stats.skew(sales)
kurt = stats.kurtosis(sales)
print(f'Skewness: {skewness:.2f}')
print(f'Kurtosis: {kurt:.2f}')
Part 4: Z-Scores and Outliers#
z_scores = stats.zscore(sales)
df['sales_zscore'] = z_scores
outliers = df[df['sales_zscore'].abs() > 3]
print(f'{len(outliers)} real orders beyond 3 standard deviations')
print(outliers[['Category', 'Sales', 'sales_zscore']].sort_values('sales_zscore', ascending=False).head(5))
Part 5: Visualizing the Real Distribution#
import matplotlib.pyplot as plt
log_sales = np.log(sales[sales > 0])
plt.hist(log_sales, bins=40, density=True, color='steelblue', edgecolor='white', alpha=0.7)
x = np.linspace(log_sales.min(), log_sales.max(), 100)
plt.plot(x, stats.norm.pdf(x, log_sales.mean(), log_sales.std()), color='darkred', linewidth=2)
plt.title('Real Log-Transformed Sales vs. a Fitted Normal Curve')
plt.xlabel('log(Sales)')
plt.show()
Wrap-Up: What You Learned#
- Central tendency: mean, median, and mode, and why they diverge on skewed real data.
- Spread: variance, standard deviation, range, and the interquartile range.
- Skewness and kurtosis, for quantifying a real distribution's shape.
- Z-scores, for standardizing values and flagging genuine statistical outliers.
- Visualizing a real distribution against a fitted normal curve, including a log transform to reduce skew.
- All of it on real Sample Superstore sales data. Video twenty-four continues with correlation, hypothesis testing, and confidence intervals, the tools for deciding whether a real pattern is likely genuine or just noise. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



