Mathew K Analytics

Lesson 23 · Data analytics zero to hero

Statistics Fundamentals Every Analyst Needs | Data Analytics #23

Video twenty-three of the 30-part series, and the start of a statistics block: the descriptive math underneath everything you've plotted and summarized so…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Data Analytics Zero to Hero, Video 23: Statistics Fundamentals#

  • Video twenty-three of the 30-part series, and the start of a statistics block: the descriptive math underneath everything you've plotted and summarized so far.
  • Real Sample Superstore sales and profit data throughout.
  • Let's jump straight in.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • Install SciPy if you haven't already: pip install scipy.
  • Place superstore_sales.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from scipy import stats

df = pd.read_csv('superstore_sales.csv')
sales = df['Sales']
print(sales.describe())
count     9994.000000
mean       229.858001
std        623.245101
min          0.444000
25%         17.280000
50%         54.490000
75%        209.940000
max      22638.480000
Name: Sales, dtype: float64

Part 1: Measures of Central Tendency#

mean_sales = sales.mean()
median_sales = sales.median()
mode_sales = sales.mode()[0]
print(f'Mean: {mean_sales:.2f}, Median: {median_sales:.2f}, Mode: {mode_sales:.2f}')
Mean: 229.86, Median: 54.49, Mode: 12.96

Part 2: Measures of Spread#

variance = sales.var()
std_dev = sales.std()
data_range = sales.max() - sales.min()
print(f'Variance: {variance:.2f}')
print(f'Std Dev: {std_dev:.2f}')
print(f'Range: {data_range:.2f}')
Variance: 388434.46
Std Dev: 623.25
Range: 22638.04
q1 = sales.quantile(0.25)
q3 = sales.quantile(0.75)
iqr = q3 - q1
print(f'Q1: {q1:.2f}, Q3: {q3:.2f}, IQR: {iqr:.2f}')
Q1: 17.28, Q3: 209.94, IQR: 192.66

Part 3: Skewness and Kurtosis#

skewness = stats.skew(sales)
kurt = stats.kurtosis(sales)
print(f'Skewness: {skewness:.2f}')
print(f'Kurtosis: {kurt:.2f}')
Skewness: 12.97
Kurtosis: 305.16

Part 4: Z-Scores and Outliers#

z_scores = stats.zscore(sales)
df['sales_zscore'] = z_scores
outliers = df[df['sales_zscore'].abs() > 3]
print(f'{len(outliers)} real orders beyond 3 standard deviations')
print(outliers[['Category', 'Sales', 'sales_zscore']].sort_values('sales_zscore', ascending=False).head(5))
127 real orders beyond 3 standard deviations
        Category      Sales  sales_zscore
2697  Technology  22638.480     35.956549
6826  Technology  17499.950     27.711339
8153  Technology  13999.960     22.095306
2623  Technology  11199.968     17.602479
4190  Technology  10499.970     16.479273

Part 5: Visualizing the Real Distribution#

import matplotlib.pyplot as plt

log_sales = np.log(sales[sales > 0])
plt.hist(log_sales, bins=40, density=True, color='steelblue', edgecolor='white', alpha=0.7)
x = np.linspace(log_sales.min(), log_sales.max(), 100)
plt.plot(x, stats.norm.pdf(x, log_sales.mean(), log_sales.std()), color='darkred', linewidth=2)
plt.title('Real Log-Transformed Sales vs. a Fitted Normal Curve')
plt.xlabel('log(Sales)')
plt.show()
No description has been provided for this image

Wrap-Up: What You Learned#

  • Central tendency: mean, median, and mode, and why they diverge on skewed real data.
  • Spread: variance, standard deviation, range, and the interquartile range.
  • Skewness and kurtosis, for quantifying a real distribution's shape.
  • Z-scores, for standardizing values and flagging genuine statistical outliers.
  • Visualizing a real distribution against a fitted normal curve, including a log transform to reduce skew.
  • All of it on real Sample Superstore sales data. Video twenty-four continues with correlation, hypothesis testing, and confidence intervals, the tools for deciding whether a real pattern is likely genuine or just noise. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.