Mathew K Analytics

Lesson 24 · Data analytics zero to hero

Correlation, P-Values & Confidence Intervals | Data Analytics #24

Video twenty-four of the 30-part series, wrapping up the statistics block: the tools for deciding whether a real pattern is likely genuine or just noise.…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Data Analytics Zero to Hero, Video 24: Correlation, Hypothesis Testing, and Confidence Intervals#

  • Video twenty-four of the 30-part series, wrapping up the statistics block: the tools for deciding whether a real pattern is likely genuine or just noise.
  • Back to the real mtcars dataset, whose real transmission-type question turns out to be a genuinely famous statistical example.
  • Let's jump straight in.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • Place mtcars.csv in the same folder as this notebook.
import pandas as pd
from scipy import stats

df = pd.read_csv('mtcars.csv')
print(df.shape)
(32, 12)

Part 1: Correlation#

corr_coef, p_value = stats.pearsonr(df['wt'], df['mpg'])
print(f'Correlation: {corr_coef:.3f}')
print(f'P-value: {p_value:.6f}')
Correlation: -0.868
P-value: 0.000000
numeric_df = df.drop(columns='model')
corr_matrix = numeric_df.corr()
print(corr_matrix['mpg'].sort_values())
wt     -0.867659
cyl    -0.852162
disp   -0.847551
hp     -0.776168
carb   -0.550925
qsec    0.418684
gear    0.480285
am      0.599832
vs      0.664039
drat    0.681172
mpg     1.000000
Name: mpg, dtype: float64

Part 2: Confidence Intervals#

mean_mpg = df['mpg'].mean()
sem_mpg = stats.sem(df['mpg'])
ci_low, ci_high = stats.t.interval(0.95, df=len(df) - 1, loc=mean_mpg, scale=sem_mpg)
print(f'Mean MPG: {mean_mpg:.2f}')
print(f'95% CI: [{ci_low:.2f}, {ci_high:.2f}]')
Mean MPG: 20.09
95% CI: [17.92, 22.26]

Part 3: Two-Sample t-test#

automatic = df[df['am'] == 0]['mpg']
manual = df[df['am'] == 1]['mpg']
print(f'Automatic mean: {automatic.mean():.2f}, n={len(automatic)}')
print(f'Manual mean: {manual.mean():.2f}, n={len(manual)}')
t_stat, p_val = stats.ttest_ind(automatic, manual)
print(f't-statistic: {t_stat:.3f}, p-value: {p_val:.4f}')
Automatic mean: 17.15, n=19
Manual mean: 24.39, n=13
t-statistic: -4.106, p-value: 0.0003

Part 4: Chi-Square Test of Independence#

contingency = pd.crosstab(df['cyl'], df['am'])
print(contingency)
chi2, p_val, dof, expected = stats.chi2_contingency(contingency)
print(f'Chi-square: {chi2:.3f}, p-value: {p_val:.4f}, dof: {dof}')
am    0  1
cyl       
4     3  8
6     4  3
8    12  2
Chi-square: 8.741, p-value: 0.0126, dof: 2

Part 5: One-Way ANOVA#

cyl4 = df[df['cyl'] == 4]['mpg']
cyl6 = df[df['cyl'] == 6]['mpg']
cyl8 = df[df['cyl'] == 8]['mpg']
f_stat, p_val = stats.f_oneway(cyl4, cyl6, cyl8)
print(f'4-cyl mean: {cyl4.mean():.2f}, 6-cyl mean: {cyl6.mean():.2f}, 8-cyl mean: {cyl8.mean():.2f}')
print(f'F-statistic: {f_stat:.3f}, p-value: {p_val:.6f}')
4-cyl mean: 26.66, 6-cyl mean: 19.74, 8-cyl mean: 15.10
F-statistic: 39.698, p-value: 0.000000

Wrap-Up: What You Learned#

  • Correlation with pearsonr, and reading a full correlation matrix.
  • Confidence intervals, for expressing real uncertainty around a sample mean instead of overstating precision.
  • Two-sample t-tests, applied to mtcars' genuinely famous automatic-versus-manual mpg question.
  • Chi-square tests, for checking association between two real categorical variables.
  • One-way ANOVA, extending a t-test to three or more real groups at once.
  • This wraps up the statistics block. Video twenty-five starts a machine learning block: your first real predictive model with scikit-learn. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.