Mathew K Analytics

Lesson 40 · Probability and Statistics in python

Comprehensive Time Series Forecasting Project: From Data Preparation to Deployment

In this lesson, we will explore probability, statistics, and forecasting using Python. We will use real datasets to learn new skills for analyzing and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Welcome: Intro to Time Series Forecasting in Python#

In this lesson, we will explore probability, statistics, and forecasting using Python.

We will use real datasets to learn new skills for analyzing and predicting time-based data.

No prior knowledge is needed. Let us get started and have fun!

# Suppress warnings for a cleaner notebook.
import warnings; warnings.filterwarnings("ignore")
 
# Import libraries for data and plotting.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
np.random.seed(42)
%matplotlib inline
# Data setup: load airline passengers monthly dataset.
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/airline-passengers.csv"
df = pd.read_csv(url)

print("Data shape:", df.shape)
df.head()
Data shape: (144, 2)
Month Passengers
0 1949-01 112
1 1949-02 118
2 1949-03 132
3 1949-04 129
4 1949-05 121
# First look: line plot of passengers over time.
plt.figure(figsize=(10,4))
plt.plot(df['Month'], df['Passengers'], marker='o')
plt.title("Monthly Airline Passengers 1949-1960")
plt.xlabel("Month")
plt.ylabel("Passengers")
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
No description has been provided for this image

What is a Time Series?#

A time series is data collected over a sequence of time points. For example, daily temperatures or monthly sales.

Time series analysis helps us to track changes, spot trends, and make predictions for the future.

# Probability in action: random coin flip simulation.
np.random.seed(42)
flips = np.random.choice(["Heads", "Tails"], size=10, p=[0.5, 0.5])
print("Coin flips:", flips)
Coin flips: ['Heads' 'Tails' 'Tails' 'Tails' 'Heads' 'Heads' 'Heads' 'Tails' 'Tails'
 'Tails']
# Descriptive statistics: look at passenger summary.
print("How many months?", len(df))
print("Average passengers:", np.mean(df['Passengers']))
print("Smallest month:", np.min(df['Passengers']))
print("Largest month:", np.max(df['Passengers']))
How many months? 144
Average passengers: 280.2986111111111
Smallest month: 104
Largest month: 622
# Handle missing values (if any): is the data clean?
print("Any missing data?", df.isnull().values.any())

# Fill or drop missing data if needed.
df_clean = df.dropna()
Any missing data? False
# Let us try a Bernoulli distribution simulation.
from scipy.stats import bernoulli
outcomes = bernoulli.rvs(p=0.7, size=12, random_state=42)
print("Bernoulli outcomes:", outcomes)
Bernoulli outcomes: [1 0 0 1 1 1 1 0 1 0 1 0]
# Binomial distribution: number of successes in many tries.
from scipy.stats import binom
trials = 15
prob_success = 0.6
binom_samples = binom.rvs(trials, prob_success, size=8, random_state=42)
print("Binomial samples:", binom_samples)
Binomial samples: [10  6  8  9 11 11 12  7]
# Continuous distribution example: normally distributed data.
normal_data = np.random.normal(loc=0, scale=1, size=1000)
plt.hist(normal_data, bins=30, alpha=0.7)
plt.title("Histogram of 1000 random values (Normal Distribution)")
plt.xlabel("Value")
plt.ylabel("Frequency")
plt.show()
No description has been provided for this image
# Sampling and the Central Limit Theorem (CLT) demo.
pop = np.random.exponential(scale=1.5, size=5000)
means = [np.mean(np.random.choice(pop, 30)) for _ in range(200)]
plt.hist(means, bins=25, alpha=0.7)
plt.title("Sample Means from Exponential Population")
plt.xlabel("Sample mean")
plt.ylabel("Count")
plt.show()
No description has been provided for this image
# Confidence interval for mean passengers.
import scipy.stats as stats
mu = np.mean(df['Passengers'])
sigma = np.std(df['Passengers'], ddof=1)
n = len(df['Passengers'])
conf_int = stats.t.interval(0.95, n-1, loc=mu, scale=sigma/np.sqrt(n))
print("95% confidence interval for mean passengers:", conf_int)
95% confidence interval for mean passengers: (np.float64(260.53723755161025), np.float64(300.0599846706119))
# Hypothesis test: Did passengers grow from 1949 to 1960?
early = df[df['Month'] < '1955']['Passengers']
late = df[df['Month'] >= '1955']['Passengers']
t_stat, p_value = stats.ttest_ind(early, late, equal_var=False)
print("P-value:", p_value)
P-value: 4.2899079396310666e-32
# Correlation and covariance: link between time and passengers.
df["MonthNum"] = np.arange(len(df))
corr = np.corrcoef(df["MonthNum"], df["Passengers"])[0,1]
cov = np.cov(df["MonthNum"], df["Passengers"])[0,1]
print("Correlation:", corr)
print("Covariance:", cov)
Correlation: 0.9239254112768995
Covariance: 4623.5
# Simple linear regression: predict next month's passengers.
from sklearn.linear_model import LinearRegression
model = LinearRegression()
X = df[["MonthNum"]]
y = df["Passengers"]
model.fit(X, y)
preds = model.predict(X)
plt.plot(df["MonthNum"], y, label="Actual")
plt.plot(df["MonthNum"], preds, label="Predicted")
plt.legend()
plt.title("Linear Regression Fit")
plt.xlabel("Month Number")
plt.ylabel("Passengers")
plt.show()
No description has been provided for this image
# Logistic regression preview: Titanic dataset survival prediction.
import seaborn as sns
titanic = sns.load_dataset("titanic")
titanic = titanic.dropna(subset=["age", "fare"])
from sklearn.linear_model import LogisticRegression
X_titanic = titanic[["age", "fare"]]
y_titanic = titanic["survived"]
logr = LogisticRegression(max_iter=200)
logr.fit(X_titanic, y_titanic)
score = logr.score(X_titanic, y_titanic)
print("Accuracy:", round(score, 2))
Accuracy: 0.66
# Time series train-test split: predict the last year.
train = df[df['Month'] < '1959-01']
test = df[df['Month'] >= '1959-01']
print("Train months:", train.shape[0])
print("Test months:", test.shape[0])
Train months: 120
Test months: 24
# Simple moving average forecast for test months.
test = test.copy()
window = 12
train_means = train['Passengers'].rolling(window=window).mean().dropna()
last_mean = train_means.iloc[-1]
test['SMA_Pred'] = last_mean
plt.plot(train['Month'], train['Passengers'], label='Train')
plt.plot(test['Month'], test['Passengers'], label='Test')
plt.plot(test['Month'], test['SMA_Pred'], label='SMA Prediction')
plt.legend()
plt.title('Moving Average Forecast Example')
plt.xlabel('Month')
plt.ylabel('Passengers')
plt.show()
No description has been provided for this image
# Bonus: Let us try a real forecasting model (ARIMA).
from statsmodels.tsa.arima.model import ARIMA
model_arima = ARIMA(train['Passengers'], order=(1,1,1))
fit_arima = model_arima.fit()
forecast = fit_arima.forecast(steps=len(test))
test = test.copy()
test['ARIMA_Pred'] = forecast.values
plt.plot(train['Month'], train['Passengers'], label='Train')
plt.plot(test['Month'], test['Passengers'], label='Test')
plt.plot(test['Month'], test['ARIMA_Pred'], label='ARIMA Forecast')
plt.legend()
plt.title('ARIMA Model Forecast')
plt.xlabel('Month')
plt.ylabel('Passengers')
plt.show()
No description has been provided for this image
# Mini-project part 1: Simulate a seasonal sales pattern.
np.random.seed(100)
months = np.arange(60)
season = 20 * np.sin(2*np.pi*months/12)
trend = 3 * months
noise = np.random.normal(0, 10, 60)
sales = 100 + trend + season + noise
plt.plot(months, sales)
plt.title("Simulated Monthly Sales (5 Years)")
plt.xlabel("Month")
plt.ylabel("Sales")
plt.show()
No description has been provided for this image
# Mini-project part 2: Analyze daily temperature real data.
url_temp = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/daily-min-temperatures.csv"
temps = pd.read_csv(url_temp)
temps['Date'] = pd.to_datetime(temps['Date'])
temps.set_index('Date', inplace=True)
temps['Temp'].plot(figsize=(10,4))
plt.title('Daily Minimum Temperatures, Melbourne')
plt.xlabel('Date')
plt.ylabel('Min Temp (C)')
plt.show()
No description has been provided for this image

Best Practices for Time Series#

  • Always check for missing dates or jumps in your data.
  • Use train-test splits that respect time order. Do not mix past and future.
  • Visualize everything to spot surprises early.
  • Try simple models before advanced ones.
# Extra tip: making features from dates.
temps['Month'] = temps.index.month
monthly_avg = temps.groupby('Month')['Temp'].mean()
print(monthly_avg)
Month
1     15.030323
2     15.373759
3     14.565484
4     12.088333
5      9.866452
6      7.278333
7      6.692581
8      7.891290
9      8.976333
10    10.309355
11    12.479667
12    13.851948
Name: Temp, dtype: float64
# Challenge exercise: Shift temperature data and compare.
temps['Yesterday'] = temps['Temp'].shift(1)
plt.scatter(temps['Yesterday'], temps['Temp'], alpha=0.5)
plt.title('Current Temp vs Previous Day')
plt.xlabel('Yesterday Temp (C)')
plt.ylabel('Today Temp (C)')
plt.show()
No description has been provided for this image

Recap: You are now ready for time series analysis!#

You learned how to:

  • Summarize and plot time-based data.
  • Build simple forecasts and advanced models.
  • Handle missing values and practical real-world datasets.

Keep experimenting. Every dataset is a new story to discover!

Thank you for joining!

Try the exercises, pause as you need, and share what you learned.

For more videos, please like, subscribe, and turn on notifications.

See you next time for more Python learning!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.