Lesson 40 · Probability and Statistics in python
Comprehensive Time Series Forecasting Project: From Data Preparation to Deployment
In this lesson, we will explore probability, statistics, and forecasting using Python. We will use real datasets to learn new skills for analyzing and…
- CourseProbability and Statistics in python
- Lesson40 of 35
- Video12 min
- FormatJupyter notebook · 22 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWelcome: Intro to Time Series Forecasting in Python#
In this lesson, we will explore probability, statistics, and forecasting using Python.
We will use real datasets to learn new skills for analyzing and predicting time-based data.
No prior knowledge is needed. Let us get started and have fun!
# Suppress warnings for a cleaner notebook.
import warnings; warnings.filterwarnings("ignore")
# Import libraries for data and plotting.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
np.random.seed(42)
%matplotlib inline
# Data setup: load airline passengers monthly dataset.
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/airline-passengers.csv"
df = pd.read_csv(url)
print("Data shape:", df.shape)
df.head()
# First look: line plot of passengers over time.
plt.figure(figsize=(10,4))
plt.plot(df['Month'], df['Passengers'], marker='o')
plt.title("Monthly Airline Passengers 1949-1960")
plt.xlabel("Month")
plt.ylabel("Passengers")
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
What is a Time Series?#
A time series is data collected over a sequence of time points. For example, daily temperatures or monthly sales.
Time series analysis helps us to track changes, spot trends, and make predictions for the future.
# Probability in action: random coin flip simulation.
np.random.seed(42)
flips = np.random.choice(["Heads", "Tails"], size=10, p=[0.5, 0.5])
print("Coin flips:", flips)
# Descriptive statistics: look at passenger summary.
print("How many months?", len(df))
print("Average passengers:", np.mean(df['Passengers']))
print("Smallest month:", np.min(df['Passengers']))
print("Largest month:", np.max(df['Passengers']))
# Handle missing values (if any): is the data clean?
print("Any missing data?", df.isnull().values.any())
# Fill or drop missing data if needed.
df_clean = df.dropna()
# Let us try a Bernoulli distribution simulation.
from scipy.stats import bernoulli
outcomes = bernoulli.rvs(p=0.7, size=12, random_state=42)
print("Bernoulli outcomes:", outcomes)
# Binomial distribution: number of successes in many tries.
from scipy.stats import binom
trials = 15
prob_success = 0.6
binom_samples = binom.rvs(trials, prob_success, size=8, random_state=42)
print("Binomial samples:", binom_samples)
# Continuous distribution example: normally distributed data.
normal_data = np.random.normal(loc=0, scale=1, size=1000)
plt.hist(normal_data, bins=30, alpha=0.7)
plt.title("Histogram of 1000 random values (Normal Distribution)")
plt.xlabel("Value")
plt.ylabel("Frequency")
plt.show()
# Sampling and the Central Limit Theorem (CLT) demo.
pop = np.random.exponential(scale=1.5, size=5000)
means = [np.mean(np.random.choice(pop, 30)) for _ in range(200)]
plt.hist(means, bins=25, alpha=0.7)
plt.title("Sample Means from Exponential Population")
plt.xlabel("Sample mean")
plt.ylabel("Count")
plt.show()
# Confidence interval for mean passengers.
import scipy.stats as stats
mu = np.mean(df['Passengers'])
sigma = np.std(df['Passengers'], ddof=1)
n = len(df['Passengers'])
conf_int = stats.t.interval(0.95, n-1, loc=mu, scale=sigma/np.sqrt(n))
print("95% confidence interval for mean passengers:", conf_int)
# Hypothesis test: Did passengers grow from 1949 to 1960?
early = df[df['Month'] < '1955']['Passengers']
late = df[df['Month'] >= '1955']['Passengers']
t_stat, p_value = stats.ttest_ind(early, late, equal_var=False)
print("P-value:", p_value)
# Correlation and covariance: link between time and passengers.
df["MonthNum"] = np.arange(len(df))
corr = np.corrcoef(df["MonthNum"], df["Passengers"])[0,1]
cov = np.cov(df["MonthNum"], df["Passengers"])[0,1]
print("Correlation:", corr)
print("Covariance:", cov)
# Simple linear regression: predict next month's passengers.
from sklearn.linear_model import LinearRegression
model = LinearRegression()
X = df[["MonthNum"]]
y = df["Passengers"]
model.fit(X, y)
preds = model.predict(X)
plt.plot(df["MonthNum"], y, label="Actual")
plt.plot(df["MonthNum"], preds, label="Predicted")
plt.legend()
plt.title("Linear Regression Fit")
plt.xlabel("Month Number")
plt.ylabel("Passengers")
plt.show()
# Logistic regression preview: Titanic dataset survival prediction.
import seaborn as sns
titanic = sns.load_dataset("titanic")
titanic = titanic.dropna(subset=["age", "fare"])
from sklearn.linear_model import LogisticRegression
X_titanic = titanic[["age", "fare"]]
y_titanic = titanic["survived"]
logr = LogisticRegression(max_iter=200)
logr.fit(X_titanic, y_titanic)
score = logr.score(X_titanic, y_titanic)
print("Accuracy:", round(score, 2))
# Time series train-test split: predict the last year.
train = df[df['Month'] < '1959-01']
test = df[df['Month'] >= '1959-01']
print("Train months:", train.shape[0])
print("Test months:", test.shape[0])
# Simple moving average forecast for test months.
test = test.copy()
window = 12
train_means = train['Passengers'].rolling(window=window).mean().dropna()
last_mean = train_means.iloc[-1]
test['SMA_Pred'] = last_mean
plt.plot(train['Month'], train['Passengers'], label='Train')
plt.plot(test['Month'], test['Passengers'], label='Test')
plt.plot(test['Month'], test['SMA_Pred'], label='SMA Prediction')
plt.legend()
plt.title('Moving Average Forecast Example')
plt.xlabel('Month')
plt.ylabel('Passengers')
plt.show()
# Bonus: Let us try a real forecasting model (ARIMA).
from statsmodels.tsa.arima.model import ARIMA
model_arima = ARIMA(train['Passengers'], order=(1,1,1))
fit_arima = model_arima.fit()
forecast = fit_arima.forecast(steps=len(test))
test = test.copy()
test['ARIMA_Pred'] = forecast.values
plt.plot(train['Month'], train['Passengers'], label='Train')
plt.plot(test['Month'], test['Passengers'], label='Test')
plt.plot(test['Month'], test['ARIMA_Pred'], label='ARIMA Forecast')
plt.legend()
plt.title('ARIMA Model Forecast')
plt.xlabel('Month')
plt.ylabel('Passengers')
plt.show()
# Mini-project part 1: Simulate a seasonal sales pattern.
np.random.seed(100)
months = np.arange(60)
season = 20 * np.sin(2*np.pi*months/12)
trend = 3 * months
noise = np.random.normal(0, 10, 60)
sales = 100 + trend + season + noise
plt.plot(months, sales)
plt.title("Simulated Monthly Sales (5 Years)")
plt.xlabel("Month")
plt.ylabel("Sales")
plt.show()
# Mini-project part 2: Analyze daily temperature real data.
url_temp = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/daily-min-temperatures.csv"
temps = pd.read_csv(url_temp)
temps['Date'] = pd.to_datetime(temps['Date'])
temps.set_index('Date', inplace=True)
temps['Temp'].plot(figsize=(10,4))
plt.title('Daily Minimum Temperatures, Melbourne')
plt.xlabel('Date')
plt.ylabel('Min Temp (C)')
plt.show()
Best Practices for Time Series#
- Always check for missing dates or jumps in your data.
- Use train-test splits that respect time order. Do not mix past and future.
- Visualize everything to spot surprises early.
- Try simple models before advanced ones.
# Extra tip: making features from dates.
temps['Month'] = temps.index.month
monthly_avg = temps.groupby('Month')['Temp'].mean()
print(monthly_avg)
# Challenge exercise: Shift temperature data and compare.
temps['Yesterday'] = temps['Temp'].shift(1)
plt.scatter(temps['Yesterday'], temps['Temp'], alpha=0.5)
plt.title('Current Temp vs Previous Day')
plt.xlabel('Yesterday Temp (C)')
plt.ylabel('Today Temp (C)')
plt.show()
Recap: You are now ready for time series analysis!#
You learned how to:
- Summarize and plot time-based data.
- Build simple forecasts and advanced models.
- Handle missing values and practical real-world datasets.
Keep experimenting. Every dataset is a new story to discover!
Thank you for joining!
Try the exercises, pause as you need, and share what you learned.
For more videos, please like, subscribe, and turn on notifications.
See you next time for more Python learning!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



