Lesson 16 · Real-World Data Analytics
Python Data Analytics #16: Backtesting a Simple Trading Strategy in Python
Video sixteen of the hundred-video real-world data analytics series. Turning the real moving-average crossover into an honest, cost-aware backtest against…
- CourseReal-World Data Analytics
- Lesson16 of 100
- Video31 min
- FormatJupyter notebook · 30 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- aapl_clean.csv58.9 KB
📓 Full notebook
Download .ipynbData Analytics 100, Video 16: Backtesting a Simple Trading Strategy#
- Video sixteen of the hundred-video real-world data analytics series.
- Turning the real moving-average crossover into an honest, cost-aware backtest against real Apple prices.
- Let's get into it.
Part 1: A Real Backtest Needs Real Honesty#
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
aapl = pd.read_csv('aapl_clean.csv', parse_dates=['Date'])
close = aapl['AAPL.Close']
Part 2: Real Strategy Rules#
sma_fast = close.rolling(10).mean()
sma_slow = close.rolling(30).mean()
signal = (sma_fast > sma_slow).astype(int)
signal.value_counts()
Part 3: Real Position Lag, Avoiding Real Lookahead Bias#
position = signal.shift(1).fillna(0)
daily_return = close.pct_change()
position.head(12)
Part 4: Real Gross Strategy Returns#
gross_strategy_return = position * daily_return
cumulative_gross = (1 + gross_strategy_return.fillna(0)).cumprod()
round(cumulative_gross.iloc[-1], 3)
Part 5: Real Transaction Costs#
position_change = position.diff().abs().fillna(0)
cost_per_trade = 0.001
transaction_costs = position_change * cost_per_trade
transaction_costs.sum()
Part 6: Real Net Strategy Returns#
net_strategy_return = gross_strategy_return - transaction_costs
cumulative_net = (1 + net_strategy_return.fillna(0)).cumprod()
round(cumulative_net.iloc[-1], 3)
Part 7: Real Strategy vs Real Buy-and-Hold#
cumulative_buyhold = (1 + daily_return.fillna(0)).cumprod()
round(cumulative_net.iloc[-1], 3), round(cumulative_buyhold.iloc[-1], 3)
Part 8: Visualizing Real Net Strategy vs Real Buy-and-Hold#
plt.figure(figsize=(11, 6))
plt.plot(aapl['Date'], cumulative_net, label='Real Net Strategy (After Costs)')
plt.plot(aapl['Date'], cumulative_gross, label='Real Gross Strategy (No Costs)', linestyle='--', alpha=0.6)
plt.plot(aapl['Date'], cumulative_buyhold, label='Real Buy and Hold')
plt.xlabel('Real Date')
plt.ylabel('Real Growth of $1 Invested')
plt.title('Real SMA Crossover Backtest, Gross vs Net vs Buy-and-Hold')
plt.legend()
plt.tight_layout()
plt.savefig('backtest_equity_curve.png', dpi=120)
plt.close()
Part 9: Real Trade-by-Trade Log#
position_diff = position.diff().fillna(0)
entry_indices = position[position_diff == 1].index.tolist()
exit_indices = position[position_diff == -1].index.tolist()
len(entry_indices), len(exit_indices)
Part 10: Real Pairing of Entries and Exits#
trades = []
for entry_idx in entry_indices:
later_exits = [x for x in exit_indices if x > entry_idx]
if later_exits:
trades.append((entry_idx, later_exits[0]))
len(trades)
Part 11: Real Trade Returns#
trade_returns = [close.iloc[exit_i] / close.iloc[entry_i] - 1 for entry_i, exit_i in trades]
trade_returns_series = pd.Series(trade_returns)
trade_returns_series.round(4)
Part 12: Real Win Rate#
real_wins = trade_returns_series[trade_returns_series > 0]
real_losses = trade_returns_series[trade_returns_series <= 0]
win_rate = len(real_wins) / len(trade_returns_series)
round(win_rate * 100, 1)
Part 13: Real Profit Factor#
total_wins_amount = real_wins.sum()
total_losses_amount = abs(real_losses.sum())
profit_factor = total_wins_amount / total_losses_amount if total_losses_amount > 0 else np.nan
round(profit_factor, 3)
Part 14: Real Average Win vs Real Average Loss#
avg_win = round(real_wins.mean() * 100, 2)
avg_loss = round(real_losses.mean() * 100, 2)
avg_win, avg_loss
Part 15: Real Strategy Sharpe Ratio#
strategy_annual_return = net_strategy_return.mean() * 252
strategy_annual_vol = net_strategy_return.std() * np.sqrt(252)
strategy_sharpe = strategy_annual_return / strategy_annual_vol
round(strategy_sharpe, 3)
Part 16: Real Buy-and-Hold Sharpe Ratio#
buyhold_annual_return = daily_return.mean() * 252
buyhold_annual_vol = daily_return.std() * np.sqrt(252)
buyhold_sharpe = buyhold_annual_return / buyhold_annual_vol
round(strategy_sharpe, 3), round(buyhold_sharpe, 3)
Part 17: Real Maximum Drawdown Comparison#
def max_drawdown(cumulative):
running_peak = cumulative.cummax()
return ((cumulative - running_peak) / running_peak).min()
round(max_drawdown(cumulative_net) * 100, 2), round(max_drawdown(cumulative_buyhold) * 100, 2)
Part 18: Real Time in the Market#
pct_time_invested = position.mean()
round(pct_time_invested * 100, 1)
Part 19: Real Walk-Forward Split#
split_point = int(len(aapl) * 0.6)
in_sample_net_return = net_strategy_return.iloc[:split_point]
out_sample_net_return = net_strategy_return.iloc[split_point:]
in_sample_growth = (1 + in_sample_net_return.fillna(0)).cumprod().iloc[-1]
out_sample_growth = (1 + out_sample_net_return.fillna(0)).cumprod().iloc[-1]
round(in_sample_growth, 3), round(out_sample_growth, 3)
Part 20: A Real Alternative Strategy, RSI Mean-Reversion#
delta = close.diff()
gain = delta.clip(lower=0)
loss = -delta.clip(upper=0)
rs = gain.rolling(14).mean() / loss.rolling(14).mean()
rsi = 100 - (100 / (1 + rs))
rsi_signal = (rsi < 30).astype(int)
Part 21: Real RSI Strategy Backtest#
rsi_position = rsi_signal.shift(1).fillna(0)
rsi_gross_return = rsi_position * daily_return
rsi_trade_flags = rsi_position.diff().abs().fillna(0)
rsi_net_return = rsi_gross_return - (rsi_trade_flags * cost_per_trade)
cumulative_rsi_net = (1 + rsi_net_return.fillna(0)).cumprod()
round(cumulative_rsi_net.iloc[-1], 3)
Part 22: Real Head-to-Head Strategy Comparison#
comparison = pd.Series({'SMA Crossover (Net)': cumulative_net.iloc[-1], 'RSI Mean-Reversion (Net)': cumulative_rsi_net.iloc[-1], 'Buy and Hold': cumulative_buyhold.iloc[-1]})
comparison.round(3).sort_values(ascending=False)
Part 23: Visualizing All Three Real Equity Curves#
plt.figure(figsize=(11, 6))
plt.plot(aapl['Date'], cumulative_net, label='Real SMA Crossover (Net)')
plt.plot(aapl['Date'], cumulative_rsi_net, label='Real RSI Mean-Reversion (Net)')
plt.plot(aapl['Date'], cumulative_buyhold, label='Real Buy and Hold')
plt.xlabel('Real Date')
plt.ylabel('Real Growth of $1 Invested')
plt.title('Real Strategy Comparison, All Net of Costs')
plt.legend()
plt.tight_layout()
plt.savefig('strategy_comparison.png', dpi=120)
plt.close()
Part 24: Real Cost Sensitivity Check#
higher_cost = 0.005
stressed_net_return = gross_strategy_return - (position_change * higher_cost)
stressed_cumulative = (1 + stressed_net_return.fillna(0)).cumprod()
round(stressed_cumulative.iloc[-1], 3)
Part 25: Real Number of Trades vs Real Cost Assumption#
n_completed_trades = len(trades)
total_cost_drag_pct = round(transaction_costs.sum() * 100, 2)
n_completed_trades, total_cost_drag_pct
Part 26: Saving the Real Trade Log#
trade_log = pd.DataFrame({'EntryDate': [aapl['Date'].iloc[e] for e, x in trades], 'ExitDate': [aapl['Date'].iloc[x] for e, x in trades], 'Return': trade_returns})
trade_log.to_csv('trade_log.csv', index=False)
reloaded_log = pd.read_csv('trade_log.csv')
reloaded_log.shape[0] == len(trades)
Part 27: Saving the Real Backtest Summary#
backtest_summary = pd.DataFrame({'Strategy': ['SMA Crossover', 'RSI Mean-Reversion', 'Buy and Hold'], 'EndingValue': [round(cumulative_net.iloc[-1], 3), round(cumulative_rsi_net.iloc[-1], 3), round(cumulative_buyhold.iloc[-1], 3)], 'MaxDrawdownPct': [round(max_drawdown(cumulative_net) * 100, 2), round(max_drawdown(cumulative_rsi_net) * 100, 2), round(max_drawdown(cumulative_buyhold) * 100, 2)]})
backtest_summary.to_csv('backtest_summary.csv', index=False)
reloaded_summary = pd.read_csv('backtest_summary.csv')
reloaded_summary.shape[0] == 3
Part 28: Real Sanity Check on Trade Pairing#
all(exit_i > entry_i for entry_i, exit_i in trades)
Part 29: Real Recap Print#
net_pct = round((cumulative_net.iloc[-1] - 1) * 100, 1)
bh_pct = round((cumulative_buyhold.iloc[-1] - 1) * 100, 1)
print(f'After real trading costs across {n_completed_trades} completed trades, the SMA crossover returned {net_pct}% versus {bh_pct}% for buy-and-hold, with a {round(win_rate*100,1)}% win rate.')
Part 30: Real Longest Winning Streak#
import itertools
streak_signs = [1 if r > 0 else -1 for r in trade_returns]
win_streaks = [sum(1 for _ in group) for key, group in itertools.groupby(streak_signs) if key == 1]
longest_win_streak = max(win_streaks) if win_streaks else 0
longest_win_streak
Wrap-Up: What You Learned#
- An honest backtest has to lag the real signal by one day and charge a real trading cost every time the position changes, otherwise it is quietly cheating.
- Reconstructing individual real trades reveals a win rate and profit factor that a single aggregate return number would hide completely.
- Comparing real Sharpe ratio and real maximum drawdown side by side, not just ending value, is what tells you whether a strategy is actually worth holding through.
- A real walk-forward split is a basic, honest check for whether a fixed rule is genuinely robust or just fit to one lucky stretch of history.
- Testing a real second, different style of strategy, mean-reversion instead of momentum, keeps you from mistaking one lucky rule for a genuinely reliable edge.
- Next video: applying these real same tools to a real cryptocurrency market, testing whether these lessons carry over to a very differently behaved real asset.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



