Mathew K Analytics

Lesson 16 · Real-World Data Analytics

Python Data Analytics #16: Backtesting a Simple Trading Strategy in Python

Video sixteen of the hundred-video real-world data analytics series. Turning the real moving-average crossover into an honest, cost-aware backtest against…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Data Analytics 100, Video 16: Backtesting a Simple Trading Strategy#

  • Video sixteen of the hundred-video real-world data analytics series.
  • Turning the real moving-average crossover into an honest, cost-aware backtest against real Apple prices.
  • Let's get into it.

Part 1: A Real Backtest Needs Real Honesty#

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
aapl = pd.read_csv('aapl_clean.csv', parse_dates=['Date'])
close = aapl['AAPL.Close']

Part 2: Real Strategy Rules#

sma_fast = close.rolling(10).mean()
sma_slow = close.rolling(30).mean()
signal = (sma_fast > sma_slow).astype(int)
signal.value_counts()
AAPL.Close
1    255
0    251
Name: count, dtype: int64

Part 3: Real Position Lag, Avoiding Real Lookahead Bias#

position = signal.shift(1).fillna(0)
daily_return = close.pct_change()
position.head(12)
0     0.0
1     0.0
2     0.0
3     0.0
4     0.0
5     0.0
6     0.0
7     0.0
8     0.0
9     0.0
10    0.0
11    0.0
Name: AAPL.Close, dtype: float64

Part 4: Real Gross Strategy Returns#

gross_strategy_return = position * daily_return
cumulative_gross = (1 + gross_strategy_return.fillna(0)).cumprod()
round(cumulative_gross.iloc[-1], 3)
np.float64(1.17)

Part 5: Real Transaction Costs#

position_change = position.diff().abs().fillna(0)
cost_per_trade = 0.001
transaction_costs = position_change * cost_per_trade
transaction_costs.sum()
np.float64(0.021)

Part 6: Real Net Strategy Returns#

net_strategy_return = gross_strategy_return - transaction_costs
cumulative_net = (1 + net_strategy_return.fillna(0)).cumprod()
round(cumulative_net.iloc[-1], 3)
np.float64(1.146)

Part 7: Real Strategy vs Real Buy-and-Hold#

cumulative_buyhold = (1 + daily_return.fillna(0)).cumprod()
round(cumulative_net.iloc[-1], 3), round(cumulative_buyhold.iloc[-1], 3)
(np.float64(1.146), np.float64(1.059))

Part 8: Visualizing Real Net Strategy vs Real Buy-and-Hold#

plt.figure(figsize=(11, 6))
plt.plot(aapl['Date'], cumulative_net, label='Real Net Strategy (After Costs)')
plt.plot(aapl['Date'], cumulative_gross, label='Real Gross Strategy (No Costs)', linestyle='--', alpha=0.6)
plt.plot(aapl['Date'], cumulative_buyhold, label='Real Buy and Hold')
plt.xlabel('Real Date')
plt.ylabel('Real Growth of $1 Invested')
plt.title('Real SMA Crossover Backtest, Gross vs Net vs Buy-and-Hold')
plt.legend()
plt.tight_layout()
plt.savefig('backtest_equity_curve.png', dpi=120)
plt.close()

Part 9: Real Trade-by-Trade Log#

position_diff = position.diff().fillna(0)
entry_indices = position[position_diff == 1].index.tolist()
exit_indices = position[position_diff == -1].index.tolist()
len(entry_indices), len(exit_indices)
(11, 10)

Part 10: Real Pairing of Entries and Exits#

trades = []
for entry_idx in entry_indices:
    later_exits = [x for x in exit_indices if x > entry_idx]
    if later_exits:
        trades.append((entry_idx, later_exits[0]))
len(trades)
10

Part 11: Real Trade Returns#

trade_returns = [close.iloc[exit_i] / close.iloc[entry_i] - 1 for entry_i, exit_i in trades]
trade_returns_series = pd.Series(trade_returns)
trade_returns_series.round(4)
0   -0.0061
1   -0.0288
2   -0.0308
3   -0.0276
4    0.0312
5    0.0332
6   -0.0270
7   -0.0514
8    0.0570
9   -0.0417
dtype: float64

Part 12: Real Win Rate#

real_wins = trade_returns_series[trade_returns_series > 0]
real_losses = trade_returns_series[trade_returns_series <= 0]
win_rate = len(real_wins) / len(trade_returns_series)
round(win_rate * 100, 1)
30.0

Part 13: Real Profit Factor#

total_wins_amount = real_wins.sum()
total_losses_amount = abs(real_losses.sum())
profit_factor = total_wins_amount / total_losses_amount if total_losses_amount > 0 else np.nan
round(profit_factor, 3)
np.float64(0.569)

Part 14: Real Average Win vs Real Average Loss#

avg_win = round(real_wins.mean() * 100, 2)
avg_loss = round(real_losses.mean() * 100, 2)
avg_win, avg_loss
(np.float64(4.05), np.float64(-3.05))

Part 15: Real Strategy Sharpe Ratio#

strategy_annual_return = net_strategy_return.mean() * 252
strategy_annual_vol = net_strategy_return.std() * np.sqrt(252)
strategy_sharpe = strategy_annual_return / strategy_annual_vol
round(strategy_sharpe, 3)
np.float64(0.559)

Part 16: Real Buy-and-Hold Sharpe Ratio#

buyhold_annual_return = daily_return.mean() * 252
buyhold_annual_vol = daily_return.std() * np.sqrt(252)
buyhold_sharpe = buyhold_annual_return / buyhold_annual_vol
round(strategy_sharpe, 3), round(buyhold_sharpe, 3)
(np.float64(0.559), np.float64(0.239))

Part 17: Real Maximum Drawdown Comparison#

def max_drawdown(cumulative):
    running_peak = cumulative.cummax()
    return ((cumulative - running_peak) / running_peak).min()
round(max_drawdown(cumulative_net) * 100, 2), round(max_drawdown(cumulative_buyhold) * 100, 2)
(np.float64(-19.03), np.float64(-32.08))

Part 18: Real Time in the Market#

pct_time_invested = position.mean()
round(pct_time_invested * 100, 1)
np.float64(50.2)

Part 19: Real Walk-Forward Split#

split_point = int(len(aapl) * 0.6)
in_sample_net_return = net_strategy_return.iloc[:split_point]
out_sample_net_return = net_strategy_return.iloc[split_point:]
in_sample_growth = (1 + in_sample_net_return.fillna(0)).cumprod().iloc[-1]
out_sample_growth = (1 + out_sample_net_return.fillna(0)).cumprod().iloc[-1]
round(in_sample_growth, 3), round(out_sample_growth, 3)
(np.float64(0.947), np.float64(1.209))

Part 20: A Real Alternative Strategy, RSI Mean-Reversion#

delta = close.diff()
gain = delta.clip(lower=0)
loss = -delta.clip(upper=0)
rs = gain.rolling(14).mean() / loss.rolling(14).mean()
rsi = 100 - (100 / (1 + rs))
rsi_signal = (rsi < 30).astype(int)

Part 21: Real RSI Strategy Backtest#

rsi_position = rsi_signal.shift(1).fillna(0)
rsi_gross_return = rsi_position * daily_return
rsi_trade_flags = rsi_position.diff().abs().fillna(0)
rsi_net_return = rsi_gross_return - (rsi_trade_flags * cost_per_trade)
cumulative_rsi_net = (1 + rsi_net_return.fillna(0)).cumprod()
round(cumulative_rsi_net.iloc[-1], 3)
np.float64(0.883)

Part 22: Real Head-to-Head Strategy Comparison#

comparison = pd.Series({'SMA Crossover (Net)': cumulative_net.iloc[-1], 'RSI Mean-Reversion (Net)': cumulative_rsi_net.iloc[-1], 'Buy and Hold': cumulative_buyhold.iloc[-1]})
comparison.round(3).sort_values(ascending=False)
SMA Crossover (Net)         1.146
Buy and Hold                1.059
RSI Mean-Reversion (Net)    0.883
dtype: float64

Part 23: Visualizing All Three Real Equity Curves#

plt.figure(figsize=(11, 6))
plt.plot(aapl['Date'], cumulative_net, label='Real SMA Crossover (Net)')
plt.plot(aapl['Date'], cumulative_rsi_net, label='Real RSI Mean-Reversion (Net)')
plt.plot(aapl['Date'], cumulative_buyhold, label='Real Buy and Hold')
plt.xlabel('Real Date')
plt.ylabel('Real Growth of $1 Invested')
plt.title('Real Strategy Comparison, All Net of Costs')
plt.legend()
plt.tight_layout()
plt.savefig('strategy_comparison.png', dpi=120)
plt.close()

Part 24: Real Cost Sensitivity Check#

higher_cost = 0.005
stressed_net_return = gross_strategy_return - (position_change * higher_cost)
stressed_cumulative = (1 + stressed_net_return.fillna(0)).cumprod()
round(stressed_cumulative.iloc[-1], 3)
np.float64(1.053)

Part 25: Real Number of Trades vs Real Cost Assumption#

n_completed_trades = len(trades)
total_cost_drag_pct = round(transaction_costs.sum() * 100, 2)
n_completed_trades, total_cost_drag_pct
(10, np.float64(2.1))

Part 26: Saving the Real Trade Log#

trade_log = pd.DataFrame({'EntryDate': [aapl['Date'].iloc[e] for e, x in trades], 'ExitDate': [aapl['Date'].iloc[x] for e, x in trades], 'Return': trade_returns})
trade_log.to_csv('trade_log.csv', index=False)
reloaded_log = pd.read_csv('trade_log.csv')
reloaded_log.shape[0] == len(trades)
True

Part 27: Saving the Real Backtest Summary#

backtest_summary = pd.DataFrame({'Strategy': ['SMA Crossover', 'RSI Mean-Reversion', 'Buy and Hold'], 'EndingValue': [round(cumulative_net.iloc[-1], 3), round(cumulative_rsi_net.iloc[-1], 3), round(cumulative_buyhold.iloc[-1], 3)], 'MaxDrawdownPct': [round(max_drawdown(cumulative_net) * 100, 2), round(max_drawdown(cumulative_rsi_net) * 100, 2), round(max_drawdown(cumulative_buyhold) * 100, 2)]})
backtest_summary.to_csv('backtest_summary.csv', index=False)
reloaded_summary = pd.read_csv('backtest_summary.csv')
reloaded_summary.shape[0] == 3
True

Part 28: Real Sanity Check on Trade Pairing#

all(exit_i > entry_i for entry_i, exit_i in trades)
True

Part 29: Real Recap Print#

net_pct = round((cumulative_net.iloc[-1] - 1) * 100, 1)
bh_pct = round((cumulative_buyhold.iloc[-1] - 1) * 100, 1)
print(f'After real trading costs across {n_completed_trades} completed trades, the SMA crossover returned {net_pct}% versus {bh_pct}% for buy-and-hold, with a {round(win_rate*100,1)}% win rate.')
After real trading costs across 10 completed trades, the SMA crossover returned 14.6% versus 5.9% for buy-and-hold, with a 30.0% win rate.

Part 30: Real Longest Winning Streak#

import itertools
streak_signs = [1 if r > 0 else -1 for r in trade_returns]
win_streaks = [sum(1 for _ in group) for key, group in itertools.groupby(streak_signs) if key == 1]
longest_win_streak = max(win_streaks) if win_streaks else 0
longest_win_streak
2

Wrap-Up: What You Learned#

  • An honest backtest has to lag the real signal by one day and charge a real trading cost every time the position changes, otherwise it is quietly cheating.
  • Reconstructing individual real trades reveals a win rate and profit factor that a single aggregate return number would hide completely.
  • Comparing real Sharpe ratio and real maximum drawdown side by side, not just ending value, is what tells you whether a strategy is actually worth holding through.
  • A real walk-forward split is a basic, honest check for whether a fixed rule is genuinely robust or just fit to one lucky stretch of history.
  • Testing a real second, different style of strategy, mean-reversion instead of momentum, keeps you from mistaking one lucky rule for a genuinely reliable edge.
  • Next video: applying these real same tools to a real cryptocurrency market, testing whether these lessons carry over to a very differently behaved real asset.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.