Mathew K Analytics

Lesson 8 · Finance and Stock Market Analytics

Numerical Analysis with NumPy for Financial Data: Training and Applications

In this lesson we solve real-world finance and stock market problems using NumPy. Numerical analysis allows you to process, analyze, and interpret large…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Numerical Analysis with NumPy for Financial Data#

  • In this lesson we solve real-world finance and stock market problems using NumPy.
  • Numerical analysis allows you to process, analyze, and interpret large quantities of financial data efficiently.
  • The ability to perform quick, accurate calculations is critical for portfolio analysis, risk management, and trading.
  • By the end, you will understand and apply NumPy techniques to manipulate, summarize, and extract insights from real financial datasets.
import warnings
warnings.filterwarnings('ignore')
import pandas as pd
import numpy as np
import yfinance as yf

Understanding Our Core Datasets#

  • We will use real stock data for five large tech companies: AAPL, MSFT, GOOGL, AMZN, and TSLA.
  • Data includes daily Open, High, Low, Close, and Volume values for each ticker, in long format.
  • Each row corresponds to a single stock on a single date.
  • Common mistakes: misinterpreting multi-ticker data shapes or flattening columns incorrectly.
  • Always check whether your data is 'wide' (columns per ticker) or 'long' (row per date-ticker pair).
tickers = ['AAPL','MSFT','GOOGL','AMZN','TSLA']
ohlcv = yf.download(tickers, period='1y', auto_adjust=True, progress=False)
ohlcv = ohlcv.stack(future_stack=True).rename_axis(['Date','Ticker']).reset_index()
ohlcv.columns.name = None
print(ohlcv.shape)
print(ohlcv.head(3))
(1255, 7)
        Date Ticker       Close        High         Low        Open    Volume
0 2025-07-14   AAPL  207.795624  210.076583  206.719890  209.100445  38840100
1 2025-07-14   AMZN  225.690002  226.660004  224.240005  225.070007  35702600
2 2025-07-14  GOOGL  181.043518  183.147516  179.168861  180.495080  32536600

Beginner Example 1: Extracting Price Arrays with NumPy#

  • Let us select Apple (AAPL) prices and convert to a NumPy array.
  • NumPy arrays allow fast computation for large numeric datasets.
  • This is the foundation for more advanced financial analytics.
aapl_prices = ohlcv[ohlcv['Ticker'] == 'AAPL']['Close'].values
print(type(aapl_prices))
print('Number of days:', aapl_prices.shape[0])
print('First 5 closing prices:', aapl_prices[:5])
<class 'numpy.ndarray'>
Number of days: 251
First 5 closing prices: [207.79562378 208.28370667 209.32954407 209.19009399 210.34550476]

Beginner Example 2: Calculating Average and Standard Deviation#

  • NumPy's mean() and std() allow us to compute key statistics instantly.
  • These numbers tell us about stock performance and risk.
aapl_mean = np.mean(aapl_prices)
aapl_std = np.std(aapl_prices)
print('AAPL 1-year average price:', round(aapl_mean,2))
print('AAPL 1-year price volatility (std):', round(aapl_std,2))
AAPL 1-year average price: 262.71
AAPL 1-year price volatility (std): 25.84

Beginner Example 3: Computing Daily Returns with NumPy#

  • One of the most important tasks in finance is calculating returns for each period.
  • Returns show us percentage change from one day to the next.
aapl_ret = np.diff(aapl_prices) / aapl_prices[:-1] * 100
print('AAPL: Daily return array shape:', aapl_ret.shape)
print('First 5 daily returns (%):', np.round(aapl_ret[:5],2))
AAPL: Daily return array shape: (250,)
First 5 daily returns (%): [ 0.23  0.5  -0.07  0.55  0.62]

Intermediate Example 1: Working with Multiple Tickers at Once#

  • NumPy and pandas let us process entire portfolioshundreds of stocks at once.
  • We will pivot the data to get a matrix of closing prices by date and ticker.
price_matrix = ohlcv.pivot(index='Date', columns='Ticker', values='Close')
print('Shape:', price_matrix.shape)
print(price_matrix.head(3))
Shape: (251, 5)
Ticker            AAPL        AMZN       GOOGL        MSFT        TSLA
Date                                                                  
2025-07-14  207.795624  225.690002  181.043518  499.033936  316.899994
2025-07-15  208.283707  226.350006  181.482269  501.811768  310.779999
2025-07-16  209.329544  223.190002  182.449524  501.613373  321.670013

Intermediate Example 2: Vectorized Calculations on Multiple Stocks#

  • NumPy makes it easy to compute statistics across the entire matrix, not just one stock.
  • Let us calculate average price and volatility for every ticker at once.
mean_prices = price_matrix.mean(axis=0)
volatility = price_matrix.std(axis=0)
print('Tickers:', list(mean_prices.index))
print('Mean prices:', np.round(mean_prices.values,2))
print('Volatility (std):', np.round(volatility.values,2))
Tickers: ['AAPL', 'AMZN', 'GOOGL', 'MSFT', 'TSLA']
Mean prices: [262.71 232.29 297.02 453.96 403.01]
Volatility (std): [25.89 17.27 57.95 52.92 43.26]

Intermediate Example 3: Correlation Matrix of Stock Returns#

  • Correlation shows how two stocks move togethera key to portfolio risk.
  • We will use NumPy and pandas to compute daily returns and the full correlation matrix.
return_matrix = price_matrix.pct_change().iloc[1:] * 100
corr_matrix = return_matrix.corr()
print('Pairwise return correlations:')
print(np.round(corr_matrix,2))
Pairwise return correlations:
Ticker  AAPL  AMZN  GOOGL  MSFT  TSLA
Ticker                               
AAPL    1.00  0.32   0.28  0.19  0.24
AMZN    0.32  1.00   0.42  0.33  0.33
GOOGL   0.28  0.42   1.00  0.10  0.37
MSFT    0.19  0.33   0.10  1.00  0.17
TSLA    0.24  0.33   0.37  0.17  1.00

Advanced Example 1: Rolling Volatility (Risk) Calculation#

  • Financial risk is often measured as rolling standard deviation of returns.
  • NumPy and pandas make it easy to calculate this over sliding windows.
aapl_returns = price_matrix['AAPL'].pct_change().iloc[1:] * 100
roll_vol = aapl_returns.rolling(21).std()
print('Last 5 rolling 21-day volatilities (%):')
print(np.round(roll_vol.tail(),2))
Last 5 rolling 21-day volatilities (%):
Date
2026-07-07    2.38
2026-07-08    2.37
2026-07-09    2.33
2026-07-10    2.16
2026-07-13    2.16
Name: AAPL, dtype: float64

Advanced Example 2: Vectorized Portfolio Return Calculation#

  • NumPy lets us compute portfolio performance from raw positions with minimal code.
  • We will use a simulated portfolio holding random share counts in real stocks.
np.random.seed(42)
tickers = ['AAPL','MSFT','GOOGL','AMZN','TSLA','NVDA','META','NFLX','JPM','JNJ']
sector_map = {'AAPL':'Tech','MSFT':'Tech','GOOGL':'Tech','AMZN':'Consumer',
              'TSLA':'Auto','NVDA':'Tech','META':'Tech','NFLX':'Media',
              'JPM':'Finance','JNJ':'Health'}
data = yf.download(tickers, period='1y', auto_adjust=True, progress=False)['Close']
rows = []
for tk in tickers:
    buy = round(float(data[tk].iloc[0]), 2)
    cur = round(float(data[tk].iloc[-1]), 2)
    rows.append({'ticker': tk, 'shares': int(np.random.randint(5, 100)),
                 'buy_price': buy, 'cur_price': cur, 'sector': sector_map[tk]})
df_port = pd.DataFrame(rows)
df_port['gain_pct'] = np.round((df_port['cur_price'] - df_port['buy_price']) / df_port['buy_price'] * 100, 2)
print(df_port[['ticker','shares','buy_price','cur_price','gain_pct']])
  ticker  shares  buy_price  cur_price  gain_pct
0   AAPL      56     207.80     317.16     52.63
1   MSFT      97     499.03     393.06    -21.24
2  GOOGL      19     181.04     355.28     96.24
3   AMZN      76     225.69     249.07     10.36
4   TSLA      65     316.90     392.92     23.99
5   NVDA      25     163.85     203.90     24.44
6   META      87     718.56     660.73     -8.05
7   NFLX      91     126.19      74.41    -41.03
8    JPM      79     283.28     334.77     18.18
9    JNJ      79     153.00     258.14     68.72

Advanced Example 3: Masked Arrays for Filtering Outliers#

  • Sometimes we want to exclude extreme price moves or data glitches in our analysis.
  • NumPy's masked arrays help us filter data before taking statistics.
from numpy import ma
masked_returns = ma.masked_outside(aapl_ret, -7, 7)
print('Masked outlier count:', masked_returns.mask.sum())
print('Mean daily return (ex-outliers):', round(masked_returns.mean(),3))
Masked outlier count: 0
Mean daily return (ex-outliers): 0.18

Error Handling Example: Safe Division in NumPy#

  • Division by zero or invalid values may happen in timeseries (for example, price stuck at zero).
  • Use NumPy's nan_to_num to handle such cases gracefully.
bad_prices = np.array([150, 152, 0, 155, 0])
returns = np.diff(bad_prices) / bad_prices[:-1]
safe_returns = np.nan_to_num(returns, nan=0, posinf=0, neginf=0)
print('Safe returns:', safe_returns)
Safe returns: [ 0.01333333 -1.          0.         -1.        ]

Debugging Example: Checking Data Shapes Early#

  • One common bug: getting arrays of shapes that do not match.
  • Use shape prints and assertions to catch mistakes before they crash your model.
arr1 = np.ones((50,))
arr2 = np.random.randn(49,)
print('arr1 shape:', arr1.shape,'arr2 shape:', arr2.shape)
try:
    arr1 + arr2
except ValueError as e:
    print('Error:', e)
arr1 shape: (50,) arr2 shape: (49,)
Error: operands could not be broadcast together with shapes (50,) (49,) 

Best Practice: Always Seed Random Number Generators#

  • Setting a seed with np.random.seed ensures reproducible results for simulations.
  • This helps others reproduce your exact numbers and debug easily.
np.random.seed(42)
rand_nums = np.random.randn(3)
print('Random array (should be same every time):', rand_nums)
Random array (should be same every time): [ 0.49671415 -0.1382643   0.64768854]

Common Pattern: Vectorized Computation for All Assets#

  • Rewriting explicit Python loops as NumPy vector operations speeds up your code drastically.
  • This makes even portfolio-level analytics efficient and clean.
daily_rets = price_matrix.pct_change().iloc[1:] * 100
mean_daily_rets = daily_rets.mean()
print('Mean daily return by stock:')
print(mean_daily_rets.round(3))
Mean daily return by stock:
Ticker
AAPL     0.180
AMZN     0.059
GOOGL    0.288
MSFT    -0.081
TSLA     0.126
dtype: float64

Mini End-to-End Problem: Portfolio Volatility in Real Time#

  • Let us combine everything: simulate a portfolio, compute its daily value, and plot volatility.
  • You will see how NumPy and pandas work together for true financial analytics from start to finish.
import matplotlib.pyplot as plt
np.random.seed(42)
weights = np.random.dirichlet(np.ones(len(price_matrix.columns)),size=1)[0]
norm_prices = price_matrix / price_matrix.iloc[0]
portfolio_val = (norm_prices * weights).sum(axis=1)
port_daily_ret = portfolio_val.pct_change().iloc[1:] * 100
port_vol = port_daily_ret.rolling(21).std()
plt.figure(figsize=(10,4))
plt.plot(port_vol.index, port_vol, label='21-day Volatility (%)')
plt.ylabel('Portfolio Volatility (%)')
plt.xlabel('Date')
plt.title('Rolling Volatility of Simulated Portfolio')
plt.legend()
plt.tight_layout()
plt.show()
No description has been provided for this image
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.