Lesson 8 · Finance and Stock Market Analytics
Numerical Analysis with NumPy for Financial Data: Training and Applications
In this lesson we solve real-world finance and stock market problems using NumPy. Numerical analysis allows you to process, analyze, and interpret large…
- CourseFinance and Stock Market Analytics
- Lesson8 of 16
- Video24 min
- FormatJupyter notebook · 17 code cells
What you'll learn
- Understanding Our Core Datasets
- Beginner Example 1: Extracting Price Arrays with NumPy
- Beginner Example 2: Calculating Average and Standard Deviation
- Beginner Example 3: Computing Daily Returns with NumPy
- Intermediate Example 1: Working with Multiple Tickers at Once
- Intermediate Example 2: Vectorized Calculations on Multiple Stocks
- Intermediate Example 3: Correlation Matrix of Stock Returns
- Advanced Example 1: Rolling Volatility (Risk) Calculation
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbNumerical Analysis with NumPy for Financial Data#
- In this lesson we solve real-world finance and stock market problems using NumPy.
- Numerical analysis allows you to process, analyze, and interpret large quantities of financial data efficiently.
- The ability to perform quick, accurate calculations is critical for portfolio analysis, risk management, and trading.
- By the end, you will understand and apply NumPy techniques to manipulate, summarize, and extract insights from real financial datasets.
import warnings
warnings.filterwarnings('ignore')
import pandas as pd
import numpy as np
import yfinance as yf
Understanding Our Core Datasets#
- We will use real stock data for five large tech companies: AAPL, MSFT, GOOGL, AMZN, and TSLA.
- Data includes daily Open, High, Low, Close, and Volume values for each ticker, in long format.
- Each row corresponds to a single stock on a single date.
- Common mistakes: misinterpreting multi-ticker data shapes or flattening columns incorrectly.
- Always check whether your data is 'wide' (columns per ticker) or 'long' (row per date-ticker pair).
tickers = ['AAPL','MSFT','GOOGL','AMZN','TSLA']
ohlcv = yf.download(tickers, period='1y', auto_adjust=True, progress=False)
ohlcv = ohlcv.stack(future_stack=True).rename_axis(['Date','Ticker']).reset_index()
ohlcv.columns.name = None
print(ohlcv.shape)
print(ohlcv.head(3))
Beginner Example 1: Extracting Price Arrays with NumPy#
- Let us select Apple (AAPL) prices and convert to a NumPy array.
- NumPy arrays allow fast computation for large numeric datasets.
- This is the foundation for more advanced financial analytics.
aapl_prices = ohlcv[ohlcv['Ticker'] == 'AAPL']['Close'].values
print(type(aapl_prices))
print('Number of days:', aapl_prices.shape[0])
print('First 5 closing prices:', aapl_prices[:5])
Beginner Example 2: Calculating Average and Standard Deviation#
- NumPy's mean() and std() allow us to compute key statistics instantly.
- These numbers tell us about stock performance and risk.
aapl_mean = np.mean(aapl_prices)
aapl_std = np.std(aapl_prices)
print('AAPL 1-year average price:', round(aapl_mean,2))
print('AAPL 1-year price volatility (std):', round(aapl_std,2))
Beginner Example 3: Computing Daily Returns with NumPy#
- One of the most important tasks in finance is calculating returns for each period.
- Returns show us percentage change from one day to the next.
aapl_ret = np.diff(aapl_prices) / aapl_prices[:-1] * 100
print('AAPL: Daily return array shape:', aapl_ret.shape)
print('First 5 daily returns (%):', np.round(aapl_ret[:5],2))
Intermediate Example 1: Working with Multiple Tickers at Once#
- NumPy and pandas let us process entire portfolioshundreds of stocks at once.
- We will pivot the data to get a matrix of closing prices by date and ticker.
price_matrix = ohlcv.pivot(index='Date', columns='Ticker', values='Close')
print('Shape:', price_matrix.shape)
print(price_matrix.head(3))
Intermediate Example 2: Vectorized Calculations on Multiple Stocks#
- NumPy makes it easy to compute statistics across the entire matrix, not just one stock.
- Let us calculate average price and volatility for every ticker at once.
mean_prices = price_matrix.mean(axis=0)
volatility = price_matrix.std(axis=0)
print('Tickers:', list(mean_prices.index))
print('Mean prices:', np.round(mean_prices.values,2))
print('Volatility (std):', np.round(volatility.values,2))
Intermediate Example 3: Correlation Matrix of Stock Returns#
- Correlation shows how two stocks move togethera key to portfolio risk.
- We will use NumPy and pandas to compute daily returns and the full correlation matrix.
return_matrix = price_matrix.pct_change().iloc[1:] * 100
corr_matrix = return_matrix.corr()
print('Pairwise return correlations:')
print(np.round(corr_matrix,2))
Advanced Example 1: Rolling Volatility (Risk) Calculation#
- Financial risk is often measured as rolling standard deviation of returns.
- NumPy and pandas make it easy to calculate this over sliding windows.
aapl_returns = price_matrix['AAPL'].pct_change().iloc[1:] * 100
roll_vol = aapl_returns.rolling(21).std()
print('Last 5 rolling 21-day volatilities (%):')
print(np.round(roll_vol.tail(),2))
Advanced Example 2: Vectorized Portfolio Return Calculation#
- NumPy lets us compute portfolio performance from raw positions with minimal code.
- We will use a simulated portfolio holding random share counts in real stocks.
np.random.seed(42)
tickers = ['AAPL','MSFT','GOOGL','AMZN','TSLA','NVDA','META','NFLX','JPM','JNJ']
sector_map = {'AAPL':'Tech','MSFT':'Tech','GOOGL':'Tech','AMZN':'Consumer',
'TSLA':'Auto','NVDA':'Tech','META':'Tech','NFLX':'Media',
'JPM':'Finance','JNJ':'Health'}
data = yf.download(tickers, period='1y', auto_adjust=True, progress=False)['Close']
rows = []
for tk in tickers:
buy = round(float(data[tk].iloc[0]), 2)
cur = round(float(data[tk].iloc[-1]), 2)
rows.append({'ticker': tk, 'shares': int(np.random.randint(5, 100)),
'buy_price': buy, 'cur_price': cur, 'sector': sector_map[tk]})
df_port = pd.DataFrame(rows)
df_port['gain_pct'] = np.round((df_port['cur_price'] - df_port['buy_price']) / df_port['buy_price'] * 100, 2)
print(df_port[['ticker','shares','buy_price','cur_price','gain_pct']])
Advanced Example 3: Masked Arrays for Filtering Outliers#
- Sometimes we want to exclude extreme price moves or data glitches in our analysis.
- NumPy's masked arrays help us filter data before taking statistics.
from numpy import ma
masked_returns = ma.masked_outside(aapl_ret, -7, 7)
print('Masked outlier count:', masked_returns.mask.sum())
print('Mean daily return (ex-outliers):', round(masked_returns.mean(),3))
Error Handling Example: Safe Division in NumPy#
- Division by zero or invalid values may happen in timeseries (for example, price stuck at zero).
- Use NumPy's nan_to_num to handle such cases gracefully.
bad_prices = np.array([150, 152, 0, 155, 0])
returns = np.diff(bad_prices) / bad_prices[:-1]
safe_returns = np.nan_to_num(returns, nan=0, posinf=0, neginf=0)
print('Safe returns:', safe_returns)
Debugging Example: Checking Data Shapes Early#
- One common bug: getting arrays of shapes that do not match.
- Use shape prints and assertions to catch mistakes before they crash your model.
arr1 = np.ones((50,))
arr2 = np.random.randn(49,)
print('arr1 shape:', arr1.shape,'arr2 shape:', arr2.shape)
try:
arr1 + arr2
except ValueError as e:
print('Error:', e)
Best Practice: Always Seed Random Number Generators#
- Setting a seed with np.random.seed ensures reproducible results for simulations.
- This helps others reproduce your exact numbers and debug easily.
np.random.seed(42)
rand_nums = np.random.randn(3)
print('Random array (should be same every time):', rand_nums)
Common Pattern: Vectorized Computation for All Assets#
- Rewriting explicit Python loops as NumPy vector operations speeds up your code drastically.
- This makes even portfolio-level analytics efficient and clean.
daily_rets = price_matrix.pct_change().iloc[1:] * 100
mean_daily_rets = daily_rets.mean()
print('Mean daily return by stock:')
print(mean_daily_rets.round(3))
Mini End-to-End Problem: Portfolio Volatility in Real Time#
- Let us combine everything: simulate a portfolio, compute its daily value, and plot volatility.
- You will see how NumPy and pandas work together for true financial analytics from start to finish.
import matplotlib.pyplot as plt
np.random.seed(42)
weights = np.random.dirichlet(np.ones(len(price_matrix.columns)),size=1)[0]
norm_prices = price_matrix / price_matrix.iloc[0]
portfolio_val = (norm_prices * weights).sum(axis=1)
port_daily_ret = portfolio_val.pct_change().iloc[1:] * 100
port_vol = port_daily_ret.rolling(21).std()
plt.figure(figsize=(10,4))
plt.plot(port_vol.index, port_vol, label='21-day Volatility (%)')
plt.ylabel('Portfolio Volatility (%)')
plt.xlabel('Date')
plt.title('Rolling Volatility of Simulated Portfolio')
plt.legend()
plt.tight_layout()
plt.show()
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



