Lesson 2 · Finance and Stock Market Analytics
Understanding Financial Markets and Instruments: Key Concepts and Analytics
In this lesson, we explore what financial markets are and how different instruments like stocks and ETFs work. Financial markets help companies raise money…
- CourseFinance and Stock Market Analytics
- Lesson2 of 16
- Video35 min
- FormatJupyter notebook · 24 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbUnderstanding Financial Markets and Instruments#
- In this lesson, we explore what financial markets are and how different instruments like stocks and ETFs work.
- Financial markets help companies raise money and let investors buy and sell investments like stocks.
- We will use real-world datasets to see how these markets operate and analyze popular financial instruments.
- By the end, you will understand core data structures, common pitfalls, and how to solve basic finance problems in Python.
- Let us get started!
import warnings
warnings.filterwarnings('ignore')
import pandas as pd
import numpy as np
import yfinance as yf
Core Data Concepts in Financial Markets#
- The data we use in finance describes transactions, prices, company fundamentals, and more.
- Stock data typically includes open, high, low, close prices, volumes, and sometimes dividends.
- Financial instruments are uniquely identified by their ticker symbols.
- Data is often structured either in a long format (one row per date/ticker) or a wide format (dates as rows, tickers as columns).
- Beginners often get confused by multi-index columns or forget to use real dates as indexes.
- Many mistakes are caused by mixing formats or manually editing dataframes instead of using built-in pandas methods.
# Beginner Example 1: Load a list of S&P 500 companies
url = 'https://raw.githubusercontent.com/datasets/s-and-p-500-companies/main/data/constituents.csv'
sp500_df = pd.read_csv(url)
print('Number of S&P 500 companies:', sp500_df.shape[0])
print(sp500_df.head(3))
# Beginner Example 2: Download daily close prices for five major tech stocks
tickers = ['AAPL','MSFT','GOOGL','AMZN','TSLA']
close_prices = yf.download(tickers, period='1y', auto_adjust=True, progress=False)['Close'].reset_index()
close_prices.columns.name = None
print(close_prices.shape)
print(close_prices.head(3))
# Beginner Example 3: Examine stock price movement for one ticker (AAPL)
aapl_only = close_prices[['Date', 'AAPL']]
print(aapl_only.head(5))
# Beginner Example 4: Load financial fundamentals for major tickers
fundamentals = []
fundamental_tickers = ['AAPL','MSFT','GOOGL','AMZN','TSLA','NVDA','META','JPM','JNJ','XOM','WMT','V']
for tk in fundamental_tickers:
info = yf.Ticker(tk).info
fundamentals.append({
'ticker': tk,
'sector': info.get('sector', 'Unknown'),
'market_cap': info.get('marketCap'),
'pe_ratio': info.get('trailingPE'),
'eps': info.get('trailingEps'),
'profit_margin': info.get('profitMargins'),
'debt_equity': info.get('debtToEquity')
})
fundamental_df = pd.DataFrame(fundamentals)
print(fundamental_df.head(3))
# Beginner Example 5: What is the average market capitalization in our list?
avg_market_cap = fundamental_df['market_cap'].mean()
print('Average market cap (USD):', round(avg_market_cap, 2))
# Beginner Example 6: Show companies in the 'Technology' sector only
tech_companies = fundamental_df[fundamental_df['sector'] == 'Technology']
print(tech_companies[['ticker','market_cap','pe_ratio']].reset_index(drop=True))
# Intermediate Example 1: Download one year of daily OHLCV multi-ticker data (long format)
tickers = ['AAPL','MSFT','GOOGL','AMZN','TSLA']
ohlcv = yf.download(tickers, period='1y', auto_adjust=True, progress=False)
ohlcv = ohlcv.stack(future_stack=True).rename_axis(['Date','Ticker']).reset_index()
ohlcv.columns.name = None
print(ohlcv.shape)
print(ohlcv.head(3))
# Intermediate Example 2: Pivot OHLCV data to see the closing price per ticker, per day
close_wide = ohlcv.pivot(index='Date', columns='Ticker', values='Close')
print(close_wide.head(5))
# Intermediate Example 3: Calculate daily returns for SPY ETF
spy_data = yf.download('SPY', period='2y', auto_adjust=True, progress=False)
spy_returns = spy_data[['Close','Volume']].reset_index()
spy_returns.columns = ['date','price','volume']
spy_returns['daily_return'] = (spy_returns['price'].pct_change() * 100).round(4)
spy_returns = spy_returns.dropna().reset_index(drop=True)
print(spy_returns.head(3))
# Intermediate Example 4: Build a basic simulated portfolio with real closing prices
tickers = ['AAPL','MSFT','GOOGL','AMZN','TSLA','NVDA','META','NFLX','JPM','JNJ']
sector_map = {'AAPL':'Tech','MSFT':'Tech','GOOGL':'Tech','AMZN':'Consumer',
'TSLA':'Auto','NVDA':'Tech','META':'Tech','NFLX':'Media',
'JPM':'Finance','JNJ':'Health'}
np.random.seed(42)
data = yf.download(tickers, period='1y', auto_adjust=True, progress=False)['Close']
rows = []
for tk in tickers:
buy = round(float(data[tk].iloc[0]), 2)
cur = round(float(data[tk].iloc[-1]), 2)
rows.append({'ticker': tk, 'shares': int(np.random.randint(5, 100)),
'buy_price': buy, 'cur_price': cur, 'sector': sector_map[tk]})
df_portfolio = pd.DataFrame(rows)
df_portfolio['gain_pct'] = np.round((df_portfolio['cur_price'] - df_portfolio['buy_price']) / df_portfolio['buy_price'] * 100, 2)
print(df_portfolio.head(3))
# Intermediate Example 5: Find the sector with the highest average percentage gain
sector_gains = df_portfolio.groupby('sector')['gain_pct'].mean().sort_values(ascending=False)
print('Average percentage gain by sector:')
print(sector_gains)
# Intermediate Example 6: Generating a simple trading signal for Apple using moving averages
aapl_signal_data = yf.download('AAPL', period='1y', auto_adjust=True, progress=False)
aapl_signal = aapl_signal_data[['Close','Volume']].reset_index()
aapl_signal.columns = ['date','close','volume']
aapl_signal['sma20'] = aapl_signal['close'].rolling(20).mean().round(2)
aapl_signal['sma50'] = aapl_signal['close'].rolling(50).mean().round(2)
aapl_signal['signal'] = np.where(aapl_signal['sma20'] > aapl_signal['sma50'], 1, -1)
print(aapl_signal.tail(3))
# Advanced Example 1: Calculate annualized volatility for each stock
close_only = close_prices.set_index('Date').dropna()
returns = close_only.pct_change().dropna()
annualized_volatility = returns.std() * np.sqrt(252)
print('Annualized volatility:')
print(annualized_volatility.round(4))
# Advanced Example 2: Simulate a random portfolio and compute total return
np.random.seed(42)
num_stocks = 5
initial_investment = 10000
weights = np.random.dirichlet(np.ones(num_stocks), 1).flatten()
start_prices = close_only.iloc[0]
end_prices = close_only.iloc[-1]
returns_vec = (end_prices.values - start_prices.values) / start_prices.values
portfolio_return = np.dot(returns_vec, weights)
final_value = initial_investment * (1 + portfolio_return)
print('Portfolio weights:', weights.round(2))
print('Total portfolio return (percent):', round(portfolio_return*100,2))
print('Final portfolio value (USD):', round(final_value, 2))
# Advanced Example 3: Detect and handle missing data in daily close prices
missing = close_prices.isnull().sum()
print('Missing values per column:')
print(missing)
# Error Handling Example 1: Try loading a dataset from an invalid URL
try:
bad_df = pd.read_csv('https://invalid.url/fakefile.csv')
except Exception as e:
print('Error loading file:', str(e))
# Error Handling Example 2: Handling missing fundamental data safely
def safe_pe(row):
try:
return float(row['pe_ratio']) if not pd.isnull(row['pe_ratio']) else None
except Exception:
return None
fundamental_df['safe_pe'] = fundamental_df.apply(safe_pe, axis=1)
print(fundamental_df[['ticker','safe_pe']])
# Error Handling Example 3: Dealing with non-trading days in time series
sample_dates = pd.date_range(start=close_prices['Date'].min(), end=close_prices['Date'].max(), freq='B')
merged = pd.DataFrame({'Date': sample_dates}).merge(close_prices, on='Date', how='left')
print('Rows after join:', merged.shape[0])
print('Missing days:', merged.isnull().any(axis=1).sum())
# Best Practice Example 1: Always check data types before analysis
print('Dtypes for close_prices:')
print(close_prices.dtypes)
# Best Practice Example 2: Always document data sources and sampling periods
print('Data for S&P 500 companies from:', url)
print('Stock price data period:', close_prices['Date'].min(), '-', close_prices['Date'].max())
# Best Practice Example 3: Detect and avoid forward-looking bias
def no_peeking_split(df, train_frac=0.7):
n = int(len(df) * train_frac)
train = df.iloc[:n]
test = df.iloc[n:]
return train, test
train, test = no_peeking_split(spy_returns)
print('Train set:', train.shape)
print('Test set:', test.shape)
# Best Practice Example 4: Always set random seed when simulating data
np.random.seed(42)
sim_shares = np.random.randint(10, 100, size=5)
print('Simulated shares:', sim_shares)
# End-to-End Example: From data load to portfolio gain calculation
portfolio_tickers = ['AAPL', 'MSFT', 'GOOGL', 'AMZN', 'TSLA']
portfolio_closes = yf.download(portfolio_tickers, period='1y', auto_adjust=True, progress=False)['Close'].reset_index()
portfolio_closes.columns.name = None
np.random.seed(42)
shares = np.random.randint(10, 100, size=len(portfolio_tickers))
buy_prices = portfolio_closes.iloc[0][1:]
end_prices = portfolio_closes.iloc[-1][1:]
total_buy = np.sum(buy_prices.values * shares)
total_end = np.sum(end_prices.values * shares)
gain_pct = round((total_end - total_buy) / total_buy * 100, 2)
print(f'Portfolio initial value: ${round(total_buy,2)}')
print(f'Portfolio final value: ${round(total_end,2)}')
print(f'Portfolio total gain: {gain_pct}%')
You Did It!#
- You have learned how to
- Load real finance data
- Understand financial instruments and prices
- Handle errors and missing data
- Run simple, intermediate, and advanced analytics
- Practice, experiment, and keep building!
- For more tips, check out YouTube for top finance analytics channels.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



