Lesson 1 · Python for Banking and Finance
Python Basics for Banking
Learn how to use Python to solve real-world problems in banking. Work with synthetic banking transaction data to understand customer activity. Discover…
- CoursePython for Banking and Finance
- Lesson1 of 24
- Video22 min
- FormatJupyter notebook · 24 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbPython Basics for Banking#
- Learn how to use Python to solve real-world problems in banking.
- Work with synthetic banking transaction data to understand customer activity.
- Discover fundamental and intermediate skills, from loading and exploring data to best practices.
- Build practical skills for banking analysis you can use at work or for projects.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
Understanding Banking Data in Python#
- Banking data includes transactions, accounts, and customer info.
- Transaction tables usually show each movement of money: customer, amount, type, when, and channel.
- Account tables record details about accounts: type, open date, and ID.
- Beginners often forget to check data types or miss relationships between tables.
- Getting familiar with structure helps avoid errors and confusion.
np.random.seed(42) # Set seed for reproducibility
n_transactions = 1000
n_customers = 200
df = pd.DataFrame({
'transaction_id': range(1, n_transactions + 1),
'customer_id': np.random.choice([f'CUST_{i:04d}' for i in range(1, n_customers + 1)], n_transactions),
'amount': np.round(np.random.normal(150, 60, n_transactions), 2),
'transaction_type': np.random.choice(['Debit', 'Credit'], n_transactions),
'channel': np.random.choice(['ATM', 'Online', 'Branch', 'POS'], n_transactions),
'date': pd.date_range(start='2024-01-01', periods=n_transactions, freq='h')
})
print(df.shape)
print(df.head(3))
customer_ids = [f'CUST_{i:04d}' for i in range(1, 201)]
customers = pd.DataFrame({
'customer_id': customer_ids,
'segment': ['Retail'] * 150 + ['Business'] * 50,
'region': ['Metro'] * 100 + ['Regional'] * 100
})
print(customers.shape)
print(customers.head(3))
np.random.seed(42) # Set seed for reproducibility
accounts = pd.DataFrame({
'account_id': [f'ACC_{i:05d}' for i in range(1, 201)],
'customer_id': customer_ids,
'account_type': np.random.choice(['Savings', 'Cheque', 'Credit'], size=200),
'open_date': pd.date_range(start='2015-01-01', periods=200, freq='30D')
})
print(accounts.shape)
print(accounts.head(3))
print(df.columns)
print(df['transaction_type'].value_counts())
avg_amount = df['amount'].mean()
print('Average transaction amount:', avg_amount)
customer_txn_counts = df['customer_id'].value_counts().head(5)
print('Top 5 customers by transaction count:')
print(customer_txn_counts)
channel_stats = df.groupby('channel')['amount'].agg(['mean', 'max', 'min'])
print('Amount stats by channel:')
print(channel_stats)
monthly_totals = df.resample('M', on='date')['amount'].sum()
print('Total transaction amount by month:')
print(monthly_totals)
result = pd.merge(df, customers, on='customer_id', how='left')
print('Merged transaction and customer data:')
print(result.head(3))
pivot = df.pivot_table('amount', index='transaction_type', columns='channel', aggfunc='mean')
print('Pivot table: mean amount by type and channel')
print(pivot)
largest_txn = df.loc[df['amount'].idxmax()]
print('Largest transaction:')
print(largest_txn)
debit_percent = df['transaction_type'].value_counts(normalize=True)['Debit'] * 100
print(f'Percentage of Debit transactions: {debit_percent:.2f}%')
print(df.describe())
summary = df.groupby(['customer_id', 'transaction_type'])['amount'].sum().unstack(fill_value=0)
summary['net_flow'] = summary['Credit'] - summary['Debit']
print('Net flow (Credit - Debit) per customer:')
print(summary.head())
debit_channels = df[df['transaction_type'] == 'Debit']['channel'].value_counts()
print('Most common channels for Debits:')
print(debit_channels)
try:
no_column = df['balance']
except KeyError as e:
print('Error:', e)
try:
negative_txn = df[df['amount'] < 0]
if negative_txn.empty:
print('All transaction amounts are positive.')
else:
print('Some transaction amounts are negative:')
print(negative_txn.head())
except Exception as ex:
print('An error occurred:', ex)
try:
idx = df[df['date'] == '2025-01-01'].index[0]
print('Row found at index:', idx)
except IndexError:
print('No transactions found for that date.')
assert df['amount'].isnull().sum() == 0, 'No missing amounts allowed!'
# Always review datatypes early
print(df.dtypes)
# Clean data: ensure all channel categories are expected
expected_channels = {'ATM', 'Online', 'Branch', 'POS'}
actual_channels = set(df['channel'].unique())
if not actual_channels.issubset(expected_channels):
print('Unexpected channel values found:', actual_channels - expected_channels)
else:
print('All channel values are as expected.')
# End-to-End: What is the most profitable business region?
merged = pd.merge(df, customers, on='customer_id')
region_profit = merged[merged['segment'] == 'Business'].groupby('region')['amount'].sum()
best_region = region_profit.idxmax()
print('Business segment profit by region:')
print(region_profit)
print('Most profitable business region is:', best_region)
Great job: What next?#
- You learned how to use Python for banking data from setup to advanced analytics.
- Practice designing your own banking DataFrame and write end-to-end solutions.
- Watch our YouTube Python banking playlist for deeper dives and real case studies.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



