Mathew K Analytics

Lesson 8 · Pandas Projects

Sports Analytics with Python Pandas: Player & Team Stats

A complete, standalone tutorial: build a synthetic season of basketball box scores, then rank, track, and compare players with pandas. No prior pandas…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Pandas for Sports Analytics: Analyzing a Basketball Season#

  • A complete, standalone tutorial: build a synthetic season of basketball box scores, then rank, track, and compare players with pandas.
  • No prior pandas experience needed. Let's jump straight in.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • If pandas isn't installed yet, open a terminal in VS Code and run: pip install pandas

Part 1: Building the Dataset#

import pandas as pd
import numpy as np
print(pd.__version__)
2.3.0

Generating a Synthetic Season#

rng = np.random.default_rng(seed=22)

teams = {
    'Comets': ['Rylan', 'Tho', 'Bea', 'Omar', 'Nia'],
    'Falcons': ['Dara', 'Mo', 'Ines', 'Kwame', 'Suri']
}
games_per_player = 20
print(teams)
{'Comets': ['Rylan', 'Tho', 'Bea', 'Omar', 'Nia'], 'Falcons': ['Dara', 'Mo', 'Ines', 'Kwame', 'Suri']}
rows = []
for team, players in teams.items():
    for player in players:
        for game_num in range(1, games_per_player + 1):
            minutes = round(float(rng.uniform(15, 38)), 1)
            points = int(rng.poisson(14))
            rebounds = int(rng.poisson(5))
            assists = int(rng.poisson(4))
            rows.append([team, player, game_num, minutes, points, rebounds, assists])

games = pd.DataFrame(rows, columns=['team', 'player', 'game_num', 'minutes', 'points', 'rebounds', 'assists'])
print(games.shape)
(200, 7)
games.to_csv('season_box_scores.csv', index=False)
games = pd.read_csv('season_box_scores.csv')
games.head()
team player game_num minutes points rebounds assists
0 Comets Rylan 1 23.4 10 6 2
1 Comets Rylan 2 22.6 13 4 7
2 Comets Rylan 3 24.5 13 3 8
3 Comets Rylan 4 36.9 21 7 5
4 Comets Rylan 5 26.9 13 9 4

Part 2: First Look at the Data#

print(games.shape)
games.info()
(200, 7)
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 200 entries, 0 to 199
Data columns (total 7 columns):
 #   Column    Non-Null Count  Dtype  
---  ------    --------------  -----  
 0   team      200 non-null    object 
 1   player    200 non-null    object 
 2   game_num  200 non-null    int64  
 3   minutes   200 non-null    float64
 4   points    200 non-null    int64  
 5   rebounds  200 non-null    int64  
 6   assists   200 non-null    int64  
dtypes: float64(1), int64(4), object(2)
memory usage: 11.1+ KB
games[['minutes', 'points', 'rebounds', 'assists']].describe()
minutes points rebounds assists
count 200.000000 200.000000 200.000000 200.000000
mean 26.581500 14.335000 4.820000 4.145000
std 6.522139 3.953597 2.243382 2.067692
min 15.000000 5.000000 0.000000 0.000000
25% 21.350000 12.000000 3.000000 3.000000
50% 26.700000 14.000000 5.000000 4.000000
75% 31.225000 17.000000 6.000000 5.000000
max 38.000000 23.000000 12.000000 11.000000
games['team'].value_counts()
team
Comets     100
Falcons    100
Name: count, dtype: int64

Part 3: Season Totals and Ranking#

season_totals = games.groupby('player').agg(
    team=('team', 'first'),
    games_played=('game_num', 'count'),
    total_points=('points', 'sum'),
    total_rebounds=('rebounds', 'sum'),
    total_assists=('assists', 'sum')
).sort_values('total_points', ascending=False)
season_totals
team games_played total_points total_rebounds total_assists
player
Nia Comets 20 317 100 83
Dara Falcons 20 301 102 76
Mo Falcons 20 296 109 90
Omar Comets 20 294 83 86
Ines Falcons 20 293 94 93
Rylan Comets 20 290 95 72
Suri Falcons 20 280 84 88
Bea Comets 20 276 99 74
Kwame Falcons 20 267 105 82
Tho Comets 20 253 93 85

Ranking with rank()#

season_totals['points_rank'] = season_totals['total_points'].rank(ascending=False).astype(int)
season_totals[['team', 'total_points', 'points_rank']]
team total_points points_rank
player
Nia Comets 317 1
Dara Falcons 301 2
Mo Falcons 296 3
Omar Comets 294 4
Ines Falcons 293 5
Rylan Comets 290 6
Suri Falcons 280 7
Bea Comets 276 8
Kwame Falcons 267 9
Tho Comets 253 10

Per-Game Averages#

per_game_avg = games.groupby('player')[['points', 'rebounds', 'assists']].mean().round(1)
per_game_avg = per_game_avg.sort_values('points', ascending=False)
per_game_avg
points rebounds assists
player
Nia 15.8 5.0 4.2
Dara 15.0 5.1 3.8
Mo 14.8 5.4 4.5
Omar 14.7 4.2 4.3
Ines 14.6 4.7 4.6
Rylan 14.5 4.8 3.6
Suri 14.0 4.2 4.4
Bea 13.8 5.0 3.7
Kwame 13.4 5.2 4.1
Tho 12.6 4.6 4.2

Part 4: A Simple Efficiency Metric#

games['efficiency'] = (games['points'] + games['rebounds'] + games['assists']) / games['minutes']
games[['player', 'game_num', 'minutes', 'efficiency']].head()
player game_num minutes efficiency
0 Rylan 1 23.4 0.769231
1 Rylan 2 22.6 1.061947
2 Rylan 3 24.5 0.979592
3 Rylan 4 36.9 0.894309
4 Rylan 5 26.9 0.966543
avg_efficiency = games.groupby('player')['efficiency'].mean().sort_values(ascending=False)
avg_efficiency.round(3)
player
Nia      1.052
Mo       1.042
Dara     0.986
Omar     0.978
Ines     0.943
Bea      0.901
Suri     0.894
Kwame    0.887
Rylan    0.859
Tho      0.837
Name: efficiency, dtype: float64

Part 5: Cumulative and Rolling Stats#

games = games.sort_values(['player', 'game_num'])
games['cumulative_points'] = games.groupby('player')['points'].cumsum()
games[games['player'] == 'Rylan'][['game_num', 'points', 'cumulative_points']].tail()
game_num points cumulative_points
15 16 11 229
16 17 12 241
17 18 15 256
18 19 20 276
19 20 14 290
games['rolling_5g_avg'] = games.groupby('player')['points'].transform(lambda s: s.rolling(window=5).mean())
games[games['player'] == 'Dara'][['game_num', 'points', 'rolling_5g_avg']].tail(8)
game_num points rolling_5g_avg
112 13 17 14.2
113 14 15 14.8
114 15 17 16.8
115 16 18 17.2
116 17 20 17.4
117 18 13 16.6
118 19 16 16.8
119 20 12 15.8

Part 6: Comparing Teams#

team_comparison = games.groupby('team').agg(
    avg_points=('points', 'mean'),
    avg_rebounds=('rebounds', 'mean'),
    avg_assists=('assists', 'mean'),
    avg_efficiency=('efficiency', 'mean')
).round(2)
team_comparison
avg_points avg_rebounds avg_assists avg_efficiency
team
Comets 14.30 4.70 4.00 0.93
Falcons 14.37 4.94 4.29 0.95

Pivot Table: Player Scoring by Game Range#

games['season_third'] = pd.cut(games['game_num'], bins=[0, 7, 14, 20], labels=['Games 1-7', 'Games 8-14', 'Games 15-20'])
scoring_trend = pd.pivot_table(games, values='points', index='player', columns='season_third', aggfunc='mean', observed=True).round(1)
scoring_trend
season_third Games 1-7 Games 8-14 Games 15-20
player
Bea 14.0 14.0 13.3
Dara 14.4 14.9 16.0
Ines 14.9 12.9 16.5
Kwame 12.9 14.1 13.0
Mo 13.4 15.4 15.7
Nia 16.0 16.6 14.8
Omar 14.9 16.9 12.0
Rylan 14.7 15.4 13.2
Suri 15.0 13.9 13.0
Tho 11.4 13.9 12.7

Part 7: Visualizing the Results#

import matplotlib
matplotlib.use('Agg')
import matplotlib.pyplot as plt

fig, ax = plt.subplots(figsize=(8, 5))
per_game_avg['points'].plot(kind='bar', ax=ax, color='darkorange')
ax.set_title('Points Per Game by Player')
ax.set_ylabel('Points Per Game')
plt.tight_layout()
plt.savefig('points_per_game.png', dpi=150)
plt.close(fig)
print('Saved points_per_game.png')
Saved points_per_game.png
rylan_games = games[games['player'] == 'Rylan']
fig, ax = plt.subplots(figsize=(9, 5))
ax.plot(rylan_games['game_num'], rylan_games['points'], color='gray', alpha=0.5, marker='o', label='Points per game')
ax.plot(rylan_games['game_num'], rylan_games['rolling_5g_avg'], color='crimson', linewidth=2, label='5-game rolling average')
ax.set_title("Rylan's Scoring Trend This Season")
ax.set_xlabel('Game Number')
ax.set_ylabel('Points')
ax.legend()
plt.tight_layout()
plt.savefig('player_scoring_trend.png', dpi=150)
plt.close(fig)
print('Saved player_scoring_trend.png')
Saved player_scoring_trend.png

Wrap-Up: What You Learned#

  • Generating a realistic synthetic season of box scores with NumPy's Poisson distribution, then saving and reloading with to_csv and read_csv.
  • First-look exploration: shape, info, describe, and value_counts.
  • Season totals with a multi-statistic agg, plus rank() for a real leaderboard.
  • Building your own derived metric, like our per-minute efficiency score.
  • Cumulative totals with cumsum, and per-player rolling averages with groupby plus transform.
  • Team-level comparisons, and a pivot_table binned by season third to spot trends over time.
  • You went from raw synthetic box scores to a full season analysis with rankings, trends, and charts. If you want the next dataset in this series to land in your feed automatically, subscribing is the move see you in the next one.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.