Mathew K Analytics

Python library centre

Mastering Polars for Efficient Big Data Analysis in Python

Polars is a fast DataFrame library for Python. It is used for data analysis and manipulation, similar to pandas. Polars is known for being very fast and…

Polarslogging

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Introduction to Polars#

  • Polars is a fast DataFrame library for Python.
  • It is used for data analysis and manipulation, similar to pandas.
  • Polars is known for being very fast and memory efficient.
  • You can work with tabular data like CSVs and big datasets easily.
  • Real world uses include big data analysis, quick data exploration, and ETL workflows.
  • Polars can handle much larger datasets than can fit into RAM.
import sys
try:
    import polars as pl
except ImportError:
    print('Polars is not installed. Installing now...')
    !{sys.executable} -m pip install polars
    import polars as pl

Core Objects in Polars#

  • Polars has two main objects: DataFrame and Series.
  • DataFrame: A table of data with columns and rows, like an Excel sheet.
  • Series: A single column of data from a DataFrame.
  • Most data operations are done on DataFrames.
df = pl.DataFrame({
    'name': ['Alice', 'Bob', 'Charlie'],
    'age': [25, 30, 35],
    'city': ['NY', 'LA', 'SF']
})
print(df)
shape: (3, 3)
┌─────────┬─────┬──────┐
│ name    ┆ age ┆ city │
│ ---     ┆ --- ┆ ---  │
│ str     ┆ i64 ┆ str  │
╞═════════╪═════╪══════╡
│ Alice   ┆ 25  ┆ NY   │
│ Bob     ┆ 30  ┆ LA   │
│ Charlie ┆ 35  ┆ SF   │
└─────────┴─────┴──────┘
print(df.columns)
print(df.shape)
['name', 'age', 'city']
(3, 3)
series = df['age']
print(series)
print(type(series))
shape: (3,)
Series: 'age' [i64]
[
	25
	30
	35
]
<class 'polars.series.series.Series'>
print(df.head(2))
shape: (2, 3)
┌───────┬─────┬──────┐
│ name  ┆ age ┆ city │
│ ---   ┆ --- ┆ ---  │
│ str   ┆ i64 ┆ str  │
╞═══════╪═════╪══════╡
│ Alice ┆ 25  ┆ NY   │
│ Bob   ┆ 30  ┆ LA   │
└───────┴─────┴──────┘
print(df.describe())
shape: (9, 4)
┌────────────┬─────────┬──────┬──────┐
│ statistic  ┆ name    ┆ age  ┆ city │
│ ---        ┆ ---     ┆ ---  ┆ ---  │
│ str        ┆ str     ┆ f64  ┆ str  │
╞════════════╪═════════╪══════╪══════╡
│ count      ┆ 3       ┆ 3.0  ┆ 3    │
│ null_count ┆ 0       ┆ 0.0  ┆ 0    │
│ mean       ┆ null    ┆ 30.0 ┆ null │
│ std        ┆ null    ┆ 5.0  ┆ null │
│ min        ┆ Alice   ┆ 25.0 ┆ LA   │
│ 25%        ┆ null    ┆ 30.0 ┆ null │
│ 50%        ┆ null    ┆ 30.0 ┆ null │
│ 75%        ┆ null    ┆ 35.0 ┆ null │
│ max        ┆ Charlie ┆ 35.0 ┆ SF   │
└────────────┴─────────┴──────┴──────┘
new_df = df.with_columns(
    (pl.col('age') + 1).alias('age_next_year')
)
print(new_df)
shape: (3, 4)
┌─────────┬─────┬──────┬───────────────┐
│ name    ┆ age ┆ city ┆ age_next_year │
│ ---     ┆ --- ┆ ---  ┆ ---           │
│ str     ┆ i64 ┆ str  ┆ i64           │
╞═════════╪═════╪══════╪═══════════════╡
│ Alice   ┆ 25  ┆ NY   ┆ 26            │
│ Bob     ┆ 30  ┆ LA   ┆ 31            │
│ Charlie ┆ 35  ┆ SF   ┆ 36            │
└─────────┴─────┴──────┴───────────────┘
filtered = new_df.filter(pl.col('age') > 28)
print(filtered)
shape: (2, 4)
┌─────────┬─────┬──────┬───────────────┐
│ name    ┆ age ┆ city ┆ age_next_year │
│ ---     ┆ --- ┆ ---  ┆ ---           │
│ str     ┆ i64 ┆ str  ┆ i64           │
╞═════════╪═════╪══════╪═══════════════╡
│ Bob     ┆ 30  ┆ LA   ┆ 31            │
│ Charlie ┆ 35  ┆ SF   ┆ 36            │
└─────────┴─────┴──────┴───────────────┘
sorted_df = new_df.sort('age', reverse=True)
print(sorted_df)
---------------------------------------------------------------------------
TypeError                                 Traceback (most recent call last)
Cell In[9], line 1
----> 1 sorted_df = new_df.sort('age', reverse=True)
      2 print(sorted_df)

TypeError: DataFrame.sort() got an unexpected keyword argument 'reverse'
csv_path = 'sample_data.csv'
new_df.write_csv(csv_path)
print('CSV saved to:', csv_path)
CSV saved to: sample_data.csv
read_df = pl.read_csv(csv_path)
print(read_df)
shape: (3, 4)
┌─────────┬─────┬──────┬───────────────┐
│ name    ┆ age ┆ city ┆ age_next_year │
│ ---     ┆ --- ┆ ---  ┆ ---           │
│ str     ┆ i64 ┆ str  ┆ i64           │
╞═════════╪═════╪══════╪═══════════════╡
│ Alice   ┆ 25  ┆ NY   ┆ 26            │
│ Bob     ┆ 30  ┆ LA   ┆ 31            │
│ Charlie ┆ 35  ┆ SF   ┆ 36            │
└─────────┴─────┴──────┴───────────────┘
grouped = new_df.groupby('city').agg([pl.col('age').mean()])
print(grouped)
---------------------------------------------------------------------------
AttributeError                            Traceback (most recent call last)
Cell In[12], line 1
----> 1 grouped = new_df.groupby('city').agg([pl.col('age').mean()])
      2 print(grouped)

AttributeError: 'DataFrame' object has no attribute 'groupby'
pivoted = new_df.pivot(values='age', index='city', columns='name')
print(pivoted)
shape: (3, 4)
┌──────┬───────┬──────┬─────────┐
│ city ┆ Alice ┆ Bob  ┆ Charlie │
│ ---  ┆ ---   ┆ ---  ┆ ---     │
│ str  ┆ i64   ┆ i64  ┆ i64     │
╞══════╪═══════╪══════╪═════════╡
│ NY   ┆ 25    ┆ null ┆ null    │
│ LA   ┆ null  ┆ 30   ┆ null    │
│ SF   ┆ null  ┆ null ┆ 35      │
└──────┴───────┴──────┴─────────┘
C:\Users\makmw\AppData\Local\Temp\ipykernel_9528\1178418076.py:1: DeprecationWarning: the argument `columns` for `DataFrame.pivot` is deprecated. It was renamed to `on` in version 1.0.0.
  pivoted = new_df.pivot(values='age', index='city', columns='name')
df2 = pl.DataFrame({'name': ['Dave'], 'age': [40], 'city': ['LA']})
combined = pl.concat([new_df, df2])
print(combined)
---------------------------------------------------------------------------
ShapeError                                Traceback (most recent call last)
Cell In[14], line 2
      1 df2 = pl.DataFrame({'name': ['Dave'], 'age': [40], 'city': ['LA']})
----> 2 combined = pl.concat([new_df, df2])
      3 print(combined)

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\polars\functions\eager.py:234, in concat(items, how, rechunk, parallel, strict)
    232 if isinstance(first, pl.DataFrame):
    233     if how == "vertical":
--> 234         out = wrap_df(plr.concat_df(elems))
    235     elif how == "vertical_relaxed":
    236         out = wrap_ldf(
    237             plr.concat_lf(
    238                 [df.lazy() for df in elems],
   (...)
    243             )
    244         ).collect(optimizations=QueryOptFlags._eager())

ShapeError: unable to append to a DataFrame of width 4 with a DataFrame of width 3
mask = new_df['age'] > 28
print(mask)
print(new_df[mask])
shape: (3,)
Series: 'age' [bool]
[
	false
	true
	true
]
---------------------------------------------------------------------------
TypeError                                 Traceback (most recent call last)
File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\polars\_utils\getitem.py:167, in get_df_item_by_key(df, key)
    166 try:
--> 167     return _select_rows(df, key)  # type: ignore[arg-type]
    168 except TypeError:

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\polars\_utils\getitem.py:319, in _select_rows(df, key)
    318 elif isinstance(key, pl.Series):
--> 319     indices = _convert_series_to_indices(key, df.height)
    320     return _select_rows_by_index(df, indices)

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\polars\_utils\getitem.py:360, in _convert_series_to_indices(s, size)
    359 if s.dtype == Boolean:
--> 360     _raise_on_boolean_mask()
    361 else:

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\polars\_utils\getitem.py:457, in _raise_on_boolean_mask()
    453 msg = (
    454     "selecting rows by passing a boolean mask to `__getitem__` is not supported"
    455     "\n\nHint: Use the `filter` method instead."
    456 )
--> 457 raise TypeError(msg)

TypeError: selecting rows by passing a boolean mask to `__getitem__` is not supported

Hint: Use the `filter` method instead.

During handling of the above exception, another exception occurred:

ValueError                                Traceback (most recent call last)
Cell In[15], line 3
      1 mask = new_df['age'] > 28
      2 print(mask)
----> 3 print(new_df[mask])

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\polars\dataframe\frame.py:1403, in DataFrame.__getitem__(self, key)
   1266 def __getitem__(
   1267     self,
   1268     key: (
   (...)
   1277     ),
   1278 ) -> DataFrame | Series | Any:
   1279     """
   1280     Get part of the DataFrame as a new DataFrame, Series, or scalar.
   1281 
   (...)
   1401     └─────┴─────┴─────┘
   1402     """
-> 1403     return get_df_item_by_key(self, key)

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\polars\_utils\getitem.py:169, in get_df_item_by_key(df, key)
    167     return _select_rows(df, key)  # type: ignore[arg-type]
    168 except TypeError:
--> 169     return _select_columns(df, key)

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\polars\_utils\getitem.py:231, in _select_columns(df, key)
    229     return _select_columns_by_index(df, key)
    230 elif dtype == Boolean:
--> 231     return _select_columns_by_mask(df, key)
    232 else:
    233     msg = f"cannot select columns using Series of type {dtype}"

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\polars\_utils\getitem.py:277, in _select_columns_by_mask(df, key)
    275 if len(key) != df.width:
    276     msg = f"expected {df.width} values when selecting columns by boolean mask, got {len(key)}"
--> 277     raise ValueError(msg)
    279 indices = (i for i, val in enumerate(key) if val)
    280 return _select_columns_by_index(df, indices)

ValueError: expected 4 values when selecting columns by boolean mask, got 3
new_df = new_df.with_row_count('row_id')
print(new_df)
shape: (3, 5)
┌────────┬─────────┬─────┬──────┬───────────────┐
│ row_id ┆ name    ┆ age ┆ city ┆ age_next_year │
│ ---    ┆ ---     ┆ --- ┆ ---  ┆ ---           │
│ u32    ┆ str     ┆ i64 ┆ str  ┆ i64           │
╞════════╪═════════╪═════╪══════╪═══════════════╡
│ 0      ┆ Alice   ┆ 25  ┆ NY   ┆ 26            │
│ 1      ┆ Bob     ┆ 30  ┆ LA   ┆ 31            │
│ 2      ┆ Charlie ┆ 35  ┆ SF   ┆ 36            │
└────────┴─────────┴─────┴──────┴───────────────┘
C:\Users\makmw\AppData\Local\Temp\ipykernel_9528\1206851682.py:1: DeprecationWarning: `DataFrame.with_row_count` is deprecated; use `with_row_index` instead. Note that the default column name has changed from 'row_nr' to 'index'.
  new_df = new_df.with_row_count('row_id')
try:
    non_existing = new_df['height']
except Exception as e:
    print('Error:', e)
Error: "height" not found
try:
    strange_df = pl.DataFrame({'a': [1, 2], 'b': [3]})
except Exception as e:
    print('Error:', e)
Error: could not create a new DataFrame: height of column 'b' (1) does not match height of column 'a' (2)
import logging
logging.basicConfig(level=logging.INFO)
logging.info('Polars logs will show up here if errors occur.')
INFO:root:Polars logs will show up here if errors occur.
# Best practice: Use lazy evaluation for big data
lazy_df = new_df.lazy()
result = lazy_df.filter(pl.col('age') > 28).collect()
print(result)
shape: (2, 5)
┌────────┬─────────┬─────┬──────┬───────────────┐
│ row_id ┆ name    ┆ age ┆ city ┆ age_next_year │
│ ---    ┆ ---     ┆ --- ┆ ---  ┆ ---           │
│ u32    ┆ str     ┆ i64 ┆ str  ┆ i64           │
╞════════╪═════════╪═════╪══════╪═══════════════╡
│ 1      ┆ Bob     ┆ 30  ┆ LA   ┆ 31            │
│ 2      ┆ Charlie ┆ 35  ┆ SF   ┆ 36            │
└────────┴─────────┴─────┴──────┴───────────────┘
# Best practice: Chain methods for clarity
clean_df = (
    new_df
    .filter(pl.col('age') > 25)
    .sort('name')
)
print(clean_df)
shape: (2, 5)
┌────────┬─────────┬─────┬──────┬───────────────┐
│ row_id ┆ name    ┆ age ┆ city ┆ age_next_year │
│ ---    ┆ ---     ┆ --- ┆ ---  ┆ ---           │
│ u32    ┆ str     ┆ i64 ┆ str  ┆ i64           │
╞════════╪═════════╪═════╪══════╪═══════════════╡
│ 1      ┆ Bob     ┆ 30  ┆ LA   ┆ 31            │
│ 2      ┆ Charlie ┆ 35  ┆ SF   ┆ 36            │
└────────┴─────────┴─────┴──────┴───────────────┘
# Mini-project: Analyze exam scores
mpath = 'exam_scores.csv'
with open(mpath, 'w') as f:
    f.write('student,score,subject\n')
    f.write('Anna,88,Math\n')
    f.write('Ben,91,Math\n')
    f.write('Cara,79,Math\n')
    f.write('Anna,82,Science\n')
    f.write('Ben,78,Science\n')
    f.write('Cara,90,Science\n')
scores_df = pl.read_csv('exam_scores.csv')
print(scores_df)
shape: (6, 3)
┌─────────┬───────┬─────────┐
│ student ┆ score ┆ subject │
│ ---     ┆ ---   ┆ ---     │
│ str     ┆ i64   ┆ str     │
╞═════════╪═══════╪═════════╡
│ Anna    ┆ 88    ┆ Math    │
│ Ben     ┆ 91    ┆ Math    │
│ Cara    ┆ 79    ┆ Math    │
│ Anna    ┆ 82    ┆ Science │
│ Ben     ┆ 78    ┆ Science │
│ Cara    ┆ 90    ┆ Science │
└─────────┴───────┴─────────┘
avg_scores = scores_df.groupby('student').agg(
    pl.col('score').mean().alias('avg_score')
)
print(avg_scores)
---------------------------------------------------------------------------
AttributeError                            Traceback (most recent call last)
Cell In[24], line 1
----> 1 avg_scores = scores_df.groupby('student').agg(
      2     pl.col('score').mean().alias('avg_score')
      3 )
      4 print(avg_scores)

AttributeError: 'DataFrame' object has no attribute 'groupby'
top_scores = scores_df.sort('score', reverse=True).head(3)
print(top_scores)
---------------------------------------------------------------------------
TypeError                                 Traceback (most recent call last)
Cell In[25], line 1
----> 1 top_scores = scores_df.sort('score', reverse=True).head(3)
      2 print(top_scores)

TypeError: DataFrame.sort() got an unexpected keyword argument 'reverse'

Thank you for learning Polars with us!#

  • Subscribe for more Python data tutorials.
  • Like and share if this lesson helped you.
from pynput.mouse import Controller, Button
import time
import os
import pyautogui
mouse = Controller()

# def perform_action(action):
#     import os, time, subprocess

#     t = action["type"]
#     target = action.get("target")
#     params = action.get("params", {})

#     try:
#         if t == "open_file":
#             time.sleep(params.get("delay", 0.3))
#             if target and os.path.exists(target):
#                 os.startfile(target)
#                 time.sleep(params.get("delay", 0.3))
#                 mouse.scroll(0, -2)
#                 # pyautogui.press("pg dn"); time.sleep(0.1)
#                 # pyautogui.press("pg dn"); time.sleep(0.1)
#                 # pyautogui.press("pg dn"); time.sleep(0.1)
#                 time.sleep(params.get("view_time", 2.0))
#                 pyautogui.press("esc"); time.sleep(1)
#                 # pyautogui.hotkey("alt","tab")
#         else:
#             raise ValueError(f"Unsupported action type: {t}")
#     except:
#         pass
# perform_action("open_file")
# # Move mouse
# mouse.position = (400, 200)
# time.sleep(1)

# # Left click
# mouse.click(Button.left, 1)

# # Right click
# mouse.click(Button.right, 1)

# # Press and hold
# mouse.press(Button.left)
# time.sleep(1)
# mouse.release(Button.left)

os.startfile("example_style.docx"); time.sleep(5)
pyautogui.hotkey("ctrl","pagedown"); time.sleep(0.2)
time.sleep(5)

# Scroll
# mouse.scroll(0, -29)
# mouse.scroll(0, -29)
# mouse.scroll(0, -29)
# mouse.scroll(0, -29)

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.