Python library centre
Mastering Polars for Efficient Big Data Analysis in Python
Polars is a fast DataFrame library for Python. It is used for data analysis and manipulation, similar to pandas. Polars is known for being very fast and…
- CoursePython library centre
- Video20 min
- FormatJupyter notebook · 26 code cells
- Data2 datasets
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- sample_data.csv73 B
- exam_scores.csv114 B
📓 Full notebook
Download .ipynbIntroduction to Polars#
- Polars is a fast DataFrame library for Python.
- It is used for data analysis and manipulation, similar to pandas.
- Polars is known for being very fast and memory efficient.
- You can work with tabular data like CSVs and big datasets easily.
- Real world uses include big data analysis, quick data exploration, and ETL workflows.
- Polars can handle much larger datasets than can fit into RAM.
import sys
try:
import polars as pl
except ImportError:
print('Polars is not installed. Installing now...')
!{sys.executable} -m pip install polars
import polars as pl
Core Objects in Polars#
- Polars has two main objects: DataFrame and Series.
- DataFrame: A table of data with columns and rows, like an Excel sheet.
- Series: A single column of data from a DataFrame.
- Most data operations are done on DataFrames.
df = pl.DataFrame({
'name': ['Alice', 'Bob', 'Charlie'],
'age': [25, 30, 35],
'city': ['NY', 'LA', 'SF']
})
print(df)
print(df.columns)
print(df.shape)
series = df['age']
print(series)
print(type(series))
print(df.head(2))
print(df.describe())
new_df = df.with_columns(
(pl.col('age') + 1).alias('age_next_year')
)
print(new_df)
filtered = new_df.filter(pl.col('age') > 28)
print(filtered)
sorted_df = new_df.sort('age', reverse=True)
print(sorted_df)
csv_path = 'sample_data.csv'
new_df.write_csv(csv_path)
print('CSV saved to:', csv_path)
read_df = pl.read_csv(csv_path)
print(read_df)
grouped = new_df.groupby('city').agg([pl.col('age').mean()])
print(grouped)
pivoted = new_df.pivot(values='age', index='city', columns='name')
print(pivoted)
df2 = pl.DataFrame({'name': ['Dave'], 'age': [40], 'city': ['LA']})
combined = pl.concat([new_df, df2])
print(combined)
mask = new_df['age'] > 28
print(mask)
print(new_df[mask])
new_df = new_df.with_row_count('row_id')
print(new_df)
try:
non_existing = new_df['height']
except Exception as e:
print('Error:', e)
try:
strange_df = pl.DataFrame({'a': [1, 2], 'b': [3]})
except Exception as e:
print('Error:', e)
import logging
logging.basicConfig(level=logging.INFO)
logging.info('Polars logs will show up here if errors occur.')
# Best practice: Use lazy evaluation for big data
lazy_df = new_df.lazy()
result = lazy_df.filter(pl.col('age') > 28).collect()
print(result)
# Best practice: Chain methods for clarity
clean_df = (
new_df
.filter(pl.col('age') > 25)
.sort('name')
)
print(clean_df)
# Mini-project: Analyze exam scores
mpath = 'exam_scores.csv'
with open(mpath, 'w') as f:
f.write('student,score,subject\n')
f.write('Anna,88,Math\n')
f.write('Ben,91,Math\n')
f.write('Cara,79,Math\n')
f.write('Anna,82,Science\n')
f.write('Ben,78,Science\n')
f.write('Cara,90,Science\n')
scores_df = pl.read_csv('exam_scores.csv')
print(scores_df)
avg_scores = scores_df.groupby('student').agg(
pl.col('score').mean().alias('avg_score')
)
print(avg_scores)
top_scores = scores_df.sort('score', reverse=True).head(3)
print(top_scores)
Thank you for learning Polars with us!#
- Subscribe for more Python data tutorials.
- Like and share if this lesson helped you.
from pynput.mouse import Controller, Button
import time
import os
import pyautogui
mouse = Controller()
# def perform_action(action):
# import os, time, subprocess
# t = action["type"]
# target = action.get("target")
# params = action.get("params", {})
# try:
# if t == "open_file":
# time.sleep(params.get("delay", 0.3))
# if target and os.path.exists(target):
# os.startfile(target)
# time.sleep(params.get("delay", 0.3))
# mouse.scroll(0, -2)
# # pyautogui.press("pg dn"); time.sleep(0.1)
# # pyautogui.press("pg dn"); time.sleep(0.1)
# # pyautogui.press("pg dn"); time.sleep(0.1)
# time.sleep(params.get("view_time", 2.0))
# pyautogui.press("esc"); time.sleep(1)
# # pyautogui.hotkey("alt","tab")
# else:
# raise ValueError(f"Unsupported action type: {t}")
# except:
# pass
# perform_action("open_file")
# # Move mouse
# mouse.position = (400, 200)
# time.sleep(1)
# # Left click
# mouse.click(Button.left, 1)
# # Right click
# mouse.click(Button.right, 1)
# # Press and hold
# mouse.press(Button.left)
# time.sleep(1)
# mouse.release(Button.left)
os.startfile("example_style.docx"); time.sleep(5)
pyautogui.hotkey("ctrl","pagedown"); time.sleep(0.2)
time.sleep(5)
# Scroll
# mouse.scroll(0, -29)
# mouse.scroll(0, -29)
# mouse.scroll(0, -29)
# mouse.scroll(0, -29)
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



