Python library centre
Comprehensive PyArrow Training for Efficient Data Processing in Python
PyArrow is a Python library for working with Apache Arrow data. It allows fast, zero-copy data exchange between different computing tools. PyArrow is useful…
- CoursePython library centre
- Video21 min
- FormatJupyter notebook · 24 code cells
- Data2 datasets
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- my_table.parquet736 B
- filtered.parquet1.1 KB
📓 Full notebook
Download .ipynbIntroduction to PyArrow#
- PyArrow is a Python library for working with Apache Arrow data.
- It allows fast, zero-copy data exchange between different computing tools.
- PyArrow is useful for working with big tabular datasets efficiently.
- Real world uses: accelerating data loading, reading Parquet files, sharing data between Python & other languages, powering machine learning, and more.
- PyArrow makes big data workflows faster and more memory efficient.
import warnings; warnings.filterwarnings("ignore")
import sys
# If you are on Windows and need to install PyArrow, run:
# pip install pyarrow
import pyarrow as pa
import pyarrow.feather as feather
import pyarrow.parquet as pq
What are Arrow Arrays and Tables?#
- PyArrow has two main data types: Arrays and Tables.
- An Arrow Array is like a column of data with a specific type.
- An Arrow Table is similar to a table in SQL or a DataFrame in pandas.
- Arrow Tables are made from multiple Arrow Arrays.
- Arrow data is stored in a fast, columnar format.
# Let us create a simple Arrow Array
arr = pa.array([1, 2, 3, 4, 5])
print(arr)
# You can make Arrow Arrays with different types.
arr_float = pa.array([1.5, 2.0, 3.2])
print(arr_float)
arr_str = pa.array(['a', 'b', 'c'])
print(arr_str)
# Now let us build an Arrow Table from columns
table = pa.table({'numbers': arr_float, 'letters': arr_str})
print(table)
Beginner: Accessing Data#
- You can select columns and rows in Arrow Tables.
- PyArrow allows you to easily inspect and slice tables.
# Access a column by name
print(table['numbers'])
# Access rows using slice notation
print(table.slice(0, 2))
# Convert Arrow Table to pandas DataFrame
import pandas as pd
df = table.to_pandas()
print(df)
Intermediate: Reading and Writing Arrow Files#
- PyArrow can save data to disk and load it back quickly.
- You can save Tables as Arrow IPC (Feather) files or as Parquet files.
# Write the table to a Feather file
pa.feather.write_feather(table, 'my_table.feather')
print('Saved my_table.feather')
# Read the table back from a Feather file
table2 = pa.feather.read_feather('my_table.feather')
print(table2)
# Save table as Parquet file
pq.write_table(table, 'my_table.parquet')
print('Saved my_table.parquet')
# Load table from Parquet file
parquet_table = pq.read_table('my_table.parquet')
print(parquet_table)
Intermediate: Arrow Schema and Data Types#
- Arrow Arrays and Tables have strong, explicit data types.
- Table schemas help ensure data consistency when sharing or saving data.
# Check the schema of a table
print(table.schema)
# Create an Arrow Array with nulls/None values
arr_with_nulls = pa.array([1, None, 3])
print(arr_with_nulls)
Advanced: Chunked Arrays and Large Datasets#
- Arrow Tables can store data in chunks, useful for big data.
- You can combine arrays or append new data efficiently.
- ChunkedArray lets you process data without loading everything at once.
# Creating a ChunkedArray from multiple Arrow Arrays
chunked_arr = pa.chunked_array([[1, 2], [3, 4, 5]])
print(chunked_arr)
print('Number of chunks:', chunked_arr.num_chunks)
# Append new rows to an Arrow Table using concat_tables
t1 = pa.table({'a': [1, 2]})
t2 = pa.table({'a': [3, 4]})
combined = pa.concat_tables([t1, t2])
print(combined)
Advanced: Zero-copy Data Sharing#
- Arrow format allows multiple languages and libraries to share data without copying.
- You can pass Arrow Arrays from NumPy or pandas efficiently (using zero-copy).
# Convert NumPy array to Arrow Array (zero-copy)
import numpy as np
np_arr = np.array([10, 20, 30])
arrow_arr = pa.array(np_arr)
print(arrow_arr)
# Convert Arrow Array back to NumPy
back_to_np = arrow_arr.to_numpy()
print(back_to_np)
Error Handling Example#
- PyArrow will raise errors if you use unsupported data.
- Let us see what happens with a bad data type.
# Trying to create an Arrow Array from a nested list (should fail)
try:
arr_bad = pa.array([[1, 2], [3, 4]])
except Exception as e:
print('Error:', e)
# Safely check if your data can be converted to Arrow Array
def safe_arrow_array(data):
try:
arr = pa.array(data)
print('Success:', arr)
except Exception as e:
print('Conversion failed:', e)
safe_arrow_array([1, 2, 3])
safe_arrow_array([{'x': 1}, {'x': 2}])
Best Practices with PyArrow#
- Prefer Arrow for tabular, column data in big-data workflow.
- Use Arrow Arrays and Tables for efficient serialization.
- Use zero-copy interfaces for integration with NumPy and pandas.
- Convert data back to pandas or NumPy only when necessary.
- Always specify schema when saving data for reliability.
# Specify schema when building Tables
myschema = pa.schema([('a', pa.int64()), ('b', pa.string())])
table_with_schema = pa.table([pa.array([7, 8]), pa.array(['x', 'y'])], schema=myschema)
print(table_with_schema)
# Always close file handles after use (for larger files)
with pa.OSFile('my_table.feather', 'rb') as source:
file_table = pa.feather.read_table(source)
print('Read using OSFile:', file_table.shape)
Mini-Project: Analyze a Small Dataset#
- Let us simulate loading, filtering, and saving a dataset using PyArrow.
- We will:
- Create sample data
- Filter rows
- Save and load it in Parquet format
# 1. Create sample data as Arrow Table
ids = pa.array([1, 2, 3, 4, 5])
values = pa.array([5.2, 6.3, 7.1, 8.4, 9.9])
scores = pa.array([50, 60, 70, 80, 90])
mini_table = pa.table({'id': ids, 'value': values, 'score': scores})
print('Sample data:')
print(mini_table)
# 2. Filter rows: Select rows with score greater than 60
mask = pa.compute.greater(mini_table['score'], 60)
filtered = mini_table.filter(mask)
print('Rows with score > 60:')
print(filtered)
# 3. Save and reload the filtered Table as Parquet
pq.write_table(filtered, 'filtered.parquet')
filtered_loaded = pq.read_table('filtered.parquet')
print('Reloaded filtered table:')
print(filtered_loaded)
YouTube: Like and Subscribe!#
- If you enjoyed this PyArrow lesson, please like and subscribe for more Python tutorials.
Want More?#
- Comment below if you want a deep dive on PyArrow or data engineering topics!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



