Mathew K Analytics

Python library centre

Comprehensive PyArrow Training for Efficient Data Processing in Python

PyArrow is a Python library for working with Apache Arrow data. It allows fast, zero-copy data exchange between different computing tools. PyArrow is useful…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Introduction to PyArrow#

  • PyArrow is a Python library for working with Apache Arrow data.
  • It allows fast, zero-copy data exchange between different computing tools.
  • PyArrow is useful for working with big tabular datasets efficiently.
  • Real world uses: accelerating data loading, reading Parquet files, sharing data between Python & other languages, powering machine learning, and more.
  • PyArrow makes big data workflows faster and more memory efficient.
import warnings; warnings.filterwarnings("ignore")
import sys
 
# If you are on Windows and need to install PyArrow, run:
# pip install pyarrow
import pyarrow as pa
import pyarrow.feather as feather
import pyarrow.parquet as pq

What are Arrow Arrays and Tables?#

  • PyArrow has two main data types: Arrays and Tables.
  • An Arrow Array is like a column of data with a specific type.
  • An Arrow Table is similar to a table in SQL or a DataFrame in pandas.
  • Arrow Tables are made from multiple Arrow Arrays.
  • Arrow data is stored in a fast, columnar format.
# Let us create a simple Arrow Array
arr = pa.array([1, 2, 3, 4, 5])
print(arr)
[
  1,
  2,
  3,
  4,
  5
]
# You can make Arrow Arrays with different types.
arr_float = pa.array([1.5, 2.0, 3.2])
print(arr_float)

arr_str = pa.array(['a', 'b', 'c'])
print(arr_str)
[
  1.5,
  2,
  3.2
]
[
  "a",
  "b",
  "c"
]
# Now let us build an Arrow Table from columns
table = pa.table({'numbers': arr_float, 'letters': arr_str})
print(table)
pyarrow.Table
numbers: double
letters: string
----
numbers: [[1.5,2,3.2]]
letters: [["a","b","c"]]

Beginner: Accessing Data#

  • You can select columns and rows in Arrow Tables.
  • PyArrow allows you to easily inspect and slice tables.
# Access a column by name
print(table['numbers'])
[
  [
    1.5,
    2,
    3.2
  ]
]
# Access rows using slice notation
print(table.slice(0, 2))
pyarrow.Table
numbers: double
letters: string
----
numbers: [[1.5,2]]
letters: [["a","b"]]
# Convert Arrow Table to pandas DataFrame
import pandas as pd
df = table.to_pandas()
print(df)
   numbers letters
0      1.5       a
1      2.0       b
2      3.2       c

Intermediate: Reading and Writing Arrow Files#

  • PyArrow can save data to disk and load it back quickly.
  • You can save Tables as Arrow IPC (Feather) files or as Parquet files.
# Write the table to a Feather file
pa.feather.write_feather(table, 'my_table.feather')
print('Saved my_table.feather')
Saved my_table.feather
# Read the table back from a Feather file
table2 = pa.feather.read_feather('my_table.feather')
print(table2)
   numbers letters
0      1.5       a
1      2.0       b
2      3.2       c
# Save table as Parquet file
pq.write_table(table, 'my_table.parquet')
print('Saved my_table.parquet')
Saved my_table.parquet
# Load table from Parquet file
parquet_table = pq.read_table('my_table.parquet')
print(parquet_table)
pyarrow.Table
numbers: double
letters: string
----
numbers: [[1.5,2,3.2]]
letters: [["a","b","c"]]

Intermediate: Arrow Schema and Data Types#

  • Arrow Arrays and Tables have strong, explicit data types.
  • Table schemas help ensure data consistency when sharing or saving data.
# Check the schema of a table
print(table.schema)
numbers: double
letters: string
# Create an Arrow Array with nulls/None values
arr_with_nulls = pa.array([1, None, 3])
print(arr_with_nulls)
[
  1,
  null,
  3
]

Advanced: Chunked Arrays and Large Datasets#

  • Arrow Tables can store data in chunks, useful for big data.
  • You can combine arrays or append new data efficiently.
  • ChunkedArray lets you process data without loading everything at once.
# Creating a ChunkedArray from multiple Arrow Arrays
chunked_arr = pa.chunked_array([[1, 2], [3, 4, 5]])
print(chunked_arr)
print('Number of chunks:', chunked_arr.num_chunks)
[
  [
    1,
    2
  ],
  [
    3,
    4,
    5
  ]
]
Number of chunks: 2
# Append new rows to an Arrow Table using concat_tables
t1 = pa.table({'a': [1, 2]})
t2 = pa.table({'a': [3, 4]})
combined = pa.concat_tables([t1, t2])
print(combined)
pyarrow.Table
a: int64
----
a: [[1,2],[3,4]]

Advanced: Zero-copy Data Sharing#

  • Arrow format allows multiple languages and libraries to share data without copying.
  • You can pass Arrow Arrays from NumPy or pandas efficiently (using zero-copy).
# Convert NumPy array to Arrow Array (zero-copy)
import numpy as np
np_arr = np.array([10, 20, 30])
arrow_arr = pa.array(np_arr)
print(arrow_arr)
[
  10,
  20,
  30
]
# Convert Arrow Array back to NumPy
back_to_np = arrow_arr.to_numpy()
print(back_to_np)
[10 20 30]

Error Handling Example#

  • PyArrow will raise errors if you use unsupported data.
  • Let us see what happens with a bad data type.
# Trying to create an Arrow Array from a nested list (should fail)
try:
    arr_bad = pa.array([[1, 2], [3, 4]])
except Exception as e:
    print('Error:', e)
# Safely check if your data can be converted to Arrow Array
def safe_arrow_array(data):
    try:
        arr = pa.array(data)
        print('Success:', arr)
    except Exception as e:
        print('Conversion failed:', e)

safe_arrow_array([1, 2, 3])
safe_arrow_array([{'x': 1}, {'x': 2}])
Success: [
  1,
  2,
  3
]
Success: -- is_valid: all not null
-- child 0 type: int64
  [
    1,
    2
  ]

Best Practices with PyArrow#

  • Prefer Arrow for tabular, column data in big-data workflow.
  • Use Arrow Arrays and Tables for efficient serialization.
  • Use zero-copy interfaces for integration with NumPy and pandas.
  • Convert data back to pandas or NumPy only when necessary.
  • Always specify schema when saving data for reliability.
# Specify schema when building Tables
myschema = pa.schema([('a', pa.int64()), ('b', pa.string())])
table_with_schema = pa.table([pa.array([7, 8]), pa.array(['x', 'y'])], schema=myschema)
print(table_with_schema)
pyarrow.Table
a: int64
b: string
----
a: [[7,8]]
b: [["x","y"]]
# Always close file handles after use (for larger files)
with pa.OSFile('my_table.feather', 'rb') as source:
    file_table = pa.feather.read_table(source)
    print('Read using OSFile:', file_table.shape)
Read using OSFile: (3, 2)

Mini-Project: Analyze a Small Dataset#

  • Let us simulate loading, filtering, and saving a dataset using PyArrow.
  • We will:
    1. Create sample data
    1. Filter rows
    1. Save and load it in Parquet format
# 1. Create sample data as Arrow Table
ids = pa.array([1, 2, 3, 4, 5])
values = pa.array([5.2, 6.3, 7.1, 8.4, 9.9])
scores = pa.array([50, 60, 70, 80, 90])
mini_table = pa.table({'id': ids, 'value': values, 'score': scores})
print('Sample data:')
print(mini_table)
Sample data:
pyarrow.Table
id: int64
value: double
score: int64
----
id: [[1,2,3,4,5]]
value: [[5.2,6.3,7.1,8.4,9.9]]
score: [[50,60,70,80,90]]
# 2. Filter rows: Select rows with score greater than 60
mask = pa.compute.greater(mini_table['score'], 60)
filtered = mini_table.filter(mask)
print('Rows with score > 60:')
print(filtered)
Rows with score > 60:
pyarrow.Table
id: int64
value: double
score: int64
----
id: [[3,4,5]]
value: [[7.1,8.4,9.9]]
score: [[70,80,90]]
# 3. Save and reload the filtered Table as Parquet
pq.write_table(filtered, 'filtered.parquet')
filtered_loaded = pq.read_table('filtered.parquet')
print('Reloaded filtered table:')
print(filtered_loaded)
Reloaded filtered table:
pyarrow.Table
id: int64
value: double
score: int64
----
id: [[3,4,5]]
value: [[7.1,8.4,9.9]]
score: [[70,80,90]]

YouTube: Like and Subscribe!#

  • If you enjoyed this PyArrow lesson, please like and subscribe for more Python tutorials.

Want More?#

  • Comment below if you want a deep dive on PyArrow or data engineering topics!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.