Mathew K Analytics

Lesson 15 · Python standard library deep dive

Python pickle & shelve Explained: Save Objects to Disk | Standard Library #15

Video fifteen of the twenty-five-part series: pickle and shelve, for serializing genuinely arbitrary Python objects, not just the JSON-safe ones. Basic…

⬇ Download notebookOpen in Colab ↗
pickle

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Python Standard Library Deep-Dive, Video 15: pickle and shelve#

  • Video fifteen of the twenty-five-part series: pickle and shelve, for serializing genuinely arbitrary Python objects, not just the JSON-safe ones.
  • Basic serialization, files, custom objects, limitations, security, and a persistent dict-like store.
  • Let's get into it.

Part 1: What pickle and shelve Offer#

import pickle
import shelve
data = {'a': 1, 'b': [1, 2, 3]}
print(pickle.dumps(data)[:20])
b'\x80\x04\x95\x19\x00\x00\x00\x00\x00\x00\x00}\x94(\x8c\x01a\x94K\x01'

Part 2: pickle.dumps() and pickle.loads()#

original = {'name': 'Ana', 'scores': [85, 90, 78], 'active': True}
serialized = pickle.dumps(original)
print(type(serialized))
restored = pickle.loads(serialized)
print(restored)
print(restored == original)
print(restored is original)
<class 'bytes'>
{'name': 'Ana', 'scores': [85, 90, 78], 'active': True}
True
False
a_tuple = (1, 'two', 3.0)
a_set = {1, 2, 3}
print(pickle.loads(pickle.dumps(a_tuple)))
print(pickle.loads(pickle.dumps(a_set)))
nested = {'users': [{'name': 'Ana'}, {'name': 'Sam'}], 'count': 2}
print(pickle.loads(pickle.dumps(nested)) == nested)
(1, 'two', 3.0)
{1, 2, 3}
True

Part 3: pickle.dump() and pickle.load() - Files#

data = {'model': 'demo', 'version': 3, 'weights': [0.1, 0.2, 0.3]}
with open('demo_data.pkl', 'wb') as f:
    pickle.dump(data, f)
with open('demo_data.pkl', 'rb') as f:
    loaded = pickle.load(f)
print(loaded)
print(loaded == data)
{'model': 'demo', 'version': 3, 'weights': [0.1, 0.2, 0.3]}
True

Part 4: Pickling Custom Objects#

class Point:
    def __init__(self, x, y):
        self.x = x
        self.y = y
    def __repr__(self):
        return f'Point({self.x}, {self.y})'
p = Point(3, 4)
serialized = pickle.dumps(p)
restored = pickle.loads(serialized)
print(restored)
print(type(restored))
print(restored.x, restored.y)
Point(3, 4)
<class '__main__.Point'>
3 4

Part 5: What Can't Be Pickled#

square = lambda n: n ** 2
try:
    pickle.dumps(square)
except (pickle.PicklingError, AttributeError, TypeError) as e:
    print(f'Caught: {type(e).__name__}')
def real_square(n):
    return n ** 2
print(pickle.loads(pickle.dumps(real_square))(5))
Caught: PicklingError
25

Part 6: Pickle Protocol Versions#

print(pickle.HIGHEST_PROTOCOL)
print(pickle.DEFAULT_PROTOCOL)
data = list(range(100))
old_style = pickle.dumps(data, protocol=0)
new_style = pickle.dumps(data, protocol=pickle.HIGHEST_PROTOCOL)
print(len(old_style) > len(new_style))
5
4
True

Part 7: Security Warning: Never Unpickle Untrusted Data#

trusted_data = {'safe': True}
safe_bytes = pickle.dumps(trusted_data)
print(pickle.loads(safe_bytes))
print('Only ever unpickle data your own application genuinely produced and controls')
{'safe': True}
Only ever unpickle data your own application genuinely produced and controls

Part 8: shelve Basics: a Persistent Dict#

with shelve.open('demo_shelf') as db:
    db['user_count'] = 150
    db['config'] = {'theme': 'dark', 'retries': 3}
    db['tags'] = ['python', 'stdlib']
with shelve.open('demo_shelf') as db:
    print(db['user_count'])
    print(db['config'])
    print(list(db.keys()))
150
{'theme': 'dark', 'retries': 3}
['user_count', 'config', 'tags']

Part 9: shelve Writeback and Best Practices#

with shelve.open('demo_shelf2') as db:
    db['items'] = [1, 2, 3]
    db['items'].append(4)
with shelve.open('demo_shelf2') as db:
    print(db['items'])
with shelve.open('demo_shelf3', writeback=True) as db:
    db['items'] = [1, 2, 3]
    db['items'].append(4)
with shelve.open('demo_shelf3') as db:
    print(db['items'])
[1, 2, 3]
[1, 2, 3, 4]

Part 10: Common Patterns#

import os
def expensive_computation(n):
    return sum(i ** 2 for i in range(n))
def cached_compute(n, cache_file='demo_cache.pkl'):
    if os.path.exists(cache_file):
        with open(cache_file, 'rb') as f:
            cache = pickle.load(f)
    else:
        cache = {}
    if n not in cache:
        cache[n] = expensive_computation(n)
        with open(cache_file, 'wb') as f:
            pickle.dump(cache, f)
    return cache[n]
print(cached_compute(1000))
print(cached_compute(1000))
332833500
332833500

Wrap-Up: What You Learned#

  • pickle serializes almost any Python object, including custom class instances, unlike json's limited type support.
  • dumps/loads for bytes, dump/load for binary files, opened with 'wb' and 'rb'.
  • Custom class instances pickle cleanly, as long as the class is importable wherever it's unpickled.
  • Lambdas and open file handles genuinely can't be pickled; ordinary named functions can.
  • Protocol versions trade compatibility for compactness; HIGHEST_PROTOCOL is the newest, most efficient one.
  • Never unpickle untrusted data; unpickling can genuinely execute arbitrary code.
  • shelve provides a persistent, dict-like store backed by pickle, with string keys only.
  • writeback=True is required for in-place mutations of shelved values to actually persist.
  • A real pattern: disk-backed caching of expensive computations with pickle.
  • That wraps up pickle and shelve. Next up: hashlib, hmac, and secrets, for hashing and cryptographic basics.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.