Lesson 5 · Python Libraries
Comprehensive Guide to Using pypdf for PDF File Management in Python
PyPDF is a Python library for working with PDF files. It is used to read, extract, merge, split, and modify PDFs. Real world uses include extracting text,…
- CoursePython Libraries
- Lesson5 of 6
- Video17 min
- FormatJupyter notebook · 20 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIntroduction to PyPDF#
- PyPDF is a Python library for working with PDF files.
- It is used to read, extract, merge, split, and modify PDFs.
- Real world uses include extracting text, combining files, and automating PDF document workflows.
- PyPDF makes handling PDFs in Python simple and flexible.
# On Windows, install with: pip install pypdf
from pypdf import PdfReader, PdfWriter
Core Concepts in PyPDF#
- PdfReader is used to open and read PDF documents.
- PdfWriter helps create and modify PDF files.
- Pages in a PDF are accessed as a list using reader.pages.
- PyPDF can extract text, images, and meta data from PDFs.
- No external PDF software is required.
# Example 1: Open a PDF file for reading
reader = PdfReader('sample.pdf')
# Example 2: Count the number of pages
num_pages = len(reader.pages)
print('Number of pages:', num_pages)
# Example 3: Extract text from the first page
page = reader.pages[0]
text = page.extract_text()
print('First page text:')
print(text)
# Example 4: Merge two PDF files
writer = PdfWriter()
for file_name in ['sample.pdf', 'second.pdf']:
input_pdf = PdfReader(file_name)
for page in input_pdf.pages:
writer.add_page(page)
writer.write('merged.pdf')
# Example 5: Split a PDF into single-page files
for page_num, page in enumerate(reader.pages):
writer = PdfWriter()
writer.add_page(page)
output_name = f'page_{page_num + 1}.pdf'
writer.write(output_name)
# Example 6: Add metadata to a PDF
writer = PdfWriter()
for page in reader.pages:
writer.add_page(page)
writer.add_metadata({'/Title': 'My Custom PDF', '/Author': 'YouTube Demo'})
writer.write('metadata_sample.pdf')
# Example 7: Decrypting an encrypted PDF (Intermediate)
protected_reader = PdfReader('protected.pdf')
if protected_reader.is_encrypted:
protected_reader.decrypt('mypassword')
text = protected_reader.pages[0].extract_text()
print('Decrypted page 1 text:')
print(text)
# Example 8: Rotate a page clockwise by 90 degrees
from copy import deepcopy
writer = PdfWriter()
rotated_page = deepcopy(reader.pages[0])
rotated_page.rotate(90)
writer.add_page(rotated_page)
writer.write('rotated.pdf')
# Example 9: Extracting all text from a PDF (Intermediate)
all_text = ''
for page in reader.pages:
page_text = page.extract_text()
if page_text:
all_text += page_text + '\n'
print('All PDF text:')
print(all_text)
# Example 10: Copy only odd pages to a new PDF (Intermediate)
writer = PdfWriter()
for page_num, page in enumerate(reader.pages):
if page_num % 2 == 0:
writer.add_page(page)
writer.write('odd_pages.pdf')
# Example 11: Extract PDF meta data (Intermediate)
meta = reader.metadata
for key, value in meta.items():
print(f'{key}: {value}')
# Example 12: Add password protection to a PDF (Advanced)
writer = PdfWriter()
for page in reader.pages:
writer.add_page(page)
writer.encrypt('mypassword')
writer.write('protected_out.pdf')
# Example 13: Linearizing a PDF for fast web view (Advanced)
# PyPDF cannot fully linearize PDFs, but we can optimize structure.
writer = PdfWriter()
for page in reader.pages:
writer.add_page(page)
writer.write('linearized_sample.pdf')
# Example 14: Catch missing file errors (Error Handling)
try:
missing_reader = PdfReader('missing.pdf')
except FileNotFoundError:
print('File missing.pdf not found! Check the file name.')
# Example 15: Warn when extracting text returns None (Error Handling)
page = reader.pages[0]
text = page.extract_text()
if text is None:
print('No text extracted. This page may be scanned or image-only.')
else:
print('Text:', text)
# Example 16: Always close writers for best practice
writer = PdfWriter()
for page in reader.pages:
writer.add_page(page)
with open('best_practice.pdf', 'wb') as out_file:
writer.write(out_file)
# Example 17: Check for encryption before extracting text (Best Practice)
if reader.is_encrypted:
print('PDF is encrypted! Cannot extract text without the password.')
else:
text = reader.pages[0].extract_text()
print('Extracted text:', text)
# Example 18: Get total size of all PDF outputs (Common Pattern)
import os
total_size = 0
for name in ['merged.pdf', 'best_practice.pdf', 'odd_pages.pdf']:
if os.path.exists(name):
total_size += os.path.getsize(name)
print('Total size in bytes:', total_size)
# Example 19: Tiny Mini-Project: Combine only first pages of several PDFs
import glob
writer = PdfWriter()
for pdf_file in glob.glob('*.pdf'):
try:
pdf = PdfReader(pdf_file)
first_page = pdf.pages[0]
writer.add_page(first_page)
print(f'Added first page from {pdf_file}')
except Exception as e:
print(f'Cannot add {pdf_file}:', e)
writer.write('first_pages_collection.pdf')
Thanks for learning PyPDF!#
- Try more examples and explore PyPDF documentation.
- Subscribe for more simple Python tutorials.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



