Mathew K Analytics

Lesson 5 · Python Libraries

Comprehensive Guide to Using pypdf for PDF File Management in Python

PyPDF is a Python library for working with PDF files. It is used to read, extract, merge, split, and modify PDFs. Real world uses include extracting text,…

⬇ Download notebookOpen in Colab ↗
pypdf

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Introduction to PyPDF#

  • PyPDF is a Python library for working with PDF files.
  • It is used to read, extract, merge, split, and modify PDFs.
  • Real world uses include extracting text, combining files, and automating PDF document workflows.
  • PyPDF makes handling PDFs in Python simple and flexible.
# On Windows, install with: pip install pypdf
from pypdf import PdfReader, PdfWriter

Core Concepts in PyPDF#

  • PdfReader is used to open and read PDF documents.
  • PdfWriter helps create and modify PDF files.
  • Pages in a PDF are accessed as a list using reader.pages.
  • PyPDF can extract text, images, and meta data from PDFs.
  • No external PDF software is required.
# Example 1: Open a PDF file for reading
reader = PdfReader('sample.pdf')
# Example 2: Count the number of pages
num_pages = len(reader.pages)
print('Number of pages:', num_pages)
Number of pages: 5
# Example 3: Extract text from the first page
page = reader.pages[0]
text = page.extract_text()
print('First page text:')
print(text)
First page text:
Sample PDF – Page 1
This is page 1 of a 5-page sample PDF document. You can use this file to test PDF splitting,
merging, or page-range extraction logic in your scripts. Each page is intentionally simple and clearly
labeled.

# Example 4: Merge two PDF files
writer = PdfWriter()
for file_name in ['sample.pdf', 'second.pdf']:
    input_pdf = PdfReader(file_name)
    for page in input_pdf.pages:
        writer.add_page(page)
writer.write('merged.pdf')
(True, <_io.FileIO [closed]>)
# Example 5: Split a PDF into single-page files
for page_num, page in enumerate(reader.pages):
    writer = PdfWriter()
    writer.add_page(page)
    output_name = f'page_{page_num + 1}.pdf'
    writer.write(output_name)
# Example 6: Add metadata to a PDF
writer = PdfWriter()
for page in reader.pages:
    writer.add_page(page)
writer.add_metadata({'/Title': 'My Custom PDF', '/Author': 'YouTube Demo'})
writer.write('metadata_sample.pdf')
(True, <_io.FileIO [closed]>)
# Example 7: Decrypting an encrypted PDF (Intermediate)
protected_reader = PdfReader('protected.pdf')
if protected_reader.is_encrypted:
    protected_reader.decrypt('mypassword')
    text = protected_reader.pages[0].extract_text()
    print('Decrypted page 1 text:')
    print(text)
Decrypted page 1 text:
Sample PDF – Page 1
This is page 1 of a 5-page sample PDF document. You can use this file to test PDF splitting,
merging, or page-range extraction logic in your scripts. Each page is intentionally simple and clearly
labeled.

# Example 8: Rotate a page clockwise by 90 degrees
from copy import deepcopy
writer = PdfWriter()
rotated_page = deepcopy(reader.pages[0])
rotated_page.rotate(90)
writer.add_page(rotated_page)
writer.write('rotated.pdf')
(True, <_io.FileIO [closed]>)
# Example 9: Extracting all text from a PDF (Intermediate)
all_text = ''
for page in reader.pages:
    page_text = page.extract_text()
    if page_text:
        all_text += page_text + '\n'
print('All PDF text:')
print(all_text)
All PDF text:
Sample PDF – Page 1
This is page 1 of a 5-page sample PDF document. You can use this file to test PDF splitting,
merging, or page-range extraction logic in your scripts. Each page is intentionally simple and clearly
labeled.

Sample PDF – Page 2
This is page 2 of a 5-page sample PDF document. You can use this file to test PDF splitting,
merging, or page-range extraction logic in your scripts. Each page is intentionally simple and clearly
labeled.

Sample PDF – Page 3
This is page 3 of a 5-page sample PDF document. You can use this file to test PDF splitting,
merging, or page-range extraction logic in your scripts. Each page is intentionally simple and clearly
labeled.

Sample PDF – Page 4
This is page 4 of a 5-page sample PDF document. You can use this file to test PDF splitting,
merging, or page-range extraction logic in your scripts. Each page is intentionally simple and clearly
labeled.

Sample PDF – Page 5
This is page 5 of a 5-page sample PDF document. You can use this file to test PDF splitting,
merging, or page-range extraction logic in your scripts. Each page is intentionally simple and clearly
labeled.


# Example 10: Copy only odd pages to a new PDF (Intermediate)
writer = PdfWriter()
for page_num, page in enumerate(reader.pages):
    if page_num % 2 == 0:
        writer.add_page(page)
writer.write('odd_pages.pdf')
(True, <_io.FileIO [closed]>)
# Example 11: Extract PDF meta data (Intermediate)
meta = reader.metadata
for key, value in meta.items():
    print(f'{key}: {value}')
/Author: (anonymous)
/CreationDate: D:20251217002804+00'00'
/Creator: (unspecified)
/Keywords: 
/ModDate: D:20251217002804+00'00'
/Producer: ReportLab PDF Library - www.reportlab.com
/Subject: (unspecified)
/Title: (anonymous)
/Trapped: /False
# Example 12: Add password protection to a PDF (Advanced)
writer = PdfWriter()
for page in reader.pages:
    writer.add_page(page)
writer.encrypt('mypassword')
writer.write('protected_out.pdf')
(True, <_io.FileIO [closed]>)
# Example 13: Linearizing a PDF for fast web view (Advanced)
# PyPDF cannot fully linearize PDFs, but we can optimize structure.
writer = PdfWriter()
for page in reader.pages:
    writer.add_page(page)
writer.write('linearized_sample.pdf')
(True, <_io.FileIO [closed]>)
# Example 14: Catch missing file errors (Error Handling)
try:
    missing_reader = PdfReader('missing.pdf')
except FileNotFoundError:
    print('File missing.pdf not found! Check the file name.')
File missing.pdf not found! Check the file name.
# Example 15: Warn when extracting text returns None (Error Handling)
page = reader.pages[0]
text = page.extract_text()
if text is None:
    print('No text extracted. This page may be scanned or image-only.')
else:
    print('Text:', text)
Text: Sample PDF – Page 1
This is page 1 of a 5-page sample PDF document. You can use this file to test PDF splitting,
merging, or page-range extraction logic in your scripts. Each page is intentionally simple and clearly
labeled.

# Example 16: Always close writers for best practice
writer = PdfWriter()
for page in reader.pages:
    writer.add_page(page)
with open('best_practice.pdf', 'wb') as out_file:
    writer.write(out_file)
# Example 17: Check for encryption before extracting text (Best Practice)
if reader.is_encrypted:
    print('PDF is encrypted! Cannot extract text without the password.')
else:
    text = reader.pages[0].extract_text()
    print('Extracted text:', text)
Extracted text: Sample PDF – Page 1
This is page 1 of a 5-page sample PDF document. You can use this file to test PDF splitting,
merging, or page-range extraction logic in your scripts. Each page is intentionally simple and clearly
labeled.

# Example 18: Get total size of all PDF outputs (Common Pattern)
import os
total_size = 0
for name in ['merged.pdf', 'best_practice.pdf', 'odd_pages.pdf']:
    if os.path.exists(name):
        total_size += os.path.getsize(name)
print('Total size in bytes:', total_size)
Total size in bytes: 35734
# Example 19: Tiny Mini-Project: Combine only first pages of several PDFs
import glob
writer = PdfWriter()
for pdf_file in glob.glob('*.pdf'):
    try:
        pdf = PdfReader(pdf_file)
        first_page = pdf.pages[0]
        writer.add_page(first_page)
        print(f'Added first page from {pdf_file}')
    except Exception as e:
        print(f'Cannot add {pdf_file}:', e)
writer.write('first_pages_collection.pdf')
Added first page from all_combined.pdf
Added first page from appendix.pdf
Added first page from best_practice.pdf
Cannot add encrypted.pdf: File has not been decrypted
Added first page from file1.pdf
Added first page from file2.pdf
Added first page from file3.pdf
Added first page from final_report.pdf
Added first page from first_page.pdf
Added first page from first_pages_collection.pdf
Added first page from first_pages_combined.pdf
Added first page from linearized_sample.pdf
Added first page from main.pdf
Added first page from merged.pdf
Added first page from meta.pdf
Added first page from metadata_sample.pdf
Added first page from odd_pages.pdf
Added first page from page_1.pdf
Added first page from page_2.pdf
Added first page from page_3.pdf
Added first page from page_4.pdf
Added first page from page_5.pdf
Cannot add protected.pdf: File has not been decrypted
Cannot add protected_out.pdf: File has not been decrypted
Added first page from rotated.pdf
Added first page from rotated_page.pdf
Added first page from sample.pdf
Added first page from second.pdf
Added first page from title.pdf
Added first page from watermark.pdf
Added first page from watermarked.pdf
Added first page from with_blank.pdf
(True, <_io.FileIO [closed]>)

Thanks for learning PyPDF!#

  • Try more examples and explore PyPDF documentation.
  • Subscribe for more simple Python tutorials.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.