Mathew K Analytics

Lesson 9 · Python for Data Analysts

Web Scraping with BeautifulSoup

Everything you need to parse real HTML and pull structured data out of it: tags, CSS selectors, tables, pagination, and scraping ethics. No prior scraping…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Web Scraping with BeautifulSoup#

  • Everything you need to parse real HTML and pull structured data out of it: tags, CSS selectors, tables, pagination, and scraping ethics.
  • No prior scraping experience needed. Let's get straight into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • If they aren't installed yet, open a terminal in VS Code and run: pip install beautifulsoup4 requests lxml

Part 1: HTML Basics for Scraping#

The Building Blocks#

  • HTML is made of nested tags, like
  • Scraping means finding the right tags, by name, attribute, or CSS class, and pulling out either their text or their attribute values.
  • Browser DevTools, right-click, Inspect, is how you'd normally find the exact tags and classes a real page uses.

Setting Up a Local Practice Site#

import threading
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from urllib.parse import urlparse, parse_qs

BOOKS = [
    {'title': 'Learning Python', 'price': 39.99, 'rating': 5},
    {'title': 'Deep Work', 'price': 18.50, 'rating': 4},
    {'title': 'The Pragmatic Programmer', 'price': 42.00, 'rating': 5},
    {'title': 'Clean Code', 'price': 34.75, 'rating': 3},
    {'title': 'Automate the Boring Stuff', 'price': 29.99, 'rating': 4},
    {'title': 'Fluent Python', 'price': 44.50, 'rating': 5},
    {'title': 'Data Science from Scratch', 'price': 36.25, 'rating': 3},
    {'title': 'Effective Python', 'price': 31.00, 'rating': 4},
    {'title': 'Python Crash Course', 'price': 27.99, 'rating': 4},
    {'title': 'Storytelling with Data', 'price': 33.40, 'rating': 5},
    {'title': 'Grokking Algorithms', 'price': 28.75, 'rating': 4},
    {'title': 'Designing Data-Intensive Applications', 'price': 49.99, 'rating': 5},
]
PAGE_SIZE = 4
def render_catalog_page(page):
    start = (page - 1) * PAGE_SIZE
    items = BOOKS[start:start + PAGE_SIZE]
    total_pages = -(-len(BOOKS) // PAGE_SIZE)

    cards = ''
    for book in items:
        cards += (
            f'<div class="product" data-rating="{book["rating"]}">'
            f'<h2 class="title">{book["title"]}</h2>'
            f'<span class="price">${book["price"]:.2f}</span>'
            f'</div>'
        )

    nav = ''
    if page > 1:
        nav += f'<a class="prev" href="/catalog?page={page - 1}">Previous</a>'
    if page < total_pages:
        nav += f'<a class="next" href="/catalog?page={page + 1}">Next</a>'

    return (
        '<html><body>'
        '<h1>Book Catalog</h1>'
        f'<div class="catalog">{cards}</div>'
        f'<div class="pagination">{nav}</div>'
        '</body></html>'
    )

def render_report_page():
    rows = ''
    for region, revenue in [('North', 12000), ('South', 9400), ('East', 15250), ('West', 11100)]:
        rows += f'<tr><td>{region}</td><td>{revenue}</td></tr>'
    return (
        '<html><body>'
        '<h1>Quarterly Sales</h1>'
        f'<table id="sales"><tr><th>Region</th><th>Revenue</th></tr>{rows}</table>'
        '</body></html>'
    )
ROBOTS_TXT = 'User-agent: *\nDisallow: /admin\nAllow: /catalog\nAllow: /report\n'

class ScrapeSiteHandler(BaseHTTPRequestHandler):
    def _send_html(self, status, html):
        body = html.encode('utf-8')
        self.send_response(status)
        self.send_header('Content-Type', 'text/html; charset=utf-8')
        self.send_header('Content-Length', str(len(body)))
        self.end_headers()
        self.wfile.write(body)

    def log_message(self, format, *args):
        pass

    def do_GET(self):
        parsed = urlparse(self.path)
        query = parse_qs(parsed.query)

        if parsed.path == '/catalog':
            page = int(query.get('page', ['1'])[0])
            self._send_html(200, render_catalog_page(page))
        elif parsed.path == '/report':
            self._send_html(200, render_report_page())
        elif parsed.path == '/robots.txt':
            self._send_html(200, ROBOTS_TXT)
        else:
            self._send_html(404, '<html><body><h1>Not Found</h1></body></html>')
server = ThreadingHTTPServer(('127.0.0.1', 0), ScrapeSiteHandler)
server_thread = threading.Thread(target=server.serve_forever, daemon=True)
server_thread.start()
BASE_URL = f'http://127.0.0.1:{server.server_port}'
print(f'Practice site running at {BASE_URL}')
Practice site running at http://127.0.0.1:54670
import requests
from bs4 import BeautifulSoup

Part 2: Parsing HTML with BeautifulSoup#

response = requests.get(f'{BASE_URL}/catalog')
soup = BeautifulSoup(response.text, 'lxml')
print(type(soup))
print(soup.h1)
<class 'bs4.BeautifulSoup'>
<h1>Book Catalog</h1>
first_title = soup.find('h2', class_='title')
print(first_title)
print(first_title.text)
<h2 class="title">Learning Python</h2>
Learning Python
all_titles = soup.find_all('h2', class_='title')
print(len(all_titles))
for title in all_titles:
    print(title.text)
4
Learning Python
Deep Work
The Pragmatic Programmer
Clean Code
all_prices = soup.find_all('span', class_='price')
prices = [p.text for p in all_prices]
print(prices)
['$39.99', '$18.50', '$42.00', '$34.75']

Part 3: CSS Selectors#

first_product = soup.select_one('div.product')
print(first_product.h2.text)
Learning Python
all_products = soup.select('div.product')
print(len(all_products))
for product in all_products:
    title = product.select_one('h2.title').text
    price = product.select_one('span.price').text
    print(f'{title}: {price}')
4
Learning Python: $39.99
Deep Work: $18.50
The Pragmatic Programmer: $42.00
Clean Code: $34.75
high_rated = soup.select('div.product[data-rating="5"]')
print(len(high_rated))
for product in high_rated:
    print(product.h2.text)
2
Learning Python
The Pragmatic Programmer

Part 4: Extracting Data and Attributes#

product = soup.select_one('div.product')
print(product['data-rating'])
print(product.get('data-rating'))
print(product.get('data-missing', 'N/A'))
5
5
N/A
price_text = product.select_one('span.price').text
price_number = float(price_text.replace('$', ''))
print(price_text, '->', price_number)
$39.99 -> 39.99
clean_data = []
for product in soup.select('div.product'):
    title = product.select_one('h2.title').text.strip()
    price = float(product.select_one('span.price').text.replace('$', ''))
    rating = int(product['data-rating'])
    clean_data.append({'title': title, 'price': price, 'rating': rating})
print(clean_data)
[{'title': 'Learning Python', 'price': 39.99, 'rating': 5}, {'title': 'Deep Work', 'price': 18.5, 'rating': 4}, {'title': 'The Pragmatic Programmer', 'price': 42.0, 'rating': 5}, {'title': 'Clean Code', 'price': 34.75, 'rating': 3}]

Part 5: Navigating the Tree#

title_tag = soup.select_one('h2.title')
print(title_tag.parent)
<div class="product" data-rating="5"><h2 class="title">Learning Python</h2><span class="price">$39.99</span></div>
product = soup.select_one('div.product')
for child in product.children:
    print(repr(child))
<h2 class="title">Learning Python</h2>
<span class="price">$39.99</span>
title_tag = soup.select_one('h2.title')
sibling = title_tag.find_next_sibling('span')
print(sibling.text)
$39.99

Part 6: Extracting Tables to pandas#

import pandas as pd
from io import StringIO

report_response = requests.get(f'{BASE_URL}/report')
tables = pd.read_html(StringIO(report_response.text))
print(len(tables))
print(tables[0])
1
  Region  Revenue
0  North    12000
1  South     9400
2   East    15250
3   West    11100
report_soup = BeautifulSoup(report_response.text, 'lxml')
table_tag = report_soup.find('table', id='sales')
rows = []
for row in table_tag.find_all('tr')[1:]:
    cells = row.find_all('td')
    rows.append({'region': cells[0].text, 'revenue': int(cells[1].text)})
manual_df = pd.DataFrame(rows)
print(manual_df)
  region  revenue
0  North    12000
1  South     9400
2   East    15250
3   West    11100

Part 7: Pagination Handling#

all_books = []
page = 1
while True:
    resp = requests.get(f'{BASE_URL}/catalog', params={'page': page})
    page_soup = BeautifulSoup(resp.text, 'lxml')
    products = page_soup.select('div.product')
    if not products:
        break
    for product in products:
        all_books.append({
            'title': product.select_one('h2.title').text.strip(),
            'price': float(product.select_one('span.price').text.replace('$', '')),
            'rating': int(product['data-rating']),
        })
    has_next = page_soup.select_one('a.next') is not None
    if not has_next:
        break
    page += 1

print(f'Collected {len(all_books)} books across {page} pages')
Collected 12 books across 3 pages

Part 8: Saving to CSV#

books_df = pd.DataFrame(all_books)
print(books_df.head())
books_df.to_csv('scraped_books.csv', index=False)
print('Saved scraped_books.csv')
                       title  price  rating
0            Learning Python  39.99       5
1                  Deep Work  18.50       4
2   The Pragmatic Programmer  42.00       5
3                 Clean Code  34.75       3
4  Automate the Boring Stuff  29.99       4
Saved scraped_books.csv
reloaded_books = pd.read_csv('scraped_books.csv')
print(reloaded_books.shape)
print(reloaded_books.sort_values('price', ascending=False).head(3))
(12, 3)
                                    title  price  rating
11  Designing Data-Intensive Applications  49.99       5
5                           Fluent Python  44.50       5
2                The Pragmatic Programmer  42.00       5

Part 9: Scraping Ethics#

robots_response = requests.get(f'{BASE_URL}/robots.txt')
print(robots_response.text)
User-agent: *
Disallow: /admin
Allow: /catalog
Allow: /report

import time

for page_num in [1, 2, 3]:
    requests.get(f'{BASE_URL}/catalog', params={'page': page_num})
    time.sleep(0.2)
print('Finished a polite, rate-limited crawl')
Finished a polite, rate-limited crawl

Capstone Project: A Reusable Scrape-to-DataFrame Pipeline#

def scrape_catalog(base_url, delay=0.1):
    records = []
    page = 1
    while True:
        response = requests.get(f'{base_url}/catalog', params={'page': page})
        response.raise_for_status()
        page_soup = BeautifulSoup(response.text, 'lxml')
        products = page_soup.select('div.product')
        if not products:
            break
        for product in products:
            records.append({
                'title': product.select_one('h2.title').text.strip(),
                'price': float(product.select_one('span.price').text.replace('$', '')),
                'rating': int(product['data-rating']),
                'page': page,
            })
        if page_soup.select_one('a.next') is None:
            break
        page += 1
        time.sleep(delay)
    return pd.DataFrame(records)
catalog_df = scrape_catalog(BASE_URL)
print(catalog_df.shape)
print(catalog_df.groupby('rating')['price'].mean())
(12, 4)
rating
3    35.500
4    27.246
5    41.976
Name: price, dtype: float64

Wrap-Up: What You Learned#

  • HTML basics for scraping, and setting up a real practice site.
  • Parsing HTML with BeautifulSoup: find, find_all, and dot-access.
  • CSS selectors with select and select_one, including scoped and attribute selectors.
  • Extracting text and attributes, and converting them into real, usable data types.
  • Navigating the tree with parent, children, and sibling lookups.
  • Extracting real HTML tables straight into pandas with read_html.
  • Handling pagination by following Next links until they run out.
  • Saving scraped data to CSV, and scraping ethics: robots.txt, rate limiting, and respecting site terms.
  • A capstone pipeline turning an entire paginated site into one clean DataFrame.
  • You went from a raw HTML tag to a full, rate-limited, paginated scraper in one sitting. If you want the next build to land in your feed automatically, subscribing is the move see you in the next one.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.