Lesson 9 · Python for Data Analysts
Web Scraping with BeautifulSoup
Everything you need to parse real HTML and pull structured data out of it: tags, CSS selectors, tables, pagination, and scraping ethics. No prior scraping…
- CoursePython for Data Analysts
- Lesson9 of 12
- Video33 min
- FormatJupyter notebook · 27 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeb Scraping with BeautifulSoup#
- Everything you need to parse real HTML and pull structured data out of it: tags, CSS selectors, tables, pagination, and scraping ethics.
- No prior scraping experience needed. Let's get straight into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- If they aren't installed yet, open a terminal in VS Code and run: pip install beautifulsoup4 requests lxml
Part 1: HTML Basics for Scraping#
The Building Blocks#
- HTML is made of nested tags, like ,
, , and
, each optionally carrying attributes like class, id, or href.
- Scraping means finding the right tags, by name, attribute, or CSS class, and pulling out either their text or their attribute values.
- Browser DevTools, right-click, Inspect, is how you'd normally find the exact tags and classes a real page uses.
Setting Up a Local Practice Site#
import threading
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from urllib.parse import urlparse, parse_qs
BOOKS = [
{'title': 'Learning Python', 'price': 39.99, 'rating': 5},
{'title': 'Deep Work', 'price': 18.50, 'rating': 4},
{'title': 'The Pragmatic Programmer', 'price': 42.00, 'rating': 5},
{'title': 'Clean Code', 'price': 34.75, 'rating': 3},
{'title': 'Automate the Boring Stuff', 'price': 29.99, 'rating': 4},
{'title': 'Fluent Python', 'price': 44.50, 'rating': 5},
{'title': 'Data Science from Scratch', 'price': 36.25, 'rating': 3},
{'title': 'Effective Python', 'price': 31.00, 'rating': 4},
{'title': 'Python Crash Course', 'price': 27.99, 'rating': 4},
{'title': 'Storytelling with Data', 'price': 33.40, 'rating': 5},
{'title': 'Grokking Algorithms', 'price': 28.75, 'rating': 4},
{'title': 'Designing Data-Intensive Applications', 'price': 49.99, 'rating': 5},
]
PAGE_SIZE = 4
def render_catalog_page(page):
start = (page - 1) * PAGE_SIZE
items = BOOKS[start:start + PAGE_SIZE]
total_pages = -(-len(BOOKS) // PAGE_SIZE)
cards = ''
for book in items:
cards += (
f'<div class="product" data-rating="{book["rating"]}">'
f'<h2 class="title">{book["title"]}</h2>'
f'<span class="price">${book["price"]:.2f}</span>'
f'</div>'
)
nav = ''
if page > 1:
nav += f'<a class="prev" href="/catalog?page={page - 1}">Previous</a>'
if page < total_pages:
nav += f'<a class="next" href="/catalog?page={page + 1}">Next</a>'
return (
'<html><body>'
'<h1>Book Catalog</h1>'
f'<div class="catalog">{cards}</div>'
f'<div class="pagination">{nav}</div>'
'</body></html>'
)
def render_report_page():
rows = ''
for region, revenue in [('North', 12000), ('South', 9400), ('East', 15250), ('West', 11100)]:
rows += f'<tr><td>{region}</td><td>{revenue}</td></tr>'
return (
'<html><body>'
'<h1>Quarterly Sales</h1>'
f'<table id="sales"><tr><th>Region</th><th>Revenue</th></tr>{rows}</table>'
'</body></html>'
)
ROBOTS_TXT = 'User-agent: *\nDisallow: /admin\nAllow: /catalog\nAllow: /report\n'
class ScrapeSiteHandler(BaseHTTPRequestHandler):
def _send_html(self, status, html):
body = html.encode('utf-8')
self.send_response(status)
self.send_header('Content-Type', 'text/html; charset=utf-8')
self.send_header('Content-Length', str(len(body)))
self.end_headers()
self.wfile.write(body)
def log_message(self, format, *args):
pass
def do_GET(self):
parsed = urlparse(self.path)
query = parse_qs(parsed.query)
if parsed.path == '/catalog':
page = int(query.get('page', ['1'])[0])
self._send_html(200, render_catalog_page(page))
elif parsed.path == '/report':
self._send_html(200, render_report_page())
elif parsed.path == '/robots.txt':
self._send_html(200, ROBOTS_TXT)
else:
self._send_html(404, '<html><body><h1>Not Found</h1></body></html>')
server = ThreadingHTTPServer(('127.0.0.1', 0), ScrapeSiteHandler)
server_thread = threading.Thread(target=server.serve_forever, daemon=True)
server_thread.start()
BASE_URL = f'http://127.0.0.1:{server.server_port}'
print(f'Practice site running at {BASE_URL}')
import requests
from bs4 import BeautifulSoup
Part 2: Parsing HTML with BeautifulSoup#
response = requests.get(f'{BASE_URL}/catalog')
soup = BeautifulSoup(response.text, 'lxml')
print(type(soup))
print(soup.h1)
first_title = soup.find('h2', class_='title')
print(first_title)
print(first_title.text)
all_titles = soup.find_all('h2', class_='title')
print(len(all_titles))
for title in all_titles:
print(title.text)
all_prices = soup.find_all('span', class_='price')
prices = [p.text for p in all_prices]
print(prices)
Part 3: CSS Selectors#
first_product = soup.select_one('div.product')
print(first_product.h2.text)
all_products = soup.select('div.product')
print(len(all_products))
for product in all_products:
title = product.select_one('h2.title').text
price = product.select_one('span.price').text
print(f'{title}: {price}')
high_rated = soup.select('div.product[data-rating="5"]')
print(len(high_rated))
for product in high_rated:
print(product.h2.text)
Part 4: Extracting Data and Attributes#
product = soup.select_one('div.product')
print(product['data-rating'])
print(product.get('data-rating'))
print(product.get('data-missing', 'N/A'))
price_text = product.select_one('span.price').text
price_number = float(price_text.replace('$', ''))
print(price_text, '->', price_number)
clean_data = []
for product in soup.select('div.product'):
title = product.select_one('h2.title').text.strip()
price = float(product.select_one('span.price').text.replace('$', ''))
rating = int(product['data-rating'])
clean_data.append({'title': title, 'price': price, 'rating': rating})
print(clean_data)
Part 5: Navigating the Tree#
title_tag = soup.select_one('h2.title')
print(title_tag.parent)
product = soup.select_one('div.product')
for child in product.children:
print(repr(child))
title_tag = soup.select_one('h2.title')
sibling = title_tag.find_next_sibling('span')
print(sibling.text)
Part 6: Extracting Tables to pandas#
import pandas as pd
from io import StringIO
report_response = requests.get(f'{BASE_URL}/report')
tables = pd.read_html(StringIO(report_response.text))
print(len(tables))
print(tables[0])
report_soup = BeautifulSoup(report_response.text, 'lxml')
table_tag = report_soup.find('table', id='sales')
rows = []
for row in table_tag.find_all('tr')[1:]:
cells = row.find_all('td')
rows.append({'region': cells[0].text, 'revenue': int(cells[1].text)})
manual_df = pd.DataFrame(rows)
print(manual_df)
Part 7: Pagination Handling#
all_books = []
page = 1
while True:
resp = requests.get(f'{BASE_URL}/catalog', params={'page': page})
page_soup = BeautifulSoup(resp.text, 'lxml')
products = page_soup.select('div.product')
if not products:
break
for product in products:
all_books.append({
'title': product.select_one('h2.title').text.strip(),
'price': float(product.select_one('span.price').text.replace('$', '')),
'rating': int(product['data-rating']),
})
has_next = page_soup.select_one('a.next') is not None
if not has_next:
break
page += 1
print(f'Collected {len(all_books)} books across {page} pages')
Part 8: Saving to CSV#
books_df = pd.DataFrame(all_books)
print(books_df.head())
books_df.to_csv('scraped_books.csv', index=False)
print('Saved scraped_books.csv')
reloaded_books = pd.read_csv('scraped_books.csv')
print(reloaded_books.shape)
print(reloaded_books.sort_values('price', ascending=False).head(3))
Part 9: Scraping Ethics#
robots_response = requests.get(f'{BASE_URL}/robots.txt')
print(robots_response.text)
import time
for page_num in [1, 2, 3]:
requests.get(f'{BASE_URL}/catalog', params={'page': page_num})
time.sleep(0.2)
print('Finished a polite, rate-limited crawl')
Capstone Project: A Reusable Scrape-to-DataFrame Pipeline#
def scrape_catalog(base_url, delay=0.1):
records = []
page = 1
while True:
response = requests.get(f'{base_url}/catalog', params={'page': page})
response.raise_for_status()
page_soup = BeautifulSoup(response.text, 'lxml')
products = page_soup.select('div.product')
if not products:
break
for product in products:
records.append({
'title': product.select_one('h2.title').text.strip(),
'price': float(product.select_one('span.price').text.replace('$', '')),
'rating': int(product['data-rating']),
'page': page,
})
if page_soup.select_one('a.next') is None:
break
page += 1
time.sleep(delay)
return pd.DataFrame(records)
catalog_df = scrape_catalog(BASE_URL)
print(catalog_df.shape)
print(catalog_df.groupby('rating')['price'].mean())
Wrap-Up: What You Learned#
- HTML basics for scraping, and setting up a real practice site.
- Parsing HTML with BeautifulSoup: find, find_all, and dot-access.
- CSS selectors with select and select_one, including scoped and attribute selectors.
- Extracting text and attributes, and converting them into real, usable data types.
- Navigating the tree with parent, children, and sibling lookups.
- Extracting real HTML tables straight into pandas with read_html.
- Handling pagination by following Next links until they run out.
- Saving scraped data to CSV, and scraping ethics: robots.txt, rate limiting, and respecting site terms.
- A capstone pipeline turning an entire paginated site into one clean DataFrame.
- You went from a raw HTML tag to a full, rate-limited, paginated scraper in one sitting. If you want the next build to land in your feed automatically, subscribing is the move see you in the next one.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



