Lesson 22 · Data analytics zero to hero
Web Scraping with Python: Step by Step | Data Analytics #22
Video twenty-two of the 30-part series: pulling real structured data straight out of a web page's HTML, for the many real sites that don't offer an API.…
- CourseData analytics zero to hero
- Lesson22 of 30
- Video10 min
- FormatJupyter notebook · 6 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbData Analytics Zero to Hero, Video 22: Web Scraping for Data Analysts#
- Video twenty-two of the 30-part series: pulling real structured data straight out of a web page's HTML, for the many real sites that don't offer an API.
- We're using books.toscrape.com, a real site built specifically and legally for practicing web scraping.
- Let's jump straight in.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- Install requests and BeautifulSoup if you haven't already:
pip install requests beautifulsoup4. - You'll need an active internet connection for this video, since every cell here fetches a real, live page.
- Always check a real site's terms of service and robots.txt before scraping it for anything beyond practice.
Part 1: Fetching and Parsing Real HTML#
import requests
from bs4 import BeautifulSoup
url = 'http://books.toscrape.com/'
response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')
print(response.status_code)
print(soup.title.text)
Part 2: Finding Real Elements#
books = soup.find_all('article', class_='product_pod')
print(len(books))
print(books[0].prettify()[:500])
first_book = books[0]
title = first_book.h3.a['title']
price = first_book.find('p', class_='price_color').text
availability = first_book.find('p', class_='instock availability').text.strip()
print(title, price, availability)
Part 3: Building a Real Dataset from One Page#
rows = []
for book in books:
rows.append({
'title': book.h3.a['title'],
'price': book.find('p', class_='price_color').text,
'rating': book.p['class'][1],
'in_stock': 'In stock' in book.find('p', class_='instock availability').text
})
print(len(rows))
print(rows[0])
Part 4: Following Real Pagination#
import time
all_rows = []
for page_num in range(1, 4):
page_url = f'http://books.toscrape.com/catalogue/page-{page_num}.html'
page_response = requests.get(page_url)
page_soup = BeautifulSoup(page_response.text, 'html.parser')
page_books = page_soup.find_all('article', class_='product_pod')
for book in page_books:
all_rows.append({
'title': book.h3.a['title'],
'price': book.find('p', class_='price_color').text
})
time.sleep(1)
print(len(all_rows))
Part 5: Turning Scraped Data into a DataFrame#
import pandas as pd
books_df = pd.DataFrame(all_rows)
books_df['price'] = books_df['price'].str.replace('', '', regex=False).astype(float)
print(books_df.head())
print(books_df['price'].describe())
Wrap-Up: What You Learned#
- Fetching real HTML with requests, and parsing it into a searchable tree with BeautifulSoup.
- find and find_all, for locating real elements by tag and class, and pulling out text or attribute values.
- Looping over matched elements to build a real dataset, one dictionary per item.
- Following real pagination across multiple pages, with a polite delay between requests.
- Cleaning and converting scraped text into a proper pandas DataFrame.
- All of it against books.toscrape.com, a real site built for practicing exactly this. Video twenty-three starts a statistics block: descriptive statistics and distributions, the math underneath everything you've plotted and summarized so far. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



