Mathew K Analytics

Lesson 22 · Data analytics zero to hero

Web Scraping with Python: Step by Step | Data Analytics #22

Video twenty-two of the 30-part series: pulling real structured data straight out of a web page's HTML, for the many real sites that don't offer an API.…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Data Analytics Zero to Hero, Video 22: Web Scraping for Data Analysts#

  • Video twenty-two of the 30-part series: pulling real structured data straight out of a web page's HTML, for the many real sites that don't offer an API.
  • We're using books.toscrape.com, a real site built specifically and legally for practicing web scraping.
  • Let's jump straight in.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • Install requests and BeautifulSoup if you haven't already: pip install requests beautifulsoup4.
  • You'll need an active internet connection for this video, since every cell here fetches a real, live page.
  • Always check a real site's terms of service and robots.txt before scraping it for anything beyond practice.

Part 1: Fetching and Parsing Real HTML#

import requests
from bs4 import BeautifulSoup

url = 'http://books.toscrape.com/'
response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')
print(response.status_code)
print(soup.title.text)
200

    All products | Books to Scrape - Sandbox

Part 2: Finding Real Elements#

books = soup.find_all('article', class_='product_pod')
print(len(books))
print(books[0].prettify()[:500])
20
<article class="product_pod">
 <div class="image_container">
  <a href="catalogue/a-light-in-the-attic_1000/index.html">
   <img alt="A Light in the Attic" class="thumbnail" src="media/cache/2c/da/2cdad67c44b002e7ead0cc35693c0e8b.jpg"/>
  </a>
 </div>
 <p class="star-rating Three">
  <i class="icon-star">
  </i>
  <i class="icon-star">
  </i>
  <i class="icon-star">
  </i>
  <i class="icon-star">
  </i>
  <i class="icon-star">
  </i>
 </p>
 <h3>
  <a href="catalogue/a-light-in-the-attic_1000/ind
first_book = books[0]
title = first_book.h3.a['title']
price = first_book.find('p', class_='price_color').text
availability = first_book.find('p', class_='instock availability').text.strip()
print(title, price, availability)
A Light in the Attic £51.77 In stock

Part 3: Building a Real Dataset from One Page#

rows = []
for book in books:
    rows.append({
        'title': book.h3.a['title'],
        'price': book.find('p', class_='price_color').text,
        'rating': book.p['class'][1],
        'in_stock': 'In stock' in book.find('p', class_='instock availability').text
    })
print(len(rows))
print(rows[0])
20
{'title': 'A Light in the Attic', 'price': '£51.77', 'rating': 'Three', 'in_stock': True}

Part 4: Following Real Pagination#

import time

all_rows = []
for page_num in range(1, 4):
    page_url = f'http://books.toscrape.com/catalogue/page-{page_num}.html'
    page_response = requests.get(page_url)
    page_soup = BeautifulSoup(page_response.text, 'html.parser')
    page_books = page_soup.find_all('article', class_='product_pod')
    for book in page_books:
        all_rows.append({
            'title': book.h3.a['title'],
            'price': book.find('p', class_='price_color').text
        })
    time.sleep(1)

print(len(all_rows))
60

Part 5: Turning Scraped Data into a DataFrame#

import pandas as pd

books_df = pd.DataFrame(all_rows)
books_df['price'] = books_df['price'].str.replace('', '', regex=False).astype(float)
print(books_df.head())
print(books_df['price'].describe())
---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
Cell In[6], line 4
      1 import pandas as pd
      3 books_df = pd.DataFrame(all_rows)
----> 4 books_df['price'] = books_df['price'].str.replace('', '', regex=False).astype(float)
      5 print(books_df.head())
      6 print(books_df['price'].describe())

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\core\generic.py:6662, in NDFrame.astype(self, dtype, copy, errors)
   6656     results = [
   6657         ser.astype(dtype, copy=copy, errors=errors) for _, ser in self.items()
   6658     ]
   6660 else:
   6661     # else, only a single dtype is given
-> 6662     new_data = self._mgr.astype(dtype=dtype, copy=copy, errors=errors)
   6663     res = self._constructor_from_mgr(new_data, axes=new_data.axes)
   6664     return res.__finalize__(self, method="astype")

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\core\internals\managers.py:430, in BaseBlockManager.astype(self, dtype, copy, errors)
    427 elif using_copy_on_write():
    428     copy = False
--> 430 return self.apply(
    431     "astype",
    432     dtype=dtype,
    433     copy=copy,
    434     errors=errors,
    435     using_cow=using_copy_on_write(),
    436 )

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\core\internals\managers.py:363, in BaseBlockManager.apply(self, f, align_keys, **kwargs)
    361         applied = b.apply(f, **kwargs)
    362     else:
--> 363         applied = getattr(b, f)(**kwargs)
    364     result_blocks = extend_blocks(applied, result_blocks)
    366 out = type(self).from_blocks(result_blocks, self.axes)

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\core\internals\blocks.py:784, in Block.astype(self, dtype, copy, errors, using_cow, squeeze)
    781         raise ValueError("Can not squeeze with more than one column.")
    782     values = values[0, :]  # type: ignore[call-overload]
--> 784 new_values = astype_array_safe(values, dtype, copy=copy, errors=errors)
    786 new_values = maybe_coerce_values(new_values)
    788 refs = None

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\core\dtypes\astype.py:237, in astype_array_safe(values, dtype, copy, errors)
    234     dtype = dtype.numpy_dtype
    236 try:
--> 237     new_values = astype_array(values, dtype, copy=copy)
    238 except (ValueError, TypeError):
    239     # e.g. _astype_nansafe can fail on object-dtype of strings
    240     #  trying to convert to float
    241     if errors == "ignore":

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\core\dtypes\astype.py:182, in astype_array(values, dtype, copy)
    179     values = values.astype(dtype, copy=copy)
    181 else:
--> 182     values = _astype_nansafe(values, dtype, copy=copy)
    184 # in pandas we don't store numpy str dtypes, so convert to object
    185 if isinstance(dtype, np.dtype) and issubclass(values.dtype.type, str):

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\core\dtypes\astype.py:133, in _astype_nansafe(arr, dtype, copy, skipna)
    129     raise ValueError(msg)
    131 if copy or arr.dtype == object or dtype == object:
    132     # Explicit copy, or required since NumPy can't view from / to object.
--> 133     return arr.astype(dtype, copy=True)
    135 return arr.astype(dtype, copy=copy)

ValueError: could not convert string to float: '£51.77'

Wrap-Up: What You Learned#

  • Fetching real HTML with requests, and parsing it into a searchable tree with BeautifulSoup.
  • find and find_all, for locating real elements by tag and class, and pulling out text or attribute values.
  • Looping over matched elements to build a real dataset, one dictionary per item.
  • Following real pagination across multiple pages, with a polite delay between requests.
  • Cleaning and converting scraped text into a proper pandas DataFrame.
  • All of it against books.toscrape.com, a real site built for practicing exactly this. Video twenty-three starts a statistics block: descriptive statistics and distributions, the math underneath everything you've plotted and summarized so far. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.