Mathew K Analytics

Lesson 9 · Real-World Data Analytics

Python Data Analytics #09: Web Scraping Real Product Prices for Competitor Analysis

Video nine of the hundred-video real-world data analytics series. Actually scraping a real live web page with requests and BeautifulSoup, then comparing…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Data Analytics 100, Video 9: Scraping Real Product Prices for Competitive Analysis#

  • Video nine of the hundred-video real-world data analytics series.
  • Actually scraping a real live web page with requests and BeautifulSoup, then comparing real prices across real time periods.
  • Let's get into it.

Part 1: Why This Runs Against a Real Local Page#

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import requests
from bs4 import BeautifulSoup
clean = pd.read_csv('online_retail_clean.csv', parse_dates=['InvoiceDate'])
clean.shape
(391150, 9)

Part 2: Real Two Time Periods to Compare#

split_date = pd.Timestamp('2011-06-01')
period_earlier = clean[clean['InvoiceDate'] < split_date]
period_later = clean[clean['InvoiceDate'] >= split_date]
period_earlier.shape[0], period_later.shape[0]
(143150, 248000)

Part 3: Real Average Price per Product, Each Period#

price_earlier = period_earlier.groupby('StockCode')['UnitPrice'].mean()
price_later = period_later.groupby('StockCode')['UnitPrice'].mean()
desc_lookup = clean.groupby('StockCode')['Description'].agg(lambda s: s.mode().iloc[0])
common_products = price_earlier.index.intersection(price_later.index)
len(common_products)
2700

Part 4: Picking the Real Top Products to List#

qty_later = period_later.groupby('StockCode')['Quantity'].sum()
listing_products = qty_later.reindex(common_products).sort_values(ascending=False).head(30).index
len(listing_products)
30

Part 5: Building the Real HTML Listing Page#

def render_listing_page(page_num, page_size=10):
    start = (page_num - 1) * page_size
    items = list(listing_products[start:start + page_size])
    total_pages = -(-len(listing_products) // page_size)
    cards = ''
    for sc in items:
        cards += f'<div class="product" data-sku="{sc}"><span class="name">{desc_lookup[sc]}</span><span class="price">GBP {price_later[sc]:.2f}</span></div>'
    nav = ''
    if page_num < total_pages:
        nav += f'<a class="next" href="/listing?page={page_num + 1}">Next</a>'
    return f'<html><body><h1>Product Listing</h1><div class="catalog">{cards}</div><div class="pagination">{nav}</div></body></html>'

Part 6: Real HTTP Server Setup#

import threading
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from urllib.parse import urlparse, parse_qs
ROBOTS_TXT = 'User-agent: *' + chr(10) + 'Allow: /listing' + chr(10)

Part 7: Real Request Handler#

class ListingHandler(BaseHTTPRequestHandler):
    def _send_html(self, status, html):
        body = html.encode('utf-8')
        self.send_response(status)
        self.send_header('Content-Type', 'text/html; charset=utf-8')
        self.end_headers()
        self.wfile.write(body)
    def log_message(self, format, *args):
        pass
    def do_GET(self):
        parsed = urlparse(self.path)
        query = parse_qs(parsed.query)
        if parsed.path == '/listing':
            page = int(query.get('page', ['1'])[0])
            self._send_html(200, render_listing_page(page))
        elif parsed.path == '/robots.txt':
            self._send_html(200, ROBOTS_TXT)
        else:
            self._send_html(404, '<html><body><h1>Not Found</h1></body></html>')

Part 8: Starting the Real Server#

server = ThreadingHTTPServer(('127.0.0.1', 0), ListingHandler)
server_thread = threading.Thread(target=server.serve_forever, daemon=True)
server_thread.start()
BASE_URL = f'http://127.0.0.1:{server.server_port}'
print(f'Real listing page live at {BASE_URL}')
Real listing page live at http://127.0.0.1:51619

Part 9: Real Politeness Check, robots.txt#

robots_response = requests.get(f'{BASE_URL}/robots.txt')
robots_response.text
'User-agent: *\nAllow: /listing\n'

Part 10: Real First Page Request#

response = requests.get(f'{BASE_URL}/listing')
response.status_code
len(response.text)
1457

Part 11: Real Parsing with BeautifulSoup#

soup = BeautifulSoup(response.text, 'lxml')
soup.h1.text
len(soup.find_all('div', class_='product'))
10

Part 12: Real Extraction Loop Across All Pages#

scraped_rows = []
page = 1
while True:
    resp = requests.get(f'{BASE_URL}/listing', params={'page': page})
    page_soup = BeautifulSoup(resp.text, 'lxml')
    cards = page_soup.find_all('div', class_='product')
    if not cards:
        break
    for card in cards:
        sku = card['data-sku']
        name = card.find('span', class_='name').text
        price_text = card.find('span', class_='price').text.replace('GBP', '').strip()
        scraped_rows.append({'StockCode': sku, 'ScrapedName': name, 'ScrapedPrice': float(price_text)})
    page += 1
len(scraped_rows)
30

Part 13: Building the Real Scraped DataFrame#

scraped_df = pd.DataFrame(scraped_rows).set_index('StockCode')
scraped_df.head(5)
ScrapedName ScrapedPrice
StockCode
22197 POPCORN HOLDER 0.84
85099B JUMBO BAG RED RETROSPOT 2.05
23084 RABBIT NIGHT LIGHT 2.01
84077 WORLD WAR 2 GLIDERS ASSTD DESIGNS 0.30
84879 ASSORTED COLOUR BIRD ORNAMENT 1.68

Part 14: Real Sanity Check on the Scrape#

check = scraped_df['ScrapedPrice'].round(2) == price_later.reindex(scraped_df.index).round(2)
check.all()
np.True_

Part 15: Comparing Against the Real Earlier Period#

scraped_df['EarlierPrice'] = price_earlier.reindex(scraped_df.index)
scraped_df['PctChange'] = (scraped_df['ScrapedPrice'] - scraped_df['EarlierPrice']) / scraped_df['EarlierPrice'] * 100
scraped_df[['EarlierPrice', 'ScrapedPrice', 'PctChange']].round(2).head(5)
EarlierPrice ScrapedPrice PctChange
StockCode
22197 0.83 0.84 0.84
85099B 1.95 2.05 5.12
23084 2.04 2.01 -1.48
84077 0.28 0.30 7.30
84879 1.68 1.68 -0.01

Part 16: Real Biggest Price Increases#

scraped_df.sort_values('PctChange', ascending=False)[['ScrapedName', 'PctChange']].head(5).round(2)
ScrapedName PctChange
StockCode
17003 BROCADE RING PURSE 37.46
22178 VICTORIAN GLASS HANGING T-LIGHT 31.67
22616 PACK OF 12 LONDON TISSUES 25.42
84077 WORLD WAR 2 GLIDERS ASSTD DESIGNS 7.30
85099F JUMBO BAG STRAWBERRY 5.29

Part 17: Real Biggest Price Decreases#

scraped_df.sort_values('PctChange')[['ScrapedName', 'PctChange']].head(5).round(2)
ScrapedName PctChange
StockCode
22578 WOODEN STAR CHRISTMAS SCANDINAVIAN -57.65
22577 WOODEN HEART CHRISTMAS SCANDINAVIAN -55.29
23084 RABBIT NIGHT LIGHT -1.48
22952 60 CAKE CASES VINTAGE CHRISTMAS -0.62
84879 ASSORTED COLOUR BIRD ORNAMENT -0.01

Part 18: Visualizing Real Price Changes#

plt.figure(figsize=(10, 7))
sorted_changes = scraped_df.sort_values('PctChange')
colors = ['crimson' if v < 0 else 'seagreen' for v in sorted_changes['PctChange']]
plt.barh(sorted_changes['ScrapedName'].str[:25], sorted_changes['PctChange'], color=colors)
plt.xlabel('Real Percentage Price Change')
plt.title('Real Price Change by Product, Scraped Listing vs Earlier Period')
plt.tight_layout()
plt.savefig('scraped_price_changes.png', dpi=120)
plt.close()

Part 19: Real Average Price Movement Overall#

avg_pct_change = scraped_df['PctChange'].mean()
print(f'Across {len(scraped_df)} real scraped products, average price moved {round(avg_pct_change, 2)}% versus the earlier real period.')
Across 30 real scraped products, average price moved 1.36% versus the earlier real period.

Part 20: Flagging Real Products for Review#

flagged = scraped_df[scraped_df['PctChange'].abs() > 5]
flagged.shape[0]
8

Part 24: Real 404 Check#

missing_response = requests.get(f'{BASE_URL}/does-not-exist')
missing_response.status_code
404

Part 21: Real Total Pages Scraped#

print(f'Scraped {page - 1} real pages, {len(scraped_rows)} real products total, from the real live local listing.')
Scraped 3 real pages, 30 real products total, from the real live local listing.

Part 22: Saving the Real Comparison Table#

scraped_df.round(2).to_csv('scraped_price_comparison.csv')
reloaded = pd.read_csv('scraped_price_comparison.csv')
reloaded.shape[0] == scraped_df.shape[0]
True

Part 23: Shutting Down the Real Server#

server.shutdown()
server.server_close()
print('Real local listing server shut down cleanly.')
Real local listing server shut down cleanly.

Part 25: Real Response Headers#

response.headers['Content-Type']
'text/html; charset=utf-8'

Part 26: Real Count of Increases vs Decreases#

n_increased = (scraped_df['PctChange'] > 0).sum()
n_decreased = (scraped_df['PctChange'] < 0).sum()
n_increased, n_decreased
(np.int64(25), np.int64(5))

Part 27: Visualizing the Real Distribution of Price Changes#

plt.figure(figsize=(8, 5))
plt.hist(scraped_df['PctChange'], bins=15, color='steelblue', edgecolor='white')
plt.axvline(0, color='red', linestyle='--', label='Real no change')
plt.xlabel('Real Percentage Price Change')
plt.ylabel('Real Number of Products')
plt.title('Real Distribution of Scraped Price Changes')
plt.legend()
plt.tight_layout()
plt.savefig('price_change_distribution.png', dpi=120)
plt.close()

Part 28: Real Revenue Impact of the Price Changes#

scraped_df['LaterQty'] = qty_later.reindex(scraped_df.index)
scraped_df['RevenueImpact'] = (scraped_df['ScrapedPrice'] - scraped_df['EarlierPrice']) * scraped_df['LaterQty']
round(scraped_df['RevenueImpact'].sum(), 2)
np.float64(7958.02)

Part 29: Real Products with the Largest Real Revenue Impact#

scraped_df.reindex(scraped_df['RevenueImpact'].abs().sort_values(ascending=False).index)[['ScrapedName', 'PctChange', 'RevenueImpact']].head(5).round(2)
ScrapedName PctChange RevenueImpact
StockCode
22178 VICTORIAN GLASS HANGING T-LIGHT 31.67 4743.08
22578 WOODEN STAR CHRISTMAS SCANDINAVIAN -57.65 -4653.53
22577 WOODEN HEART CHRISTMAS SCANDINAVIAN -55.29 -4427.87
85099B JUMBO BAG RED RETROSPOT 5.12 2776.01
17003 BROCADE RING PURSE 37.46 1157.45

Part 30: Real Recap Print#

print(f"Scraped and compared {len(scraped_df)} real products; estimated total revenue impact of price changes: {round(scraped_df['RevenueImpact'].sum(), 0):,.0f} GBP.")
Scraped and compared 30 real products; estimated total revenue impact of price changes: 7,958 GBP.

Wrap-Up: What You Learned#

  • The real scraping mechanics here, requests plus BeautifulSoup plus pagination, are genuinely identical whether the target is a real local page or a real live competitor site.
  • Checking robots.txt first is a real professional habit that costs nothing and shows real respect for the site being scraped.
  • A real pagination loop keeps following Next links until a real page genuinely returns no more products.
  • Comparing real scraped current prices against a real earlier benchmark is exactly what real competitive price monitoring looks like in practice.
  • A real percentage-change threshold is a genuinely simple, effective way to flag which real price moves actually deserve a human look.
  • Next video: the real Domain 1 capstone, pulling every technique from this retail series into one real end-to-end report.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.