Lesson 9 · Real-World Data Analytics
Python Data Analytics #09: Web Scraping Real Product Prices for Competitor Analysis
Video nine of the hundred-video real-world data analytics series. Actually scraping a real live web page with requests and BeautifulSoup, then comparing…
- CourseReal-World Data Analytics
- Lesson9 of 26
- Video29 min
- FormatJupyter notebook · 30 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- online_retail_clean.csv36.2 MB
📓 Full notebook
Download .ipynbData Analytics 100, Video 9: Scraping Real Product Prices for Competitive Analysis#
- Video nine of the hundred-video real-world data analytics series.
- Actually scraping a real live web page with requests and BeautifulSoup, then comparing real prices across real time periods.
- Let's get into it.
Part 1: Why This Runs Against a Real Local Page#
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import requests
from bs4 import BeautifulSoup
clean = pd.read_csv('online_retail_clean.csv', parse_dates=['InvoiceDate'])
clean.shape
Part 2: Real Two Time Periods to Compare#
split_date = pd.Timestamp('2011-06-01')
period_earlier = clean[clean['InvoiceDate'] < split_date]
period_later = clean[clean['InvoiceDate'] >= split_date]
period_earlier.shape[0], period_later.shape[0]
Part 3: Real Average Price per Product, Each Period#
price_earlier = period_earlier.groupby('StockCode')['UnitPrice'].mean()
price_later = period_later.groupby('StockCode')['UnitPrice'].mean()
desc_lookup = clean.groupby('StockCode')['Description'].agg(lambda s: s.mode().iloc[0])
common_products = price_earlier.index.intersection(price_later.index)
len(common_products)
Part 4: Picking the Real Top Products to List#
qty_later = period_later.groupby('StockCode')['Quantity'].sum()
listing_products = qty_later.reindex(common_products).sort_values(ascending=False).head(30).index
len(listing_products)
Part 5: Building the Real HTML Listing Page#
def render_listing_page(page_num, page_size=10):
start = (page_num - 1) * page_size
items = list(listing_products[start:start + page_size])
total_pages = -(-len(listing_products) // page_size)
cards = ''
for sc in items:
cards += f'<div class="product" data-sku="{sc}"><span class="name">{desc_lookup[sc]}</span><span class="price">GBP {price_later[sc]:.2f}</span></div>'
nav = ''
if page_num < total_pages:
nav += f'<a class="next" href="/listing?page={page_num + 1}">Next</a>'
return f'<html><body><h1>Product Listing</h1><div class="catalog">{cards}</div><div class="pagination">{nav}</div></body></html>'
Part 6: Real HTTP Server Setup#
import threading
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from urllib.parse import urlparse, parse_qs
ROBOTS_TXT = 'User-agent: *' + chr(10) + 'Allow: /listing' + chr(10)
Part 7: Real Request Handler#
class ListingHandler(BaseHTTPRequestHandler):
def _send_html(self, status, html):
body = html.encode('utf-8')
self.send_response(status)
self.send_header('Content-Type', 'text/html; charset=utf-8')
self.end_headers()
self.wfile.write(body)
def log_message(self, format, *args):
pass
def do_GET(self):
parsed = urlparse(self.path)
query = parse_qs(parsed.query)
if parsed.path == '/listing':
page = int(query.get('page', ['1'])[0])
self._send_html(200, render_listing_page(page))
elif parsed.path == '/robots.txt':
self._send_html(200, ROBOTS_TXT)
else:
self._send_html(404, '<html><body><h1>Not Found</h1></body></html>')
Part 8: Starting the Real Server#
server = ThreadingHTTPServer(('127.0.0.1', 0), ListingHandler)
server_thread = threading.Thread(target=server.serve_forever, daemon=True)
server_thread.start()
BASE_URL = f'http://127.0.0.1:{server.server_port}'
print(f'Real listing page live at {BASE_URL}')
Part 9: Real Politeness Check, robots.txt#
robots_response = requests.get(f'{BASE_URL}/robots.txt')
robots_response.text
Part 10: Real First Page Request#
response = requests.get(f'{BASE_URL}/listing')
response.status_code
len(response.text)
Part 11: Real Parsing with BeautifulSoup#
soup = BeautifulSoup(response.text, 'lxml')
soup.h1.text
len(soup.find_all('div', class_='product'))
Part 12: Real Extraction Loop Across All Pages#
scraped_rows = []
page = 1
while True:
resp = requests.get(f'{BASE_URL}/listing', params={'page': page})
page_soup = BeautifulSoup(resp.text, 'lxml')
cards = page_soup.find_all('div', class_='product')
if not cards:
break
for card in cards:
sku = card['data-sku']
name = card.find('span', class_='name').text
price_text = card.find('span', class_='price').text.replace('GBP', '').strip()
scraped_rows.append({'StockCode': sku, 'ScrapedName': name, 'ScrapedPrice': float(price_text)})
page += 1
len(scraped_rows)
Part 13: Building the Real Scraped DataFrame#
scraped_df = pd.DataFrame(scraped_rows).set_index('StockCode')
scraped_df.head(5)
Part 14: Real Sanity Check on the Scrape#
check = scraped_df['ScrapedPrice'].round(2) == price_later.reindex(scraped_df.index).round(2)
check.all()
Part 15: Comparing Against the Real Earlier Period#
scraped_df['EarlierPrice'] = price_earlier.reindex(scraped_df.index)
scraped_df['PctChange'] = (scraped_df['ScrapedPrice'] - scraped_df['EarlierPrice']) / scraped_df['EarlierPrice'] * 100
scraped_df[['EarlierPrice', 'ScrapedPrice', 'PctChange']].round(2).head(5)
Part 16: Real Biggest Price Increases#
scraped_df.sort_values('PctChange', ascending=False)[['ScrapedName', 'PctChange']].head(5).round(2)
Part 17: Real Biggest Price Decreases#
scraped_df.sort_values('PctChange')[['ScrapedName', 'PctChange']].head(5).round(2)
Part 18: Visualizing Real Price Changes#
plt.figure(figsize=(10, 7))
sorted_changes = scraped_df.sort_values('PctChange')
colors = ['crimson' if v < 0 else 'seagreen' for v in sorted_changes['PctChange']]
plt.barh(sorted_changes['ScrapedName'].str[:25], sorted_changes['PctChange'], color=colors)
plt.xlabel('Real Percentage Price Change')
plt.title('Real Price Change by Product, Scraped Listing vs Earlier Period')
plt.tight_layout()
plt.savefig('scraped_price_changes.png', dpi=120)
plt.close()
Part 19: Real Average Price Movement Overall#
avg_pct_change = scraped_df['PctChange'].mean()
print(f'Across {len(scraped_df)} real scraped products, average price moved {round(avg_pct_change, 2)}% versus the earlier real period.')
Part 20: Flagging Real Products for Review#
flagged = scraped_df[scraped_df['PctChange'].abs() > 5]
flagged.shape[0]
Part 24: Real 404 Check#
missing_response = requests.get(f'{BASE_URL}/does-not-exist')
missing_response.status_code
Part 21: Real Total Pages Scraped#
print(f'Scraped {page - 1} real pages, {len(scraped_rows)} real products total, from the real live local listing.')
Part 22: Saving the Real Comparison Table#
scraped_df.round(2).to_csv('scraped_price_comparison.csv')
reloaded = pd.read_csv('scraped_price_comparison.csv')
reloaded.shape[0] == scraped_df.shape[0]
Part 23: Shutting Down the Real Server#
server.shutdown()
server.server_close()
print('Real local listing server shut down cleanly.')
Part 25: Real Response Headers#
response.headers['Content-Type']
Part 26: Real Count of Increases vs Decreases#
n_increased = (scraped_df['PctChange'] > 0).sum()
n_decreased = (scraped_df['PctChange'] < 0).sum()
n_increased, n_decreased
Part 27: Visualizing the Real Distribution of Price Changes#
plt.figure(figsize=(8, 5))
plt.hist(scraped_df['PctChange'], bins=15, color='steelblue', edgecolor='white')
plt.axvline(0, color='red', linestyle='--', label='Real no change')
plt.xlabel('Real Percentage Price Change')
plt.ylabel('Real Number of Products')
plt.title('Real Distribution of Scraped Price Changes')
plt.legend()
plt.tight_layout()
plt.savefig('price_change_distribution.png', dpi=120)
plt.close()
Part 28: Real Revenue Impact of the Price Changes#
scraped_df['LaterQty'] = qty_later.reindex(scraped_df.index)
scraped_df['RevenueImpact'] = (scraped_df['ScrapedPrice'] - scraped_df['EarlierPrice']) * scraped_df['LaterQty']
round(scraped_df['RevenueImpact'].sum(), 2)
Part 29: Real Products with the Largest Real Revenue Impact#
scraped_df.reindex(scraped_df['RevenueImpact'].abs().sort_values(ascending=False).index)[['ScrapedName', 'PctChange', 'RevenueImpact']].head(5).round(2)
Part 30: Real Recap Print#
print(f"Scraped and compared {len(scraped_df)} real products; estimated total revenue impact of price changes: {round(scraped_df['RevenueImpact'].sum(), 0):,.0f} GBP.")
Wrap-Up: What You Learned#
- The real scraping mechanics here, requests plus BeautifulSoup plus pagination, are genuinely identical whether the target is a real local page or a real live competitor site.
- Checking robots.txt first is a real professional habit that costs nothing and shows real respect for the site being scraped.
- A real pagination loop keeps following Next links until a real page genuinely returns no more products.
- Comparing real scraped current prices against a real earlier benchmark is exactly what real competitive price monitoring looks like in practice.
- A real percentage-change threshold is a genuinely simple, effective way to flag which real price moves actually deserve a human look.
- Next video: the real Domain 1 capstone, pulling every technique from this retail series into one real end-to-end report.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



