Mathew K Analytics

Lesson 5 · Data Science Projects

Web Scraping and Visualization Techniques for Real Estate Data Analysis

Introduction to scraping real estate data using Python. See how to extract and visualize property data from a sample website. Learn why web scraping is a…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Real Estate Property Scraping and Visualization#

  • Introduction to scraping real estate data using Python.
  • See how to extract and visualize property data from a sample website.
  • Learn why web scraping is a valuable skill for data mining.
  • We will use pandas, requests, and beautifulsoup libraries.
  • By the end, you will be able to collect and plot property information.

What is web scraping?#

  • Web scraping means collecting data from web pages automatically.
  • It helps extract useful information from websites for analysis.
  • This is often used in real estate, finance, news, and many other industries.
  • The Python libraries we use will make scraping easier.
  • Let us see how it is done with real examples.
# Suppress warnings for a cleaner output
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
# Data setup
import requests
from bs4 import BeautifulSoup
url = 'https://webscraper.io/test-sites/e-commerce/static/computers'
html = requests.get(url).text
soup = BeautifulSoup(html, 'html.parser')
cards = soup.select('.thumbnail')
records = [{'title': c.select_one('.title').get_text(strip=True), 'price': c.select_one('.price').get_text(strip=True)} for c in cards]
print(records[:5])
[{'title': 'Asus ROG Strix...', 'price': '$1769'}, {'title': 'Acer Aspire 3...', 'price': '$408.98'}, {'title': 'Lenovo IdeaPad...', 'price': '$1212.16'}]
# Convert records to pandas DataFrame for easy analysis
import pandas as pd
df = pd.DataFrame(records)
print(df.shape)
print(df.head())
(3, 2)
               title     price
0  Asus ROG Strix...     $1769
1   Acer Aspire 3...   $408.98
2  Lenovo IdeaPad...  $1212.16

Quick look at the data#

  • Let us review what columns we have.
  • Each row is a property or product, almost like a real estate listing.
  • Columns help us filter, sort, and plot data later.
# Check column names and missing values
print(df.columns)
print(df.isnull().sum())
Index(['title', 'price'], dtype='object')
title    0
price    0
dtype: int64
# Clean price column (remove currency signs and convert to float)
df['price'] = df['price'].str.replace('$', '').astype(float)

Why preprocess the data?#

  • Clean data avoids mistakes in charts and calculations.
  • It is a key step in any data mining process.
  • Always check types and fix problems early.
# Show the basic statistics for the price
print(df['price'].describe())
count       3.000000
mean     1130.046667
std       683.718180
min       408.980000
25%       810.570000
50%      1212.160000
75%      1490.580000
max      1769.000000
Name: price, dtype: float64
# Plot a simple price distribution
import matplotlib.pyplot as plt
plt.figure(figsize=(8, 4))
plt.hist(df['price'], bins=10, color='skyblue', edgecolor='black')
plt.title('Distribution of Scraped Property Prices')
plt.xlabel('Price')
plt.ylabel('Count')
plt.show()
No description has been provided for this image

Explore listings with the highest prices#

  • Sometimes it is useful to focus on properties with the highest value.
  • This can show what luxury features or locations appear in pricier listings.
# Display top 5 most expensive listings
print(df.sort_values('price', ascending=False).head())
               title    price
0  Asus ROG Strix...  1769.00
2  Lenovo IdeaPad...  1212.16
1   Acer Aspire 3...   408.98
# Search for a property keyword using input
keyword = input("Enter a keyword to search titles: ")
matches = df[df['title'].str.lower().str.contains(keyword.lower())]
print(matches)
Empty DataFrame
Columns: [title, price]
Index: []

Data Mining: What insights can we extract?#

  • Web scraping lets us answer questions such as:
  • Which product types are most expensive or cheapest?
  • How many listings mention a certain brand or keyword?
  • What is the typical price range?
  • These basic insights help investors, buyers, or analysts make decisions.
# Count how many listings contain the word 'laptop'
laptop_count = df['title'].str.lower().str.contains('laptop').sum()
print(f"Number of listings with 'laptop': {laptop_count}")
Number of listings with 'laptop': 0
# Create a bar plot of the top 5 most common words in titles
from collections import Counter
words = " ".join(df['title']).lower().split()
stopwords = set(['and','or','the','with','for','of','by','to','in'])
filtered = [w for w in words if w not in stopwords]
top5 = Counter(filtered).most_common(5)
labels, counts = zip(*top5)
plt.figure(figsize=(7,4))
plt.bar(labels, counts, color='orange')
plt.title('Top 5 Words in Property Titles')
plt.ylabel('Count')
plt.show()
No description has been provided for this image

Mini project: Try scraping another site#

  • Find another e-commerce or listing site with a safe terms of use.
  • Try changing the url variable to a new site and adapt the code.
  • Can you scrape product names and prices from a different category?
  • Remember, always check the website rules before scraping.
# Export your cleaned data for later use
df.to_csv('real_estate_scraped.csv', index=False)
print('Saved as real_estate_scraped.csv')
Saved as real_estate_scraped.csv

What else can you do next?#

  • Try grouping by keywords or sorting differently.
  • Experiment with more advanced plots like scatter charts.
  • The more you explore, the more insights you find.
  • Subscribe to the channel for more hands on Python data mining lessons.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.