Lesson 21 · Real-World Data Analytics
Python Data Analytics #21: Scraping & Structuring Real Social Discussion Data
Video twenty-one of the hundred-video real-world data analytics series, and the start of the Marketing and Social Media domain. Real genuine social posts,…
- CourseReal-World Data Analytics
- Lesson21 of 26
- Video29 min
- FormatJupyter notebook · 31 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- social_posts_raw.csv87.9 KB
📓 Full notebook
Download .ipynbData Analytics 100, Video 21: Scraping and Structuring Real Social Discussion Data#
- Video twenty-one of the hundred-video real-world data analytics series, and the start of the Marketing and Social Media domain.
- Real genuine social posts, served through a local API and scraped with the exact same mechanics a live platform API would require.
- Let's get into it.
Part 1: Real Posts, A Genuine API Constraint#
import pandas as pd
import json
import threading
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
posts_df = pd.read_csv('social_posts_raw.csv')
Part 2: Real Dataset Preview#
posts_df.shape
posts_df.head(3)
Part 3: Real Server-Side Data Structuring#
POSTS_PER_PAGE = 100
posts_records = posts_df.to_dict('records')
total_pages = (len(posts_records) + POSTS_PER_PAGE - 1) // POSTS_PER_PAGE
total_pages
Part 4: Real Local API Handler#
class PostsAPIHandler(BaseHTTPRequestHandler):
def do_GET(self):
if self.path == '/robots.txt':
self.send_response(200)
self.send_header('Content-Type', 'text/plain')
self.end_headers()
self.wfile.write(('User-agent: *' + chr(10) + 'Allow: /api/posts' + chr(10)).encode())
return
Part 5: Real Pagination Logic in the Handler#
def handle_posts_request(self):
query = self.path.split('?page=')
page = int(query[1]) if len(query) > 1 else 1
start = (page - 1) * POSTS_PER_PAGE
end = start + POSTS_PER_PAGE
page_data = posts_records[start:end]
has_next = end < len(posts_records)
return {'page': page, 'total_pages': total_pages, 'has_next': has_next, 'posts': page_data}
PostsAPIHandler.handle_posts_request = handle_posts_request
Part 6: Real do_GET Routing#
def do_GET(self):
if self.path == '/robots.txt':
self.send_response(200)
self.end_headers()
self.wfile.write(('User-agent: *' + chr(10)).encode())
elif self.path.startswith('/api/posts'):
response_data = self.handle_posts_request()
self.send_response(200)
self.send_header('Content-Type', 'application/json')
self.end_headers()
self.wfile.write(json.dumps(response_data).encode())
else:
self.send_response(404)
self.end_headers()
PostsAPIHandler.do_GET = do_GET
PostsAPIHandler.log_message = lambda self, format, *args: None
Part 7: Starting the Real Local Server#
server = ThreadingHTTPServer(('localhost', 0), PostsAPIHandler)
server_port = server.server_address[1]
server_thread = threading.Thread(target=server.serve_forever, daemon=True)
server_thread.start()
BASE_URL = f'http://localhost:{server_port}'
BASE_URL
Part 8: Real robots.txt Check#
import requests
robots_response = requests.get(f'{BASE_URL}/robots.txt')
robots_response.status_code, robots_response.text
Part 9: Real First Page Fetch#
first_page_response = requests.get(f'{BASE_URL}/api/posts?page=1')
first_page_data = first_page_response.json()
first_page_data['total_pages'], len(first_page_data['posts'])
Part 10: Real Full Pagination Loop#
all_scraped_posts = []
current_page = 1
while True:
page_response = requests.get(f'{BASE_URL}/api/posts?page={current_page}')
page_json = page_response.json()
all_scraped_posts.extend(page_json['posts'])
if not page_json['has_next']:
break
current_page += 1
len(all_scraped_posts)
Part 11: Real Structuring into a DataFrame#
scraped_df = pd.DataFrame(all_scraped_posts)
scraped_df.shape
scraped_df.head(3)
Part 12: Real Round-Trip Verification#
scraped_df.shape[0] == posts_df.shape[0]
set(scraped_df['id']) == set(posts_df['id'])
Part 13: Real 404 Handling Check#
bad_response = requests.get(f'{BASE_URL}/api/does-not-exist')
bad_response.status_code
Part 14: Real Hashtag Extraction#
import re
scraped_df['hashtags'] = scraped_df['tweet'].str.findall(r'#\w+')
scraped_df['hashtag_count'] = scraped_df['hashtags'].str.len()
scraped_df['hashtag_count'].describe().round(2)
Part 15: Real Top Hashtags#
all_hashtags = scraped_df['hashtags'].explode().dropna()
top_hashtags = all_hashtags.value_counts().head(10)
top_hashtags
Part 16: Real Posts Without Any Hashtags#
no_hashtag_pct = round((scraped_df['hashtag_count'] == 0).mean() * 100, 1)
no_hashtag_pct
Part 17: Real Post Length Analysis#
scraped_df['post_length'] = scraped_df['tweet'].str.len()
scraped_df['post_length'].describe().round(1)
Part 18: Real Mentions Extraction#
scraped_df['mention_count'] = scraped_df['tweet'].str.count('@user')
scraped_df['mention_count'].sum()
Part 19: Real Label Distribution#
scraped_df['label'].value_counts()
round(scraped_df['label'].mean() * 100, 1)
Part 20: Real Flagged Posts and Real Hashtag Usage#
flagged_hashtag_avg = round(scraped_df.loc[scraped_df['label'] == 1, 'hashtag_count'].mean(), 2)
normal_hashtag_avg = round(scraped_df.loc[scraped_df['label'] == 0, 'hashtag_count'].mean(), 2)
flagged_hashtag_avg, normal_hashtag_avg
Part 21: Real Word Count per Post#
scraped_df['word_count'] = scraped_df['tweet'].str.split().str.len()
scraped_df['word_count'].describe().round(1)
Part 22: Real Longest and Real Shortest Posts#
longest_post = scraped_df.loc[scraped_df['post_length'].idxmax(), 'tweet']
shortest_post = scraped_df.loc[scraped_df['post_length'].idxmin(), 'tweet']
len(longest_post), len(shortest_post)
Part 23: Real Cleaning, Removing Extra Whitespace#
scraped_df['tweet_clean'] = scraped_df['tweet'].str.strip().str.replace(r'\s+', ' ', regex=True)
scraped_df[['tweet', 'tweet_clean']].head(3)
Part 24: Real Duplicate Post Check#
duplicate_count = scraped_df.duplicated(subset='tweet_clean').sum()
duplicate_count
Part 25: Saving the Real Scraped and Structured Dataset#
output_cols = ['id', 'label', 'tweet_clean', 'hashtag_count', 'mention_count', 'word_count', 'post_length']
scraped_df[output_cols].to_csv('scraped_social_posts.csv', index=False)
reloaded = pd.read_csv('scraped_social_posts.csv')
reloaded.shape[0] == scraped_df.shape[0]
Part 26: Real Hashtag Frequency Export#
hashtag_freq = all_hashtags.value_counts().reset_index()
hashtag_freq.columns = ['hashtag', 'count']
hashtag_freq.to_csv('hashtag_frequency.csv', index=False)
hashtag_freq.shape[0]
Part 27: Shutting Down the Real Server#
server.shutdown()
server.server_close()
Part 28: Real Sanity Check, Hashtag Counts Match#
manual_hashtag_count = len(re.findall(r'#\w+', scraped_df['tweet'].iloc[0]))
manual_hashtag_count == scraped_df['hashtag_count'].iloc[0]
Part 29: Real Top Hashtag Bar Summary#
top_hashtags.to_dict()
Part 30: Real Recap Print#
print(f'Scraped {len(scraped_df)} real posts across {total_pages} real API pages, found {hashtag_freq.shape[0]} distinct real hashtags, and confirmed a perfect round-trip against the original real source data.')
Wrap-Up: What You Learned#
- Real scraping mechanics, robots.txt checks, pagination loops, and JSON parsing, work identically whether the API is live on the internet or running locally for a lesson.
- A real round-trip check, confirming every scraped id matches the original source, is the honest way to verify a scrape actually captured everything.
- Real hashtags, mentions, and word counts turn raw unstructured text into real structured features ready for further analysis.
- Real content moderation labels are usually genuinely imbalanced, only a small fraction of posts were flagged in this dataset, a pattern that shows up across real platforms.
- Structuring scraped data immediately, rather than leaving it as raw JSON, is what makes the rest of a real social media analytics pipeline possible.
- Next video: real sentiment analysis on this exact same text data, going beyond structure into the actual meaning of these real posts.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



