Mathew K Analytics

Lesson 21 · Real-World Data Analytics

Python Data Analytics #21: Scraping & Structuring Real Social Discussion Data

Video twenty-one of the hundred-video real-world data analytics series, and the start of the Marketing and Social Media domain. Real genuine social posts,…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Data Analytics 100, Video 21: Scraping and Structuring Real Social Discussion Data#

  • Video twenty-one of the hundred-video real-world data analytics series, and the start of the Marketing and Social Media domain.
  • Real genuine social posts, served through a local API and scraped with the exact same mechanics a live platform API would require.
  • Let's get into it.

Part 1: Real Posts, A Genuine API Constraint#

import pandas as pd
import json
import threading
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
posts_df = pd.read_csv('social_posts_raw.csv')

Part 2: Real Dataset Preview#

posts_df.shape
posts_df.head(3)
id label tweet
0 1 0 @user when a father is dysfunctional and is s...
1 2 0 @user @user thanks for #lyft credit i can't us...
2 3 0 bihday your majesty

Part 3: Real Server-Side Data Structuring#

POSTS_PER_PAGE = 100
posts_records = posts_df.to_dict('records')
total_pages = (len(posts_records) + POSTS_PER_PAGE - 1) // POSTS_PER_PAGE
total_pages
10

Part 4: Real Local API Handler#

class PostsAPIHandler(BaseHTTPRequestHandler):
    def do_GET(self):
        if self.path == '/robots.txt':
            self.send_response(200)
            self.send_header('Content-Type', 'text/plain')
            self.end_headers()
            self.wfile.write(('User-agent: *' + chr(10) + 'Allow: /api/posts' + chr(10)).encode())
            return

Part 5: Real Pagination Logic in the Handler#

def handle_posts_request(self):
        query = self.path.split('?page=')
        page = int(query[1]) if len(query) > 1 else 1
        start = (page - 1) * POSTS_PER_PAGE
        end = start + POSTS_PER_PAGE
        page_data = posts_records[start:end]
        has_next = end < len(posts_records)
        return {'page': page, 'total_pages': total_pages, 'has_next': has_next, 'posts': page_data}
PostsAPIHandler.handle_posts_request = handle_posts_request

Part 6: Real do_GET Routing#

def do_GET(self):
    if self.path == '/robots.txt':
        self.send_response(200)
        self.end_headers()
        self.wfile.write(('User-agent: *' + chr(10)).encode())
    elif self.path.startswith('/api/posts'):
        response_data = self.handle_posts_request()
        self.send_response(200)
        self.send_header('Content-Type', 'application/json')
        self.end_headers()
        self.wfile.write(json.dumps(response_data).encode())
    else:
        self.send_response(404)
        self.end_headers()
PostsAPIHandler.do_GET = do_GET
PostsAPIHandler.log_message = lambda self, format, *args: None

Part 7: Starting the Real Local Server#

server = ThreadingHTTPServer(('localhost', 0), PostsAPIHandler)
server_port = server.server_address[1]
server_thread = threading.Thread(target=server.serve_forever, daemon=True)
server_thread.start()
BASE_URL = f'http://localhost:{server_port}'
BASE_URL
'http://localhost:49679'

Part 8: Real robots.txt Check#

import requests
robots_response = requests.get(f'{BASE_URL}/robots.txt')
robots_response.status_code, robots_response.text
(200, 'User-agent: *\n')

Part 9: Real First Page Fetch#

first_page_response = requests.get(f'{BASE_URL}/api/posts?page=1')
first_page_data = first_page_response.json()
first_page_data['total_pages'], len(first_page_data['posts'])
(10, 100)

Part 10: Real Full Pagination Loop#

all_scraped_posts = []
current_page = 1
while True:
    page_response = requests.get(f'{BASE_URL}/api/posts?page={current_page}')
    page_json = page_response.json()
    all_scraped_posts.extend(page_json['posts'])
    if not page_json['has_next']:
        break
    current_page += 1
len(all_scraped_posts)
944

Part 11: Real Structuring into a DataFrame#

scraped_df = pd.DataFrame(all_scraped_posts)
scraped_df.shape
scraped_df.head(3)
id label tweet
0 1 0 @user when a father is dysfunctional and is s...
1 2 0 @user @user thanks for #lyft credit i can't us...
2 3 0 bihday your majesty

Part 12: Real Round-Trip Verification#

scraped_df.shape[0] == posts_df.shape[0]
set(scraped_df['id']) == set(posts_df['id'])
True

Part 13: Real 404 Handling Check#

bad_response = requests.get(f'{BASE_URL}/api/does-not-exist')
bad_response.status_code
404

Part 14: Real Hashtag Extraction#

import re
scraped_df['hashtags'] = scraped_df['tweet'].str.findall(r'#\w+')
scraped_df['hashtag_count'] = scraped_df['hashtags'].str.len()
scraped_df['hashtag_count'].describe().round(2)
count    944.00
mean       2.35
std        2.46
min        0.00
25%        0.00
50%        2.00
75%        4.00
max       20.00
Name: hashtag_count, dtype: float64

Part 15: Real Top Hashtags#

all_hashtags = scraped_df['hashtags'].explode().dropna()
top_hashtags = all_hashtags.value_counts().head(10)
top_hashtags
hashtags
#love             46
#positive         28
#thankful         19
#healthy          17
#model            16
#gold             15
#silver           15
#blog             15
#forex            13
#altwaystoheal    12
Name: count, dtype: int64

Part 16: Real Posts Without Any Hashtags#

no_hashtag_pct = round((scraped_df['hashtag_count'] == 0).mean() * 100, 1)
no_hashtag_pct
np.float64(26.5)

Part 17: Real Post Length Analysis#

scraped_df['post_length'] = scraped_df['tweet'].str.len()
scraped_df['post_length'].describe().round(1)
count    944.0
mean      84.6
std       28.7
min       18.0
25%       63.0
50%       86.0
75%      107.0
max      151.0
Name: post_length, dtype: float64

Part 18: Real Mentions Extraction#

scraped_df['mention_count'] = scraped_df['tweet'].str.count('@user')
scraped_df['mention_count'].sum()
np.int64(515)

Part 19: Real Label Distribution#

scraped_df['label'].value_counts()
round(scraped_df['label'].mean() * 100, 1)
np.float64(7.5)

Part 20: Real Flagged Posts and Real Hashtag Usage#

flagged_hashtag_avg = round(scraped_df.loc[scraped_df['label'] == 1, 'hashtag_count'].mean(), 2)
normal_hashtag_avg = round(scraped_df.loc[scraped_df['label'] == 0, 'hashtag_count'].mean(), 2)
flagged_hashtag_avg, normal_hashtag_avg
(np.float64(2.15), np.float64(2.37))

Part 21: Real Word Count per Post#

scraped_df['word_count'] = scraped_df['tweet'].str.split().str.len()
scraped_df['word_count'].describe().round(1)
count    944.0
mean      13.1
std        5.4
min        3.0
25%        9.0
50%       12.0
75%       17.0
max       30.0
Name: word_count, dtype: float64

Part 22: Real Longest and Real Shortest Posts#

longest_post = scraped_df.loc[scraped_df['post_length'].idxmax(), 'tweet']
shortest_post = scraped_df.loc[scraped_df['post_length'].idxmin(), 'tweet']
len(longest_post), len(shortest_post)
(151, 18)

Part 23: Real Cleaning, Removing Extra Whitespace#

scraped_df['tweet_clean'] = scraped_df['tweet'].str.strip().str.replace(r'\s+', ' ', regex=True)
scraped_df[['tweet', 'tweet_clean']].head(3)
tweet tweet_clean
0 @user when a father is dysfunctional and is s... @user when a father is dysfunctional and is so...
1 @user @user thanks for #lyft credit i can't us... @user @user thanks for #lyft credit i can't us...
2 bihday your majesty bihday your majesty

Part 24: Real Duplicate Post Check#

duplicate_count = scraped_df.duplicated(subset='tweet_clean').sum()
duplicate_count
np.int64(23)

Part 25: Saving the Real Scraped and Structured Dataset#

output_cols = ['id', 'label', 'tweet_clean', 'hashtag_count', 'mention_count', 'word_count', 'post_length']
scraped_df[output_cols].to_csv('scraped_social_posts.csv', index=False)
reloaded = pd.read_csv('scraped_social_posts.csv')
reloaded.shape[0] == scraped_df.shape[0]
True

Part 26: Real Hashtag Frequency Export#

hashtag_freq = all_hashtags.value_counts().reset_index()
hashtag_freq.columns = ['hashtag', 'count']
hashtag_freq.to_csv('hashtag_frequency.csv', index=False)
hashtag_freq.shape[0]
1514

Part 27: Shutting Down the Real Server#

server.shutdown()
server.server_close()

Part 28: Real Sanity Check, Hashtag Counts Match#

manual_hashtag_count = len(re.findall(r'#\w+', scraped_df['tweet'].iloc[0]))
manual_hashtag_count == scraped_df['hashtag_count'].iloc[0]
np.True_

Part 29: Real Top Hashtag Bar Summary#

top_hashtags.to_dict()
{'#love': 46,
 '#positive': 28,
 '#thankful': 19,
 '#healthy': 17,
 '#model': 16,
 '#gold': 15,
 '#silver': 15,
 '#blog': 15,
 '#forex': 13,
 '#altwaystoheal': 12}

Part 30: Real Recap Print#

print(f'Scraped {len(scraped_df)} real posts across {total_pages} real API pages, found {hashtag_freq.shape[0]} distinct real hashtags, and confirmed a perfect round-trip against the original real source data.')
Scraped 944 real posts across 10 real API pages, found 1514 distinct real hashtags, and confirmed a perfect round-trip against the original real source data.

Wrap-Up: What You Learned#

  • Real scraping mechanics, robots.txt checks, pagination loops, and JSON parsing, work identically whether the API is live on the internet or running locally for a lesson.
  • A real round-trip check, confirming every scraped id matches the original source, is the honest way to verify a scrape actually captured everything.
  • Real hashtags, mentions, and word counts turn raw unstructured text into real structured features ready for further analysis.
  • Real content moderation labels are usually genuinely imbalanced, only a small fraction of posts were flagged in this dataset, a pattern that shows up across real platforms.
  • Structuring scraped data immediately, rather than leaving it as raw JSON, is what makes the rest of a real social media analytics pipeline possible.
  • Next video: real sentiment analysis on this exact same text data, going beyond structure into the actual meaning of these real posts.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.