Lesson 26 · Real-World Data Analytics
Python Data Analytics #26: Analysing Real Hashtags & Topic Trends in Python
Video twenty-six of the hundred-video real-world data analytics series. Returning to the real nine hundred forty-four social posts, this time hunting for…
- CourseReal-World Data Analytics
- Lesson26 of 100
- Video27 min
- FormatJupyter notebook · 30 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- social_posts_raw.csv87.9 KB
📓 Full notebook
Download .ipynbData Analytics 100, Video 26: Real Hashtag and Topic Trend Analysis#
- Video twenty-six of the hundred-video real-world data analytics series.
- Returning to the real nine hundred forty-four social posts, this time hunting for real hashtag patterns and clusters.
- Let's get into it.
Part 1: Real Hashtags, Real Patterns#
import pandas as pd
import re
from collections import Counter
from itertools import combinations
import matplotlib.pyplot as plt
posts = pd.read_csv('social_posts_raw.csv')
posts.shape
Part 2: Real Hashtag Extraction#
def extract_hashtags(text):
return re.findall(r'#(\w+)', str(text).lower())
posts['hashtags'] = posts['tweet'].apply(extract_hashtags)
posts[['tweet', 'hashtags']].head(3)
Part 3: Real Hashtag Volume#
all_hashtags = [tag for tags in posts['hashtags'] for tag in tags]
total_mentions = len(all_hashtags)
distinct_hashtags = len(set(all_hashtags))
total_mentions, distinct_hashtags
Part 4: Real Top Twenty Hashtags#
hashtag_counts = Counter(all_hashtags)
top20 = hashtag_counts.most_common(20)
top20
Part 5: Visualizing Real Top Hashtags#
top15_df = pd.DataFrame(hashtag_counts.most_common(15), columns=['hashtag', 'count'])
plt.figure(figsize=(9, 6))
plt.barh(top15_df['hashtag'], top15_df['count'], color='mediumvioletred')
plt.xlabel('Real Mention Count')
plt.title('Real Top 15 Hashtags')
plt.gca().invert_yaxis()
plt.tight_layout()
plt.savefig('top_hashtags.png', dpi=120)
plt.close()
Part 6: Real Hashtag Usage by Content Label#
posts['has_hashtag'] = posts['hashtags'].apply(len) > 0
hashtag_rate_by_label = posts.groupby('label')['has_hashtag'].mean().round(3) * 100
hashtag_rate_by_label
Part 7: Real Hashtags per Post Distribution#
posts['hashtag_count'] = posts['hashtags'].apply(len)
posts['hashtag_count'].describe().round(2)
Part 8: Real Hashtag Co-occurrence#
cooccurrence = Counter()
for tags in posts['hashtags']:
unique_tags = sorted(set(tags))
for tag_a, tag_b in combinations(unique_tags, 2):
cooccurrence[(tag_a, tag_b)] += 1
len(cooccurrence)
Part 9: Real Top Co-occurring Pairs#
top_pairs = cooccurrence.most_common(10)
top_pairs
Part 10: Real Topic Cluster Identification#
financial_tags = {'gold', 'silver', 'forex', 'blog'}
wellness_tags = {'positive', 'thankful', 'affirmation', 'i_am'}
financial_posts = posts[posts['hashtags'].apply(lambda tags: bool(set(tags) & financial_tags))]
wellness_posts = posts[posts['hashtags'].apply(lambda tags: bool(set(tags) & wellness_tags))]
len(financial_posts), len(wellness_posts)
Part 11: Real Cluster Overlap Check#
overlap_ids = set(financial_posts['id']) & set(wellness_posts['id'])
len(overlap_ids)
Part 12: Real Single-Use Hashtags#
single_use = [tag for tag, count in hashtag_counts.items() if count == 1]
single_use_pct = round(len(single_use) / distinct_hashtags * 100, 1)
single_use_pct
Part 13: Real Hashtag Length Analysis#
hashtag_lengths = pd.Series([len(tag) for tag in all_hashtags])
hashtag_lengths.describe().round(2)
Part 14: Real Longest Hashtags Used#
longest_hashtags = sorted(set(all_hashtags), key=len, reverse=True)[:5]
longest_hashtags
Part 15: Real Word Count vs Hashtag Count Correlation#
posts['word_count'] = posts['tweet'].apply(lambda t: len(str(t).split()))
wordcount_hashtag_corr = posts['word_count'].corr(posts['hashtag_count'])
round(wordcount_hashtag_corr, 3)
Part 16: Real Zero-Hashtag Posts#
zero_hashtag_posts = posts[posts['hashtag_count'] == 0]
zero_hashtag_pct = round(len(zero_hashtag_posts) / len(posts) * 100, 1)
zero_hashtag_pct
Part 17: Real Heavy Hashtag Users#
heavy_hashtag_posts = posts[posts['hashtag_count'] >= 5]
len(heavy_hashtag_posts)
heavy_hashtag_posts[['tweet', 'hashtag_count']].head(3)
Part 18: Real Hashtag Diversity by Label#
flagged_tags = [tag for tags in posts.loc[posts['label']==1, 'hashtags'] for tag in tags]
normal_tags = [tag for tags in posts.loc[posts['label']==0, 'hashtags'] for tag in tags]
len(set(flagged_tags)), len(set(normal_tags))
Part 19: Real Top Hashtags in Flagged Posts#
flagged_top = Counter(flagged_tags).most_common(10)
flagged_top
Part 20: Real Hashtags Unique to Each Group#
flagged_only = set(flagged_tags) - set(normal_tags)
normal_only = set(normal_tags) - set(flagged_tags)
len(flagged_only), len(normal_only)
Part 21: Real Trend Across Post Order#
posts_sorted = posts.sort_values('id').reset_index(drop=True)
posts_sorted['batch'] = pd.cut(posts_sorted.index, bins=10, labels=False)
hashtag_rate_by_batch = posts_sorted.groupby('batch')['has_hashtag'].mean().round(3) * 100
hashtag_rate_by_batch
Part 22: Visualizing the Real Batch Trend#
plt.figure(figsize=(9, 5))
hashtag_rate_by_batch.plot(kind='line', marker='o', color='darkcyan')
plt.xlabel('Real Sequential Batch')
plt.ylabel('Real Hashtag Usage Rate (%)')
plt.title('Real Hashtag Usage Rate Across the Dataset')
plt.tight_layout()
plt.savefig('hashtag_trend.png', dpi=120)
plt.close()
Part 23: Real Mention Extraction for Comparison#
posts['mentions'] = posts['tweet'].apply(lambda t: re.findall(r'@(\w+)', str(t).lower()))
posts['mention_count'] = posts['mentions'].apply(len)
hashtag_mention_corr = posts['hashtag_count'].corr(posts['mention_count'])
round(hashtag_mention_corr, 3)
Part 24: Real Hashtag Vocabulary Growth#
seen = set()
new_tag_counts = []
for tags in posts_sorted['hashtags']:
new_tags = set(tags) - seen
new_tag_counts.append(len(new_tags))
seen.update(tags)
sum(new_tag_counts), len(seen)
Part 25: Saving the Real Hashtag Frequency Table#
hashtag_freq_table = pd.DataFrame(hashtag_counts.most_common(), columns=['hashtag', 'count'])
hashtag_freq_table.to_csv('hashtag_trend_frequency.csv', index=False)
reloaded_freq = pd.read_csv('hashtag_trend_frequency.csv')
reloaded_freq.shape[0] == len(hashtag_counts)
Part 26: Saving the Real Co-occurrence Table#
cooccur_table = pd.DataFrame([(a, b, c) for (a, b), c in cooccurrence.most_common()], columns=['hashtag_a', 'hashtag_b', 'count'])
cooccur_table.to_csv('hashtag_cooccurrence.csv', index=False)
cooccur_table.shape
Part 27: Real Sanity Check, Counts are Consistent#
sum(hashtag_counts.values()) == total_mentions
Part 28: Real Sanity Check, Rates in Range#
all((0 <= hashtag_rate_by_label) & (hashtag_rate_by_label <= 100))
Part 29: Real Top Cluster Summary#
cluster_summary = pd.DataFrame({'cluster': ['Financial', 'Wellness'], 'post_count': [len(financial_posts), len(wellness_posts)]})
cluster_summary
Part 30: Real Recap Print#
print(f'Across {len(posts)} real posts, we found {distinct_hashtags} distinct real hashtags, with the top pair {top_pairs[0][0]} genuinely co-occurring {top_pairs[0][1]} times.')
Wrap-Up: What You Learned#
- Regular expressions are the real simplest reliable way to pull hashtags and mentions out of raw real social text.
- Co-occurrence counting, not just frequency counting, is what actually reveals real topic clusters hiding in hashtag data.
- Real hashtag vocabularies follow a real long-tail pattern, a handful of tags dominate while most are used only once.
- Comparing real hashtag usage across a content label can surface real differences in how genuine and flagged content gets tagged.
- Without a real timestamp field, sequential post order is an honest, if imperfect, real proxy for tracking trend movement.
- Next video: real web traffic analytics and attribution, tracing how real visitors reach a site in the first place.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



