Mathew K Analytics

Lesson 26 · Real-World Data Analytics

Python Data Analytics #26: Analysing Real Hashtags & Topic Trends in Python

Video twenty-six of the hundred-video real-world data analytics series. Returning to the real nine hundred forty-four social posts, this time hunting for…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Data Analytics 100, Video 26: Real Hashtag and Topic Trend Analysis#

  • Video twenty-six of the hundred-video real-world data analytics series.
  • Returning to the real nine hundred forty-four social posts, this time hunting for real hashtag patterns and clusters.
  • Let's get into it.

Part 1: Real Hashtags, Real Patterns#

import pandas as pd
import re
from collections import Counter
from itertools import combinations
import matplotlib.pyplot as plt
posts = pd.read_csv('social_posts_raw.csv')
posts.shape
(944, 3)

Part 2: Real Hashtag Extraction#

def extract_hashtags(text):
    return re.findall(r'#(\w+)', str(text).lower())
posts['hashtags'] = posts['tweet'].apply(extract_hashtags)
posts[['tweet', 'hashtags']].head(3)
tweet hashtags
0 @user when a father is dysfunctional and is s... [run]
1 @user @user thanks for #lyft credit i can't us... [lyft, disapointed, getthanked]
2 bihday your majesty []

Part 3: Real Hashtag Volume#

all_hashtags = [tag for tags in posts['hashtags'] for tag in tags]
total_mentions = len(all_hashtags)
distinct_hashtags = len(set(all_hashtags))
total_mentions, distinct_hashtags
(2220, 1514)

Part 4: Real Top Twenty Hashtags#

hashtag_counts = Counter(all_hashtags)
top20 = hashtag_counts.most_common(20)
top20
[('love', 46),
 ('positive', 28),
 ('thankful', 19),
 ('healthy', 17),
 ('model', 16),
 ('blog', 15),
 ('silver', 15),
 ('gold', 15),
 ('forex', 13),
 ('altwaystoheal', 12),
 ('orlando', 11),
 ('i_am', 11),
 ('affirmation', 11),
 ('fun', 11),
 ('smile', 11),
 ('fathersday', 10),
 ('weekend', 9),
 ('summer', 9),
 ('happiness', 9),
 ('healing', 9)]

Part 5: Visualizing Real Top Hashtags#

top15_df = pd.DataFrame(hashtag_counts.most_common(15), columns=['hashtag', 'count'])
plt.figure(figsize=(9, 6))
plt.barh(top15_df['hashtag'], top15_df['count'], color='mediumvioletred')
plt.xlabel('Real Mention Count')
plt.title('Real Top 15 Hashtags')
plt.gca().invert_yaxis()
plt.tight_layout()
plt.savefig('top_hashtags.png', dpi=120)
plt.close()

Part 6: Real Hashtag Usage by Content Label#

posts['has_hashtag'] = posts['hashtags'].apply(len) > 0
hashtag_rate_by_label = posts.groupby('label')['has_hashtag'].mean().round(3) * 100
hashtag_rate_by_label
label
0    73.1
1    78.9
Name: has_hashtag, dtype: float64

Part 7: Real Hashtags per Post Distribution#

posts['hashtag_count'] = posts['hashtags'].apply(len)
posts['hashtag_count'].describe().round(2)
count    944.00
mean       2.35
std        2.46
min        0.00
25%        0.00
50%        2.00
75%        4.00
max       20.00
Name: hashtag_count, dtype: float64

Part 8: Real Hashtag Co-occurrence#

cooccurrence = Counter()
for tags in posts['hashtags']:
    unique_tags = sorted(set(tags))
    for tag_a, tag_b in combinations(unique_tags, 2):
        cooccurrence[(tag_a, tag_b)] += 1
len(cooccurrence)
3991

Part 9: Real Top Co-occurring Pairs#

top_pairs = cooccurrence.most_common(10)
top_pairs
[(('positive', 'thankful'), 16),
 (('blog', 'silver'), 15),
 (('blog', 'gold'), 14),
 (('gold', 'silver'), 14),
 (('blog', 'forex'), 13),
 (('forex', 'gold'), 13),
 (('forex', 'silver'), 13),
 (('affirmation', 'i_am'), 11),
 (('affirmation', 'positive'), 11),
 (('i_am', 'positive'), 11)]

Part 10: Real Topic Cluster Identification#

financial_tags = {'gold', 'silver', 'forex', 'blog'}
wellness_tags = {'positive', 'thankful', 'affirmation', 'i_am'}
financial_posts = posts[posts['hashtags'].apply(lambda tags: bool(set(tags) & financial_tags))]
wellness_posts = posts[posts['hashtags'].apply(lambda tags: bool(set(tags) & wellness_tags))]
len(financial_posts), len(wellness_posts)
(15, 31)

Part 11: Real Cluster Overlap Check#

overlap_ids = set(financial_posts['id']) & set(wellness_posts['id'])
len(overlap_ids)
0

Part 12: Real Single-Use Hashtags#

single_use = [tag for tag, count in hashtag_counts.items() if count == 1]
single_use_pct = round(len(single_use) / distinct_hashtags * 100, 1)
single_use_pct
83.6

Part 13: Real Hashtag Length Analysis#

hashtag_lengths = pd.Series([len(tag) for tag in all_hashtags])
hashtag_lengths.describe().round(2)
count    2220.00
mean        7.37
std         3.69
min         1.00
25%         5.00
50%         7.00
75%         9.00
max        27.00
dtype: float64

Part 14: Real Longest Hashtags Used#

longest_hashtags = sorted(set(all_hashtags), key=len, reverse=True)[:5]
longest_hashtags
['golfstrengthandconditioning',
 'internationaldayofyoga2016',
 'thinkbigsundaywithmarshað',
 'whenrealtorscompeteyouwin',
 'wheresallthenaturalphotos']

Part 15: Real Word Count vs Hashtag Count Correlation#

posts['word_count'] = posts['tweet'].apply(lambda t: len(str(t).split()))
wordcount_hashtag_corr = posts['word_count'].corr(posts['hashtag_count'])
round(wordcount_hashtag_corr, 3)
np.float64(-0.078)

Part 16: Real Zero-Hashtag Posts#

zero_hashtag_posts = posts[posts['hashtag_count'] == 0]
zero_hashtag_pct = round(len(zero_hashtag_posts) / len(posts) * 100, 1)
zero_hashtag_pct
26.5

Part 17: Real Heavy Hashtag Users#

heavy_hashtag_posts = posts[posts['hashtag_count'] >= 5]
len(heavy_hashtag_posts)
heavy_hashtag_posts[['tweet', 'hashtag_count']].head(3)
tweet hashtag_count
7 the next school year is the year for exams.ðŸ˜... 7
8 we won!!! love the land!!! #allin #cavs #champ... 5
10 ↝ #ireland consumer price index (mom) climb... 5

Part 18: Real Hashtag Diversity by Label#

flagged_tags = [tag for tags in posts.loc[posts['label']==1, 'hashtags'] for tag in tags]
normal_tags = [tag for tags in posts.loc[posts['label']==0, 'hashtags'] for tag in tags]
len(set(flagged_tags)), len(set(normal_tags))
(124, 1406)

Part 19: Real Top Hashtags in Flagged Posts#

flagged_top = Counter(flagged_tags).most_common(10)
flagged_top
[('trump', 6),
 ('sjw', 4),
 ('libtard', 3),
 ('liberal', 3),
 ('politics', 3),
 ('hate', 3),
 ('bigotry', 3),
 ('tcot', 2),
 ('helpcovedolphins', 2),
 ('race', 2)]

Part 20: Real Hashtags Unique to Each Group#

flagged_only = set(flagged_tags) - set(normal_tags)
normal_only = set(normal_tags) - set(flagged_tags)
len(flagged_only), len(normal_only)
(108, 1390)

Part 21: Real Trend Across Post Order#

posts_sorted = posts.sort_values('id').reset_index(drop=True)
posts_sorted['batch'] = pd.cut(posts_sorted.index, bins=10, labels=False)
hashtag_rate_by_batch = posts_sorted.groupby('batch')['has_hashtag'].mean().round(3) * 100
hashtag_rate_by_batch
batch
0    73.7
1    78.7
2    75.5
3    71.6
4    79.8
5    70.2
6    72.6
7    73.4
8    75.5
9    64.2
Name: has_hashtag, dtype: float64

Part 22: Visualizing the Real Batch Trend#

plt.figure(figsize=(9, 5))
hashtag_rate_by_batch.plot(kind='line', marker='o', color='darkcyan')
plt.xlabel('Real Sequential Batch')
plt.ylabel('Real Hashtag Usage Rate (%)')
plt.title('Real Hashtag Usage Rate Across the Dataset')
plt.tight_layout()
plt.savefig('hashtag_trend.png', dpi=120)
plt.close()

Part 23: Real Mention Extraction for Comparison#

posts['mentions'] = posts['tweet'].apply(lambda t: re.findall(r'@(\w+)', str(t).lower()))
posts['mention_count'] = posts['mentions'].apply(len)
hashtag_mention_corr = posts['hashtag_count'].corr(posts['mention_count'])
round(hashtag_mention_corr, 3)
np.float64(-0.144)

Part 24: Real Hashtag Vocabulary Growth#

seen = set()
new_tag_counts = []
for tags in posts_sorted['hashtags']:
    new_tags = set(tags) - seen
    new_tag_counts.append(len(new_tags))
    seen.update(tags)
sum(new_tag_counts), len(seen)
(1514, 1514)

Part 25: Saving the Real Hashtag Frequency Table#

hashtag_freq_table = pd.DataFrame(hashtag_counts.most_common(), columns=['hashtag', 'count'])
hashtag_freq_table.to_csv('hashtag_trend_frequency.csv', index=False)
reloaded_freq = pd.read_csv('hashtag_trend_frequency.csv')
reloaded_freq.shape[0] == len(hashtag_counts)
True

Part 26: Saving the Real Co-occurrence Table#

cooccur_table = pd.DataFrame([(a, b, c) for (a, b), c in cooccurrence.most_common()], columns=['hashtag_a', 'hashtag_b', 'count'])
cooccur_table.to_csv('hashtag_cooccurrence.csv', index=False)
cooccur_table.shape
(3991, 3)

Part 27: Real Sanity Check, Counts are Consistent#

sum(hashtag_counts.values()) == total_mentions
True

Part 28: Real Sanity Check, Rates in Range#

all((0 <= hashtag_rate_by_label) & (hashtag_rate_by_label <= 100))
True

Part 29: Real Top Cluster Summary#

cluster_summary = pd.DataFrame({'cluster': ['Financial', 'Wellness'], 'post_count': [len(financial_posts), len(wellness_posts)]})
cluster_summary
cluster post_count
0 Financial 15
1 Wellness 31

Part 30: Real Recap Print#

print(f'Across {len(posts)} real posts, we found {distinct_hashtags} distinct real hashtags, with the top pair {top_pairs[0][0]} genuinely co-occurring {top_pairs[0][1]} times.')
Across 944 real posts, we found 1514 distinct real hashtags, with the top pair ('positive', 'thankful') genuinely co-occurring 16 times.

Wrap-Up: What You Learned#

  • Regular expressions are the real simplest reliable way to pull hashtags and mentions out of raw real social text.
  • Co-occurrence counting, not just frequency counting, is what actually reveals real topic clusters hiding in hashtag data.
  • Real hashtag vocabularies follow a real long-tail pattern, a handful of tags dominate while most are used only once.
  • Comparing real hashtag usage across a content label can surface real differences in how genuine and flagged content gets tagged.
  • Without a real timestamp field, sequential post order is an honest, if imperfect, real proxy for tracking trend movement.
  • Next video: real web traffic analytics and attribution, tracing how real visitors reach a site in the first place.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.