Lesson 48 · Social Media Content Analytics
Detecting Viral Spikes and Trends in Social Media Data
In this lesson, we explore how to use Python to detect viral spikes and trending content across social media platforms. Understanding these patterns helps…
- CourseSocial Media Content Analytics
- Lesson48 of 41
- Video31 min
- FormatJupyter notebook · 21 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbDetecting Viral Spikes and Trends in Social Media Data#
- In this lesson, we explore how to use Python to detect viral spikes and trending content across social media platforms.
- Understanding these patterns helps content creators and businesses adapt strategies for maximum audience impact.
- You will learn to analyze time series data, engagement metrics, and discover techniques to identify viral moments and sustained trends.
- By the end, you will be able to spot both sudden spikes and long-term growth patterns in real social media datasets.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')
Core Concepts: Social Media Analytics for Viral Detection#
- Social media datasets contain data on posts, videos, audience engagements, and timestamps.
- Key metrics include views, likes, comments, shares, click-through rate (CTR), and watch time.
- Engagement rates show how active your audience is compared to total views.
- A spike in engagement or views often signals viral or trending content.
- Beginners often mistake total engagement for high performance without considering posting time or audience size.
- Always contextualize spikes with timelines and audience reach.
# Beginner Example 1: Load a simple synthetic social media time series dataset
np.random.seed(42)
n_days = 30
dates = pd.date_range('2023-01-01', periods=n_days, freq='D')
views = np.random.randint(1000, 20000, n_days)
likes = (views * np.random.uniform(0.04, 0.09, n_days)).astype(int)
df_ts = pd.DataFrame({
'date': dates,
'views': views,
'likes': likes,
'engagement_rate': np.round(likes / views * 100, 2)
})
print(df_ts.head(3))
# Beginner Example 2: Visualize views over time to spot spikes
plt.figure(figsize=(9,4))
plt.plot(df_ts['date'], df_ts['views'], marker='o')
plt.title('Daily Views Over Time')
plt.xlabel('Date')
plt.ylabel('Views')
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
# Beginner Example 3: Calculate and visualize engagement rate trends
plt.figure(figsize=(9,4))
plt.plot(df_ts['date'], df_ts['engagement_rate'], color='orange', marker='o')
plt.title('Daily Engagement Rate (%)')
plt.xlabel('Date')
plt.ylabel('Engagement Rate (%)')
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
# Beginner Example 4: Identify the day with highest views (potential viral spike)
max_views_row = df_ts.loc[df_ts['views'].idxmax()]
print('Day with highest views:')
print(max_views_row)
# Intermediate Example 1: Load a real YouTube Trending Videos dataset or synthetic fallback
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request
SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
def get_yt_service():
api_key = os.environ.get('YOUTUBE_API_KEY')
if api_key:
return build('youtube', 'v3', developerKey=api_key)
if Path('client_secret.json').exists():
creds = None
if Path('token_ro.pickle').exists():
with open('token_ro.pickle', 'rb') as f:
creds = pickle.load(f)
if not creds or not creds.valid:
if creds and creds.expired and creds.refresh_token:
creds.refresh(Request())
else:
flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
creds = flow.run_local_server(port=0)
with open('token_ro.pickle', 'wb') as f:
pickle.dump(creds, f)
return build('youtube', 'v3', credentials=creds)
raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
def fetch_yt_trending(max_results=200, region='US'):
youtube = get_yt_service()
records, token = [], None
while len(records) < max_results:
resp = youtube.videos().list(
part='snippet,statistics',
chart='mostPopular',
regionCode=region,
maxResults=min(50, max_results - len(records)),
pageToken=token
).execute()
for item in resp.get('items', []):
s = item['snippet']; st = item.get('statistics', {})
records.append({
'video_id': item['id'],
'trending_date': pd.Timestamp.today().date(),
'title': s.get('title', ''),
'channel_title': s.get('channelTitle', ''),
'category_id': s.get('categoryId', ''),
'views': int(st.get('viewCount', 0)),
'likes': int(st.get('likeCount', 0)),
'comment_count': int(st.get('commentCount', 0)),
})
token = resp.get('nextPageToken')
if not token: break
return pd.DataFrame(records)
try:
df_trend = fetch_yt_trending()
print('Loaded real trending data:', df_trend.shape)
except Exception as e:
print(f'Falling back to synthetic: {e}')
np.random.seed(42)
n = 1000
df_trend = pd.DataFrame({
'video_id': [f'vid{i}' for i in range(n)],
'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
'title': [f'Video Title {i}' for i in range(n)],
'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
'category_id': np.random.choice([1,2,10,22,24,28], n),
'views': np.random.randint(10000, 5000000, n),
'likes': np.random.randint(100, 200000, n),
'comment_count': np.random.randint(10, 50000, n),
})
print('Synthetic fallback:', df_trend.shape)
print(df_trend.head(3))
# Intermediate Example 2: Find top 5 trending videos with most views
top_videos = df_trend.sort_values('views', ascending=False).head(5)
print('Top 5 trending videos by view count:')
print(top_videos[['title', 'channel_title', 'views', 'likes', 'comment_count']])
# Intermediate Example 3: Visualize trending video view distribution
plt.figure(figsize=(8,4))
sns.histplot(df_trend['views'], bins=30, color='skyblue')
plt.title('Distribution of Trending Video Views')
plt.xlabel('Views per Video')
plt.ylabel('Count of Videos')
plt.tight_layout()
plt.show()
# Intermediate Example 4: Calculate engagement rate for each trending video
df_trend['engagement_rate'] = np.where(
df_trend['views'] > 0,
(df_trend['likes'] + df_trend['comment_count']) / df_trend['views'] * 100,
np.nan
)
print('Engagement rate stats:')
print(df_trend['engagement_rate'].describe())
# Intermediate Example 5: Identify potential viral videos using engagement and views
viral_videos = df_trend[(df_trend['views'] > df_trend['views'].quantile(0.95)) &
(df_trend['engagement_rate'] > df_trend['engagement_rate'].quantile(0.95))]
print('Potential viral videos (top 5% for both views and engagement rate):')
print(viral_videos[['title', 'channel_title', 'views', 'likes', 'comment_count', 'engagement_rate']])
# Advanced Example 1: Detect viral spikes in a full year social media time series
np.random.seed(42)
n_days_full = 365
dates_full = pd.date_range('2023-01-01', periods=n_days_full)
views_full = np.random.randint(1000, 50000, n_days_full)
spikes = np.random.choice([0, 0, 0, 1], n_days_full, p=[0.96, 0.01, 0.01, 0.02])
views_full += spikes * np.random.randint(80000, 200000, n_days_full)
likes_full = (views_full * np.random.uniform(0.03, 0.12, n_days_full)).astype(int)
df_full = pd.DataFrame({
'date': dates_full,
'views': views_full,
'likes': likes_full,
'engagement_rate': np.round(likes_full / views_full * 100, 2)
})
plt.figure(figsize=(14,5))
plt.plot(df_full['date'], df_full['views'], label='Views', color='blue')
plt.title('Full Year Daily Views with Synthetic Viral Spikes')
plt.xlabel('Date')
plt.ylabel('Views')
plt.tight_layout()
plt.show()
# Advanced Example 2: Automatic viral spike detection algorithm
threshold = df_full['views'].mean() + 3 * df_full['views'].std()
viral_days = df_full[df_full['views'] > threshold]
print(f'Days identified as viral (spikes > mean + 3*std): {viral_days.shape[0]}')
print(viral_days[['date', 'views', 'likes', 'engagement_rate']])
# Advanced Example 3: Explore trends across content categories in trending videos
category_trends = df_trend.groupby('category_id').agg({'views': 'mean', 'likes': 'mean', 'engagement_rate': 'mean'}).sort_values('views', ascending=False)
print('Average views, likes, and engagement rate by category:')
print(category_trends)
# Advanced Example 4: Visualize engagement trend with a moving average
df_full['views_ma7'] = df_full['views'].rolling(window=7).mean()
plt.figure(figsize=(12,5))
plt.plot(df_full['date'], df_full['views'], alpha=0.4, label='Daily Views')
plt.plot(df_full['date'], df_full['views_ma7'], color='red', linewidth=2, label='7-Day Moving Average')
plt.title('Viral Trendlines: Daily Views and 7-Day Moving Average')
plt.xlabel('Date')
plt.ylabel('Views')
plt.legend()
plt.tight_layout()
plt.show()
# Error Handling Example 1: Missing values in engagement columns
df_trend_corrupt = df_trend.copy()
df_trend_corrupt.loc[0, 'likes'] = np.nan
df_trend_corrupt.loc[1, 'comment_count'] = np.nan
df_trend_corrupt['engagement_fixed'] = (
df_trend_corrupt['likes'].fillna(0) + df_trend_corrupt['comment_count'].fillna(0)
) / df_trend_corrupt['views'] * 100
print('Engagement rate with missing values (rows 0 and 1):')
print(df_trend_corrupt[['likes', 'comment_count', 'views', 'engagement_fixed']].head(3))
# Error Handling Example 2: Incorrect metric aggregation
grouped = df_trend.groupby('category_id').sum(numeric_only=True)
print('Incorrect metric aggregation (should not sum engagement rates):')
print(grouped[['views', 'likes', 'comment_count', 'engagement_rate']].head())
# Error Handling Example 3: Misinterpreting ratios like CTR
np.random.seed(42)
dummy_views = np.random.randint(1, 20, 10)
dummy_clicks = np.random.randint(0, 5, 10)
ctr = dummy_clicks / dummy_views * 100
print('Example of click-through rates (CTR) correctly calculated:')
print('Views:', dummy_views)
print('Clicks:', dummy_clicks)
print('CTR (%):', np.round(ctr,2))
# Error Handling Example 4: Wrong grouping logic for content categories
wrong_group = df_trend.groupby(['channel_title', 'category_id']).mean(numeric_only=True)
print('Grouped by BOTH channel and category - can make small groups:')
print(wrong_group.head(7))
Best Practices for Social Media Trend Analytics#
- Always visualize your data before calculating spikes or trends.
- Use moving averages to smooth out daily or weekly fluctuations.
- Benchmark performance relative to similar content, not only across your own posts.
- Segment audience or content by reasonable categories (e.g., genre, time, channel).
- Define viral thresholds clearly (e.g., above percentile, or multiple of std).
- Check all metric calculations: use means for rates, not sums.
- Document anomalies and check for missing values before reporting results.
# Analytics Pattern: Finding best posting days of the week
df_full['weekday'] = df_full['date'].dt.day_name()
weekday_means = df_full.groupby('weekday')['views'].mean().reindex([
'Monday', 'Tuesday', 'Wednesday', 'Thursday', 'Friday', 'Saturday', 'Sunday'
])
print('Average views by weekday:')
print(weekday_means)
# Analytics Pattern: Consistent metric definition using a helper function
def calc_engagement(df):
return (df['likes'] + df.get('comment_count', 0)) / df['views'] * 100
df_trend['practice_engagement'] = calc_engagement(df_trend)
print('First five engagement calculations:')
print(df_trend[['likes', 'comment_count', 'views', 'practice_engagement']].head())
# End-to-end Problem: Identify, explain, and recommend content strategy for a detected viral spike
detected_viral = df_full.loc[df_full['views'] > threshold]
for idx, row in detected_viral.iterrows():
print(f'Date: {row.date.date()} | Views: {row.views} | Engagement Rate: {row.engagement_rate:.2f}%')
print('---')
if not detected_viral.empty:
rec_day = detected_viral.iloc[0]['date'].day_name()
print(f'Strategy: On {rec_day}s, consider scheduling more high-impact content, and analyze what drove engagement on these spike dates.')
else:
print('No strong viral spikes found: Review your content plan for new tactics.')
Wrap-Up: Viral Trend Detection Cheatsheet#
- Use time series plots to visually spot spikes.
- Compute percentiles or std deviations to quantify viral moments.
- Combine views with engagement rate to confirm real popularity.
- Avoid common analysis mistakes (wrong grouping, missing values, incorrect metric sums).
- Test your strategy with both synthetic and real data.
- Try explaining every detected spike in plain language.
- Next: Download this notebook or watch our full video course on YouTube!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



