Lesson 62 · Social Media Content Analytics
Viral Video Detection: A Social Media Case Study
In this lesson, we will learn to detect viral videos using real-world YouTube trending data. Viral videos drive massive attention, growth, and business…
- CourseSocial Media Content Analytics
- Lesson62 of 41
- Video24 min
- FormatJupyter notebook · 18 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbViral Video Detection: A Social Media Case Study#
In this lesson, we will learn to detect viral videos using real-world YouTube trending data.
Viral videos drive massive attention, growth, and business results for creators and brands.
You will analyze, visualize, and identify what makes a video 'go viral' and how to spot them early.
We focus on essential metrics like views, likes, and engagement to build actionable insights.
By the end, you can use these methods to benchmark and optimize your content for virality.
import pandas as pd
import numpy as np
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request
import warnings
warnings.filterwarnings('ignore')
SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
def get_yt_service():
api_key = os.environ.get('YOUTUBE_API_KEY')
if api_key:
return build('youtube', 'v3', developerKey=api_key)
if Path('client_secret.json').exists():
creds = None
if Path('token_ro.pickle').exists():
with open('token_ro.pickle', 'rb') as f:
creds = pickle.load(f)
if not creds or not creds.valid:
if creds and creds.expired and creds.refresh_token:
creds.refresh(Request())
else:
flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
creds = flow.run_local_server(port=0)
with open('token_ro.pickle', 'wb') as f:
pickle.dump(creds, f)
return build('youtube', 'v3', credentials=creds)
raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
def fetch_yt_trending(max_results=200, region='US'):
youtube = get_yt_service()
records, token = [], None
while len(records) < max_results:
resp = youtube.videos().list(
part='snippet,statistics',
chart='mostPopular',
regionCode=region,
maxResults=min(50, max_results - len(records)),
pageToken=token
).execute()
for item in resp.get('items', []):
s = item['snippet']; st = item.get('statistics', {})
records.append({
'video_id': item['id'],
'trending_date': pd.Timestamp.today().date(),
'title': s.get('title', ''),
'channel_title': s.get('channelTitle', ''),
'category_id': s.get('categoryId', ''),
'views': int(st.get('viewCount', 0)),
'likes': int(st.get('likeCount', 0)),
'comment_count': int(st.get('commentCount', 0)),
})
token = resp.get('nextPageToken')
if not token: break
return pd.DataFrame(records)
try:
df = fetch_yt_trending()
print('Live trending data:', df.shape)
except Exception as e:
print(f'Falling back to synthetic: {e}')
np.random.seed(42)
n = 1000
df = pd.DataFrame({
'video_id': [f'vid{i}' for i in range(n)],
'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
'title': [f'Video Title {i}' for i in range(n)],
'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
'category_id': np.random.choice([1,2,10,22,24,28], n),
'views': np.random.randint(10000, 5000000, n),
'likes': np.random.randint(100, 200000, n),
'comment_count': np.random.randint(10, 50000, n),
})
print('Synthetic fallback:', df.shape)
print(df.head(3))
What is in a social media trending dataset?#
Each row represents a trending YouTube video, including:
- Unique IDs, trending date, video title, channel, category
- Content performance: views, likes, and comments
These metrics help us measure engagement and popularity.
High values can indicate viral performancebut context matters.
Mistakes to avoid:
- Not accounting for category differences (music, education, etc.)
- Focusing only on raw counts (views) instead of rates or ratios
- Ignoring how fast a video gained popularity
Core social media analytics concepts#
Views: how many times a video was watched.
Likes: indicates audience approval or connection.
Comments: shows two-way audience interaction.
Engagement rate: a metric (likes + comments) divided by views.
Trending date: when a video appeared in top viral lists.
Categories: group similar types of content for fair comparison.
Common mistakes:
- Comparing videos across very different categories
- Thinking all high-view content is truly viral
- Not recognizing different types of engagement (likes vs comments)
# Beginner Example 1: Calculate engagement rate for each trending video
df['engagement_rate'] = (df['likes'] + df['comment_count']) / df['views'] * 100
print(df[['video_id', 'title', 'views', 'likes', 'comment_count', 'engagement_rate']].head())
# Beginner Example 2: Find the single most viewed trending video
top_view = df.loc[df['views'].idxmax()]
print('Top viewed video:', top_view['title'], '| Views:', top_view['views'])
# Beginner Example 3: Average engagement rate across all trending videos
avg_engagement = df['engagement_rate'].mean()
print(f'Average engagement rate: {avg_engagement:.2f}%')
# Intermediate Example 1: Find videos with engagement rate above 10%
viral_candidates = df[df['engagement_rate'] > 10]
print('Number of high-engagement videos:', len(viral_candidates))
print(viral_candidates[['title', 'views', 'engagement_rate']].head(3))
# Intermediate Example 2: Which channels appear most often in trending videos?
top_channels = df['channel_title'].value_counts().head(3)
print('Most featured channels on trending list:')
print(top_channels)
# Intermediate Example 3: Analyze engagement rate by category
category_stats = df.groupby('category_id')['engagement_rate'].mean().sort_values(ascending=False)
print('Average engagement rate by category:')
print(category_stats)
# Advanced Example 1: Time-series analysis - engagement trends over time
df['trending_date'] = pd.to_datetime(df['trending_date'])
daily_trend = df.groupby('trending_date')['engagement_rate'].mean()
import matplotlib.pyplot as plt
plt.figure(figsize=(10,4))
plt.plot(daily_trend.index, daily_trend.values, label='Avg Engagement Rate')
plt.xlabel('Trending Date')
plt.ylabel('Avg Engagement Rate (%)')
plt.title('Engagement Rate Trend Over Time')
plt.legend()
plt.tight_layout()
plt.show()
# Advanced Example 2: Detecting outlier viral videos using z-score
from scipy.stats import zscore
df['views_z'] = zscore(df['views'])
viral_outliers = df[df['views_z'] > 3]
print('Videos with extreme virality (z > 3):')
print(viral_outliers[['title', 'views', 'engagement_rate', 'views_z']].head())
# Advanced Example 3: Calculate and rank videos by viral index (views * engagement rate)
df['viral_index'] = df['views'] * df['engagement_rate']
top_viral = df.nlargest(5, 'viral_index')
print('Top 5 potential viral videos by viral index:')
print(top_viral[['title', 'views', 'engagement_rate', 'viral_index']])
# Error handling: missing or zero engagement values
has_engagement_issue = df['views'] == 0
if has_engagement_issue.any():
print('Warning: Some videos have zero views (impossible on trending)!')
print(df[has_engagement_issue])
else:
print('All trending videos have at least one view. Data integrity looks good.')
# Debugging: What if aggregation logic is incorrect?
wrong_sum = df.groupby('category_id')['views'].sum().sum()
true_sum = df['views'].sum()
if wrong_sum != true_sum:
print('Check: Grouped sum matches total sum OK')
else:
print('Possible mistake: double-counted or missing data in aggregation!')
# Error: Interpreting engagement rate wrongly (likes/views instead of (likes+comments)/views)
df['wrong_rate'] = df['likes'] / df['views'] * 100
mean_wrong = df['wrong_rate'].mean()
mean_right = df['engagement_rate'].mean()
print(f'Wrong formula avg: {mean_wrong:.2f}% | Correct: {mean_right:.2f}%')
# Error: Wrong grouping logic (should group by trending date and category together)
df['date_cat'] = df['trending_date'].astype(str) + '-' + df['category_id'].astype(str)
grouped = df.groupby('date_cat')['views'].sum()
print('Sample combined date/category group:')
print(grouped.head(2))
Best practices for social media content analytics#
Benchmark performance within content categories, not just globally.
Segment audiences and analyze by channel for deeper insights.
Track trends over timelook for patterns, not just peaks.
Use consistent and clearly defined metrics (explain your math!).
Set and revisit benchmarks using up-to-date, comparable data.
Optimize content by learning from what works (and what does not work) on trending lists.
Common analytics patterns for viral video detection#
Identify videos that outperform average by large margins.
Use engagement rates and 'viral index' for deeper ranking.
Combine visualization and statistics for better insights.
Regularly review and adapt thresholds as platforms evolve.
Document when and why you flag a video as viral.
# End-to-end Example: From raw data to viral video recommendation
viral = df[df['views_z'] > 3]
if not viral.empty:
result = viral.sort_values('viral_index', ascending=False).iloc[0]
print('Recommend doubling down on content like:', result['title'])
print(f'Viral signals: {result["views"]} views and {result["engagement_rate"]:.2f}% engagement rate')
else:
print('No strongly viral videos detected. Recommend analyzing trends and trying again later.')
Congratulations on completing the viral video detection case study!#
You analyzed trends, engagement rates, and detected outliers in trending data.
Try these skills on your own content or with new datasets.
For more tutorials, see the official YouTube Data API and analytics guides.
Want to boost your analytics superpowers? Subscribe on YouTube to stay updated.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



