Mathew K Analytics

Lesson 62 · Social Media Content Analytics

Viral Video Detection: A Social Media Case Study

In this lesson, we will learn to detect viral videos using real-world YouTube trending data. Viral videos drive massive attention, growth, and business…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Viral Video Detection: A Social Media Case Study#

  • In this lesson, we will learn to detect viral videos using real-world YouTube trending data.

  • Viral videos drive massive attention, growth, and business results for creators and brands.

  • You will analyze, visualize, and identify what makes a video 'go viral' and how to spot them early.

  • We focus on essential metrics like views, likes, and engagement to build actionable insights.

  • By the end, you can use these methods to benchmark and optimize your content for virality.

import pandas as pd
import numpy as np
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request
import warnings
warnings.filterwarnings('ignore')
SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']

def get_yt_service():
    api_key = os.environ.get('YOUTUBE_API_KEY')
    if api_key:
        return build('youtube', 'v3', developerKey=api_key)
    if Path('client_secret.json').exists():
        creds = None
        if Path('token_ro.pickle').exists():
            with open('token_ro.pickle', 'rb') as f:
                creds = pickle.load(f)
        if not creds or not creds.valid:
            if creds and creds.expired and creds.refresh_token:
                creds.refresh(Request())
            else:
                flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
                creds = flow.run_local_server(port=0)
            with open('token_ro.pickle', 'wb') as f:
                pickle.dump(creds, f)
        return build('youtube', 'v3', credentials=creds)
    raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
def fetch_yt_trending(max_results=200, region='US'):
    youtube = get_yt_service()
    records, token = [], None
    while len(records) < max_results:
        resp = youtube.videos().list(
            part='snippet,statistics',
            chart='mostPopular',
            regionCode=region,
            maxResults=min(50, max_results - len(records)),
            pageToken=token
        ).execute()
        for item in resp.get('items', []):
            s = item['snippet']; st = item.get('statistics', {})
            records.append({
                'video_id':      item['id'],
                'trending_date': pd.Timestamp.today().date(),
                'title':         s.get('title', ''),
                'channel_title': s.get('channelTitle', ''),
                'category_id':   s.get('categoryId', ''),
                'views':         int(st.get('viewCount', 0)),
                'likes':         int(st.get('likeCount', 0)),
                'comment_count': int(st.get('commentCount', 0)),
            })
        token = resp.get('nextPageToken')
        if not token: break
    return pd.DataFrame(records)
try:
    df = fetch_yt_trending()
    print('Live trending data:', df.shape)
except Exception as e:
    print(f'Falling back to synthetic: {e}')
    np.random.seed(42)
    n = 1000
    df = pd.DataFrame({
        'video_id':      [f'vid{i}' for i in range(n)],
        'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
        'title':         [f'Video Title {i}' for i in range(n)],
        'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
        'category_id':   np.random.choice([1,2,10,22,24,28], n),
        'views':         np.random.randint(10000, 5000000, n),
        'likes':         np.random.randint(100, 200000, n),
        'comment_count': np.random.randint(10, 50000, n),
    })
    print('Synthetic fallback:', df.shape)

print(df.head(3))
Live trending data: (199, 8)
      video_id trending_date  \
0  82-jTNka3uc    2026-06-03   
1  3oB9AxspVow    2026-06-03   
2  Aq9EJW9XqjQ    2026-06-03   

                                               title     channel_title  \
0  Ariana Grande - hate that i made you love me (...  ArianaGrandeVevo   
1           The End of Oak Street | Official Trailer      Warner Bros.   
2               I Shouldn’t Have Moved To This Farm…            CaseOh   

  category_id    views   likes  comment_count  
0          10  4405520  543459          30366  
1           1  4522744   46629           3597  
2          20   320318   11573            899  

What is in a social media trending dataset?#

  • Each row represents a trending YouTube video, including:

    • Unique IDs, trending date, video title, channel, category
    • Content performance: views, likes, and comments
  • These metrics help us measure engagement and popularity.

  • High values can indicate viral performancebut context matters.

  • Mistakes to avoid:

    • Not accounting for category differences (music, education, etc.)
    • Focusing only on raw counts (views) instead of rates or ratios
    • Ignoring how fast a video gained popularity

Core social media analytics concepts#

  • Views: how many times a video was watched.

  • Likes: indicates audience approval or connection.

  • Comments: shows two-way audience interaction.

  • Engagement rate: a metric (likes + comments) divided by views.

  • Trending date: when a video appeared in top viral lists.

  • Categories: group similar types of content for fair comparison.

  • Common mistakes:

    • Comparing videos across very different categories
    • Thinking all high-view content is truly viral
    • Not recognizing different types of engagement (likes vs comments)
# Beginner Example 1: Calculate engagement rate for each trending video
df['engagement_rate'] = (df['likes'] + df['comment_count']) / df['views'] * 100
print(df[['video_id', 'title', 'views', 'likes', 'comment_count', 'engagement_rate']].head())
      video_id                                              title    views  \
0  82-jTNka3uc  Ariana Grande - hate that i made you love me (...  4405520   
1  3oB9AxspVow           The End of Oak Street | Official Trailer  4522744   
2  Aq9EJW9XqjQ               I Shouldn’t Have Moved To This Farm…   320318   
3  vrY1THC_NQE  IShowSpeed - World Cup (Champions) [Official M...  4889332   
4  XXxUqLHq1xg  Cocktail 2 Official Trailer | Shahid Kapoor, K...  4135852   

    likes  comment_count  engagement_rate  
0  543459          30366        13.025137  
1   46629           3597         1.110521  
2   11573            899         3.893631  
3  724283          59345        16.027302  
4   44348           2833         1.140781  
# Beginner Example 2: Find the single most viewed trending video
top_view = df.loc[df['views'].idxmax()]
print('Top viewed video:', top_view['title'], '| Views:', top_view['views'])
Top viewed video: IShowSpeed - World Cup (Champions) [Official Music Video] | Views: 4889332
# Beginner Example 3: Average engagement rate across all trending videos
avg_engagement = df['engagement_rate'].mean()
print(f'Average engagement rate: {avg_engagement:.2f}%')
Average engagement rate: 5.81%
# Intermediate Example 1: Find videos with engagement rate above 10%
viral_candidates = df[df['engagement_rate'] > 10]
print('Number of high-engagement videos:', len(viral_candidates))
print(viral_candidates[['title', 'views', 'engagement_rate']].head(3))
Number of high-engagement videos: 25
                                               title    views  engagement_rate
0  Ariana Grande - hate that i made you love me (...  4405520        13.025137
3  IShowSpeed - World Cup (Champions) [Official M...  4889332        16.027302
6  Blin blin- Dany Ome y Kevincito El 13 ft Ovi X...    93879        11.600038
# Intermediate Example 2: Which channels appear most often in trending videos?
top_channels = df['channel_title'].value_counts().head(3)
print('Most featured channels on trending list:')
print(top_channels)
Most featured channels on trending list:
channel_title
ArianaGrandeVevo    1
Warner Bros.        1
CaseOh              1
Name: count, dtype: int64
# Intermediate Example 3: Analyze engagement rate by category
category_stats = df.groupby('category_id')['engagement_rate'].mean().sort_values(ascending=False)
print('Average engagement rate by category:')
print(category_stats)
Average engagement rate by category:
category_id
22    7.968168
10    7.659610
17    6.063859
24    6.057500
20    5.345862
28    4.471588
1     4.176650
Name: engagement_rate, dtype: float64
# Advanced Example 1: Time-series analysis - engagement trends over time
df['trending_date'] = pd.to_datetime(df['trending_date'])
daily_trend = df.groupby('trending_date')['engagement_rate'].mean()
import matplotlib.pyplot as plt
plt.figure(figsize=(10,4))
plt.plot(daily_trend.index, daily_trend.values, label='Avg Engagement Rate')
plt.xlabel('Trending Date')
plt.ylabel('Avg Engagement Rate (%)')
plt.title('Engagement Rate Trend Over Time')
plt.legend()
plt.tight_layout()
plt.show()
No description has been provided for this image
# Advanced Example 2: Detecting outlier viral videos using z-score
from scipy.stats import zscore
df['views_z'] = zscore(df['views'])
viral_outliers = df[df['views_z'] > 3]
print('Videos with extreme virality (z > 3):')
print(viral_outliers[['title', 'views', 'engagement_rate', 'views_z']].head())
Videos with extreme virality (z > 3):
                                               title    views  \
0  Ariana Grande - hate that i made you love me (...  4405520   
1           The End of Oak Street | Official Trailer  4522744   
3  IShowSpeed - World Cup (Champions) [Official M...  4889332   
4  Cocktail 2 Official Trailer | Shahid Kapoor, K...  4135852   
9                       MEOVV(미야오) - ‘DDI RO RI’ M/V  3532687   

   engagement_rate   views_z  
0        13.025137  5.488728  
1         1.110521  5.646492  
3        16.027302  6.139858  
4         1.140781  5.125801  
9         0.297337  4.314042  
# Advanced Example 3: Calculate and rank videos by viral index (views * engagement rate)
df['viral_index'] = df['views'] * df['engagement_rate']
top_viral = df.nlargest(5, 'viral_index')
print('Top 5 potential viral videos by viral index:')
print(top_viral[['title', 'views', 'engagement_rate', 'viral_index']])
Top 5 potential viral videos by viral index:
                                                title    views  \
3   IShowSpeed - World Cup (Champions) [Official M...  4889332   
0   Ariana Grande - hate that i made you love me (...  4405520   
13  No Peace Amongst the Stars | Warhammer 40,000 ...  2599959   
17            1000 VS 1000 Player Minecraft Civil War  2594460   
79  Hydroneer Tried to Optimize, So I Double Overl...  1873938   

    engagement_rate  viral_index  
3         16.027302   78362800.0  
0         13.025137   57382500.0  
13         6.718067   17466700.0  
17         6.166370   15998400.0  
79         5.273494    9882200.0  
# Error handling: missing or zero engagement values
has_engagement_issue = df['views'] == 0
if has_engagement_issue.any():
    print('Warning: Some videos have zero views (impossible on trending)!')
    print(df[has_engagement_issue])
else:
    print('All trending videos have at least one view. Data integrity looks good.')
All trending videos have at least one view. Data integrity looks good.
# Debugging: What if aggregation logic is incorrect?
wrong_sum = df.groupby('category_id')['views'].sum().sum()
true_sum = df['views'].sum()
if wrong_sum != true_sum:
    print('Check: Grouped sum matches total sum  OK')
else:
    print('Possible mistake: double-counted or missing data in aggregation!')
Possible mistake: double-counted or missing data in aggregation!
# Error: Interpreting engagement rate wrongly (likes/views instead of (likes+comments)/views)
df['wrong_rate'] = df['likes'] / df['views'] * 100
mean_wrong = df['wrong_rate'].mean()
mean_right = df['engagement_rate'].mean()
print(f'Wrong formula avg: {mean_wrong:.2f}% | Correct: {mean_right:.2f}%')
Wrong formula avg: 5.30% | Correct: 5.81%
# Error: Wrong grouping logic (should group by trending date and category together)
df['date_cat'] = df['trending_date'].astype(str) + '-' + df['category_id'].astype(str)
grouped = df.groupby('date_cat')['views'].sum()
print('Sample combined date/category group:')
print(grouped.head(2))
Sample combined date/category group:
date_cat
2026-06-03-1      6951398
2026-06-03-10    13958719
Name: views, dtype: int64

Best practices for social media content analytics#

  • Benchmark performance within content categories, not just globally.

  • Segment audiences and analyze by channel for deeper insights.

  • Track trends over timelook for patterns, not just peaks.

  • Use consistent and clearly defined metrics (explain your math!).

  • Set and revisit benchmarks using up-to-date, comparable data.

  • Optimize content by learning from what works (and what does not work) on trending lists.

Common analytics patterns for viral video detection#

  • Identify videos that outperform average by large margins.

  • Use engagement rates and 'viral index' for deeper ranking.

  • Combine visualization and statistics for better insights.

  • Regularly review and adapt thresholds as platforms evolve.

  • Document when and why you flag a video as viral.

# End-to-end Example: From raw data to viral video recommendation
viral = df[df['views_z'] > 3]
if not viral.empty:
    result = viral.sort_values('viral_index', ascending=False).iloc[0]
    print('Recommend doubling down on content like:', result['title'])
    print(f'Viral signals: {result["views"]} views and {result["engagement_rate"]:.2f}% engagement rate')
else:
    print('No strongly viral videos detected. Recommend analyzing trends and trying again later.')
Recommend doubling down on content like: IShowSpeed - World Cup (Champions) [Official Music Video]
Viral signals: 4889332 views and 16.03% engagement rate

Congratulations on completing the viral video detection case study!#

  • You analyzed trends, engagement rates, and detected outliers in trending data.

  • Try these skills on your own content or with new datasets.

  • For more tutorials, see the official YouTube Data API and analytics guides.

  • Want to boost your analytics superpowers? Subscribe on YouTube to stay updated.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.