Mathew K Analytics

Lesson 48 · Social Media Content Analytics

Detecting Viral Spikes and Trends in Social Media Data

In this lesson, we explore how to use Python to detect viral spikes and trending content across social media platforms. Understanding these patterns helps…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Detecting Viral Spikes and Trends in Social Media Data#

  • In this lesson, we explore how to use Python to detect viral spikes and trending content across social media platforms.
  • Understanding these patterns helps content creators and businesses adapt strategies for maximum audience impact.
  • You will learn to analyze time series data, engagement metrics, and discover techniques to identify viral moments and sustained trends.
  • By the end, you will be able to spot both sudden spikes and long-term growth patterns in real social media datasets.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')

Core Concepts: Social Media Analytics for Viral Detection#

  • Social media datasets contain data on posts, videos, audience engagements, and timestamps.
  • Key metrics include views, likes, comments, shares, click-through rate (CTR), and watch time.
  • Engagement rates show how active your audience is compared to total views.
  • A spike in engagement or views often signals viral or trending content.
  • Beginners often mistake total engagement for high performance without considering posting time or audience size.
  • Always contextualize spikes with timelines and audience reach.
# Beginner Example 1: Load a simple synthetic social media time series dataset
np.random.seed(42)
n_days = 30
dates = pd.date_range('2023-01-01', periods=n_days, freq='D')
views = np.random.randint(1000, 20000, n_days)
likes = (views * np.random.uniform(0.04, 0.09, n_days)).astype(int)
df_ts = pd.DataFrame({
    'date': dates,
    'views': views,
    'likes': likes,
    'engagement_rate': np.round(likes / views * 100, 2)
})
print(df_ts.head(3))
        date  views  likes  engagement_rate
0 2023-01-01  16795   1393             8.29
1 2023-01-02   1860    137             7.37
2 2023-01-03   6390    399             6.24
# Beginner Example 2: Visualize views over time to spot spikes
plt.figure(figsize=(9,4))
plt.plot(df_ts['date'], df_ts['views'], marker='o')
plt.title('Daily Views Over Time')
plt.xlabel('Date')
plt.ylabel('Views')
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
No description has been provided for this image
# Beginner Example 3: Calculate and visualize engagement rate trends
plt.figure(figsize=(9,4))
plt.plot(df_ts['date'], df_ts['engagement_rate'], color='orange', marker='o')
plt.title('Daily Engagement Rate (%)')
plt.xlabel('Date')
plt.ylabel('Engagement Rate (%)')
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
No description has been provided for this image
# Beginner Example 4: Identify the day with highest views (potential viral spike)
max_views_row = df_ts.loc[df_ts['views'].idxmax()]
print('Day with highest views:')
print(max_views_row)
Day with highest views:
date               2023-01-22 00:00:00
views                            19942
likes                             1637
engagement_rate                   8.21
Name: 21, dtype: object
# Intermediate Example 1: Load a real YouTube Trending Videos dataset or synthetic fallback
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request

SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']

def get_yt_service():
    api_key = os.environ.get('YOUTUBE_API_KEY')
    if api_key:
        return build('youtube', 'v3', developerKey=api_key)
    if Path('client_secret.json').exists():
        creds = None
        if Path('token_ro.pickle').exists():
            with open('token_ro.pickle', 'rb') as f:
                creds = pickle.load(f)
        if not creds or not creds.valid:
            if creds and creds.expired and creds.refresh_token:
                creds.refresh(Request())
            else:
                flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
                creds = flow.run_local_server(port=0)
            with open('token_ro.pickle', 'wb') as f:
                pickle.dump(creds, f)
        return build('youtube', 'v3', credentials=creds)
    raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')

def fetch_yt_trending(max_results=200, region='US'):
    youtube = get_yt_service()
    records, token = [], None
    while len(records) < max_results:
        resp = youtube.videos().list(
            part='snippet,statistics',
            chart='mostPopular',
            regionCode=region,
            maxResults=min(50, max_results - len(records)),
            pageToken=token
        ).execute()
        for item in resp.get('items', []):
            s = item['snippet']; st = item.get('statistics', {})
            records.append({
                'video_id':      item['id'],
                'trending_date': pd.Timestamp.today().date(),
                'title':         s.get('title', ''),
                'channel_title': s.get('channelTitle', ''),
                'category_id':   s.get('categoryId', ''),
                'views':         int(st.get('viewCount', 0)),
                'likes':         int(st.get('likeCount', 0)),
                'comment_count': int(st.get('commentCount', 0)),
            })
        token = resp.get('nextPageToken')
        if not token: break
    return pd.DataFrame(records)

try:
    df_trend = fetch_yt_trending()
    print('Loaded real trending data:', df_trend.shape)
except Exception as e:
    print(f'Falling back to synthetic: {e}')
    np.random.seed(42)
    n = 1000
    df_trend = pd.DataFrame({
        'video_id':      [f'vid{i}' for i in range(n)],
        'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
        'title':         [f'Video Title {i}' for i in range(n)],
        'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
        'category_id':   np.random.choice([1,2,10,22,24,28], n),
        'views':         np.random.randint(10000, 5000000, n),
        'likes':         np.random.randint(100, 200000, n),
        'comment_count': np.random.randint(10, 50000, n),
    })
    print('Synthetic fallback:', df_trend.shape)

print(df_trend.head(3))
Loaded real trending data: (199, 8)
      video_id trending_date  \
0  82-jTNka3uc    2026-06-02   
1  3oB9AxspVow    2026-06-02   
2  l-cyT28MFyk    2026-06-02   

                                               title     channel_title  \
0  Ariana Grande - hate that i made you love me (...  ArianaGrandeVevo   
1           The End of Oak Street | Official Trailer      Warner Bros.   
2        PLAYING HARDCORE MINECRAFT UNTIL WE BEAT IT            Jynxzi   

  category_id    views   likes  comment_count  
0          10  2847627  443017          27525  
1           1  2831863   35019           2876  
2          24   742329   16302            373  
# Intermediate Example 2: Find top 5 trending videos with most views
top_videos = df_trend.sort_values('views', ascending=False).head(5)
print('Top 5 trending videos by view count:')
print(top_videos[['title', 'channel_title', 'views', 'likes', 'comment_count']])
Top 5 trending videos by view count:
                                                 title     channel_title  \
38                               TREASURE - ‘IF I’ M/V    TREASURE (트레저)   
163          I Went to WAR on a Hardcore Minecraft SMP            Wemmbu   
3    IShowSpeed - World Cup (Champions) [Official M...        IShowSpeed   
0    Ariana Grande - hate that i made you love me (...  ArianaGrandeVevo   
1             The End of Oak Street | Official Trailer      Warner Bros.   

       views   likes  comment_count  
38   5206077  245042          46031  
163  3686297  198342          33773  
3    2872891  506606          44844  
0    2847627  443017          27525  
1    2831863   35019           2876  
# Intermediate Example 3: Visualize trending video view distribution
plt.figure(figsize=(8,4))
sns.histplot(df_trend['views'], bins=30, color='skyblue')
plt.title('Distribution of Trending Video Views')
plt.xlabel('Views per Video')
plt.ylabel('Count of Videos')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Intermediate Example 4: Calculate engagement rate for each trending video
df_trend['engagement_rate'] = np.where(
    df_trend['views'] > 0,
    (df_trend['likes'] + df_trend['comment_count']) / df_trend['views'] * 100,
    np.nan
)
print('Engagement rate stats:')
print(df_trend['engagement_rate'].describe())
Engagement rate stats:
count    199.000000
mean       5.969339
std        4.517802
min        0.080163
25%        3.165754
50%        5.508154
75%        7.061947
max       24.658397
Name: engagement_rate, dtype: float64
# Intermediate Example 5: Identify potential viral videos using engagement and views
viral_videos = df_trend[(df_trend['views'] > df_trend['views'].quantile(0.95)) &
                       (df_trend['engagement_rate'] > df_trend['engagement_rate'].quantile(0.95))]
print('Potential viral videos (top 5% for both views and engagement rate):')
print(viral_videos[['title', 'channel_title', 'views', 'likes', 'comment_count', 'engagement_rate']])
Potential viral videos (top 5% for both views and engagement rate):
                                               title channel_title    views  \
3  IShowSpeed - World Cup (Champions) [Official M...    IShowSpeed  2872891   

    likes  comment_count  engagement_rate  
3  506606          44844         19.19495  
# Advanced Example 1: Detect viral spikes in a full year social media time series
np.random.seed(42)
n_days_full = 365
dates_full = pd.date_range('2023-01-01', periods=n_days_full)
views_full = np.random.randint(1000, 50000, n_days_full)
spikes = np.random.choice([0, 0, 0, 1], n_days_full, p=[0.96, 0.01, 0.01, 0.02])
views_full += spikes * np.random.randint(80000, 200000, n_days_full)
likes_full = (views_full * np.random.uniform(0.03, 0.12, n_days_full)).astype(int)
df_full = pd.DataFrame({
    'date': dates_full,
    'views': views_full,
    'likes': likes_full,
    'engagement_rate': np.round(likes_full / views_full * 100, 2)
})
plt.figure(figsize=(14,5))
plt.plot(df_full['date'], df_full['views'], label='Views', color='blue')
plt.title('Full Year Daily Views with Synthetic Viral Spikes')
plt.xlabel('Date')
plt.ylabel('Views')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Advanced Example 2: Automatic viral spike detection algorithm
threshold = df_full['views'].mean() + 3 * df_full['views'].std()
viral_days = df_full[df_full['views'] > threshold]
print(f'Days identified as viral (spikes > mean + 3*std): {viral_days.shape[0]}')
print(viral_days[['date', 'views', 'likes', 'engagement_rate']])
Days identified as viral (spikes > mean + 3*std): 10
          date   views  likes  engagement_rate
9   2023-01-10  128246   9066             7.07
138 2023-05-19  207105   9312             4.50
189 2023-07-09  189464  14029             7.40
194 2023-07-14  133527  12485             9.35
223 2023-08-12  186301  16083             8.63
247 2023-09-05  163644   6536             3.99
266 2023-09-24  116675   9941             8.52
279 2023-10-07  175290  10782             6.15
280 2023-10-08  128789   4534             3.52
329 2023-11-26  142083  15352            10.80
# Advanced Example 3: Explore trends across content categories in trending videos
category_trends = df_trend.groupby('category_id').agg({'views': 'mean', 'likes': 'mean', 'engagement_rate': 'mean'}).sort_values('views', ascending=False)
print('Average views, likes, and engagement rate by category:')
print(category_trends)
Average views, likes, and engagement rate by category:
                    views         likes  engagement_rate
category_id                                             
1            1.118028e+06  11147.500000         2.347557
24           5.849731e+05  24374.692308         5.430105
22           4.669356e+05  17405.142857         8.320702
10           4.624271e+05  27549.206897         6.737327
28           2.698350e+05  12291.000000         4.908185
20           2.433250e+05  14109.545455         5.841199
17           1.596715e+05   9855.500000         7.044903
# Advanced Example 4: Visualize engagement trend with a moving average
df_full['views_ma7'] = df_full['views'].rolling(window=7).mean()
plt.figure(figsize=(12,5))
plt.plot(df_full['date'], df_full['views'], alpha=0.4, label='Daily Views')
plt.plot(df_full['date'], df_full['views_ma7'], color='red', linewidth=2, label='7-Day Moving Average')
plt.title('Viral Trendlines: Daily Views and 7-Day Moving Average')
plt.xlabel('Date')
plt.ylabel('Views')
plt.legend()
plt.tight_layout()
plt.show()
No description has been provided for this image
# Error Handling Example 1: Missing values in engagement columns
df_trend_corrupt = df_trend.copy()
df_trend_corrupt.loc[0, 'likes'] = np.nan
df_trend_corrupt.loc[1, 'comment_count'] = np.nan
df_trend_corrupt['engagement_fixed'] = (
    df_trend_corrupt['likes'].fillna(0) + df_trend_corrupt['comment_count'].fillna(0)
) / df_trend_corrupt['views'] * 100
print('Engagement rate with missing values (rows 0 and 1):')
print(df_trend_corrupt[['likes', 'comment_count', 'views', 'engagement_fixed']].head(3))
Engagement rate with missing values (rows 0 and 1):
     likes  comment_count    views  engagement_fixed
0      NaN        27525.0  2847627          0.966594
1  35019.0            NaN  2831863          1.236606
2  16302.0          373.0   742329          2.246309
# Error Handling Example 2: Incorrect metric aggregation
grouped = df_trend.groupby('category_id').sum(numeric_only=True)
print('Incorrect metric aggregation (should not sum engagement rates):')
print(grouped[['views', 'likes', 'comment_count', 'engagement_rate']].head())
Incorrect metric aggregation (should not sum engagement rates):
                views    likes  comment_count  engagement_rate
category_id                                                   
1             4472114    44590           3479         9.390227
10           13410385   798927          52890       195.382482
17             319343    19711           2133        14.089806
20           34795476  2017665         199042       835.291410
22            3268549   121836           6246        58.244917
# Error Handling Example 3: Misinterpreting ratios like CTR
np.random.seed(42)
dummy_views = np.random.randint(1, 20, 10)
dummy_clicks = np.random.randint(0, 5, 10)
ctr = dummy_clicks / dummy_views * 100
print('Example of click-through rates (CTR) correctly calculated:')
print('Views:', dummy_views)
print('Clicks:', dummy_clicks)
print('CTR (%):', np.round(ctr,2))
Example of click-through rates (CTR) correctly calculated:
Views: [ 7 15 11  8  7 19 11 11  4  8]
Clicks: [2 4 1 3 1 3 4 0 3 1]
CTR (%): [28.57 26.67  9.09 37.5  14.29 15.79 36.36  0.   75.   12.5 ]
# Error Handling Example 4: Wrong grouping logic for content categories
wrong_group = df_trend.groupby(['channel_title', 'category_id']).mean(numeric_only=True)
print('Grouped by BOTH channel and category - can make small groups:')
print(wrong_group.head(7))
Grouped by BOTH channel and category - can make small groups:
                                         views   likes  comment_count  \
channel_title            category_id                                    
3C Films                 24            49786.0  4561.0          522.0   
AR12Gaming               20            32498.0  1974.0          191.0   
Affirmation Club - Topic 10           217892.0  5590.0            0.0   
AloisNL                  20            89257.0  3952.0          124.0   
Angelazz Brookhaven      20           108026.0  4612.0          473.0   
AngryJoeShow             20            21947.0  1369.0          209.0   
Aphmau                   20           116806.0  5053.0          897.0   

                                      engagement_rate  
channel_title            category_id                   
3C Films                 24                 10.209698  
AR12Gaming               20                  6.661948  
Affirmation Club - Topic 10                  2.565491  
AloisNL                  20                  4.566589  
Angelazz Brookhaven      20                  4.707200  
AngryJoeShow             20                  7.190049  
Aphmau                   20                  5.093916  

Best Practices for Social Media Trend Analytics#

  • Always visualize your data before calculating spikes or trends.
  • Use moving averages to smooth out daily or weekly fluctuations.
  • Benchmark performance relative to similar content, not only across your own posts.
  • Segment audience or content by reasonable categories (e.g., genre, time, channel).
  • Define viral thresholds clearly (e.g., above percentile, or multiple of std).
  • Check all metric calculations: use means for rates, not sums.
  • Document anomalies and check for missing values before reporting results.
# Analytics Pattern: Finding best posting days of the week
df_full['weekday'] = df_full['date'].dt.day_name()
weekday_means = df_full.groupby('weekday')['views'].mean().reindex([
    'Monday', 'Tuesday', 'Wednesday', 'Thursday', 'Friday', 'Saturday', 'Sunday'
])
print('Average views by weekday:')
print(weekday_means)
Average views by weekday:
weekday
Monday       21718.019231
Tuesday      29737.480769
Wednesday    30303.057692
Thursday     22731.346154
Friday       31813.538462
Saturday     31121.403846
Sunday       32534.037736
Name: views, dtype: float64
# Analytics Pattern: Consistent metric definition using a helper function
def calc_engagement(df):
    return (df['likes'] + df.get('comment_count', 0)) / df['views'] * 100
df_trend['practice_engagement'] = calc_engagement(df_trend)
print('First five engagement calculations:')
print(df_trend[['likes', 'comment_count', 'views', 'practice_engagement']].head())
First five engagement calculations:
    likes  comment_count    views  practice_engagement
0  443017          27525  2847627            16.524004
1   35019           2876  2831863             1.338165
2   16302            373   742329             2.246309
3  506606          44844  2872891            19.194950
4  153901           7463  2216419             7.280392
# End-to-end Problem: Identify, explain, and recommend content strategy for a detected viral spike
detected_viral = df_full.loc[df_full['views'] > threshold]
for idx, row in detected_viral.iterrows():
    print(f'Date: {row.date.date()} | Views: {row.views} | Engagement Rate: {row.engagement_rate:.2f}%')
print('---')
if not detected_viral.empty:
    rec_day = detected_viral.iloc[0]['date'].day_name()
    print(f'Strategy: On {rec_day}s, consider scheduling more high-impact content, and analyze what drove engagement on these spike dates.')
else:
    print('No strong viral spikes found: Review your content plan for new tactics.')
Date: 2023-01-10 | Views: 128246 | Engagement Rate: 7.07%
Date: 2023-05-19 | Views: 207105 | Engagement Rate: 4.50%
Date: 2023-07-09 | Views: 189464 | Engagement Rate: 7.40%
Date: 2023-07-14 | Views: 133527 | Engagement Rate: 9.35%
Date: 2023-08-12 | Views: 186301 | Engagement Rate: 8.63%
Date: 2023-09-05 | Views: 163644 | Engagement Rate: 3.99%
Date: 2023-09-24 | Views: 116675 | Engagement Rate: 8.52%
Date: 2023-10-07 | Views: 175290 | Engagement Rate: 6.15%
Date: 2023-10-08 | Views: 128789 | Engagement Rate: 3.52%
Date: 2023-11-26 | Views: 142083 | Engagement Rate: 10.80%
---
Strategy: On Tuesdays, consider scheduling more high-impact content, and analyze what drove engagement on these spike dates.

Wrap-Up: Viral Trend Detection Cheatsheet#

  • Use time series plots to visually spot spikes.
  • Compute percentiles or std deviations to quantify viral moments.
  • Combine views with engagement rate to confirm real popularity.
  • Avoid common analysis mistakes (wrong grouping, missing values, incorrect metric sums).
  • Test your strategy with both synthetic and real data.
  • Try explaining every detected spike in plain language.
  • Next: Download this notebook or watch our full video course on YouTube!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.