Mathew K Analytics

Lesson 35 · Social Media Content Analytics

Detecting High-Impact Content Features

In this lesson, we will solve real social media analytics problems by finding which content features drive high engagement. This is important because…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Detecting High-Impact Content Features#

  • In this lesson, we will solve real social media analytics problems by finding which content features drive high engagement.
  • This is important because creators and brands need to know what makes content perform better to grow their audience.
  • You will learn how to analyze likes, views, comments, watch time, and more across posts and videos.
  • By the end, you will be able to pinpoint what makes some content trend, and suggest data-driven improvements for future posts.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Core Social Media Analytics Concepts#

  • Social media datasets represent real content like videos or posts on platforms such as YouTube, Instagram, and TikTok.
  • Each post or video typically has metrics like views, likes, comments, shares, watch time, and click-through rate (CTR).
  • High engagement usually means users are liking, sharing, or commenting on content more than average.
  • It is important to remember that a high view count is not always equal to high impactengagement rate (engagement divided by views) shows true audience reaction.
  • Beginners often make mistakes by looking only at total numbers, ignoring ratios like engagement rate and CTR.
# Load the YouTube Trending Videos dataset (with fallback if live fetch fails)
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request

SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']

def get_yt_service():
    api_key = os.environ.get('YOUTUBE_API_KEY')
    if api_key:
        return build('youtube', 'v3', developerKey=api_key)
    if Path('client_secret.json').exists():
        creds = None
        if Path('token_ro.pickle').exists():
            with open('token_ro.pickle', 'rb') as f:
                creds = pickle.load(f)
        if not creds or not creds.valid:
            if creds and creds.expired and creds.refresh_token:
                creds.refresh(Request())
            else:
                flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
                creds = flow.run_local_server(port=0)
            with open('token_ro.pickle', 'wb') as f:
                pickle.dump(creds, f)
        return build('youtube', 'v3', credentials=creds)
    raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')

def fetch_yt_trending(max_results=200, region='US'):
    youtube = get_yt_service()
    records, token = [], None
    while len(records) < max_results:
        resp = youtube.videos().list(
            part='snippet,statistics',
            chart='mostPopular',
            regionCode=region,
            maxResults=min(50, max_results - len(records)),
            pageToken=token
        ).execute()
        for item in resp.get('items', []):
            s = item['snippet']; st = item.get('statistics', {})
            records.append({
                'video_id':      item['id'],
                'trending_date': pd.Timestamp.today().date(),
                'title':         s.get('title', ''),
                'channel_title': s.get('channelTitle', ''),
                'category_id':   s.get('categoryId', ''),
                'views':         int(st.get('viewCount', 0)),
                'likes':         int(st.get('likeCount', 0)),
                'comment_count': int(st.get('commentCount', 0)),
            })
        token = resp.get('nextPageToken')
        if not token: break
    return pd.DataFrame(records)

try:
    df = fetch_yt_trending()
    print('Live trending data:', df.shape)
except Exception as e:
    print(f'Falling back to synthetic: {e}')
    np.random.seed(42)
    n = 1000
    df = pd.DataFrame({
        'video_id':      [f'vid{i}' for i in range(n)],
        'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
        'title':         [f'Video Title {i}' for i in range(n)],
        'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
        'category_id':   np.random.choice([1,2,10,22,24,28], n),
        'views':         np.random.randint(10000, 5000000, n),
        'likes':         np.random.randint(100, 200000, n),
        'comment_count': np.random.randint(10, 50000, n),
    })
    print('Synthetic fallback:', df.shape)

print(df.head(3))
Live trending data: (199, 8)
      video_id trending_date  \
0  82-jTNka3uc    2026-06-02   
1  3oB9AxspVow    2026-06-02   
2  Vagb9BqdX8g    2026-06-02   

                                               title     channel_title  \
0  Ariana Grande - hate that i made you love me (...  ArianaGrandeVevo   
1           The End of Oak Street | Official Trailer      Warner Bros.   
2                 We Almost Lost EVERYTHING Gambling        SMii7Yplus   

  category_id    views   likes  comment_count  
0          10   762209  272120          21830  
1           1   360910   21411           1834  
2          20  1224820   61634           1661  
# Preview basic dataset info and column types
print('Dataset columns:', df.columns.tolist())
print('Data types:')
print(df.dtypes)
Dataset columns: ['video_id', 'trending_date', 'title', 'channel_title', 'category_id', 'views', 'likes', 'comment_count']
Data types:
video_id         object
trending_date    object
title            object
channel_title    object
category_id      object
views             int64
likes             int64
comment_count     int64
dtype: object
# Beginner: Find the video with the highest view count
top_view = df.loc[df['views'].idxmax()]
print('Video with the most views:')
print(top_view[['video_id', 'title', 'views']])
Video with the most views:
video_id                                  _hyFGrcVv5U
title       I Went to WAR on a Hardcore Minecraft SMP
views                                         3550828
Name: 73, dtype: object
# Beginner: Calculate like rate for each video
df['like_rate'] = df['likes'] / df['views'] * 100
print('Like rates (first 3 rows):')
print(df[['title','views','likes','like_rate']].head(3))
Like rates (first 3 rows):
                                               title    views   likes  \
0  Ariana Grande - hate that i made you love me (...   762209  272120   
1           The End of Oak Street | Official Trailer   360910   21411   
2                 We Almost Lost EVERYTHING Gambling  1224820   61634   

   like_rate  
0  35.701494  
1   5.932504  
2   5.032086  
# Beginner: Calculate comment rate for each video
df['comment_rate'] = df['comment_count'] / df['views'] * 100
print('Comment rates (first 3 rows):')
print(df[['title','views','comment_count','comment_rate']].head(3))
Comment rates (first 3 rows):
                                               title    views  comment_count  \
0  Ariana Grande - hate that i made you love me (...   762209          21830   
1           The End of Oak Street | Official Trailer   360910           1834   
2                 We Almost Lost EVERYTHING Gambling  1224820           1661   

   comment_rate  
0      2.864044  
1      0.508160  
2      0.135612  
# Intermediate: Which categories have the highest average like rate?
cat_like = df.groupby('category_id')['like_rate'].mean().sort_values(ascending=False)
print('Average like rate by category:')
print(cat_like.head(5))
Average like rate by category:
category_id
25    14.787044
1      7.674104
10     6.421286
28     5.653068
23     5.103319
Name: like_rate, dtype: float64
# Intermediate: Find top 5 videos by like rate (minimum 10,000 views for fairness)
df10k = df[df['views'] >= 10000]
top5_like = df10k.sort_values('like_rate', ascending=False).head(5)
print('Top 5 content by like rate (min 10k views):')
print(top5_like[['title','views','likes','like_rate']])
Top 5 content by like rate (min 10k views):
                                                title    views   likes  \
25                                 Fortnite | Runners    26979    9803   
0   Ariana Grande - hate that i made you love me (...   762209  272120   
6                               SHINee 샤이니 'Atmos' MV   249461   46340   
64              CORTIS (코르티스) 'Blue Lips' Official MV  2283257  376038   
69  The Phan Relationship is moving too fast - Tom...   214784   32930   

    like_rate  
25  36.335668  
0   35.701494  
6   18.576050  
64  16.469368  
69  15.331682  
# Intermediate: Correlation between likes and comments
corr_likes_comments = df['likes'].corr(df['comment_count'])
print(f'Correlation between likes and comments: {corr_likes_comments:.2f}')
Correlation between likes and comments: 0.85
# Intermediate: Which channels consistently produce high like rate videos?
chan_like = df.groupby('channel_title')['like_rate'].mean().sort_values(ascending=False)
print('Top channels by average like rate:')
print(chan_like.head(5))
Top channels by average like rate:
channel_title
Fortnite            36.335668
ArianaGrandeVevo    35.701494
SMTOWN              18.576050
HYBE LABELS         16.469368
Dan and Phil        15.331682
Name: like_rate, dtype: float64
# Intermediate: Identify viral outlier videos (like rate >2 std above mean and 100k+ views)
mean_like = df['like_rate'].mean()
std_like = df['like_rate'].std()
viral = df[(df['like_rate'] > mean_like + 2*std_like) & (df['views'] > 100000)]
print('Potential viral videos:')
print(viral[['title', 'views', 'likes', 'like_rate']].head(10))
Potential viral videos:
                                                title    views   likes  \
0   Ariana Grande - hate that i made you love me (...   762209  272120   
6                               SHINee 샤이니 'Atmos' MV   249461   46340   
64              CORTIS (코르티스) 'Blue Lips' Official MV  2283257  376038   
69  The Phan Relationship is moving too fast - Tom...   214784   32930   

    like_rate  
0   35.701494  
6   18.576050  
64  16.469368  
69  15.331682  
# Advanced: Daily trend of average like rate over time
trend = df.groupby('trending_date')['like_rate'].mean()
print('First 5 days of average like rate trend:')
print(trend.head(5))
First 5 days of average like rate trend:
trending_date
2026-06-02    5.122689
Name: like_rate, dtype: float64
# Advanced: Find if video titles with strong words (e.g. 'BEST', 'AMAZING') get higher like rates
strong_words = ['best', 'amazing', 'ultimate', 'must', 'top', 'incredible']
df['strong_word_title'] = df['title'].str.lower().apply(lambda t: any(w in t for w in strong_words))
avg_like_strong = df.groupby('strong_word_title')['like_rate'].mean()
print('Average like rate by strong word presence in title:')
print(avg_like_strong)
Average like rate by strong word presence in title:
strong_word_title
False    5.176363
True     4.108257
Name: like_rate, dtype: float64
# Advanced: Export top 20 high-impact videos (by like rate & comment rate) to CSV
hi_impact = df10k.copy()
hi_impact['impact_score'] = (hi_impact['like_rate'] + hi_impact['comment_rate']) / 2
top20 = hi_impact.sort_values('impact_score', ascending=False).head(20)
top20.to_csv('top20_high_impact.csv', index=False)
print('Exported top 20 high-impact videos to CSV.')
Exported top 20 high-impact videos to CSV.
# Error Handling: Handle missing values in engagement metrics
df_missing = df.copy()
df_missing.loc[0, 'likes'] = np.nan
df_missing.loc[1, 'views'] = np.nan
filled = df_missing[['likes', 'views']].fillna(-1)
print(filled.head(3))
     likes      views
0     -1.0   762209.0
1  21411.0       -1.0
2  61634.0  1224820.0
# Debugging: Accidental aggregation bug - summing like rates by channel (WRONG!)
wrong_agg = df.groupby('channel_title')['like_rate'].sum()
print(wrong_agg.head(3))
channel_title
3FS           8.445692
AR12Gaming    4.700222
Abnaze        2.604007
Name: like_rate, dtype: float64
# Debugging: Misinterpreting ratios (like rate vs. raw likes)
by_views = df[['views','likes','like_rate']].sort_values('likes', ascending=False).head(5)
by_rate = df[['views','likes','like_rate']].sort_values('like_rate', ascending=False).head(5)
print('Most liked videos (raw):')
print(by_views)
print('Most engaging videos (by rate):')
print(by_rate)
Most liked videos (raw):
      views   likes  like_rate
64  2283257  376038  16.469368
0    762209  272120  35.701494
11  2628576  217711   8.282469
73  3550828  195775   5.513503
4   1738462  133402   7.673564
Most engaging videos (by rate):
      views   likes  like_rate
25    26979    9803  36.335668
0    762209  272120  35.701494
6    249461   46340  18.576050
64  2283257  376038  16.469368
69   214784   32930  15.331682
# Debugging: Wrong grouping can hide high-impact outliers
overall_avg = df['like_rate'].mean()
cat_avg = df.groupby('category_id')['like_rate'].mean()
mask = cat_avg > overall_avg + 2*df['like_rate'].std()
print('Categories with unusually high engagement:')
print(cat_avg[mask])
Categories with unusually high engagement:
category_id
25    14.787044
Name: like_rate, dtype: float64

Social Media Analytics Best Practices#

  • Always check for and handle missing or outlier values before analysis.
  • Benchmark each content item against similar categories or audience segments instead of a flat average.
  • Analyze both rates (like rate, comment rate, engagement rate) and raw numbers to avoid misleading results.
  • Track content trends over time, not just across topics.
  • Experiment with content features like titles, timing, or thumbnailsand compare engagement rates before and after changes.

Common Analytics Patterns#

  • Benchmark content against top performers in the same topic or format.
  • Segment your audience and look for patterns unique to different groups.
  • Use time series and moving averages to detect sustained growth or sudden drops.
  • Always define engagement metrics clearly for meaningful comparison.
  • Test content changes and collect before-and-after metrics to guide creative decisions.
# End-to-end: Raw content analytics to practical recommendation
df10k = df[df['views'] >= 10000].copy() # Use only videos with significant reach
df10k['impact_score'] = (df10k['like_rate'] + df10k['comment_rate']) / 2
top_row = df10k.sort_values('impact_score', ascending=False).iloc[0]
recommendation = f"Publish more content similar to: '{top_row['title']}' in category {top_row['category_id']} (Impact Score: {top_row['impact_score']:.2f})."
print('Content Strategy Recommendation:')
print(recommendation)
Content Strategy Recommendation:
Publish more content similar to: 'Fortnite | Runners' in category 20 (Impact Score: 20.02).

Congratulations! You have detected high-impact content features.#

  • Try applying these analyses to your own YouTube, Instagram, or TikTok data.
  • Continue practicing as social media platforms evolveaudience tastes can change rapidly!
  • Enjoyed this? Subscribe to our channel and check out our YouTube playlist for more real-world content analytics lessons.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.