Lesson 35 · Social Media Content Analytics
Detecting High-Impact Content Features
In this lesson, we will solve real social media analytics problems by finding which content features drive high engagement. This is important because…
- CourseSocial Media Content Analytics
- Lesson35 of 41
- Video23 min
- FormatJupyter notebook · 19 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbDetecting High-Impact Content Features#
- In this lesson, we will solve real social media analytics problems by finding which content features drive high engagement.
- This is important because creators and brands need to know what makes content perform better to grow their audience.
- You will learn how to analyze likes, views, comments, watch time, and more across posts and videos.
- By the end, you will be able to pinpoint what makes some content trend, and suggest data-driven improvements for future posts.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
Core Social Media Analytics Concepts#
- Social media datasets represent real content like videos or posts on platforms such as YouTube, Instagram, and TikTok.
- Each post or video typically has metrics like views, likes, comments, shares, watch time, and click-through rate (CTR).
- High engagement usually means users are liking, sharing, or commenting on content more than average.
- It is important to remember that a high view count is not always equal to high impactengagement rate (engagement divided by views) shows true audience reaction.
- Beginners often make mistakes by looking only at total numbers, ignoring ratios like engagement rate and CTR.
# Load the YouTube Trending Videos dataset (with fallback if live fetch fails)
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request
SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
def get_yt_service():
api_key = os.environ.get('YOUTUBE_API_KEY')
if api_key:
return build('youtube', 'v3', developerKey=api_key)
if Path('client_secret.json').exists():
creds = None
if Path('token_ro.pickle').exists():
with open('token_ro.pickle', 'rb') as f:
creds = pickle.load(f)
if not creds or not creds.valid:
if creds and creds.expired and creds.refresh_token:
creds.refresh(Request())
else:
flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
creds = flow.run_local_server(port=0)
with open('token_ro.pickle', 'wb') as f:
pickle.dump(creds, f)
return build('youtube', 'v3', credentials=creds)
raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
def fetch_yt_trending(max_results=200, region='US'):
youtube = get_yt_service()
records, token = [], None
while len(records) < max_results:
resp = youtube.videos().list(
part='snippet,statistics',
chart='mostPopular',
regionCode=region,
maxResults=min(50, max_results - len(records)),
pageToken=token
).execute()
for item in resp.get('items', []):
s = item['snippet']; st = item.get('statistics', {})
records.append({
'video_id': item['id'],
'trending_date': pd.Timestamp.today().date(),
'title': s.get('title', ''),
'channel_title': s.get('channelTitle', ''),
'category_id': s.get('categoryId', ''),
'views': int(st.get('viewCount', 0)),
'likes': int(st.get('likeCount', 0)),
'comment_count': int(st.get('commentCount', 0)),
})
token = resp.get('nextPageToken')
if not token: break
return pd.DataFrame(records)
try:
df = fetch_yt_trending()
print('Live trending data:', df.shape)
except Exception as e:
print(f'Falling back to synthetic: {e}')
np.random.seed(42)
n = 1000
df = pd.DataFrame({
'video_id': [f'vid{i}' for i in range(n)],
'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
'title': [f'Video Title {i}' for i in range(n)],
'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
'category_id': np.random.choice([1,2,10,22,24,28], n),
'views': np.random.randint(10000, 5000000, n),
'likes': np.random.randint(100, 200000, n),
'comment_count': np.random.randint(10, 50000, n),
})
print('Synthetic fallback:', df.shape)
print(df.head(3))
# Preview basic dataset info and column types
print('Dataset columns:', df.columns.tolist())
print('Data types:')
print(df.dtypes)
# Beginner: Find the video with the highest view count
top_view = df.loc[df['views'].idxmax()]
print('Video with the most views:')
print(top_view[['video_id', 'title', 'views']])
# Beginner: Calculate like rate for each video
df['like_rate'] = df['likes'] / df['views'] * 100
print('Like rates (first 3 rows):')
print(df[['title','views','likes','like_rate']].head(3))
# Beginner: Calculate comment rate for each video
df['comment_rate'] = df['comment_count'] / df['views'] * 100
print('Comment rates (first 3 rows):')
print(df[['title','views','comment_count','comment_rate']].head(3))
# Intermediate: Which categories have the highest average like rate?
cat_like = df.groupby('category_id')['like_rate'].mean().sort_values(ascending=False)
print('Average like rate by category:')
print(cat_like.head(5))
# Intermediate: Find top 5 videos by like rate (minimum 10,000 views for fairness)
df10k = df[df['views'] >= 10000]
top5_like = df10k.sort_values('like_rate', ascending=False).head(5)
print('Top 5 content by like rate (min 10k views):')
print(top5_like[['title','views','likes','like_rate']])
# Intermediate: Correlation between likes and comments
corr_likes_comments = df['likes'].corr(df['comment_count'])
print(f'Correlation between likes and comments: {corr_likes_comments:.2f}')
# Intermediate: Which channels consistently produce high like rate videos?
chan_like = df.groupby('channel_title')['like_rate'].mean().sort_values(ascending=False)
print('Top channels by average like rate:')
print(chan_like.head(5))
# Intermediate: Identify viral outlier videos (like rate >2 std above mean and 100k+ views)
mean_like = df['like_rate'].mean()
std_like = df['like_rate'].std()
viral = df[(df['like_rate'] > mean_like + 2*std_like) & (df['views'] > 100000)]
print('Potential viral videos:')
print(viral[['title', 'views', 'likes', 'like_rate']].head(10))
# Advanced: Daily trend of average like rate over time
trend = df.groupby('trending_date')['like_rate'].mean()
print('First 5 days of average like rate trend:')
print(trend.head(5))
# Advanced: Find if video titles with strong words (e.g. 'BEST', 'AMAZING') get higher like rates
strong_words = ['best', 'amazing', 'ultimate', 'must', 'top', 'incredible']
df['strong_word_title'] = df['title'].str.lower().apply(lambda t: any(w in t for w in strong_words))
avg_like_strong = df.groupby('strong_word_title')['like_rate'].mean()
print('Average like rate by strong word presence in title:')
print(avg_like_strong)
# Advanced: Export top 20 high-impact videos (by like rate & comment rate) to CSV
hi_impact = df10k.copy()
hi_impact['impact_score'] = (hi_impact['like_rate'] + hi_impact['comment_rate']) / 2
top20 = hi_impact.sort_values('impact_score', ascending=False).head(20)
top20.to_csv('top20_high_impact.csv', index=False)
print('Exported top 20 high-impact videos to CSV.')
# Error Handling: Handle missing values in engagement metrics
df_missing = df.copy()
df_missing.loc[0, 'likes'] = np.nan
df_missing.loc[1, 'views'] = np.nan
filled = df_missing[['likes', 'views']].fillna(-1)
print(filled.head(3))
# Debugging: Accidental aggregation bug - summing like rates by channel (WRONG!)
wrong_agg = df.groupby('channel_title')['like_rate'].sum()
print(wrong_agg.head(3))
# Debugging: Misinterpreting ratios (like rate vs. raw likes)
by_views = df[['views','likes','like_rate']].sort_values('likes', ascending=False).head(5)
by_rate = df[['views','likes','like_rate']].sort_values('like_rate', ascending=False).head(5)
print('Most liked videos (raw):')
print(by_views)
print('Most engaging videos (by rate):')
print(by_rate)
# Debugging: Wrong grouping can hide high-impact outliers
overall_avg = df['like_rate'].mean()
cat_avg = df.groupby('category_id')['like_rate'].mean()
mask = cat_avg > overall_avg + 2*df['like_rate'].std()
print('Categories with unusually high engagement:')
print(cat_avg[mask])
Social Media Analytics Best Practices#
- Always check for and handle missing or outlier values before analysis.
- Benchmark each content item against similar categories or audience segments instead of a flat average.
- Analyze both rates (like rate, comment rate, engagement rate) and raw numbers to avoid misleading results.
- Track content trends over time, not just across topics.
- Experiment with content features like titles, timing, or thumbnailsand compare engagement rates before and after changes.
Common Analytics Patterns#
- Benchmark content against top performers in the same topic or format.
- Segment your audience and look for patterns unique to different groups.
- Use time series and moving averages to detect sustained growth or sudden drops.
- Always define engagement metrics clearly for meaningful comparison.
- Test content changes and collect before-and-after metrics to guide creative decisions.
# End-to-end: Raw content analytics to practical recommendation
df10k = df[df['views'] >= 10000].copy() # Use only videos with significant reach
df10k['impact_score'] = (df10k['like_rate'] + df10k['comment_rate']) / 2
top_row = df10k.sort_values('impact_score', ascending=False).iloc[0]
recommendation = f"Publish more content similar to: '{top_row['title']}' in category {top_row['category_id']} (Impact Score: {top_row['impact_score']:.2f})."
print('Content Strategy Recommendation:')
print(recommendation)
Congratulations! You have detected high-impact content features.#
- Try applying these analyses to your own YouTube, Instagram, or TikTok data.
- Continue practicing as social media platforms evolveaudience tastes can change rapidly!
- Enjoyed this? Subscribe to our channel and check out our YouTube playlist for more real-world content analytics lessons.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



