Lesson 31 · Social Media Content Analytics
Identifying Top-Performing Content in Social Media Analytics
In this lesson, we will learn how to identify which social media posts or videos perform the best using real analytics data. Knowing what content works…
- CourseSocial Media Content Analytics
- Lesson31 of 41
- Video31 min
- FormatJupyter notebook · 24 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIdentifying Top-Performing Content in Social Media Analytics#
- In this lesson, we will learn how to identify which social media posts or videos perform the best using real analytics data.
- Knowing what content works helps creators and businesses grow their audience and make smarter decisions.
- We will analyze engagement data, spot top performers, and gain insights that lead to stronger content strategies.
- By the end, you will recognize high-impact content and avoid beginner mistakes when interpreting metrics.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
Core Concepts: Social Media Content and Performance Metrics#
- Social media content can include videos, posts, comments, and engagement data from platforms like YouTube, Instagram, and TikTok.
- Engagement metrics such as views, likes, comments, shares, CTR (click-through rate), and watch time signal content performance.
- High views do not always mean high engagementratios and context matter.
- Beginner mistake: Focusing on a single metric and ignoring how audiences interact with each content type.
- It is important to calculate engagement rates and recognize patterns across multiple metrics.
# Example 1: Load a sample social media content dataset
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df = pd.DataFrame({
'post_id': range(1, n_posts+1),
'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
'views': views,
'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df.head(3))
# Example 2: Calculate engagement rate for each post
df['engagement_rate'] = ((df['likes'] + df['comments'] + df['shares']) / df['views'] * 100).round(2)
print(df[['post_id', 'platform', 'views', 'likes', 'comments', 'shares', 'engagement_rate']].head())
# Example 3: Identify the single top-performing post by engagement rate
top_post = df.sort_values('engagement_rate', ascending=False).iloc[0]
print(f"Top Post ID: {top_post['post_id']}")
print(f"Platform: {top_post['platform']}")
print(f"Views: {top_post['views']}")
print(f"Engagement Rate: {top_post['engagement_rate']}%")
# Example 4: List Top 5 Posts by Absolute Engagement (Total likes + comments + shares)
df['total_engagement'] = df['likes'] + df['comments'] + df['shares']
top5_abs = df.sort_values('total_engagement', ascending=False).head(5)
print(top5_abs[['post_id', 'platform', 'views', 'likes', 'comments', 'shares', 'total_engagement', 'engagement_rate']])
# Example 5: Analyze top-performing content by platform
platform_top = df.groupby('platform')['engagement_rate'].mean().sort_values(ascending=False)
print(platform_top)
# Example 6: Visualize the distribution of engagement rates using a histogram
import matplotlib.pyplot as plt
plt.figure(figsize=(8,4))
plt.hist(df['engagement_rate'], bins=30, color='skyblue', edgecolor='black')
plt.title('Distribution of Engagement Rates')
plt.xlabel('Engagement Rate (%)')
plt.ylabel('Number of Posts')
plt.show()
# Example 7: Find the best post per platform by engagement rate
best_per_platform = df.loc[df.groupby('platform')['engagement_rate'].idxmax()][['platform', 'post_id', 'engagement_rate', 'views', 'likes', 'comments', 'shares']]
print(best_per_platform)
# Example 8: Average engagement rate trend over time
df['date'] = pd.to_datetime(df['date'])
daily_rates = df.groupby(df['date'].dt.date)['engagement_rate'].mean()
plt.figure(figsize=(10,4))
plt.plot(daily_rates.index, daily_rates.values, marker='o', linewidth=2)
plt.title('Average Daily Engagement Rate Over Time')
plt.xlabel('Date')
plt.ylabel('Avg Engagement Rate (%)')
plt.tight_layout()
plt.show()
# Example 9: Correlation between views and engagement rate
corr = df['views'].corr(df['engagement_rate'])
print(f"Correlation between views and engagement rate: {corr:.2f}")
# Example 10: Spotlight - Most commented post
most_comments = df.loc[df['comments'].idxmax()]
print(f"Most Commented Post ID: {most_comments['post_id']}, Platform: {most_comments['platform']}, Comments: {most_comments['comments']}, Engagement Rate: {most_comments['engagement_rate']}%")
# Example 11: Calculate median engagement rate per platform
median_rates = df.groupby('platform')['engagement_rate'].median()
print(median_rates)
# Example 12: Detect possible viral posts (outliers)
q3 = df['engagement_rate'].quantile(0.75)
iqr = q3 - df['engagement_rate'].quantile(0.25)
viral_threshold = q3 + 1.5*iqr
viral_posts = df[df['engagement_rate'] > viral_threshold]
print(f"Number of viral posts detected: {len(viral_posts)}")
print(viral_posts[['post_id', 'platform', 'engagement_rate']].head())
# Example 13: Debugging - What if likes or comments are missing?
df_missing = df.copy()
df_missing.loc[df_missing.sample(frac=0.05, random_state=42).index, 'likes'] = np.nan
df_missing['engagement_rate_fixed'] = ((df_missing['likes'].fillna(0) + df_missing['comments'] + df_missing['shares']) / df_missing['views'] * 100).round(2)
print(df_missing[['likes', 'comments', 'engagement_rate', 'engagement_rate_fixed']].head(8))
# Example 14: Debugging - Incorrect aggregation: Are we double-counting engagement?
df_bad = df.copy()
df_bad['bad_total'] = df_bad['likes'] + df_bad['comments'] + df_bad['likes'] # Do not count likes twice!
print(df_bad[['likes', 'comments', 'bad_total']].head())
# Example 15: Debugging - Misinterpreting engagement rate: Watch denominator!
df_broken = df.copy()
df_broken['engagement_wrong'] = ((df_broken['likes'] + df_broken['comments'] + df_broken['shares']) / (df_broken['likes'] + 1) * 100).round(2)
print(df_broken[['likes', 'comments', 'shares', 'views', 'engagement_rate', 'engagement_wrong']].head())
# Example 16: Debugging - Group by wrong field
by_post = df.groupby('post_id')['engagement_rate'].max().head()
print(by_post)
# Correct: group by platform
by_platform = df.groupby('platform')['engagement_rate'].max()
print(by_platform)
Best Practices: Social Media Content Analytics#
- Benchmark performance by comparing current content to historical averages.
- Segment your audience and content types to reveal hidden strengths.
- Analyze performance trends over time, not just for a single post.
- Always define your metrics clearlybe consistent in how you compute and report them.
- Use outlier detection to spot viral or underperforming content early.
- Regularly review and refine your analytics strategy based on your content goals.
# Advanced Example 1: Top-performing content by week and platform
df['week'] = df['date'].dt.isocalendar().week
weekly_platform = df.groupby(['platform', 'week'])['engagement_rate'].mean().reset_index()
top_weekly = weekly_platform.sort_values(['week', 'engagement_rate'], ascending=[True, False]).groupby('week').head(1)
print(top_weekly[['week', 'platform', 'engagement_rate']].head())
# Advanced Example 2: Analyze YouTube trending video data for top performance
try:
yt_df = fetch_yt_trending(max_results=200, region='US')
print('Loaded live YouTube trending data.')
except Exception as e:
np.random.seed(42)
n = 1000
yt_df = pd.DataFrame({
'video_id': [f'vid{i}' for i in range(n)],
'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
'title': [f'Video Title {i}' for i in range(n)],
'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
'category_id': np.random.choice([1,2,10,22,24,28], n),
'views': np.random.randint(10000, 5000000, n),
'likes': np.random.randint(100, 200000, n),
'comment_count': np.random.randint(10, 50000, n),
})
yt_df['engagement_rate'] = ((yt_df['likes'] + yt_df['comment_count']) / yt_df['views'] * 100).round(2)
yt_top = yt_df.sort_values('engagement_rate', ascending=False).head(5)
print(yt_top[['video_id', 'title', 'channel_title', 'views', 'likes', 'comment_count', 'engagement_rate']])
# Advanced Example 3: Save the top-performing content analysis as a report CSV
top10 = df.sort_values('engagement_rate', ascending=False).head(10)
top10.to_csv('top_content_report.csv', index=False)
print('Top 10 posts written to top_content_report.csv')
End-to-End Problem: Find Content to Boost Next Month#
- Your client wants to know what kind of social posts they should double down on next month.
- You will use this month's top-performing posts to inform a data-driven content strategy.
- We will identify patterns, surface winning posts, and extract actionable recommendations.
# End-to-End: Step 1 - Filter to last 30 days of data
latest_date = df['date'].max()
cutoff = latest_date - pd.Timedelta(days=30)
last_month_df = df[df['date'] > cutoff].copy()
print(f"Posts in last 30 days: {len(last_month_df)}")
# End-to-End: Step 2 - Analyze which platforms won last month
top_platforms = last_month_df.groupby('platform')['engagement_rate'].mean().sort_values(ascending=False)
print('Average engagement rate by platform (last 30 days):')
print(top_platforms)
# End-to-End: Step 3 - Highlight top 5 post topics or types
last_month_df['post_type'] = np.where(last_month_df['likes'] > last_month_df['comments'] + last_month_df['shares'], 'Like-focused', 'Discussion/Share-focused')
topic_summary = last_month_df.groupby('post_type')['engagement_rate'].mean()
print('Average engagement by post type:')
print(topic_summary)
top5_last = last_month_df.sort_values('engagement_rate', ascending=False).head(5)[['post_id', 'platform', 'date', 'engagement_rate', 'post_type']]
print('Top 5 recent posts:')
print(top5_last)
# End-to-End: Step 4 - Write recommendations to text file
with open('content_strategy_recommendation.txt', 'w') as f:
f.write('Content Strategy Recommendations for Next Month\n')
f.write('------------------------------------------\n')
best_plat = top_platforms.index[0]
f.write(f'- Focus on the {best_plat} platform; it delivered the highest engagement rate.\n')
main_type = topic_summary.idxmax()
f.write(f'- Prioritize {main_type} posts, as they outperformed others.\n')
f.write('- Consider posting more frequently during the days/weeks where engagement peaked.\n')
f.write('- Use the content from the top 5 posts as templates for upcoming campaigns.\n')
print('Recommendations written to content_strategy_recommendation.txt')
Lesson Recap: Key Takeaways on Top-Performing Content#
- Top-performing content is not just the post with the most views, but those with high engagement relative to their reach.
- Comparing engagement rate across platforms and weeks reveals real insights for growing your audience.
- Robust analytics means handling missing data, avoiding mistakes in calculations, and leveraging multiple metrics.
- Best practices help make smarter, faster decisions about what content to prioritize and promote.
- Go experiment: Repeat these analyses on your own data for maximum learning.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



