Mathew K Analytics

Lesson 31 · Social Media Content Analytics

Identifying Top-Performing Content in Social Media Analytics

In this lesson, we will learn how to identify which social media posts or videos perform the best using real analytics data. Knowing what content works…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Identifying Top-Performing Content in Social Media Analytics#

  • In this lesson, we will learn how to identify which social media posts or videos perform the best using real analytics data.
  • Knowing what content works helps creators and businesses grow their audience and make smarter decisions.
  • We will analyze engagement data, spot top performers, and gain insights that lead to stronger content strategies.
  • By the end, you will recognize high-impact content and avoid beginner mistakes when interpreting metrics.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Core Concepts: Social Media Content and Performance Metrics#

  • Social media content can include videos, posts, comments, and engagement data from platforms like YouTube, Instagram, and TikTok.
  • Engagement metrics such as views, likes, comments, shares, CTR (click-through rate), and watch time signal content performance.
  • High views do not always mean high engagementratios and context matter.
  • Beginner mistake: Focusing on a single metric and ignoring how audiences interact with each content type.
  • It is important to calculate engagement rates and recognize patterns across multiple metrics.
# Example 1: Load a sample social media content dataset
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df = pd.DataFrame({
    'post_id': range(1, n_posts+1),
    'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
    'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
    'views': views,
    'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
    'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
    'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df.head(3))
   post_id   platform                date  views  likes  comments  shares
0        1  Instagram 2023-01-01 00:00:00  15895   2356       116     253
1        2  Instagram 2023-01-01 06:00:00    960     94        16      26
2        3  Instagram 2023-01-01 12:00:00  76920   3910      2879    1185
# Example 2: Calculate engagement rate for each post
df['engagement_rate'] = ((df['likes'] + df['comments'] + df['shares']) / df['views'] * 100).round(2)
print(df[['post_id', 'platform', 'views', 'likes', 'comments', 'shares', 'engagement_rate']].head())
   post_id   platform  views  likes  comments  shares  engagement_rate
0        1  Instagram  15895   2356       116     253            17.14
1        2  Instagram    960     94        16      26            14.17
2        3  Instagram  76920   3910      2879    1185            10.37
3        4    YouTube  54986   1827       488    1637             7.19
4        5  Instagram   6365    253       261     163            10.64
# Example 3: Identify the single top-performing post by engagement rate
top_post = df.sort_values('engagement_rate', ascending=False).iloc[0]
print(f"Top Post ID: {top_post['post_id']}")
print(f"Platform: {top_post['platform']}")
print(f"Views: {top_post['views']}")
print(f"Engagement Rate: {top_post['engagement_rate']}%")
Top Post ID: 352
Platform: Instagram
Views: 74643
Engagement Rate: 21.78%
# Example 4: List Top 5 Posts by Absolute Engagement (Total likes + comments + shares)
df['total_engagement'] = df['likes'] + df['comments'] + df['shares']
top5_abs = df.sort_values('total_engagement', ascending=False).head(5)
print(top5_abs[['post_id', 'platform', 'views', 'likes', 'comments', 'shares', 'total_engagement', 'engagement_rate']])
     post_id   platform  views  likes  comments  shares  total_engagement  \
392      393     TikTok  99813  12361      4424    2837             19622   
192      193  Instagram  97604  13347      3566    2493             19406   
276      277  Instagram  96701  13427      3435    2410             19272   
74        75  Instagram  99399  13323      1894    2655             17872   
422      423    YouTube  93948  12615      2380    2571             17566   

     engagement_rate  
392            19.66  
192            19.88  
276            19.93  
74             17.98  
422            18.70  
# Example 5: Analyze top-performing content by platform
platform_top = df.groupby('platform')['engagement_rate'].mean().sort_values(ascending=False)
print(platform_top)
platform
Instagram    13.044444
TikTok       12.690850
YouTube      12.163027
Name: engagement_rate, dtype: float64
# Example 6: Visualize the distribution of engagement rates using a histogram
import matplotlib.pyplot as plt
plt.figure(figsize=(8,4))
plt.hist(df['engagement_rate'], bins=30, color='skyblue', edgecolor='black')
plt.title('Distribution of Engagement Rates')
plt.xlabel('Engagement Rate (%)')
plt.ylabel('Number of Posts')
plt.show()
No description has been provided for this image
# Example 7: Find the best post per platform by engagement rate
best_per_platform = df.loc[df.groupby('platform')['engagement_rate'].idxmax()][['platform', 'post_id', 'engagement_rate', 'views', 'likes', 'comments', 'shares']]
print(best_per_platform)
      platform  post_id  engagement_rate  views  likes  comments  shares
351  Instagram      352            21.78  74643  10653      3569    2036
383     TikTok      384            20.93  36731   4965      1769     952
245    YouTube      246            19.74  51763   6535      2192    1489
# Example 8: Average engagement rate trend over time
df['date'] = pd.to_datetime(df['date'])
daily_rates = df.groupby(df['date'].dt.date)['engagement_rate'].mean()
plt.figure(figsize=(10,4))
plt.plot(daily_rates.index, daily_rates.values, marker='o', linewidth=2)
plt.title('Average Daily Engagement Rate Over Time')
plt.xlabel('Date')
plt.ylabel('Avg Engagement Rate (%)')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Example 9: Correlation between views and engagement rate
corr = df['views'].corr(df['engagement_rate'])
print(f"Correlation between views and engagement rate: {corr:.2f}")
Correlation between views and engagement rate: 0.07
# Example 10: Spotlight - Most commented post
most_comments = df.loc[df['comments'].idxmax()]
print(f"Most Commented Post ID: {most_comments['post_id']}, Platform: {most_comments['platform']}, Comments: {most_comments['comments']}, Engagement Rate: {most_comments['engagement_rate']}%")
Most Commented Post ID: 78, Platform: TikTok, Comments: 4590, Engagement Rate: 13.72%
# Example 11: Calculate median engagement rate per platform
median_rates = df.groupby('platform')['engagement_rate'].median()
print(median_rates)
platform
Instagram    12.79
TikTok       12.90
YouTube      12.01
Name: engagement_rate, dtype: float64
# Example 12: Detect possible viral posts (outliers)
q3 = df['engagement_rate'].quantile(0.75)
iqr = q3 - df['engagement_rate'].quantile(0.25)
viral_threshold = q3 + 1.5*iqr
viral_posts = df[df['engagement_rate'] > viral_threshold]
print(f"Number of viral posts detected: {len(viral_posts)}")
print(viral_posts[['post_id', 'platform', 'engagement_rate']].head())
Number of viral posts detected: 0
Empty DataFrame
Columns: [post_id, platform, engagement_rate]
Index: []
# Example 13: Debugging - What if likes or comments are missing?
df_missing = df.copy()
df_missing.loc[df_missing.sample(frac=0.05, random_state=42).index, 'likes'] = np.nan
df_missing['engagement_rate_fixed'] = ((df_missing['likes'].fillna(0) + df_missing['comments'] + df_missing['shares']) / df_missing['views'] * 100).round(2)
print(df_missing[['likes', 'comments', 'engagement_rate', 'engagement_rate_fixed']].head(8))
    likes  comments  engagement_rate  engagement_rate_fixed
0  2356.0       116            17.14                  17.14
1    94.0        16            14.17                  14.17
2  3910.0      2879            10.37                  10.37
3  1827.0       488             7.19                   7.19
4   253.0       261            10.64                  10.64
5  4287.0      3445            10.08                  10.08
6  1524.0       964             9.47                   9.47
7  3876.0       115             4.99                   4.99
# Example 14: Debugging - Incorrect aggregation: Are we double-counting engagement?
df_bad = df.copy()
df_bad['bad_total'] = df_bad['likes'] + df_bad['comments'] + df_bad['likes']  # Do not count likes twice!
print(df_bad[['likes', 'comments', 'bad_total']].head())
   likes  comments  bad_total
0   2356       116       4828
1     94        16        204
2   3910      2879      10699
3   1827       488       4142
4    253       261        767
# Example 15: Debugging - Misinterpreting engagement rate: Watch denominator!
df_broken = df.copy()
df_broken['engagement_wrong'] = ((df_broken['likes'] + df_broken['comments'] + df_broken['shares']) / (df_broken['likes'] + 1) * 100).round(2)
print(df_broken[['likes', 'comments', 'shares', 'views', 'engagement_rate', 'engagement_wrong']].head())
   likes  comments  shares  views  engagement_rate  engagement_wrong
0   2356       116     253  15895            17.14            115.61
1     94        16      26    960            14.17            143.16
2   3910      2879    1185  76920            10.37            203.89
3   1827       488    1637  54986             7.19            216.19
4    253       261     163   6365            10.64            266.54
# Example 16: Debugging - Group by wrong field
by_post = df.groupby('post_id')['engagement_rate'].max().head()
print(by_post)
# Correct: group by platform
by_platform = df.groupby('platform')['engagement_rate'].max()
print(by_platform)
post_id
1    17.14
2    14.17
3    10.37
4     7.19
5    10.64
Name: engagement_rate, dtype: float64
platform
Instagram    21.78
TikTok       20.93
YouTube      19.74
Name: engagement_rate, dtype: float64

Best Practices: Social Media Content Analytics#

  • Benchmark performance by comparing current content to historical averages.
  • Segment your audience and content types to reveal hidden strengths.
  • Analyze performance trends over time, not just for a single post.
  • Always define your metrics clearlybe consistent in how you compute and report them.
  • Use outlier detection to spot viral or underperforming content early.
  • Regularly review and refine your analytics strategy based on your content goals.
# Advanced Example 1: Top-performing content by week and platform
df['week'] = df['date'].dt.isocalendar().week
weekly_platform = df.groupby(['platform', 'week'])['engagement_rate'].mean().reset_index()
top_weekly = weekly_platform.sort_values(['week', 'engagement_rate'], ascending=[True, False]).groupby('week').head(1)
print(top_weekly[['week', 'platform', 'engagement_rate']].head())
    week   platform  engagement_rate
0      1  Instagram        14.026923
1      2  Instagram        13.943333
2      3  Instagram        14.520000
3      4  Instagram        13.148000
41     5    YouTube        13.827273
# Advanced Example 2: Analyze YouTube trending video data for top performance
try:
    yt_df = fetch_yt_trending(max_results=200, region='US')
    print('Loaded live YouTube trending data.')
except Exception as e:
    np.random.seed(42)
    n = 1000
    yt_df = pd.DataFrame({
        'video_id':      [f'vid{i}' for i in range(n)],
        'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
        'title':         [f'Video Title {i}' for i in range(n)],
        'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
        'category_id':   np.random.choice([1,2,10,22,24,28], n),
        'views':         np.random.randint(10000, 5000000, n),
        'likes':         np.random.randint(100, 200000, n),
        'comment_count': np.random.randint(10, 50000, n),
    })
yt_df['engagement_rate'] = ((yt_df['likes'] + yt_df['comment_count']) / yt_df['views'] * 100).round(2)
yt_top = yt_df.sort_values('engagement_rate', ascending=False).head(5)
print(yt_top[['video_id', 'title', 'channel_title', 'views', 'likes', 'comment_count', 'engagement_rate']])
    video_id            title channel_title  views   likes  comment_count  \
675   vid675  Video Title 675      ChannelC  39725  199875          27974   
155   vid155  Video Title 155      ChannelC  38625  110255          23746   
821   vid821  Video Title 821      ChannelA  29675   93674           1639   
71     vid71   Video Title 71      ChannelB  71367  154555          43008   
301   vid301  Video Title 301      ChannelA  59104  100032          48702   

     engagement_rate  
675           573.57  
155           346.93  
821           321.19  
71            276.83  
301           251.65  
# Advanced Example 3: Save the top-performing content analysis as a report CSV
top10 = df.sort_values('engagement_rate', ascending=False).head(10)
top10.to_csv('top_content_report.csv', index=False)
print('Top 10 posts written to top_content_report.csv')
Top 10 posts written to top_content_report.csv

End-to-End Problem: Find Content to Boost Next Month#

  • Your client wants to know what kind of social posts they should double down on next month.
  • You will use this month's top-performing posts to inform a data-driven content strategy.
  • We will identify patterns, surface winning posts, and extract actionable recommendations.
# End-to-End: Step 1 - Filter to last 30 days of data
latest_date = df['date'].max()
cutoff = latest_date - pd.Timedelta(days=30)
last_month_df = df[df['date'] > cutoff].copy()
print(f"Posts in last 30 days: {len(last_month_df)}")
Posts in last 30 days: 120
# End-to-End: Step 2 - Analyze which platforms won last month
top_platforms = last_month_df.groupby('platform')['engagement_rate'].mean().sort_values(ascending=False)
print('Average engagement rate by platform (last 30 days):')
print(top_platforms)
Average engagement rate by platform (last 30 days):
platform
TikTok       13.753111
Instagram    13.673590
YouTube      12.177222
Name: engagement_rate, dtype: float64
# End-to-End: Step 3 - Highlight top 5 post topics or types
last_month_df['post_type'] = np.where(last_month_df['likes'] > last_month_df['comments'] + last_month_df['shares'], 'Like-focused', 'Discussion/Share-focused')
topic_summary = last_month_df.groupby('post_type')['engagement_rate'].mean()
print('Average engagement by post type:')
print(topic_summary)
top5_last = last_month_df.sort_values('engagement_rate', ascending=False).head(5)[['post_id', 'platform', 'date', 'engagement_rate', 'post_type']]
print('Top 5 recent posts:')
print(top5_last)
Average engagement by post type:
post_type
Discussion/Share-focused     8.419286
Like-focused                13.893113
Name: engagement_rate, dtype: float64
Top 5 recent posts:
     post_id   platform                date  engagement_rate     post_type
383      384     TikTok 2023-04-06 18:00:00            20.93  Like-focused
423      424     TikTok 2023-04-16 18:00:00            20.88  Like-focused
429      430     TikTok 2023-04-18 06:00:00            20.09  Like-focused
484      485  Instagram 2023-05-02 00:00:00            19.70  Like-focused
392      393     TikTok 2023-04-09 00:00:00            19.66  Like-focused
# End-to-End: Step 4 - Write recommendations to text file
with open('content_strategy_recommendation.txt', 'w') as f:
    f.write('Content Strategy Recommendations for Next Month\n')
    f.write('------------------------------------------\n')
    best_plat = top_platforms.index[0]
    f.write(f'- Focus on the {best_plat} platform; it delivered the highest engagement rate.\n')
    main_type = topic_summary.idxmax()
    f.write(f'- Prioritize {main_type} posts, as they outperformed others.\n')
    f.write('- Consider posting more frequently during the days/weeks where engagement peaked.\n')
    f.write('- Use the content from the top 5 posts as templates for upcoming campaigns.\n')
print('Recommendations written to content_strategy_recommendation.txt')
Recommendations written to content_strategy_recommendation.txt

Lesson Recap: Key Takeaways on Top-Performing Content#

  • Top-performing content is not just the post with the most views, but those with high engagement relative to their reach.
  • Comparing engagement rate across platforms and weeks reveals real insights for growing your audience.
  • Robust analytics means handling missing data, avoiding mistakes in calculations, and leveraging multiple metrics.
  • Best practices help make smarter, faster decisions about what content to prioritize and promote.
  • Go experiment: Repeat these analyses on your own data for maximum learning.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.