Lesson 23 · Social Media Content Analytics
Engagement Rate Analysis: Likes, Comments, Shares
We will learn to measure social media engagement using likes, comments, and shares. Evaluating engagement rates helps creators and brands understand what…
- CourseSocial Media Content Analytics
- Lesson23 of 41
- Video25 min
- FormatJupyter notebook · 22 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbEngagement Rate Analysis: Likes, Comments, Shares#
- We will learn to measure social media engagement using likes, comments, and shares.
- Evaluating engagement rates helps creators and brands understand what audiences love.
- You will analyze real and realistic datasets to uncover what drives the most interaction.
- By the end, you will use engagement insights to make better content decisions.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
Core Concepts in Social Media Engagement Analytics#
- Social media datasets can describe videos, posts, dates, and their metrics.
- Metrics like views, likes, comments, shares, and click-through rates track performance.
- Likes show appreciation, comments reflect conversation, shares signal viral reach.
- Engagement rate is usually the sum of likes, comments, and shares, divided by views.
- Beginners sometimes misread engagement rates or forget that high followers may mean lower rates.
# Example 1: Load a simple synthetic social media dataset
np.random.seed(42)
n_posts = 10
views = np.random.randint(500, 15000, n_posts)
df = pd.DataFrame({
'post_id': range(1, n_posts+1),
'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
'views': views,
'likes': (views * np.random.uniform(0.03, 0.16, n_posts)).astype(int),
'comments': (views * np.random.uniform(0.002, 0.03, n_posts)).astype(int),
'shares': (views * np.random.uniform(0.001, 0.02, n_posts)).astype(int)
})
print(df.head())
# Example 2: Calculate basic engagement rate for each post
df['engagement_rate'] = ((df['likes'] + df['comments'] + df['shares']) / df['views']) * 100
print(df[['post_id', 'platform', 'views', 'likes', 'comments', 'shares', 'engagement_rate']])
# Example 3: Find the post with the highest engagement rate
top_post = df.loc[df['engagement_rate'].idxmax()]
print('Top engaging post:')
print(top_post)
# Example 4: Group by platform and calculate average engagement rate
platform_eng = df.groupby('platform')['engagement_rate'].mean().sort_values(ascending=False)
print('Average engagement rate by platform:')
print(platform_eng)
# Example 5: Load a larger, more realistic dataset (social_media_content template)
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df = pd.DataFrame({
'post_id': range(1, n_posts+1),
'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
'views': views,
'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df.shape)
print(df.head(3))
# Example 6: Add an engagement rate column to the new dataset
df['engagement_rate'] = ((df['likes'] + df['comments'] + df['shares']) / df['views'] * 100).round(2)
print(df[['platform', 'views', 'likes', 'comments', 'shares', 'engagement_rate']].head(5))
# Example 7: Find top 5 posts with highest engagement rate
top5 = df.nlargest(5, 'engagement_rate')
print('Top 5 posts by engagement rate:')
print(top5[['post_id', 'platform', 'date', 'views', 'likes', 'comments', 'shares', 'engagement_rate']])
# Example 8: Calculate average engagement rate per platform
platform_means = df.groupby('platform')['engagement_rate'].mean().sort_values(ascending=False)
print('Average engagement rate (by platform):')
print(platform_means.round(2))
# Example 9: Plot engagement rate distribution across platforms
import matplotlib.pyplot as plt
plt.figure(figsize=(10,6))
for platform in df['platform'].unique():
plt.hist(df[df['platform']==platform]['engagement_rate'], bins=30, alpha=0.6, label=platform)
plt.xlabel('Engagement Rate (%)')
plt.ylabel('Post Count')
plt.title('Engagement Rate Distribution by Platform')
plt.legend()
plt.tight_layout()
plt.savefig('engagement_histogram.png')
plt.show()
# Example 10: Analyze engagement rate trends over time (all posts)
df['date'] = pd.to_datetime(df['date'])
daily_trend = df.groupby(df['date'].dt.date)['engagement_rate'].mean()
daily_trend = daily_trend.rolling(7, min_periods=1).mean() # Smooth trend
plt.figure(figsize=(12,5))
plt.plot(daily_trend.index, daily_trend.values)
plt.title('7-Day Average Engagement Rate Over Time')
plt.xlabel('Date')
plt.ylabel('Engagement Rate (%)')
plt.tight_layout()
plt.savefig('daily_engagement_trend.png')
plt.show()
# Example 11: Compare engagement rate on weekends vs weekdays
df['weekday'] = df['date'].dt.dayofweek
df['is_weekend'] = df['weekday'] >= 5
mean_weekend = df[df['is_weekend']]['engagement_rate'].mean()
mean_weekday = df[~df['is_weekend']]['engagement_rate'].mean()
print(f'Average engagement rate - Weekends: {mean_weekend:.2f}%')
print(f'Average engagement rate - Weekdays: {mean_weekday:.2f}%')
# Example 12: Identify the best time of day to post for highest engagement
df['hour'] = df['date'].dt.hour
hourly_means = df.groupby('hour')['engagement_rate'].mean()
best_hour = hourly_means.idxmax()
print('Average engagement rate by hour:')
print(hourly_means.round(2))
print(f'Best posting hour: {best_hour}:00')
# Example 13: Calculate median engagement rate for each platform
medians = df.groupby('platform')['engagement_rate'].median()
print('Median engagement rate by platform:')
print(medians.round(2))
# Example 14: Find posts with engagement rate above 15%
high_engagement = df[df['engagement_rate'] > 15]
print(f'Posts with engagement rate > 15%: {len(high_engagement)}')
print(high_engagement[['post_id', 'platform', 'date', 'engagement_rate']].head())
# Example 15: Error handling - What if engagement columns have missing data?
df_missing = df.copy()
df_missing.loc[2:4, 'likes'] = np.nan
df_missing['engagement_rate'] = ((df_missing['likes'].fillna(0) +
df_missing['comments'].fillna(0) +
df_missing['shares'].fillna(0)) /
df_missing['views']) * 100
print('Handled missing likes:')
print(df_missing[['likes', 'comments', 'shares', 'engagement_rate']].iloc[2:6])
# Example 16: Error handling - Incorrect aggregation of engagement metrics
wrong_grouping = df.groupby('platform')[['likes', 'comments', 'shares']].sum()
tot_views = df.groupby('platform')['views'].sum()
# ERROR: Summing likes/comments/shares and dividing by summed views can mislead.
wrong_rate = ((wrong_grouping['likes'] + wrong_grouping['comments'] + wrong_grouping['shares']) / tot_views) * 100
print('Incorrect aggregated engagement rate:')
print(wrong_rate)
# Example 17: Error handling - Misinterpreting ratio metrics
sample_post = df.iloc[0]
ratio = sample_post['likes'] / sample_post['views'] if sample_post['views'] > 0 else 0
print(f"Engagement ratio (likes/views) for Post {sample_post['post_id']}: {ratio:.4f}")
# Example 18: Error handling - Wrong grouping by irrelevant columns
wrong_group = df.groupby('date')['engagement_rate'].mean().tail()
print('Mean daily engagement rate (last 5 days):')
print(wrong_group)
Best Practices and Common Patterns in Content Analytics#
- Benchmark content with averages, medians, and percentiles for context.
- Segment your audience by platform, time, or post type to reveal differences.
- Track trends over weeks or months, not just days, for lasting insight.
- Always define engagement rate clearly before sharing results.
- Optimize posts based on what brings the highest engagement per audience.
# Example 19: Content benchmarking by engagement rate deciles
df['decile'] = pd.qcut(df['engagement_rate'], 10, labels=False)
benchmarks = df.groupby('decile')['engagement_rate'].agg(['count', 'mean', 'min', 'max'])
print('Engagement rate decile benchmarks:')
print(benchmarks)
# Example 20: Platform-specific weekly trend analysis
df['week'] = df['date'].dt.isocalendar().week
weekly_trends = df.groupby(['platform', 'week'])['engagement_rate'].mean().reset_index()
print('Engagement rate by platform and week (first 10):')
print(weekly_trends.head(10))
# Example 21: End-to-end: Create a mini content strategy recommendation
recent_week = df['week'].max()
report = []
for platform in df['platform'].unique():
df_p = df[(df['platform']==platform) & (df['week']==recent_week)]
if len(df_p) == 0: continue
avg_eng = df_p['engagement_rate'].mean()
top_post = df_p.loc[df_p['engagement_rate'].idxmax()]
report.append({
'platform': platform,
'posts': len(df_p),
'avg_engagement': round(avg_eng,2),
'top_post_id': int(top_post['post_id']),
'top_engagement': round(top_post['engagement_rate'],2),
'best_hour': int(df_p.loc[df_p['engagement_rate'].idxmax()]['hour'])
})
strategy = pd.DataFrame(report)
print('Content Strategy Recommendation - Most Recent Week:')
print(strategy)
End of Engagement Rate Analysis Lesson#
- You measured and compared engagement rates using likes, comments, and shares.
- You identified top-performing posts and audience patterns by platform and time.
- You learned to handle errors and benchmark content fairly.
- Use these patterns to guide your next round of content creation.
- For more walkthroughs, check out our YouTube channel.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



