Mathew K Analytics

Lesson 23 · Social Media Content Analytics

Engagement Rate Analysis: Likes, Comments, Shares

We will learn to measure social media engagement using likes, comments, and shares. Evaluating engagement rates helps creators and brands understand what…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Engagement Rate Analysis: Likes, Comments, Shares#

  • We will learn to measure social media engagement using likes, comments, and shares.
  • Evaluating engagement rates helps creators and brands understand what audiences love.
  • You will analyze real and realistic datasets to uncover what drives the most interaction.
  • By the end, you will use engagement insights to make better content decisions.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Core Concepts in Social Media Engagement Analytics#

  • Social media datasets can describe videos, posts, dates, and their metrics.
  • Metrics like views, likes, comments, shares, and click-through rates track performance.
  • Likes show appreciation, comments reflect conversation, shares signal viral reach.
  • Engagement rate is usually the sum of likes, comments, and shares, divided by views.
  • Beginners sometimes misread engagement rates or forget that high followers may mean lower rates.
# Example 1: Load a simple synthetic social media dataset
np.random.seed(42)
n_posts = 10
views = np.random.randint(500, 15000, n_posts)
df = pd.DataFrame({
    'post_id': range(1, n_posts+1),
    'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
    'views': views,
    'likes': (views * np.random.uniform(0.03, 0.16, n_posts)).astype(int),
    'comments': (views * np.random.uniform(0.002, 0.03, n_posts)).astype(int),
    'shares': (views * np.random.uniform(0.001, 0.02, n_posts)).astype(int)
})
print(df.head())
   post_id platform  views  likes  comments  shares
0        1   TikTok   7770    233        66     146
1        2   TikTok   1360    216         6      15
2        3   TikTok   5890    649       113      49
3        4  YouTube  13918   1524       176      18
4        5   TikTok   5691    175       168      30
# Example 2: Calculate basic engagement rate for each post
df['engagement_rate'] = ((df['likes'] + df['comments'] + df['shares']) / df['views']) * 100
print(df[['post_id', 'platform', 'views', 'likes', 'comments', 'shares', 'engagement_rate']])
   post_id   platform  views  likes  comments  shares  engagement_rate
0        1     TikTok   7770    233        66     146         5.727156
1        2     TikTok   1360    216         6      15        17.426471
2        3     TikTok   5890    649       113      49        13.769100
3        4    YouTube  13918   1524       176      18        12.343728
4        5     TikTok   5691    175       168      30         6.554208
5        6  Instagram  12464    411       187      69         5.351412
6        7    YouTube  11784   1157       307     164        13.815343
7        8  Instagram   6234    511       131      78        11.549567
8        9  Instagram   6765    243        98     113         6.711013
9       10  Instagram    966    151         2       4        16.252588
# Example 3: Find the post with the highest engagement rate
top_post = df.loc[df['engagement_rate'].idxmax()]
print('Top engaging post:')
print(top_post)
Top engaging post:
post_id                    2
platform              TikTok
views                   1360
likes                    216
comments                   6
shares                    15
engagement_rate    17.426471
Name: 1, dtype: object
# Example 4: Group by platform and calculate average engagement rate
platform_eng = df.groupby('platform')['engagement_rate'].mean().sort_values(ascending=False)
print('Average engagement rate by platform:')
print(platform_eng)
Average engagement rate by platform:
platform
YouTube      13.079535
TikTok       10.869234
Instagram     9.966145
Name: engagement_rate, dtype: float64
# Example 5: Load a larger, more realistic dataset (social_media_content template)
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df = pd.DataFrame({
    'post_id': range(1, n_posts+1),
    'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
    'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
    'views': views,
    'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
    'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
    'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df.shape)
print(df.head(3))
(500, 7)
   post_id   platform                date  views  likes  comments  shares
0        1  Instagram 2023-01-01 00:00:00  15895   2356       116     253
1        2  Instagram 2023-01-01 06:00:00    960     94        16      26
2        3  Instagram 2023-01-01 12:00:00  76920   3910      2879    1185
# Example 6: Add an engagement rate column to the new dataset
df['engagement_rate'] = ((df['likes'] + df['comments'] + df['shares']) / df['views'] * 100).round(2)
print(df[['platform', 'views', 'likes', 'comments', 'shares', 'engagement_rate']].head(5))
    platform  views  likes  comments  shares  engagement_rate
0  Instagram  15895   2356       116     253            17.14
1  Instagram    960     94        16      26            14.17
2  Instagram  76920   3910      2879    1185            10.37
3    YouTube  54986   1827       488    1637             7.19
4  Instagram   6365    253       261     163            10.64
# Example 7: Find top 5 posts with highest engagement rate
top5 = df.nlargest(5, 'engagement_rate')
print('Top 5 posts by engagement rate:')
print(top5[['post_id', 'platform', 'date', 'views', 'likes', 'comments', 'shares', 'engagement_rate']])
Top 5 posts by engagement rate:
     post_id   platform                date  views  likes  comments  shares  \
351      352  Instagram 2023-03-29 18:00:00  74643  10653      3569    2036   
10        11  Instagram 2023-01-03 12:00:00  16123   2202       791     426   
383      384     TikTok 2023-04-06 18:00:00  36731   4965      1769     952   
423      424     TikTok 2023-04-16 18:00:00  53021   7572      2619     881   
87        88     TikTok 2023-01-22 18:00:00  82898  11426      3823    1910   

     engagement_rate  
351            21.78  
10             21.21  
383            20.93  
423            20.88  
87             20.70  
# Example 8: Calculate average engagement rate per platform
platform_means = df.groupby('platform')['engagement_rate'].mean().sort_values(ascending=False)
print('Average engagement rate (by platform):')
print(platform_means.round(2))
Average engagement rate (by platform):
platform
Instagram    13.04
TikTok       12.69
YouTube      12.16
Name: engagement_rate, dtype: float64
# Example 9: Plot engagement rate distribution across platforms
import matplotlib.pyplot as plt
plt.figure(figsize=(10,6))
for platform in df['platform'].unique():
    plt.hist(df[df['platform']==platform]['engagement_rate'], bins=30, alpha=0.6, label=platform)
plt.xlabel('Engagement Rate (%)')
plt.ylabel('Post Count')
plt.title('Engagement Rate Distribution by Platform')
plt.legend()
plt.tight_layout()
plt.savefig('engagement_histogram.png')
plt.show()
No description has been provided for this image
# Example 10: Analyze engagement rate trends over time (all posts)
df['date'] = pd.to_datetime(df['date'])
daily_trend = df.groupby(df['date'].dt.date)['engagement_rate'].mean()
daily_trend = daily_trend.rolling(7, min_periods=1).mean() # Smooth trend
plt.figure(figsize=(12,5))
plt.plot(daily_trend.index, daily_trend.values)
plt.title('7-Day Average Engagement Rate Over Time')
plt.xlabel('Date')
plt.ylabel('Engagement Rate (%)')
plt.tight_layout()
plt.savefig('daily_engagement_trend.png')
plt.show()
No description has been provided for this image
# Example 11: Compare engagement rate on weekends vs weekdays
df['weekday'] = df['date'].dt.dayofweek
df['is_weekend'] = df['weekday'] >= 5
mean_weekend = df[df['is_weekend']]['engagement_rate'].mean()
mean_weekday = df[~df['is_weekend']]['engagement_rate'].mean()
print(f'Average engagement rate - Weekends: {mean_weekend:.2f}%')
print(f'Average engagement rate - Weekdays: {mean_weekday:.2f}%')
Average engagement rate - Weekends: 13.06%
Average engagement rate - Weekdays: 12.44%
# Example 12: Identify the best time of day to post for highest engagement
df['hour'] = df['date'].dt.hour
hourly_means = df.groupby('hour')['engagement_rate'].mean()
best_hour = hourly_means.idxmax()
print('Average engagement rate by hour:')
print(hourly_means.round(2))
print(f'Best posting hour: {best_hour}:00')
Average engagement rate by hour:
hour
0     12.63
6     11.98
12    13.24
18    12.59
Name: engagement_rate, dtype: float64
Best posting hour: 12:00
# Example 13: Calculate median engagement rate for each platform
medians = df.groupby('platform')['engagement_rate'].median()
print('Median engagement rate by platform:')
print(medians.round(2))
Median engagement rate by platform:
platform
Instagram    12.79
TikTok       12.90
YouTube      12.01
Name: engagement_rate, dtype: float64
# Example 14: Find posts with engagement rate above 15%
high_engagement = df[df['engagement_rate'] > 15]
print(f'Posts with engagement rate > 15%: {len(high_engagement)}')
print(high_engagement[['post_id', 'platform', 'date', 'engagement_rate']].head())
Posts with engagement rate > 15%: 161
    post_id   platform                date  engagement_rate
0         1  Instagram 2023-01-01 00:00:00            17.14
10       11  Instagram 2023-01-03 12:00:00            21.21
14       15  Instagram 2023-01-04 12:00:00            18.30
17       18  Instagram 2023-01-05 06:00:00            19.03
18       19     TikTok 2023-01-05 12:00:00            18.12
# Example 15: Error handling - What if engagement columns have missing data?
df_missing = df.copy()
df_missing.loc[2:4, 'likes'] = np.nan
df_missing['engagement_rate'] = ((df_missing['likes'].fillna(0) +
                                  df_missing['comments'].fillna(0) +
                                  df_missing['shares'].fillna(0)) /
                                 df_missing['views']) * 100
print('Handled missing likes:')
print(df_missing[['likes', 'comments', 'shares', 'engagement_rate']].iloc[2:6])
Handled missing likes:
    likes  comments  shares  engagement_rate
2     NaN      2879    1185         5.283411
3     NaN       488    1637         3.864620
4     NaN       261     163         6.661430
5  4287.0      3445     581        10.078074
# Example 16: Error handling - Incorrect aggregation of engagement metrics
wrong_grouping = df.groupby('platform')[['likes', 'comments', 'shares']].sum()
tot_views = df.groupby('platform')['views'].sum()
# ERROR: Summing likes/comments/shares and dividing by summed views can mislead.
wrong_rate = ((wrong_grouping['likes'] + wrong_grouping['comments'] + wrong_grouping['shares']) / tot_views) * 100
print('Incorrect aggregated engagement rate:')
print(wrong_rate)
Incorrect aggregated engagement rate:
platform
Instagram    12.984037
TikTok       13.035533
YouTube      12.358658
dtype: float64
# Example 17: Error handling - Misinterpreting ratio metrics
sample_post = df.iloc[0]
ratio = sample_post['likes'] / sample_post['views'] if sample_post['views'] > 0 else 0
print(f"Engagement ratio (likes/views) for Post {sample_post['post_id']}: {ratio:.4f}")
Engagement ratio (likes/views) for Post 1: 0.1482
# Example 18: Error handling - Wrong grouping by irrelevant columns
wrong_group = df.groupby('date')['engagement_rate'].mean().tail()
print('Mean daily engagement rate (last 5 days):')
print(wrong_group)
Mean daily engagement rate (last 5 days):
date
2023-05-04 18:00:00     7.39
2023-05-05 00:00:00    16.14
2023-05-05 06:00:00    16.78
2023-05-05 12:00:00     6.94
2023-05-05 18:00:00    17.76
Name: engagement_rate, dtype: float64

Best Practices and Common Patterns in Content Analytics#

  • Benchmark content with averages, medians, and percentiles for context.
  • Segment your audience by platform, time, or post type to reveal differences.
  • Track trends over weeks or months, not just days, for lasting insight.
  • Always define engagement rate clearly before sharing results.
  • Optimize posts based on what brings the highest engagement per audience.
# Example 19: Content benchmarking by engagement rate deciles
df['decile'] = pd.qcut(df['engagement_rate'], 10, labels=False)
benchmarks = df.groupby('decile')['engagement_rate'].agg(['count', 'mean', 'min', 'max'])
print('Engagement rate decile benchmarks:')
print(benchmarks)
Engagement rate decile benchmarks:
        count       mean    min    max
decile                                
0          51   5.848627   3.08   7.25
1          49   7.971020   7.27   8.63
2          50   9.361800   8.65  10.06
3          50  10.495400  10.07  10.91
4          50  11.835400  10.96  12.69
5          50  13.402200  12.70  14.17
6          50  14.702000  14.18  15.29
7          50  15.796800  15.31  16.58
8          50  17.405800  16.59  18.24
9          50  19.324600  18.28  21.78
# Example 20: Platform-specific weekly trend analysis
df['week'] = df['date'].dt.isocalendar().week
weekly_trends = df.groupby(['platform', 'week'])['engagement_rate'].mean().reset_index()
print('Engagement rate by platform and week (first 10):')
print(weekly_trends.head(10))
Engagement rate by platform and week (first 10):
    platform  week  engagement_rate
0  Instagram     1        14.026923
1  Instagram     2        13.943333
2  Instagram     3        14.520000
3  Instagram     4        13.148000
4  Instagram     5        12.605714
5  Instagram     6        14.071000
6  Instagram     7        13.528750
7  Instagram     8        11.671429
8  Instagram     9        11.601000
9  Instagram    10        13.646667
# Example 21: End-to-end: Create a mini content strategy recommendation
recent_week = df['week'].max()
report = []
for platform in df['platform'].unique():
    df_p = df[(df['platform']==platform) & (df['week']==recent_week)]
    if len(df_p) == 0: continue
    avg_eng = df_p['engagement_rate'].mean()
    top_post = df_p.loc[df_p['engagement_rate'].idxmax()]
    report.append({
        'platform': platform,
        'posts': len(df_p),
        'avg_engagement': round(avg_eng,2),
        'top_post_id': int(top_post['post_id']),
        'top_engagement': round(top_post['engagement_rate'],2),
        'best_hour': int(df_p.loc[df_p['engagement_rate'].idxmax()]['hour'])
    })
strategy = pd.DataFrame(report)
print('Content Strategy Recommendation - Most Recent Week:')
print(strategy)
Content Strategy Recommendation - Most Recent Week:
    platform  posts  avg_engagement  top_post_id  top_engagement  best_hour
0  Instagram      3           13.89            1           17.14          0
1    YouTube      1            7.19            4            7.19         18

End of Engagement Rate Analysis Lesson#

  • You measured and compared engagement rates using likes, comments, and shares.
  • You identified top-performing posts and audience patterns by platform and time.
  • You learned to handle errors and benchmark content fairly.
  • Use these patterns to guide your next round of content creation.
  • For more walkthroughs, check out our YouTube channel.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.