Mathew K Analytics

Lesson 2 · Social Media Content Analytics

Understanding Social Media Platforms and Metrics

In this lesson, we solve real-world social media and content analytics problems with Python. We focus on how to analyze engagement across social platforms…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Understanding Social Media Platforms and Metrics#

  • In this lesson, we solve real-world social media and content analytics problems with Python.
  • We focus on how to analyze engagement across social platforms like YouTube, Instagram, and TikTok.
  • Understanding metrics like views, likes, comments, shares, and watch time is essential for content creators and businesses.
  • By the end, you will know how to interpret platform data and draw actionable insights for improving your content strategy.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Core Concepts: Social Media Platforms and Metrics#

  • Social media platforms are online spaces for sharing content and connecting with audiences.
  • Datasets typically represent posts, videos, or comments with columns for metrics like views, likes, and comments.
  • Engagement metrics measure user interaction, such as likes, comments, shares, and watch time.
  • Click-Through Rate (CTR) shows how often viewers click after seeing a link or thumbnail.
  • Common mistakes include treating views alone as success and forgetting to account for audience size or platform differences.
np.random.seed(42)
n_posts = 12
views = np.random.randint(100, 100000, n_posts)
df = pd.DataFrame({
    'post_id': range(1, n_posts+1),
    'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
    'date': pd.date_range('2023-01-01', periods=n_posts, freq='1D'),
    'views': views,
    'likes': (views * np.random.uniform(0.04, 0.12, n_posts)).astype(int),
    'comments': (views * np.random.uniform(0.002, 0.03, n_posts)).astype(int),
    'shares': (views * np.random.uniform(0.001, 0.02, n_posts)).astype(int)
})
print(df.head())
   post_id   platform       date  views  likes  comments  shares
0        1  Instagram 2023-01-01  15895   1006       107     290
1        2    YouTube 2023-01-02    960     85         3       5
2        3  Instagram 2023-01-03  76920   3935      2197    1045
3        4  Instagram 2023-01-04  54986   3484      1596     380
4        5  Instagram 2023-01-05   6365    441       156      69
df['engagement_rate'] = ((df['likes'] + df['comments'] + df['shares']) / df['views']).round(3)
print(df[['platform','views','likes','comments','shares','engagement_rate']].head())
    platform  views  likes  comments  shares  engagement_rate
0  Instagram  15895   1006       107     290            0.088
1    YouTube    960     85         3       5            0.097
2  Instagram  76920   3935      2197    1045            0.093
3  Instagram  54986   3484      1596     380            0.099
4  Instagram   6365    441       156      69            0.105
top = df.sort_values('engagement_rate', ascending=False).head(3)
print('Top 3 posts by engagement rate:')
print(top[['post_id', 'platform', 'engagement_rate']])
Top 3 posts by engagement rate:
   post_id   platform  engagement_rate
9       10  Instagram            0.112
6        7    YouTube            0.112
8        9  Instagram            0.111
by_platform = df.groupby('platform').agg({'views':'mean','likes':'mean','comments':'mean','shares':'mean','engagement_rate':'mean'})
print('Average metrics by platform:')
print(by_platform.round(2))
Average metrics by platform:
              views    likes  comments  shares  engagement_rate
platform                                                       
Instagram  48749.43  3433.86    840.43  650.71              0.1
YouTube    36633.00  2635.00    481.80  534.60              0.1
recent = df[df['date']>='2023-01-06']
trend = recent.groupby('platform').agg({'engagement_rate':'mean','views':'sum'})
print('Engagement rate and total views for recent posts:')
print(trend.round(3))
Engagement rate and total views for recent posts:
           engagement_rate   views
platform                          
Instagram            0.107  187080
YouTube              0.098  182205
df['likes_per_comment'] = (df['likes'] / df['comments']).replace([np.inf, -np.inf], np.nan).round(2)
print(df[['platform','likes','comments','likes_per_comment']].head(7))
    platform  likes  comments  likes_per_comment
0  Instagram   1006       107               9.40
1    YouTube     85         3              28.33
2  Instagram   3935      2197               1.79
3  Instagram   3484      1596               2.18
4  Instagram    441       156               2.83
5  Instagram   6308       868               7.27
6    YouTube   3834       176              21.78

What is Click-Through Rate (CTR)?#

  • CTR is the ratio of clicks to impressions (how many viewers clicked after seeing the post or video).
  • For videos, it often measures how effectively your title and thumbnail attract clicks.
  • High CTR indicates strong interest but does not always mean high engagement after clicking.
  • CTR is crucial for evaluating video performance, especially on platforms like YouTube.
# Load YouTube video analytics data
np.random.seed(42)
n_videos = 10
views = np.random.randint(150, 50000, n_videos)
df_yt = pd.DataFrame({
    'video_id': range(1, n_videos+1),
    'publish_date': pd.date_range('2023-01-01', periods=n_videos, freq='D'),
    'views': views,
    'watch_time': np.random.randint(1000, 20000, n_videos),
    'likes': (views * np.random.uniform(0.01, 0.09, n_videos)).astype(int),
    'comments': (views * np.random.uniform(0.002, 0.018, n_videos)).astype(int),
    'ctr': np.round(np.random.uniform(2, 11, n_videos), 2)
})
print(df_yt.head())
   video_id publish_date  views  watch_time  likes  comments   ctr
0         1   2023-01-01  15945       12363    547        82  4.74
1         2   2023-01-02   1010       17023     52        10  2.88
2         3   2023-01-03  38308        9322   1706       439  8.16
3         4   2023-01-04  44882        2685   1494       123  5.96
4         5   2023-01-05  11434        1769    674       134  3.10
df_yt['engagement_rate'] = ((df_yt['likes'] + df_yt['comments']) / df_yt['views']).round(3)
print('YouTube engagement rate vs CTR:')
print(df_yt[['video_id','ctr','engagement_rate']])
YouTube engagement rate vs CTR:
   video_id    ctr  engagement_rate
0         1   4.74            0.039
1         2   2.88            0.061
2         3   8.16            0.056
3         4   5.96            0.036
4         5   3.10            0.071
5         6   6.46            0.026
6         7   2.31            0.036
7         8  10.18            0.056
8         9   4.33            0.064
9        10   7.96            0.088
viral = df_yt[(df_yt['views'] > 30000) & (df_yt['ctr'] > 8)]
print('Viral candidates:')
print(viral[['video_id','views','ctr','engagement_rate']])
Viral candidates:
   video_id  views    ctr  engagement_rate
2         3  38308   8.16            0.056
7         8  37344  10.18            0.056
# Simulate missing engagement data for one video
df_yt.loc[3, 'likes'] = np.nan
df_yt.loc[3, 'comments'] = np.nan
df_yt['engagement_rate'] = ((df_yt['likes'] + df_yt['comments']) / df_yt['views']).round(3)
print(df_yt[['video_id','views','likes','comments','engagement_rate']])
   video_id  views   likes  comments  engagement_rate
0         1  15945   547.0      82.0            0.039
1         2   1010    52.0      10.0            0.061
2         3  38308  1706.0     439.0            0.056
3         4  44882     NaN       NaN              NaN
4         5  11434   674.0     134.0            0.071
5         6   6415   135.0      30.0            0.026
6         7  17000   567.0      51.0            0.036
7         8  37344  1467.0     641.0            0.056
8         9  22112  1027.0     385.0            0.064
9        10  47341  3447.0     707.0            0.088
# Example: Incorrect aggregation mistake
totals_wrong = df_yt.groupby('publish_date')['likes'].mean()
print('INCORRECT: average likes by date (should sum or median for daily recap):')
print(totals_wrong.head())
INCORRECT: average likes by date (should sum or median for daily recap):
publish_date
2023-01-01     547.0
2023-01-02      52.0
2023-01-03    1706.0
2023-01-04       NaN
2023-01-05     674.0
Name: likes, dtype: float64
# Example: Misinterpreting CTR calculation
impressions = np.random.randint(5000, 100000, n_videos)
df_yt['computed_ctr'] = (df_yt['views'] / impressions * 100).round(2)
print(df_yt[['video_id','views','computed_ctr','ctr']].head())
   video_id  views  computed_ctr   ctr
0         1  15945         65.20  4.74
1         2   1010          1.41  2.88
2         3  38308         46.61  8.16
3         4  44882         53.46  5.96
4         5  11434         19.72  3.10
# Example: Wrong grouping logic for engagement by category
df_yt['category'] = np.random.choice(['Education','Entertainment','News'], n_videos)
wrong_group = df_yt.groupby('video_id')['engagement_rate'].mean()
correct_group = df_yt.groupby('category')['engagement_rate'].mean()
print('Wrong: Grouping by video_id gives no summary.')
print(wrong_group.head(3))
print('Correct: Grouping by category for overall engagement.')
print(correct_group.round(3))
Wrong: Grouping by video_id gives no summary.
video_id
1    0.039
2    0.061
3    0.056
Name: engagement_rate, dtype: float64
Correct: Grouping by category for overall engagement.
category
Education        0.064
Entertainment    0.056
News             0.051
Name: engagement_rate, dtype: float64

Best Practices and Analytics Patterns#

  • Always define your engagement metric clearly and consistently across datasets.
  • Benchmark content performance by comparing posts or videos to platform averages.
  • Segment your audience or content by topic, platform, or posting time for deeper insights.
  • Analyze trends over time, not just static values, to spot growth or decline.
  • Regularly review top and low-performing content to shape your strategy.
# Benchmarking content performance
mean_eng = df_yt['engagement_rate'].mean()
top_videos = df_yt[df_yt['engagement_rate'] > mean_eng]
print(f'Average engagement rate: {mean_eng:.3f}')
print(f'Number of above-average videos: {len(top_videos)}')
Average engagement rate: 0.055
Number of above-average videos: 6
# Segment by content publish timing
df_yt['weekday'] = df_yt['publish_date'].dt.day_name()
weekday_eng = df_yt.groupby('weekday')['engagement_rate'].mean().sort_values(ascending=False)
print('Average engagement rate by day of week:')
print(weekday_eng.round(3))
Average engagement rate by day of week:
weekday
Tuesday      0.072
Thursday     0.071
Monday       0.062
Sunday       0.048
Saturday     0.036
Friday       0.026
Wednesday      NaN
Name: engagement_rate, dtype: float64
# Analyze trends and growth
df_yt = df_yt.sort_values('publish_date')
rolling = df_yt['views'].rolling(window=3, min_periods=1).mean()
print('Three-video rolling average for views:')
print(rolling.round(1).values)
Three-video rolling average for views:
[15945.   8477.5 18421.  28066.7 31541.3 20910.3 11616.3 20253.  25485.3
 35599. ]
# End-to-end example: Recommend posting strategy
weekday_avg = df_yt.groupby('weekday')['engagement_rate'].mean()
best_day = weekday_avg.idxmax()
print(f'Recommendation: Based on current data, posting on {best_day} yields the highest engagement rate.')
Recommendation: Based on current data, posting on Tuesday yields the highest engagement rate.
# Optional: Save results to CSV
df_yt.to_csv('yt_engagement_summary.csv', index=False)
print('Analytics summary saved as yt_engagement_summary.csv.')
Analytics summary saved as yt_engagement_summary.csv.

Lesson Wrap-Up#

  • You have learned to interpret key metrics on major social media platforms.
  • You can now benchmark posts, analyze trends, and recommend actionable strategies.
  • Practice with your own data and keep exploring analytics to improve content success.
  • Like and subscribe to our YouTube channel for more hands-on analytics lessons!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.