Mathew K Analytics

Lesson 3 · Social Media Content Analytics

Types of Social Media Data: Views, Likes, CTR, and Watch Time

In this lesson, we will explore the main data types in social media and content analytics. Understanding these metrics helps content creators and businesses…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Types of Social Media Data: Views, Likes, CTR, and Watch Time#

  • In this lesson, we will explore the main data types in social media and content analytics.
  • Understanding these metrics helps content creators and businesses measure performance.
  • We will learn how to work with views, likes, click-through rate (CTR), and watch time data.
  • By the end, you will be able to analyze engagement and spot top content effectively.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Understanding Social Media Analytics Data#

  • Social media datasets hold information about individual posts, videos, or campaigns.
  • Each row often represents a piece of content, with metrics like views, likes, or comments.
  • Views count how many times a post or video was seen.
  • Likes show direct positive engagement from users.
  • CTR (Click-Through Rate) measures how often viewers clicked when shown a thumbnail or link.
  • Watch time is the total time users spent viewing your content.
  • Beginners sometimes confuse high views with high engagement.
  • Analyzing all key metrics helps avoid misleading conclusions.
# Beginner Example 1: Load a sample Social Media Content dataset
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
content_df = pd.DataFrame({
    'post_id': range(1, n_posts+1),
    'platform': np.random.choice(['YouTube', 'Instagram', 'TikTok'], n_posts),
    'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
    'views': views,
    'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
    'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
    'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(content_df.shape)
print(content_df.head())
(500, 7)
   post_id   platform                date  views  likes  comments  shares
0        1  Instagram 2023-01-01 00:00:00  15895   2356       116     253
1        2  Instagram 2023-01-01 06:00:00    960     94        16      26
2        3  Instagram 2023-01-01 12:00:00  76920   3910      2879    1185
3        4    YouTube 2023-01-01 18:00:00  54986   1827       488    1637
4        5  Instagram 2023-01-02 00:00:00   6365    253       261     163
# Beginner Example 2: View summary statistics for engagement metrics
summary = content_df[['views', 'likes', 'comments', 'shares']].describe()
print(summary)
             views         likes     comments      shares
count    500.00000    500.000000   500.000000   500.00000
mean   50730.11200   4385.632000  1314.766000   780.04600
std    29391.16711   3382.711272  1132.814209   659.79979
min      306.00000     22.000000     1.000000     3.00000
25%    23966.75000   1581.000000   353.750000   237.50000
50%    52353.50000   3619.000000   964.500000   612.00000
75%    75629.00000   6636.000000  2042.500000  1176.50000
max    99813.00000  13839.000000  4590.000000  2837.00000
# Beginner Example 3: Calculate engagement rate for each post
content_df['engagement_rate'] = ((content_df['likes'] + content_df['comments'] + content_df['shares']) / content_df['views']) * 100
print(content_df[['post_id', 'platform', 'views', 'likes', 'comments', 'shares', 'engagement_rate']].head())
   post_id   platform  views  likes  comments  shares  engagement_rate
0        1  Instagram  15895   2356       116     253        17.143756
1        2  Instagram    960     94        16      26        14.166667
2        3  Instagram  76920   3910      2879    1185        10.366615
3        4    YouTube  54986   1827       488    1637         7.187284
4        5  Instagram   6365    253       261     163        10.636292
# Beginner Example 4: Find the post with the highest engagement rate
top_post = content_df.sort_values('engagement_rate', ascending=False).iloc[0]
print(f"Top engaging post has ID {top_post['post_id']}, platform {top_post['platform']}, and engagement rate {top_post['engagement_rate']:.2f}%.")
Top engaging post has ID 352, platform Instagram, and engagement rate 21.78%.
# Beginner Example 5: List top 5 posts by view count
top_views = content_df.sort_values('views', ascending=False).head(5)
print(top_views[['post_id', 'platform', 'views', 'likes', 'comments', 'shares']])
     post_id   platform  views  likes  comments  shares
392      393     TikTok  99813  12361      4424    2837
324      325     TikTok  99622   8409      1888     332
74        75  Instagram  99399  13323      1894    2655
421      422  Instagram  99257   7373      2990    1434
138      139     TikTok  98906  13839      3522     133
# Intermediate Example 1: Group posts by platform and compute mean engagement
platform_group = content_df.groupby('platform').agg({'views':'mean', 'likes':'mean', 'comments':'mean', 'shares':'mean', 'engagement_rate':'mean'}).reset_index()
print(platform_group)
    platform         views        likes     comments      shares  \
0  Instagram  52754.950617  4702.827160  1334.827160  812.067901   
1     TikTok  50206.411765  4491.143791  1269.261438  784.267974   
2    YouTube  49390.124324  4020.610811  1334.832432  748.513514   

   engagement_rate  
0        13.044316  
1        12.690745  
2        12.163010  
# Intermediate Example 2: Identify posts with above average engagement rate for their platform
is_above_avg = []
for idx, row in content_df.iterrows():
    platform = row['platform']
    avg_rate = platform_group[platform_group['platform'] == platform]['engagement_rate'].values[0]
    is_above_avg.append(row['engagement_rate'] > avg_rate)
content_df['above_platform_avg'] = is_above_avg
print(content_df[['post_id', 'platform', 'engagement_rate', 'above_platform_avg']].head(10))
   post_id   platform  engagement_rate  above_platform_avg
0        1  Instagram        17.143756                True
1        2  Instagram        14.166667                True
2        3  Instagram        10.366615               False
3        4    YouTube         7.187284               False
4        5  Instagram        10.636292               False
5        6  Instagram        10.078074               False
6        7    YouTube         9.468011               False
7        8     TikTok         4.993265               False
8        9    YouTube         9.678732               False
9       10  Instagram         8.578102               False
# Intermediate Example 3: Visualize engagement rate distribution per platform
import matplotlib.pyplot as plt
plt.figure(figsize=(7,5))
for plat in content_df['platform'].unique():
    plt.hist(content_df[content_df['platform'] == plat]['engagement_rate'], bins=20, alpha=0.5, label=plat)
plt.xlabel('Engagement Rate (%)')
plt.ylabel('Number of Posts')
plt.title('Engagement Rate Distribution by Platform')
plt.legend()
plt.show()
No description has been provided for this image
# Intermediate Example 4: Load YouTube Analytics data with CTR and watch time fields
np.random.seed(42)
n_videos = 300
views_y = np.random.randint(100, 500000, n_videos)
yt_analytics = pd.DataFrame({
    'video_id': range(1, n_videos+1),
    'publish_date': pd.date_range('2022-01-01', periods=n_videos, freq='D'),
    'views': views_y,
    'watch_time': np.random.randint(1000, 500000, n_videos),
    'likes': (views_y * np.random.uniform(0.01, 0.08, n_videos)).astype(int),
    'comments': (views_y * np.random.uniform(0.001, 0.02, n_videos)).astype(int),
    'ctr': np.round(np.random.uniform(2, 10, n_videos), 2)
})
print(yt_analytics.head())
   video_id publish_date   views  watch_time  likes  comments   ctr
0         1   2022-01-01  122058      158381   1890      1975  4.14
1         2   2022-01-02  146967      481671   1730      1900  6.99
2         3   2022-01-03  132032      195806  10217       337  5.28
3         4   2022-01-04  365938       71467  25073      6439  6.42
4         5   2022-01-05  259278      184734  15224      4795  5.49
# Intermediate Example 5: Analyze correlation between CTR and watch time
correlation = yt_analytics['ctr'].corr(yt_analytics['watch_time'])
print(f'Correlation between CTR and watch time: {correlation:.2f}')
Correlation between CTR and watch time: -0.04
# Intermediate Example 6: Compare videos with top 10% CTR vs overall watch time
ctr_threshold = np.percentile(yt_analytics['ctr'], 90)
top_ctr_videos = yt_analytics[yt_analytics['ctr'] >= ctr_threshold]
mean_watch_time_top = top_ctr_videos['watch_time'].mean()
mean_watch_time_all = yt_analytics['watch_time'].mean()
print(f'Average watch time (top 10% CTR): {mean_watch_time_top:.1f}')
print(f'Average watch time (all videos): {mean_watch_time_all:.1f}')
Average watch time (top 10% CTR): 249367.5
Average watch time (all videos): 247090.2
# Advanced Example 1: Time-based analysis of daily total views and watch time
daily_stats = yt_analytics.groupby('publish_date').agg({'views':'sum', 'watch_time':'sum', 'ctr':'mean'}).reset_index()
print(daily_stats.head())
  publish_date   views  watch_time   ctr
0   2022-01-01  122058      158381  4.14
1   2022-01-02  146967      481671  6.99
2   2022-01-03  132032      195806  5.28
3   2022-01-04  365938       71467  6.42
4   2022-01-05  259278      184734  5.49
# Advanced Example 2: Identify potential viral videos using CTR and watch time thresholds
viral_threshold_ctr = yt_analytics['ctr'].quantile(0.95)
viral_threshold_watch = yt_analytics['watch_time'].quantile(0.95)
possible_viral = yt_analytics[(yt_analytics['ctr'] >= viral_threshold_ctr) & (yt_analytics['watch_time'] >= viral_threshold_watch)]
print(f'Number of possible viral videos: {len(possible_viral)}')
print(possible_viral[['video_id', 'views', 'watch_time', 'ctr']])
Number of possible viral videos: 0
Empty DataFrame
Columns: [video_id, views, watch_time, ctr]
Index: []
# Advanced Example 3: Compute engagement rate across categories (simulating with content data)
content_df['category'] = np.where(content_df['platform'] == 'YouTube', 'Video', 'Photo/Short')
category_grouped = content_df.groupby('category').agg({'views':'mean', 'engagement_rate':'mean'}).reset_index()
print(category_grouped)
      category         views  engagement_rate
0  Photo/Short  51517.088889        12.872581
1        Video  49390.124324        12.163010
# Error Handling Example 1: Handling missing engagement values
content_df_missing = content_df.copy()
content_df_missing.loc[10:14, 'likes'] = np.nan
content_df_missing['likes'].fillna(0, inplace=True)
print(content_df_missing.loc[10:14, ['post_id', 'likes']])
    post_id  likes
10       11    0.0
11       12    0.0
12       13    0.0
13       14    0.0
14       15    0.0
# Error Handling Example 2: Incorrect aggregation of metrics (summing vs averaging)
# INCORRECT: Summing engagement rates across posts
incorrect_total = content_df['engagement_rate'].sum()
print(f'Incorrect total engagement rate: {incorrect_total:.2f}%')
# CORRECT: Calculate mean engagement rate
correct_mean = content_df['engagement_rate'].mean()
print(f'Correct mean engagement rate: {correct_mean:.2f}%')
Incorrect total engagement rate: 6305.02%
Correct mean engagement rate: 12.61%
# Error Handling Example 3: Misinterpreting ratios like CTR or engagement rate
example = yt_analytics.iloc[0]
computed_ctr = (example['views'] / max(example['views'], 1)) * 100
print(f"Displayed CTR: {example['ctr']}%, Computed CTR: {computed_ctr:.2f}%")
Displayed CTR: 4.14%, Computed CTR: 100.00%
# Error Handling Example 4: Wrong grouping logic for content categories
wrong_group = content_df.groupby('platform').size().sum()
correct_group = len(content_df)
print(f"Grouped count: {wrong_group}, True post count: {correct_group}")
Grouped count: 500, True post count: 500

Best Practices in Social Media Analytics#

  • Always segment data by relevant dimensions before comparing content.
  • Use mean rates for engagement, not totals.
  • Track trends over time to identify growth or drops.
  • Benchmark against similar content types and platforms.
  • Build consistent metric definitions and formulas.
  • Use visualizations to validate and present findings.
# Analytics Pattern Example: Benchmark content performance vs. platform average
content_df['vs_platform_benchmark'] = content_df['engagement_rate'] - content_df.groupby('platform')['engagement_rate'].transform('mean')
print(content_df[['post_id', 'platform', 'engagement_rate', 'vs_platform_benchmark']].head())
   post_id   platform  engagement_rate  vs_platform_benchmark
0        1  Instagram        17.143756               4.099440
1        2  Instagram        14.166667               1.122351
2        3  Instagram        10.366615              -2.677701
3        4    YouTube         7.187284              -4.975726
4        5  Instagram        10.636292              -2.408024
# Analytics Pattern Example: Simple audience segmentation by engagement tier
content_df['engagement_tier'] = pd.cut(content_df['engagement_rate'], bins=[-np.inf, 5, 10, 20, np.inf], labels=['Low', 'Medium', 'High', 'Viral'])
tier_counts = content_df['engagement_tier'].value_counts()
print(tier_counts)
engagement_tier
High      345
Medium    132
Low        14
Viral       9
Name: count, dtype: int64
# Analytics Pattern Example: Performance trend over time (7-day rolling average)
daily_trend = content_df.set_index('date').resample('D')['engagement_rate'].mean().rolling(7).mean()
print(daily_trend.dropna().head(10))
date
2023-01-07    12.724476
2023-01-08    12.450668
2023-01-09    12.594242
2023-01-10    12.721137
2023-01-11    11.789672
2023-01-12    11.439773
2023-01-13    11.572576
2023-01-14    11.560803
2023-01-15    12.345517
2023-01-16    12.451993
Freq: D, Name: engagement_rate, dtype: float64
# End-to-End Social Media Analytics Mini-Project: Identify top-performing videos by engagement
top_videos = yt_analytics.sort_values('ctr', ascending=False).head(10)
print(top_videos[['video_id', 'views', 'likes', 'comments', 'ctr', 'watch_time']])
     video_id   views  likes  comments   ctr  watch_time
298       299  125757   7335      1032  9.98       29251
143       144  164331   2922       599  9.98      400111
193       194  245410  14216      2223  9.96      166656
177       178  158438  12125      1549  9.93      132373
163       164   77473   3636      1251  9.90      188563
100       101  202383  15483      2492  9.90      301504
21         22  263013   2895      2819  9.88      419400
150       151  289098  16028      2001  9.87      434006
260       261  459515  29802      2823  9.85      450877
285       286  225381  12785      1480  9.79      433950
# End-to-End: Synthesize content strategy recommendations
median_ctr = yt_analytics['ctr'].median()
median_watch = yt_analytics['watch_time'].median()
recommend = []
for idx, row in top_videos.iterrows():
    if row['ctr'] > median_ctr and row['watch_time'] > median_watch:
        recommend.append('Promote')
    else:
        recommend.append('Monitor')
top_videos['strategy'] = recommend
print(top_videos[['video_id', 'ctr', 'watch_time', 'strategy']])
     video_id   ctr  watch_time strategy
298       299  9.98       29251  Monitor
143       144  9.98      400111  Promote
193       194  9.96      166656  Monitor
177       178  9.93      132373  Monitor
163       164  9.90      188563  Monitor
100       101  9.90      301504  Promote
21         22  9.88      419400  Promote
150       151  9.87      434006  Promote
260       261  9.85      450877  Promote
285       286  9.79      433950  Promote
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.