Lesson 3 · Social Media Content Analytics
Types of Social Media Data: Views, Likes, CTR, and Watch Time
In this lesson, we will explore the main data types in social media and content analytics. Understanding these metrics helps content creators and businesses…
- CourseSocial Media Content Analytics
- Lesson3 of 41
- Video25 min
- FormatJupyter notebook · 25 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbTypes of Social Media Data: Views, Likes, CTR, and Watch Time#
- In this lesson, we will explore the main data types in social media and content analytics.
- Understanding these metrics helps content creators and businesses measure performance.
- We will learn how to work with views, likes, click-through rate (CTR), and watch time data.
- By the end, you will be able to analyze engagement and spot top content effectively.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
Understanding Social Media Analytics Data#
- Social media datasets hold information about individual posts, videos, or campaigns.
- Each row often represents a piece of content, with metrics like views, likes, or comments.
- Views count how many times a post or video was seen.
- Likes show direct positive engagement from users.
- CTR (Click-Through Rate) measures how often viewers clicked when shown a thumbnail or link.
- Watch time is the total time users spent viewing your content.
- Beginners sometimes confuse high views with high engagement.
- Analyzing all key metrics helps avoid misleading conclusions.
# Beginner Example 1: Load a sample Social Media Content dataset
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
content_df = pd.DataFrame({
'post_id': range(1, n_posts+1),
'platform': np.random.choice(['YouTube', 'Instagram', 'TikTok'], n_posts),
'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
'views': views,
'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(content_df.shape)
print(content_df.head())
# Beginner Example 2: View summary statistics for engagement metrics
summary = content_df[['views', 'likes', 'comments', 'shares']].describe()
print(summary)
# Beginner Example 3: Calculate engagement rate for each post
content_df['engagement_rate'] = ((content_df['likes'] + content_df['comments'] + content_df['shares']) / content_df['views']) * 100
print(content_df[['post_id', 'platform', 'views', 'likes', 'comments', 'shares', 'engagement_rate']].head())
# Beginner Example 4: Find the post with the highest engagement rate
top_post = content_df.sort_values('engagement_rate', ascending=False).iloc[0]
print(f"Top engaging post has ID {top_post['post_id']}, platform {top_post['platform']}, and engagement rate {top_post['engagement_rate']:.2f}%.")
# Beginner Example 5: List top 5 posts by view count
top_views = content_df.sort_values('views', ascending=False).head(5)
print(top_views[['post_id', 'platform', 'views', 'likes', 'comments', 'shares']])
# Intermediate Example 1: Group posts by platform and compute mean engagement
platform_group = content_df.groupby('platform').agg({'views':'mean', 'likes':'mean', 'comments':'mean', 'shares':'mean', 'engagement_rate':'mean'}).reset_index()
print(platform_group)
# Intermediate Example 2: Identify posts with above average engagement rate for their platform
is_above_avg = []
for idx, row in content_df.iterrows():
platform = row['platform']
avg_rate = platform_group[platform_group['platform'] == platform]['engagement_rate'].values[0]
is_above_avg.append(row['engagement_rate'] > avg_rate)
content_df['above_platform_avg'] = is_above_avg
print(content_df[['post_id', 'platform', 'engagement_rate', 'above_platform_avg']].head(10))
# Intermediate Example 3: Visualize engagement rate distribution per platform
import matplotlib.pyplot as plt
plt.figure(figsize=(7,5))
for plat in content_df['platform'].unique():
plt.hist(content_df[content_df['platform'] == plat]['engagement_rate'], bins=20, alpha=0.5, label=plat)
plt.xlabel('Engagement Rate (%)')
plt.ylabel('Number of Posts')
plt.title('Engagement Rate Distribution by Platform')
plt.legend()
plt.show()
# Intermediate Example 4: Load YouTube Analytics data with CTR and watch time fields
np.random.seed(42)
n_videos = 300
views_y = np.random.randint(100, 500000, n_videos)
yt_analytics = pd.DataFrame({
'video_id': range(1, n_videos+1),
'publish_date': pd.date_range('2022-01-01', periods=n_videos, freq='D'),
'views': views_y,
'watch_time': np.random.randint(1000, 500000, n_videos),
'likes': (views_y * np.random.uniform(0.01, 0.08, n_videos)).astype(int),
'comments': (views_y * np.random.uniform(0.001, 0.02, n_videos)).astype(int),
'ctr': np.round(np.random.uniform(2, 10, n_videos), 2)
})
print(yt_analytics.head())
# Intermediate Example 5: Analyze correlation between CTR and watch time
correlation = yt_analytics['ctr'].corr(yt_analytics['watch_time'])
print(f'Correlation between CTR and watch time: {correlation:.2f}')
# Intermediate Example 6: Compare videos with top 10% CTR vs overall watch time
ctr_threshold = np.percentile(yt_analytics['ctr'], 90)
top_ctr_videos = yt_analytics[yt_analytics['ctr'] >= ctr_threshold]
mean_watch_time_top = top_ctr_videos['watch_time'].mean()
mean_watch_time_all = yt_analytics['watch_time'].mean()
print(f'Average watch time (top 10% CTR): {mean_watch_time_top:.1f}')
print(f'Average watch time (all videos): {mean_watch_time_all:.1f}')
# Advanced Example 1: Time-based analysis of daily total views and watch time
daily_stats = yt_analytics.groupby('publish_date').agg({'views':'sum', 'watch_time':'sum', 'ctr':'mean'}).reset_index()
print(daily_stats.head())
# Advanced Example 2: Identify potential viral videos using CTR and watch time thresholds
viral_threshold_ctr = yt_analytics['ctr'].quantile(0.95)
viral_threshold_watch = yt_analytics['watch_time'].quantile(0.95)
possible_viral = yt_analytics[(yt_analytics['ctr'] >= viral_threshold_ctr) & (yt_analytics['watch_time'] >= viral_threshold_watch)]
print(f'Number of possible viral videos: {len(possible_viral)}')
print(possible_viral[['video_id', 'views', 'watch_time', 'ctr']])
# Advanced Example 3: Compute engagement rate across categories (simulating with content data)
content_df['category'] = np.where(content_df['platform'] == 'YouTube', 'Video', 'Photo/Short')
category_grouped = content_df.groupby('category').agg({'views':'mean', 'engagement_rate':'mean'}).reset_index()
print(category_grouped)
# Error Handling Example 1: Handling missing engagement values
content_df_missing = content_df.copy()
content_df_missing.loc[10:14, 'likes'] = np.nan
content_df_missing['likes'].fillna(0, inplace=True)
print(content_df_missing.loc[10:14, ['post_id', 'likes']])
# Error Handling Example 2: Incorrect aggregation of metrics (summing vs averaging)
# INCORRECT: Summing engagement rates across posts
incorrect_total = content_df['engagement_rate'].sum()
print(f'Incorrect total engagement rate: {incorrect_total:.2f}%')
# CORRECT: Calculate mean engagement rate
correct_mean = content_df['engagement_rate'].mean()
print(f'Correct mean engagement rate: {correct_mean:.2f}%')
# Error Handling Example 3: Misinterpreting ratios like CTR or engagement rate
example = yt_analytics.iloc[0]
computed_ctr = (example['views'] / max(example['views'], 1)) * 100
print(f"Displayed CTR: {example['ctr']}%, Computed CTR: {computed_ctr:.2f}%")
# Error Handling Example 4: Wrong grouping logic for content categories
wrong_group = content_df.groupby('platform').size().sum()
correct_group = len(content_df)
print(f"Grouped count: {wrong_group}, True post count: {correct_group}")
Best Practices in Social Media Analytics#
- Always segment data by relevant dimensions before comparing content.
- Use mean rates for engagement, not totals.
- Track trends over time to identify growth or drops.
- Benchmark against similar content types and platforms.
- Build consistent metric definitions and formulas.
- Use visualizations to validate and present findings.
# Analytics Pattern Example: Benchmark content performance vs. platform average
content_df['vs_platform_benchmark'] = content_df['engagement_rate'] - content_df.groupby('platform')['engagement_rate'].transform('mean')
print(content_df[['post_id', 'platform', 'engagement_rate', 'vs_platform_benchmark']].head())
# Analytics Pattern Example: Simple audience segmentation by engagement tier
content_df['engagement_tier'] = pd.cut(content_df['engagement_rate'], bins=[-np.inf, 5, 10, 20, np.inf], labels=['Low', 'Medium', 'High', 'Viral'])
tier_counts = content_df['engagement_tier'].value_counts()
print(tier_counts)
# Analytics Pattern Example: Performance trend over time (7-day rolling average)
daily_trend = content_df.set_index('date').resample('D')['engagement_rate'].mean().rolling(7).mean()
print(daily_trend.dropna().head(10))
# End-to-End Social Media Analytics Mini-Project: Identify top-performing videos by engagement
top_videos = yt_analytics.sort_values('ctr', ascending=False).head(10)
print(top_videos[['video_id', 'views', 'likes', 'comments', 'ctr', 'watch_time']])
# End-to-End: Synthesize content strategy recommendations
median_ctr = yt_analytics['ctr'].median()
median_watch = yt_analytics['watch_time'].median()
recommend = []
for idx, row in top_videos.iterrows():
if row['ctr'] > median_ctr and row['watch_time'] > median_watch:
recommend.append('Promote')
else:
recommend.append('Monitor')
top_videos['strategy'] = recommend
print(top_videos[['video_id', 'ctr', 'watch_time', 'strategy']])
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



