Mathew K Analytics

Lesson 36 · Social Media Content Analytics

Understanding Click-Through Rate (CTR) in Social Media and Content Analytics

We will explore what Click-Through Rate (CTR) means and why it matters for content creators and marketers. You will learn how to calculate and interpret CTR…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Understanding Click-Through Rate (CTR) in Social Media and Content Analytics#

  • We will explore what Click-Through Rate (CTR) means and why it matters for content creators and marketers.
  • You will learn how to calculate and interpret CTR using real and simulated YouTube data.
  • We will identify top-performing videos, common analysis mistakes, and actionable content insights.
  • This lesson will help you make smarter content decisions based on real engagement analytics.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
# Load a YouTube analytics synthetic dataset for CTR analysis
np.random.seed(42)
n_videos = 300
views = np.random.randint(100, 500000, n_videos)
df = pd.DataFrame({
    'video_id': range(1, n_videos+1),
    'publish_date': pd.date_range('2022-01-01', periods=n_videos, freq='D'),
    'views': views,
    'watch_time': np.random.randint(1000, 500000, n_videos),
    'likes': (views * np.random.uniform(0.01, 0.08, n_videos)).astype(int),
    'comments': (views * np.random.uniform(0.001, 0.02, n_videos)).astype(int),
    'ctr': np.round(np.random.uniform(2, 10, n_videos), 2)
})
print(df.shape)
print(df.head(3))
(300, 7)
   video_id publish_date   views  watch_time  likes  comments   ctr
0         1   2022-01-01  122058      158381   1890      1975  4.14
1         2   2022-01-02  146967      481671   1730      1900  6.99
2         3   2022-01-03  132032      195806  10217       337  5.28

Core Metrics in Social Media Analytics#

  • Each row in our dataset represents a single YouTube video and its performance.
  • Key metrics include views (how many people watched), likes, comments, and CTR (click-through rate).
  • CTR shows how often people who see a video thumbnail actually click to watch the video.
  • Watch time reveals how engaging the content is after a user clicks.
  • Beginners often confuse total views with CTR or ignore the impact of low impressions on CTR value.
  • Correctly interpreting CTR helps you spot high-potential videos and improve content strategy.
# Beginner Example 1: Calculate average CTR for all videos
average_ctr = df['ctr'].mean()
print(f'Average Click-Through Rate across all videos: {average_ctr:.2f}%')
Average Click-Through Rate across all videos: 6.17%
# Beginner Example 2: Find videos with CTR above 8%
high_ctr_videos = df[df['ctr'] > 8]
print(f'Number of videos with CTR above 8%: {len(high_ctr_videos)}')
print(high_ctr_videos[['video_id', 'views', 'ctr']].head(3))
Number of videos with CTR above 8%: 86
   video_id   views   ctr
6         7  110368  9.59
7         8  207992  8.11
9        10  137437  8.95
# Beginner Example 3: Show videos with the lowest CTR
lowest_ctr = df.nsmallest(5, 'ctr')
print('Videos with the lowest CTR:')
print(lowest_ctr[['video_id', 'views', 'ctr']])
Videos with the lowest CTR:
     video_id   views   ctr
257       258  491414  2.05
34         35  465248  2.09
162       163  297466  2.14
25         26  252809  2.15
138       139  124475  2.17

Interpreting Click-Through Rate in Context#

  • High CTR does not always mean high total views.
  • A video can have a high CTR but low total impressions if few people see the thumbnail.
  • Look for patterns: Are videos with high watch time also high in CTR?
  • Beginner mistake: Optimizing only for CTR can ignore what happens after the click (like watch time).
# Intermediate Example 1: Relationship between CTR and watch time
import matplotlib.pyplot as plt
plt.scatter(df['ctr'], df['watch_time'], alpha=0.6)
plt.xlabel('Click-Through Rate (%)')
plt.ylabel('Watch Time (minutes)')
plt.title('CTR vs Watch Time for YouTube Videos')
plt.show()
No description has been provided for this image
# Intermediate Example 2: Find top 10 videos by CTR AND by views
top_by_ctr = df.nlargest(10, 'ctr')
top_by_views = df.nlargest(10, 'views')
print('Top 10 by CTR:')
print(top_by_ctr[['video_id', 'ctr', 'views']])
print('Top 10 by Views:')
print(top_by_views[['video_id', 'ctr', 'views']])
Top 10 by CTR:
     video_id   ctr   views
143       144  9.98  164331
298       299  9.98  125757
193       194  9.96  245410
177       178  9.93  158438
100       101  9.90  202383
163       164  9.90   77473
21         22  9.88  263013
150       151  9.87  289098
260       261  9.85  459515
285       286  9.79  225381
Top 10 by Views:
     video_id   ctr   views
249       250  8.61  499146
286       287  9.75  496357
257       258  2.05  491414
173       174  9.56  491334
166       167  7.44  489670
60         61  9.49  489592
265       266  8.24  489155
224       225  2.61  488319
107       108  9.32  487979
36         37  9.20  486332
# Intermediate Example 3: Calculate average watch time for videos with high vs low CTR
high_ctr_watch_time = df[df['ctr'] > 6]['watch_time'].mean()
low_ctr_watch_time = df[df['ctr'] <= 6]['watch_time'].mean()
print(f'Avg watch time (CTR > 6%): {high_ctr_watch_time:.1f} minutes')
print(f'Avg watch time (CTR <= 6%): {low_ctr_watch_time:.1f} minutes')
Avg watch time (CTR > 6%): 243768.3 minutes
Avg watch time (CTR <= 6%): 250937.7 minutes
# Advanced Example 1: Analyze CTR distribution by month of publication
df['month'] = df['publish_date'].dt.to_period('M')
monthly_ctr = df.groupby('month')['ctr'].mean()
monthly_ctr.plot(kind='bar', figsize=(12,4), color='skyblue')
plt.title('Average CTR by Month of Publication')
plt.xlabel('Month')
plt.ylabel('Average CTR (%)')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Advanced Example 2: Correlate likes, comments, and CTR to spot viral signals
correlations = df[['likes','comments','ctr']].corr()
print('Correlation matrix:')
print(correlations)
Correlation matrix:
             likes  comments       ctr
likes     1.000000  0.511144  0.008264
comments  0.511144  1.000000 -0.007806
ctr       0.008264 -0.007806  1.000000
# Advanced Example 3: Detect outliers in CTR
q1 = df['ctr'].quantile(0.25)
q3 = df['ctr'].quantile(0.75)
iqr = q3 - q1
upper_bound = q3 + 1.5 * iqr
outliers = df[df['ctr'] > upper_bound]
print(f'Videos with CTR outliers (above {upper_bound:.2f}%):')
print(outliers[['video_id', 'ctr', 'views']])
Videos with CTR outliers (above 14.51%):
Empty DataFrame
Columns: [video_id, ctr, views]
Index: []

Common Errors and Debugging with CTR in Analytics#

  • Watch out for missing values in CTR calculationsnever divide by zero impressions.
  • Make sure you do not mix up totals: groupings must fit your analysis goal.
  • Always check definitions: CTR is not the same as engagement rate or view rate.
  • Wrong groupings (like summing CTR) can distort your finding.
  • Misinterpretation of low CTR is a trap: Maybe your audience is just small, not uninterested.
# Error Handling Example 1: Handle missing or zero views
test_df = df.copy()
test_df.loc[0, 'views'] = 0
try:
    ctr_recalc = 100 * test_df.loc[0, 'likes'] / test_df.loc[0, 'views']
except ZeroDivisionError:
    ctr_recalc = None
print('Recalculated CTR with zero views:', ctr_recalc)
Recalculated CTR with zero views: inf
# Error Handling Example 2: Avoid summing CTR directly
sum_ctr_wrong = df['ctr'].sum()
corrected_avg_ctr = np.average(df['ctr'])
print(f'Wrong: Summing all CTRs = {sum_ctr_wrong:.2f}')
print(f'Right: Average CTR = {corrected_avg_ctr:.2f}%')
Wrong: Summing all CTRs = 1850.99
Right: Average CTR = 6.17%
# Error Handling Example 3: Wrong grouping logic in CTR aggregation
monthly_total_views = df.groupby('month')['views'].sum()
monthly_ctr_wrong = df.groupby('month')['ctr'].sum()
monthly_ctr_right = df.groupby('month')['ctr'].mean()
print('CTR by summing per month (incorrect):')
print(monthly_ctr_wrong.head())
print('CTR by averaging per month (correct):')
print(monthly_ctr_right.head())
CTR by summing per month (incorrect):
month
2022-01    182.16
2022-02    158.87
2022-03    179.79
2022-04    208.91
2022-05    195.74
Freq: M, Name: ctr, dtype: float64
CTR by averaging per month (correct):
month
2022-01    5.876129
2022-02    5.673929
2022-03    5.799677
2022-04    6.963667
2022-05    6.314194
Freq: M, Name: ctr, dtype: float64

Best Practices for Analyzing and Using CTR#

  • Always compare CTR alongside other engagement metrics for best content insights.
  • Segment your audience or videos for deeper pattern detection.
  • Track CTR changes over time to see what boosts engagement.
  • Use outlier checks to find unusually effective or misleading videos.
  • Document and stick to precise metric definitions in reports.
# Best Practice Example 1: Benchmark videos with above-average CTR
avg_ctr = df['ctr'].mean()
benchmarked = df[df['ctr'] > avg_ctr]
print(f'Total videos above average CTR ({avg_ctr:.2f}%): {benchmarked.shape[0]}')
print(benchmarked[['video_id', 'views', 'ctr']].head(3))
Total videos above average CTR (6.17%): 158
   video_id   views   ctr
1         2  146967  6.99
3         4  365938  6.42
6         7  110368  9.59
# Best Practice Example 2: Segment performance by top/bottom quartile of views
q25 = df['views'].quantile(0.25)
q75 = df['views'].quantile(0.75)
low_views = df[df['views'] <= q25]
high_views = df[df['views'] >= q75]
print(f'Average CTR for bottom quartile (low views): {low_views['ctr'].mean():.2f}%')
print(f'Average CTR for top quartile (high views): {high_views['ctr'].mean():.2f}%')
Average CTR for bottom quartile (low views): 5.94%
Average CTR for top quartile (high views): 6.03%
# Best Practice Example 3: Content optimization based on CTR and engagement
df['like_rate'] = 100 * df['likes'] / df['views']
df['engagement_score'] = (df['like_rate'] + df['ctr']) / 2
optimized = df.nlargest(5, 'engagement_score')
print('Top 5 optimized videos by combined CTR and like rate:')
print(optimized[['video_id', 'ctr', 'like_rate', 'engagement_score']])
Top 5 optimized videos by combined CTR and like rate:
     video_id   ctr  like_rate  engagement_score
177       178  9.93   7.652836          8.791418
100       101  9.90   7.650346          8.775173
101       102  9.55   7.653821          8.601910
211       212  8.90   7.789341          8.344671
85         86  9.34   7.008559          8.174280

Tiny End-to-End Social Media Analytics Challenge#

  • From this real (or simulated) YouTube data, identify one clear content strategy recommendation.
  • Question: Which factor seems MOST related to high CTR: publish month, view count, or like rate?
  • Suggest one action a channel manager could take after analyzing these patterns.
# Strategy: Identify publish months with highest average CTR
ctr_by_month = df.groupby('month')['ctr'].mean()
best_month = ctr_by_month.idxmax()
print(f'Month with highest avg CTR: {best_month} ({ctr_by_month.max():.2f}%)')
Month with highest avg CTR: 2022-04 (6.96%)
# Strategy: Find traits of top CTR videos
top_ctr_videos = df.nlargest(10, 'ctr')
avg_like_rate_top = top_ctr_videos['like_rate'].mean()
avg_watch_time_top = top_ctr_videos['watch_time'].mean()
print('Top 10 CTR videos characteristics:')
print(f'Average like rate: {avg_like_rate_top:.2f}%')
print(f'Average watch time: {avg_watch_time_top:.1f} minutes')
Top 10 CTR videos characteristics:
Average like rate: 5.22%
Average watch time: 295669.1 minutes
# End-to-end recommendation: If high-CTR and high like rate cluster in certain months, focus content then
if best_month is not None:
    print(f'Recommend posting more high-quality videos targeting {str(best_month)} when audience CTR peaks.')
else:
    print('No strong monthly pattern detected. Focus on optimizing thumbnails and titles for higher CTR.')
Recommend posting more high-quality videos targeting 2022-04 when audience CTR peaks.
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.