Lesson 37 · Social Media Content Analytics
Thumbnail and Title Performance Analysis
In this lesson, we are going to learn how to analyze the impact of YouTube video thumbnails and titles on content performance. This problem is critical…
- CourseSocial Media Content Analytics
- Lesson37 of 41
- Video27 min
- FormatJupyter notebook · 20 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbThumbnail and Title Performance Analysis#
- In this lesson, we are going to learn how to analyze the impact of YouTube video thumbnails and titles on content performance.
- This problem is critical because compelling thumbnails and titles can drive higher engagement, more clicks, and ultimately greater reach for creators and brands.
- Learners will explore engagement metrics, detect which videos perform best, and extract insights to optimize their content strategy.
- We will build intuition to distinguish genuinely high-performing content from average videos using real-world social media analytics.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')
# Load YouTube Trending Videos data or create fallback data
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request
SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
def get_yt_service():
api_key = os.environ.get('YOUTUBE_API_KEY')
if api_key:
return build('youtube', 'v3', developerKey=api_key)
if Path('client_secret.json').exists():
creds = None
if Path('token_ro.pickle').exists():
with open('token_ro.pickle', 'rb') as f:
creds = pickle.load(f)
if not creds or not creds.valid:
if creds and creds.expired and creds.refresh_token:
creds.refresh(Request())
else:
flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
creds = flow.run_local_server(port=0)
with open('token_ro.pickle', 'wb') as f:
pickle.dump(creds, f)
return build('youtube', 'v3', credentials=creds)
raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
def fetch_yt_trending(max_results=200, region='US'):
youtube = get_yt_service()
records, token = [], None
while len(records) < max_results:
resp = youtube.videos().list(
part='snippet,statistics',
chart='mostPopular',
regionCode=region,
maxResults=min(50, max_results - len(records)),
pageToken=token
).execute()
for item in resp.get('items', []):
s = item['snippet']; st = item.get('statistics', {})
records.append({
'video_id': item['id'],
'trending_date': pd.Timestamp.today().date(),
'title': s.get('title', ''),
'channel_title': s.get('channelTitle', ''),
'category_id': s.get('categoryId', ''),
'views': int(st.get('viewCount', 0)),
'likes': int(st.get('likeCount', 0)),
'comment_count': int(st.get('commentCount', 0)),
})
token = resp.get('nextPageToken')
if not token: break
return pd.DataFrame(records)
try:
df = fetch_yt_trending()
print('Live trending data:', df.shape)
except Exception as e:
print(f'Falling back to synthetic: {e}')
np.random.seed(42)
n = 1000
df = pd.DataFrame({
'video_id': [f'vid{i}' for i in range(n)],
'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
'title': [f'Video Title {i}' for i in range(n)],
'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
'category_id': np.random.choice([1,2,10,22,24,28], n),
'views': np.random.randint(10000, 5000000, n),
'likes': np.random.randint(100, 200000, n),
'comment_count': np.random.randint(10, 50000, n),
})
print('Synthetic fallback:', df.shape)
print(df.head(3))
Key Concepts for Thumbnail and Title Analytics#
- Each video row represents one trending YouTube video.
- Important engagement metrics include: views, likes, and comment count.
- Thumbnails and titles act as a video's first impression (the hook for viewers).
- Click-through rate (CTR) is critical, but not always available in public datasets.
- High engagement often signals strong visual or textual hooks, but context matters.
- Common mistakes: forgetting to normalize metrics for video age, misinterpreting 'viral' vs. 'popular'.
# Beginner Example 1: Basic dataset summary
print('Total videos:', len(df))
print('Column names:', df.columns.tolist())
print('First trending dates:', df.trending_date.min(), 'to', df.trending_date.max())
# Beginner Example 2: Calculate engagement rate (likes/views * 100)
df['engagement_rate'] = (df['likes'] / df['views']) * 100
print('Engagement rate stats:')
print(df['engagement_rate'].describe())
# Beginner Example 3: Visualize the distribution of engagement rates
plt.figure(figsize=(7,3))
sns.histplot(df['engagement_rate'], bins=30, color='tomato')
plt.title('Distribution of Video Engagement Rates (%)')
plt.xlabel('Engagement Rate (%)')
plt.ylabel('Number of Videos')
plt.tight_layout()
plt.show()
# Intermediate Example 1: Find top 10 videos by engagement rate
top10_engaged = df.sort_values('engagement_rate', ascending=False).head(10)
print(top10_engaged[['title', 'views', 'likes', 'engagement_rate']])
# Intermediate Example 2: Find top 10 videos by absolute likes
top10_likes = df.sort_values('likes', ascending=False).head(10)
print(top10_likes[['title', 'views', 'likes', 'likes', 'engagement_rate']])
# Intermediate Example 3: Plot relationship between views and engagement rate
plt.figure(figsize=(7,4))
sns.scatterplot(data=df, x='views', y='engagement_rate', alpha=0.5)
plt.title('Views vs Engagement Rate')
plt.xlabel('Views')
plt.ylabel('Engagement Rate (%)')
plt.tight_layout()
plt.show()
# Intermediate Example 4: Analyze common words in top titles
from collections import Counter
import re
# Combine all top 100 titles by engagement rate
titles = df.sort_values('engagement_rate', ascending=False).head(100)['title']
words = []
for title in titles:
words.extend(re.findall(r'\w+', title.lower()))
word_counts = Counter(words)
most_common_words = word_counts.most_common(10)
print('Most common words in titles with high engagement:', most_common_words)
# Intermediate Example 5: Group by channel and compute average engagement rate
eng_by_channel = df.groupby('channel_title')['engagement_rate'].mean().sort_values(ascending=False)
print('Top channels by average engagement rate:')
print(eng_by_channel.head(5))
# Advanced Example 1: Detect outlier videos with exceptionally high or low engagement
q1, q3 = df['engagement_rate'].quantile([0.25, 0.75])
iqr = q3 - q1
lower = q1 - 1.5 * iqr
upper = q3 + 1.5 * iqr
outliers = df[(df['engagement_rate'] < lower) | (df['engagement_rate'] > upper)]
print('Outlier videos by engagement rate:', outliers.shape[0])
print(outliers[['title', 'views', 'likes', 'engagement_rate']].head())
# Advanced Example 2: Analyze engagement by content category
cat_engagement = df.groupby('category_id')['engagement_rate'].mean().sort_values(ascending=False)
plt.figure(figsize=(7,2.5))
sns.barplot(x=cat_engagement.index.astype(str), y=cat_engagement.values, palette='magma')
plt.title('Average Engagement Rate by Content Category')
plt.xlabel('Category ID')
plt.ylabel('Avg Engagement Rate (%)')
plt.tight_layout()
plt.show()
# Advanced Example 3: Are longer or shorter titles associated with better engagement?
df['title_length'] = df['title'].str.len()
plt.figure(figsize=(7,3))
sns.scatterplot(data=df, x='title_length', y='engagement_rate', alpha=0.4)
plt.title('Title Length vs Engagement Rate')
plt.xlabel('Title Length (characters)')
plt.ylabel('Engagement Rate (%)')
plt.tight_layout()
plt.show()
# Advanced Example 4: Analyze all-caps or question mark titles
def is_question(title):
return '?' in title
def is_all_caps(title):
w = re.sub(r'[^A-Za-z]', '', title)
return w.isupper() if w else False
df['is_question'] = df['title'].apply(is_question)
df['is_all_caps'] = df['title'].apply(is_all_caps)
mean_question = df[df['is_question']]['engagement_rate'].mean()
mean_caps = df[df['is_all_caps']]['engagement_rate'].mean()
print(f"Avg engagement (question titles): {mean_question:.2f}%")
print(f"Avg engagement (all-caps titles): {mean_caps:.2f}%")
# Error Handling Example 1: Missing engagement values
df_missing = df.copy()
mask = np.random.rand(len(df_missing)) < 0.01
df_missing.loc[mask, 'likes'] = np.nan
missing_count = df_missing['likes'].isnull().sum()
print(f'Missing likes values introduced: {missing_count}')
df_missing['engagement_rate'] = (df_missing['likes'] / df_missing['views']) * 100
print('Rows with NaN engagement_rate:', df_missing["engagement_rate"].isnull().sum())
# Error Handling Example 2: Incorrect metric aggregation (mean vs sum)
chan_sum = df.groupby('channel_title')['likes'].sum().sort_values(ascending=False)
chan_mean = df.groupby('channel_title')['likes'].mean().sort_values(ascending=False)
print('Channel with highest total likes:', chan_sum.idxmax())
print('Channel with highest avg likes per video:', chan_mean.idxmax())
# Error Handling Example 3: Misinterpreting ratios (engagement_rate vs total likes)
df_sorted = df.sort_values('engagement_rate', ascending=False)
print('Top by engagement:', df_sorted[['title','views','likes','engagement_rate']].head(3))
df_sorted2 = df.sort_values('likes', ascending=False)
print('Top by likes:', df_sorted2[['title','views','likes','engagement_rate']].head(3))
# Error Handling Example 4: Wrong grouping for content categories
try:
wrong_group = df.groupby('views')['likes'].mean()
print(wrong_group.head())
except Exception as e:
print('Grouping error:', e)
# Correct grouping
correct_group = df.groupby('category_id')['likes'].mean()
print('Likes per content category (correct groupby):')
print(correct_group.head())
Best Practices for Thumbnail & Title Analytics#
- Always benchmark against both absolute and relative metrics (like total likes vs. engagement rate).
- Segment analyses by content category, channel, or post type for better context.
- Track changes over time to recognize growth or sudden jumps (possible virality).
- Consider engagement metrics alongside metadata: title length, use of questions, or special characters.
- Use findings to optimize future thumbnails and titles based on real, historical audience behavior.
End-to-End Analytics: From Data to Strategy#
- Let us find the top 5 videos with the highest engagement rate and summarize what makes their titles unique.
- Based on these insights, we will recommend specific strategies for crafting high-performing video titles.
# Step 1: Identify top 5 videos by engagement rate
top5 = df.sort_values('engagement_rate', ascending=False).head(5)
print('Top 5 high-engagement video titles:')
print(top5['title'].to_list())
# Step 2: Word analysis and recommendation
all_words = []
for t in top5['title']:
all_words.extend(re.findall(r'\w+', t.lower()))
print('Words used in top titles:', set(all_words))
recommendations = []
if any('how' in w for w in all_words):
recommendations.append("Title includes 'How': Use actionable phrases to promise value.")
if any('2023' in w or '2024' in w for w in all_words):
recommendations.append("Title includes year: Timely or trending keywords can boost clicks.")
if any(w.endswith('!') for w in top5['title']):
recommendations.append("Title uses exclamation: Emotional punctuation can capture attention.")
if not recommendations:
recommendations.append('Use curiosity, lists, or questions to improve engagement.')
print('Strategy recommendations:')
for r in recommendations:
print('-', r)
YouTube Call-to-Action#
- Like this approach? Subscribe for more tips on mastering content analytics!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



