Mathew K Analytics

Lesson 37 · Social Media Content Analytics

Thumbnail and Title Performance Analysis

In this lesson, we are going to learn how to analyze the impact of YouTube video thumbnails and titles on content performance. This problem is critical…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Thumbnail and Title Performance Analysis#

  • In this lesson, we are going to learn how to analyze the impact of YouTube video thumbnails and titles on content performance.
  • This problem is critical because compelling thumbnails and titles can drive higher engagement, more clicks, and ultimately greater reach for creators and brands.
  • Learners will explore engagement metrics, detect which videos perform best, and extract insights to optimize their content strategy.
  • We will build intuition to distinguish genuinely high-performing content from average videos using real-world social media analytics.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')
# Load YouTube Trending Videos data or create fallback data
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request

SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']

def get_yt_service():
    api_key = os.environ.get('YOUTUBE_API_KEY')
    if api_key:
        return build('youtube', 'v3', developerKey=api_key)
    if Path('client_secret.json').exists():
        creds = None
        if Path('token_ro.pickle').exists():
            with open('token_ro.pickle', 'rb') as f:
                creds = pickle.load(f)
        if not creds or not creds.valid:
            if creds and creds.expired and creds.refresh_token:
                creds.refresh(Request())
            else:
                flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
                creds = flow.run_local_server(port=0)
            with open('token_ro.pickle', 'wb') as f:
                pickle.dump(creds, f)
        return build('youtube', 'v3', credentials=creds)
    raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')

def fetch_yt_trending(max_results=200, region='US'):
    youtube = get_yt_service()
    records, token = [], None
    while len(records) < max_results:
        resp = youtube.videos().list(
            part='snippet,statistics',
            chart='mostPopular',
            regionCode=region,
            maxResults=min(50, max_results - len(records)),
            pageToken=token
        ).execute()
        for item in resp.get('items', []):
            s = item['snippet']; st = item.get('statistics', {})
            records.append({
                'video_id':      item['id'],
                'trending_date': pd.Timestamp.today().date(),
                'title':         s.get('title', ''),
                'channel_title': s.get('channelTitle', ''),
                'category_id':   s.get('categoryId', ''),
                'views':         int(st.get('viewCount', 0)),
                'likes':         int(st.get('likeCount', 0)),
                'comment_count': int(st.get('commentCount', 0)),
            })
        token = resp.get('nextPageToken')
        if not token: break
    return pd.DataFrame(records)

try:
    df = fetch_yt_trending()
    print('Live trending data:', df.shape)
except Exception as e:
    print(f'Falling back to synthetic: {e}')
    np.random.seed(42)
    n = 1000
    df = pd.DataFrame({
        'video_id':      [f'vid{i}' for i in range(n)],
        'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
        'title':         [f'Video Title {i}' for i in range(n)],
        'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
        'category_id':   np.random.choice([1,2,10,22,24,28], n),
        'views':         np.random.randint(10000, 5000000, n),
        'likes':         np.random.randint(100, 200000, n),
        'comment_count': np.random.randint(10, 50000, n),
    })
    print('Synthetic fallback:', df.shape)

print(df.head(3))
Live trending data: (199, 8)
      video_id trending_date  \
0  82-jTNka3uc    2026-06-02   
1  3oB9AxspVow    2026-06-02   
2  Vagb9BqdX8g    2026-06-02   

                                               title     channel_title  \
0  Ariana Grande - hate that i made you love me (...  ArianaGrandeVevo   
1           The End of Oak Street | Official Trailer      Warner Bros.   
2                 We Almost Lost EVERYTHING Gambling        SMii7Yplus   

  category_id    views   likes  comment_count  
0          10  1229856  311806          23343  
1           1   489638   23888           2048  
2          20  1275299   63388           1680  

Key Concepts for Thumbnail and Title Analytics#

  • Each video row represents one trending YouTube video.
  • Important engagement metrics include: views, likes, and comment count.
  • Thumbnails and titles act as a video's first impression (the hook for viewers).
  • Click-through rate (CTR) is critical, but not always available in public datasets.
  • High engagement often signals strong visual or textual hooks, but context matters.
  • Common mistakes: forgetting to normalize metrics for video age, misinterpreting 'viral' vs. 'popular'.
# Beginner Example 1: Basic dataset summary
print('Total videos:', len(df))
print('Column names:', df.columns.tolist())
print('First trending dates:', df.trending_date.min(), 'to', df.trending_date.max())
Total videos: 199
Column names: ['video_id', 'trending_date', 'title', 'channel_title', 'category_id', 'views', 'likes', 'comment_count']
First trending dates: 2026-06-02 to 2026-06-02
# Beginner Example 2: Calculate engagement rate (likes/views * 100)
df['engagement_rate'] = (df['likes'] / df['views']) * 100
print('Engagement rate stats:')
print(df['engagement_rate'].describe())
Engagement rate stats:
count    199.000000
mean       4.874369
std        3.663670
min        0.000000
25%        2.405276
50%        4.423015
75%        6.392554
max       25.353049
Name: engagement_rate, dtype: float64
# Beginner Example 3: Visualize the distribution of engagement rates
plt.figure(figsize=(7,3))
sns.histplot(df['engagement_rate'], bins=30, color='tomato')
plt.title('Distribution of Video Engagement Rates (%)')
plt.xlabel('Engagement Rate (%)')
plt.ylabel('Number of Videos')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Intermediate Example 1: Find top 10 videos by engagement rate
top10_engaged = df.sort_values('engagement_rate', ascending=False).head(10)
print(top10_engaged[['title', 'views', 'likes', 'engagement_rate']])
                                                 title    views   likes  \
0    Ariana Grande - hate that i made you love me (...  1229856  311806   
194  The Deep-Sea Deception Story Time with Valeria...    20160    3848   
24             IEM Cologne Major 2026 Official Trailer    52701    9526   
6                                SHINee 샤이니 'Atmos' MV   272446   48461   
71               CORTIS (코르티스) 'Blue Lips' Official MV  2298227  376682   
72   The Phan Relationship is moving too fast - Tom...   222318   33589   
117  Netanyahu Warns of Nuclear Threats to American...    23598    2838   
13                                  Fortnite | Runners   100785   11927   
75   Monaleo - Everythang Pinka (feat. Teezo Touchd...   215117   24062   
135                      Rec Room AMA - the final AMA!    81739    8991   

     engagement_rate  
0          25.353049  
194        19.087302  
24         18.075558  
6          17.787378  
71         16.390113  
72         15.108538  
117        12.026443  
13         11.834102  
75         11.185541  
135        10.999645  
# Intermediate Example 2: Find top 10 videos by absolute likes
top10_likes = df.sort_values('likes', ascending=False).head(10)
print(top10_likes[['title', 'views', 'likes', 'likes', 'engagement_rate']])
                                                title    views   likes  \
71              CORTIS (코르티스) 'Blue Lips' Official MV  2298227  376682   
0   Ariana Grande - hate that i made you love me (...  1229856  311806   
14                              TREASURE - ‘IF I’ M/V  2934711  220451   
79          I Went to WAR on a Hardcore Minecraft SMP  3575769  196313   
4   No Peace Amongst the Stars | Warhammer 40,000 ...  1805955  136875   
5             1000 VS 1000 Player Minecraft Civil War  1745365  114197   
59  Jay Wheeler, Omar Courtz - De Lejitos (Remix) ...  1751019  111290   
84          SIDEMEN AMONG US: HARRY POTTER CHAOS MODE  2545568   90376   
16  Hydroneer Tried to Optimize, So I Double Overl...  1477971   82618   
89                       One of the Games of All Time  1135461   68102   

     likes  engagement_rate  
71  376682        16.390113  
0   311806        25.353049  
14  220451         7.511847  
79  196313         5.490092  
4   136875         7.579093  
5   114197         6.542872  
59  111290         6.355728  
84   90376         3.550327  
16   82618         5.589961  
89   68102         5.997740  
# Intermediate Example 3: Plot relationship between views and engagement rate
plt.figure(figsize=(7,4))
sns.scatterplot(data=df, x='views', y='engagement_rate', alpha=0.5)
plt.title('Views vs Engagement Rate')
plt.xlabel('Views')
plt.ylabel('Engagement Rate (%)')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Intermediate Example 4: Analyze common words in top titles
from collections import Counter
import re
# Combine all top 100 titles by engagement rate
titles = df.sort_values('engagement_rate', ascending=False).head(100)['title']
words = []
for title in titles:
    words.extend(re.findall(r'\w+', title.lower()))
word_counts = Counter(words)
most_common_words = word_counts.most_common(10)
print('Most common words in titles with high engagement:', most_common_words)
Most common words in titles with high engagement: [('the', 27), ('i', 15), ('of', 12), ('to', 12), ('official', 11), ('is', 11), ('a', 11), ('s', 9), ('in', 9), ('trailer', 8)]
# Intermediate Example 5: Group by channel and compute average engagement rate
eng_by_channel = df.groupby('channel_title')['engagement_rate'].mean().sort_values(ascending=False)
print('Top channels by average engagement rate:')
print(eng_by_channel.head(5))
Top channels by average engagement rate:
channel_title
ArianaGrandeVevo      25.353049
PlayOverwatch         19.087302
ESL Counter-Strike    18.075558
SMTOWN                17.787378
HYBE LABELS           16.390113
Name: engagement_rate, dtype: float64
# Advanced Example 1: Detect outlier videos with exceptionally high or low engagement
q1, q3 = df['engagement_rate'].quantile([0.25, 0.75])
iqr = q3 - q1
lower = q1 - 1.5 * iqr
upper = q3 + 1.5 * iqr
outliers = df[(df['engagement_rate'] < lower) | (df['engagement_rate'] > upper)]
print('Outlier videos by engagement rate:', outliers.shape[0])
print(outliers[['title', 'views', 'likes', 'engagement_rate']].head())
Outlier videos by engagement rate: 6
                                                title    views   likes  \
0   Ariana Grande - hate that i made you love me (...  1229856  311806   
6                               SHINee 샤이니 'Atmos' MV   272446   48461   
24            IEM Cologne Major 2026 Official Trailer    52701    9526   
71              CORTIS (코르티스) 'Blue Lips' Official MV  2298227  376682   
72  The Phan Relationship is moving too fast - Tom...   222318   33589   

    engagement_rate  
0         25.353049  
6         17.787378  
24        18.075558  
71        16.390113  
72        15.108538  
# Advanced Example 2: Analyze engagement by content category
cat_engagement = df.groupby('category_id')['engagement_rate'].mean().sort_values(ascending=False)
plt.figure(figsize=(7,2.5))
sns.barplot(x=cat_engagement.index.astype(str), y=cat_engagement.values, palette='magma')
plt.title('Average Engagement Rate by Content Category')
plt.xlabel('Category ID')
plt.ylabel('Avg Engagement Rate (%)')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Advanced Example 3: Are longer or shorter titles associated with better engagement?
df['title_length'] = df['title'].str.len()
plt.figure(figsize=(7,3))
sns.scatterplot(data=df, x='title_length', y='engagement_rate', alpha=0.4)
plt.title('Title Length vs Engagement Rate')
plt.xlabel('Title Length (characters)')
plt.ylabel('Engagement Rate (%)')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Advanced Example 4: Analyze all-caps or question mark titles
def is_question(title):
    return '?' in title
def is_all_caps(title):
    w = re.sub(r'[^A-Za-z]', '', title)
    return w.isupper() if w else False
df['is_question'] = df['title'].apply(is_question)
df['is_all_caps'] = df['title'].apply(is_all_caps)
mean_question = df[df['is_question']]['engagement_rate'].mean()
mean_caps = df[df['is_all_caps']]['engagement_rate'].mean()
print(f"Avg engagement (question titles): {mean_question:.2f}%")
print(f"Avg engagement (all-caps titles): {mean_caps:.2f}%")
Avg engagement (question titles): 4.38%
Avg engagement (all-caps titles): 3.77%
# Error Handling Example 1: Missing engagement values
df_missing = df.copy()
mask = np.random.rand(len(df_missing)) < 0.01
df_missing.loc[mask, 'likes'] = np.nan
missing_count = df_missing['likes'].isnull().sum()
print(f'Missing likes values introduced: {missing_count}')

df_missing['engagement_rate'] = (df_missing['likes'] / df_missing['views']) * 100
print('Rows with NaN engagement_rate:', df_missing["engagement_rate"].isnull().sum())
Missing likes values introduced: 2
Rows with NaN engagement_rate: 2
# Error Handling Example 2: Incorrect metric aggregation (mean vs sum)
chan_sum = df.groupby('channel_title')['likes'].sum().sort_values(ascending=False)
chan_mean = df.groupby('channel_title')['likes'].mean().sort_values(ascending=False)
print('Channel with highest total likes:', chan_sum.idxmax())
print('Channel with highest avg likes per video:', chan_mean.idxmax())
Channel with highest total likes: HYBE LABELS
Channel with highest avg likes per video: HYBE LABELS
# Error Handling Example 3: Misinterpreting ratios (engagement_rate vs total likes)
df_sorted = df.sort_values('engagement_rate', ascending=False)
print('Top by engagement:', df_sorted[['title','views','likes','engagement_rate']].head(3))

df_sorted2 = df.sort_values('likes', ascending=False)
print('Top by likes:', df_sorted2[['title','views','likes','engagement_rate']].head(3))
Top by engagement:                                                  title    views   likes  \
0    Ariana Grande - hate that i made you love me (...  1229856  311806   
194  The Deep-Sea Deception Story Time with Valeria...    20160    3848   
24             IEM Cologne Major 2026 Official Trailer    52701    9526   

     engagement_rate  
0          25.353049  
194        19.087302  
24         18.075558  
Top by likes:                                                 title    views   likes  \
71              CORTIS (코르티스) 'Blue Lips' Official MV  2298227  376682   
0   Ariana Grande - hate that i made you love me (...  1229856  311806   
14                              TREASURE - ‘IF I’ M/V  2934711  220451   

    engagement_rate  
71        16.390113  
0         25.353049  
14         7.511847  
# Error Handling Example 4: Wrong grouping for content categories
try:
    wrong_group = df.groupby('views')['likes'].mean()
    print(wrong_group.head())
except Exception as e:
    print('Grouping error:', e)

# Correct grouping
correct_group = df.groupby('category_id')['likes'].mean()
print('Likes per content category (correct groupby):')
print(correct_group.head())
views
6326       11.0
12286    1214.0
12670     961.0
12747     120.0
13678    1503.0
Name: likes, dtype: float64
Likes per content category (correct groupby):
category_id
1     13535.750000
10    34392.066667
17     9353.500000
20    12420.805755
22    18063.000000
Name: likes, dtype: float64

Best Practices for Thumbnail & Title Analytics#

  • Always benchmark against both absolute and relative metrics (like total likes vs. engagement rate).
  • Segment analyses by content category, channel, or post type for better context.
  • Track changes over time to recognize growth or sudden jumps (possible virality).
  • Consider engagement metrics alongside metadata: title length, use of questions, or special characters.
  • Use findings to optimize future thumbnails and titles based on real, historical audience behavior.

End-to-End Analytics: From Data to Strategy#

  • Let us find the top 5 videos with the highest engagement rate and summarize what makes their titles unique.
  • Based on these insights, we will recommend specific strategies for crafting high-performing video titles.
# Step 1: Identify top 5 videos by engagement rate
top5 = df.sort_values('engagement_rate', ascending=False).head(5)
print('Top 5 high-engagement video titles:')
print(top5['title'].to_list())
Top 5 high-engagement video titles:
['Ariana Grande - hate that i made you love me (official music video)', 'The Deep-Sea Deception Story Time with Valeria Rodriguez as Venture | Overwatch', 'IEM Cologne Major 2026 Official Trailer', "SHINee 샤이니 'Atmos' MV", "CORTIS (코르티스) 'Blue Lips' Official MV"]
# Step 2: Word analysis and recommendation
all_words = []
for t in top5['title']:
    all_words.extend(re.findall(r'\w+', t.lower()))
print('Words used in top titles:', set(all_words))

recommendations = []
if any('how' in w for w in all_words):
    recommendations.append("Title includes 'How': Use actionable phrases to promise value.")
if any('2023' in w or '2024' in w for w in all_words):
    recommendations.append("Title includes year: Timely or trending keywords can boost clicks.")
if any(w.endswith('!') for w in top5['title']):
    recommendations.append("Title uses exclamation: Emotional punctuation can capture attention.")
if not recommendations:
    recommendations.append('Use curiosity, lists, or questions to improve engagement.')

print('Strategy recommendations:')
for r in recommendations:
    print('-', r)
Words used in top titles: {'rodriguez', 'iem', 'made', '2026', 'story', 'overwatch', 'grande', 'valeria', 'cologne', 'with', 'i', 'sea', 'trailer', 'blue', 'official', 'the', 'deception', '샤이니', 'atmos', 'ariana', 'hate', 'as', 'lips', 'deep', 'love', 'that', 'me', 'video', 'time', 'mv', 'music', '코르티스', 'you', 'major', 'shinee', 'cortis', 'venture'}
Strategy recommendations:
- Use curiosity, lists, or questions to improve engagement.

YouTube Call-to-Action#

  • Like this approach? Subscribe for more tips on mastering content analytics!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.