Mathew K Analytics

Lesson 12 · Social Media Content Analytics

Understanding Video Metadata: Title, Tags, Category

In this lesson, we will learn how to analyze video metadata for social media and content analytics. We will focus on titles, tags, and categories using…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Understanding Video Metadata: Title, Tags, Category#

  • In this lesson, we will learn how to analyze video metadata for social media and content analytics.
  • We will focus on titles, tags, and categories using real-world YouTube data.
  • These elements are key for discoverability, reach, and audience engagement.
  • Mastering metadata helps creators and businesses optimize their videos for better performance.
  • By the end, you will be able to extract insights that guide smarter content strategies.
import pandas as pd
import numpy as np
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request
import warnings
warnings.filterwarnings('ignore')
# Setup: Load YouTube Trending Videos data
SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
def get_yt_service():
    api_key = os.environ.get('YOUTUBE_API_KEY')
    if api_key:
        return build('youtube', 'v3', developerKey=api_key)
    if Path('client_secret.json').exists():
        creds = None
        if Path('token_ro.pickle').exists():
            with open('token_ro.pickle', 'rb') as f:
                creds = pickle.load(f)
        if not creds or not creds.valid:
            if creds and creds.expired and creds.refresh_token:
                creds.refresh(Request())
            else:
                flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
                creds = flow.run_local_server(port=0)
            with open('token_ro.pickle', 'wb') as f:
                pickle.dump(creds, f)
        return build('youtube', 'v3', credentials=creds)
    raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')

def fetch_yt_trending(max_results=200, region='US'):
    youtube = get_yt_service()
    records, token = [], None
    while len(records) < max_results:
        resp = youtube.videos().list(
            part='snippet,statistics',
            chart='mostPopular',
            regionCode=region,
            maxResults=min(50, max_results - len(records)),
            pageToken=token
        ).execute()
        for item in resp.get('items', []):
            s = item['snippet']; st = item.get('statistics', {})
            records.append({
                'video_id':      item['id'],
                'trending_date': pd.Timestamp.today().date(),
                'title':         s.get('title', ''),
                'channel_title': s.get('channelTitle', ''),
                'category_id':   s.get('categoryId', ''),
                'views':         int(st.get('viewCount', 0)),
                'likes':         int(st.get('likeCount', 0)),
                'comment_count': int(st.get('commentCount', 0)),
            })
        token = resp.get('nextPageToken')
        if not token: break
    return pd.DataFrame(records)

try:
    df = fetch_yt_trending()
    print('Live trending data:', df.shape)
except Exception as e:
    print(f'Falling back to synthetic: {e}')
    np.random.seed(42)
    n = 1000
    df = pd.DataFrame({
        'video_id':      [f'vid{i}' for i in range(n)],
        'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
        'title':         [f'Video Title {i}' for i in range(n)],
        'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
        'category_id':   np.random.choice([1,2,10,22,24,28], n),
        'views':         np.random.randint(10000, 5000000, n),
        'likes':         np.random.randint(100, 200000, n),
        'comment_count': np.random.randint(10, 50000, n),
    })
    print('Synthetic fallback:', df.shape)
print(df.head(3))
Falling back to synthetic: <HttpError 403 when requesting https://youtube.googleapis.com/youtube/v3/videos?part=snippet%2Cstatistics&chart=mostPopular&regionCode=US&maxResults=50&key=YOUR_GOOGLE_API_KEY&alt=json returned "The request cannot be completed because you have exceeded your <a href="/youtube/v3/getting-started#quota">quota</a>.". Details: "[{'message': 'The request cannot be completed because you have exceeded your <a href="/youtube/v3/getting-started#quota">quota</a>.', 'domain': 'youtube.quota', 'reason': 'quotaExceeded'}]">
Synthetic fallback: (1000, 8)
  video_id trending_date          title channel_title  category_id    views  \
0     vid0    2023-01-01  Video Title 0      ChannelC           10  3958242   
1     vid1    2023-01-02  Video Title 1      ChannelA           10  1227060   
2     vid2    2023-01-03  Video Title 2      ChannelC           24  1041519   

    likes  comment_count  
0  144402          18611  
1  160181           4977  
2   33966          18763  

Social Media Analytics: Key Concepts and Pitfalls#

  • Social media datasets track videos, posts, and user engagement.
  • Metrics like views, likes, comments show what content attracts people.
  • Titles, tags, and categories affect whether people discover your video.
  • CTR and watch time help reveal real audience interest.
  • Beginners often misinterpret engagement by ignoring reach, context, or category size.
  • Always compare similar content types and watch for missing or misleading data.
# Beginner Example 1: Viewing core metadata columns
print(df[['video_id', 'title', 'category_id']].head(5))
  video_id          title  category_id
0     vid0  Video Title 0           10
1     vid1  Video Title 1           10
2     vid2  Video Title 2           24
3     vid3  Video Title 3           22
4     vid4  Video Title 4           28
# Beginner Example 2: Checking for missing metadata
missing_titles = df['title'].isnull().sum()
missing_categories = df['category_id'].isnull().sum()
print(f'Missing titles: {missing_titles}, Missing categories: {missing_categories}')
Missing titles: 0, Missing categories: 0
# Beginner Example 3: Counting videos by category
category_counts = df['category_id'].value_counts()
print('Videos per category:')
print(category_counts.head(10))
Videos per category:
category_id
10    173
28    170
1     170
24    167
2     167
22    153
Name: count, dtype: int64
# Beginner Example 4: Find the longest and shortest video titles
df['title_length'] = df['title'].apply(len)
print('Longest title:', df.loc[df['title_length'].idxmax(), 'title'])
print('Shortest title:', df.loc[df['title_length'].idxmin(), 'title'])
Longest title: Video Title 100
Shortest title: Video Title 0
# Beginner Example 5: Average engagement (likes + comments) by category
df['engagement'] = df['likes'] + df['comment_count']
avg_engagement = df.groupby('category_id')['engagement'].mean().sort_values(ascending=False)
print('Average engagement by category:')
print(avg_engagement.head(10))
Average engagement by category:
category_id
10    131101.063584
22    129581.352941
28    127511.258824
1     125956.305882
24    123509.652695
2     116761.239521
Name: engagement, dtype: float64
# Intermediate Example 1: Top 5 most liked videos with their titles and category
top_liked = df.nlargest(5, 'likes')[['title', 'likes', 'category_id']]
print('Top 5 Most Liked Videos:')
print(top_liked)
Top 5 Most Liked Videos:
               title   likes  category_id
675  Video Title 675  199875           10
198  Video Title 198  199508            1
602  Video Title 602  198800           28
145  Video Title 145  198752           10
478  Video Title 478  198668           10
# Intermediate Example 2: Distribution of video title lengths
title_length_stats = df['title_length'].describe()
print('Statistics for video title lengths:')
print(title_length_stats)
Statistics for video title lengths:
count    1000.000000
mean       14.890000
std         0.343538
min        13.000000
25%        15.000000
50%        15.000000
75%        15.000000
max        15.000000
Name: title_length, dtype: float64
# Intermediate Example 3: Average views per category
cat_views = df.groupby('category_id')['views'].mean().sort_values(ascending=False)
print('Average views by category:')
print(cat_views.head(8))
Average views by category:
category_id
1     2.675815e+06
28    2.661432e+06
22    2.597320e+06
24    2.524280e+06
2     2.498725e+06
10    2.423112e+06
Name: views, dtype: float64
# Intermediate Example 4: Engagement rate by category (engagement per view)
df['engagement_rate'] = df['engagement'] / df['views']
rate_cat = df.groupby('category_id')['engagement_rate'].mean().sort_values(ascending=False)
print('Average engagement rate by category:')
print(rate_cat.head(10))
Average engagement rate by category:
category_id
24    0.155854
10    0.135900
2     0.121091
1     0.104556
28    0.104179
22    0.090596
Name: engagement_rate, dtype: float64
# Intermediate Example 5: Channel with most trending videos
channel_counts = df['channel_title'].value_counts()
top_channel = channel_counts.idxmax()
top_count = channel_counts.max()
print(f'Top trending channel: {top_channel} with {top_count} trending videos')
Top trending channel: ChannelA with 355 trending videos
# Intermediate Example 6: Calculate median views for videos with short titles (<20 characters)
short_titles = df[df['title_length'] < 20]
median_views_short = short_titles['views'].median()
print(f'Median views for short-titled videos: {median_views_short}')
Median views for short-titled videos: 2600985.5
# Advanced Example 1: Most common words in trending video titles
from collections import Counter
all_title_words = ' '.join(df['title']).lower().split()
common_words = Counter(all_title_words).most_common(10)
print('Top 10 most common words in titles:')
for word, count in common_words:
    print(f'{word}: {count}')
Top 10 most common words in titles:
video: 1000
title: 1000
0: 1
1: 1
2: 1
3: 1
4: 1
5: 1
6: 1
7: 1
# Advanced Example 2: Outlier detection for unusually high engagement rate
q3 = df['engagement_rate'].quantile(0.75)
iqr = df['engagement_rate'].quantile(0.75) - df['engagement_rate'].quantile(0.25)
outlier_threshold = q3 + 1.5 * iqr
outliers = df[df['engagement_rate'] > outlier_threshold]
print(f'Videos with unusually high engagement rate: {len(outliers)}')
print(outliers[['title', 'engagement_rate']].head(3))
Videos with unusually high engagement rate: 109
             title  engagement_rate
18  Video Title 18         0.245402
32  Video Title 32         0.430007
41  Video Title 41         0.244217
# Advanced Example 3: Assigning text-based tags based on title keywords
def assign_tag(title):
    title = title.lower()
    if 'challenge' in title:
        return 'challenge'
    elif 'tutorial' in title or 'how to' in title:
        return 'education'
    elif 'review' in title or 'unboxing' in title:
        return 'review'
    elif 'reaction' in title:
        return 'reaction'
    else:
        return 'other'
df['auto_tag'] = df['title'].apply(assign_tag)
print(df[['title', 'auto_tag']].head(8))
           title auto_tag
0  Video Title 0    other
1  Video Title 1    other
2  Video Title 2    other
3  Video Title 3    other
4  Video Title 4    other
5  Video Title 5    other
6  Video Title 6    other
7  Video Title 7    other
# Error Handling Example 1: Videos with zero or negative views
zero_views = df[df['views'] <= 0]
print(f'Videos with zero or negative views: {len(zero_views)}')
if not zero_views.empty:
    print(zero_views.head())
Videos with zero or negative views: 0
# Error Handling Example 2: Check for inconsistent aggregation
category_total_likes = df.groupby('category_id')['likes'].sum().sum()
dataset_total_likes = df['likes'].sum()
if category_total_likes != dataset_total_likes:
    print('Warning: Aggregation mismatch in likes!')
else:
    print('Aggregation OK: Total likes match.')
Aggregation OK: Total likes match.
# Error Handling Example 3: Handle missing category_id when grouping
if df['category_id'].isnull().any():
    clean_df = df.dropna(subset=['category_id'])
else:
    clean_df = df
category_mean = clean_df.groupby('category_id')['views'].mean()
print('Mean views per category (clean data):')
print(category_mean.head())
Mean views per category (clean data):
category_id
1     2.675815e+06
2     2.498725e+06
10    2.423112e+06
22    2.597320e+06
24    2.524280e+06
Name: views, dtype: float64
# Error Handling Example 4: Misinterpreting engagement ratios
def safe_engagement_rate(row):
    if row['views'] > 0:
        return row['engagement'] / row['views']
    else:
        return np.nan
df['safe_engagement_rate'] = df.apply(safe_engagement_rate, axis=1)
print('Example safe engagement rates:')
print(df.loc[df['views'] <= 20000, ['views', 'engagement', 'safe_engagement_rate']].head(5))
Example safe engagement rates:
Empty DataFrame
Columns: [views, engagement, safe_engagement_rate]
Index: []

Best Practices for Video Metadata Analytics#

  • Benchmark content by always using engagement rate, not just raw counts.
  • Segment your audience by category to spot niche opportunities.
  • Analyze trends in title, tags, and timing for growth.
  • Optimize by testing different metadata approaches and measuring results.
  • Be consistent in how you define key metrics over time.
# Best Practice Example 1: Benchmarking engagement for videos above and below median views
median_views = df['views'].median()
high_perf = df[df['views'] >= median_views]
low_perf = df[df['views'] < median_views]
mean_engagement_high = high_perf['engagement_rate'].mean()
mean_engagement_low = low_perf['engagement_rate'].mean()
print(f'Average engagement rate - High View Videos: {mean_engagement_high:.3f}')
print(f'Average engagement rate - Low View Videos: {mean_engagement_low:.3f}')
Average engagement rate - High View Videos: 0.035
Average engagement rate - Low View Videos: 0.203
# Best Practice Example 2: Finding rising trends in auto-generated tags
tag_counts = df['auto_tag'].value_counts()
print('Auto-generated tag frequencies:')
print(tag_counts)
Auto-generated tag frequencies:
auto_tag
other    1000
Name: count, dtype: int64
# Best Practice Example 3: Consistent metric definition - Create a clean engagement_rate column
def compute_engagement_rate(row):
    if row['views'] > 0:
        return row['engagement'] / row['views']
    return np.nan
df['engagement_rate'] = df.apply(compute_engagement_rate, axis=1)
print('Engagement rate column updated and cleaned!')
Engagement rate column updated and cleaned!

Mini Project: Find Top-Performing Category and Video#

  • Let us use everything we learned to identify:
    • The category with the highest average engagement rate.
    • The single most engaging video overall.
    • A simple actionable recommendation for content strategy.
# Find category with highest mean engagement rate
best_cat_id = df.groupby('category_id')['engagement_rate'].mean().idxmax()
best_cat_rate = df.groupby('category_id')['engagement_rate'].mean().max()
print(f'Best-performing category (ID): {best_cat_id}')
print(f'Highest average engagement rate: {best_cat_rate:.3f}')
# Find top single video
top_video = df.loc[df['engagement_rate'].idxmax()]
print('\nTop overall video:')
print('Title:', top_video['title'])
print('Channel:', top_video['channel_title'])
print('Engagement rate:', top_video['engagement_rate'])
# High-level recommendation
print('\nRecommendation: Focus on creating videos in the highest-engagement category, with titles like our top video!')
Best-performing category (ID): 24
Highest average engagement rate: 0.156

Top overall video:
Title: Video Title 675
Channel: ChannelC
Engagement rate: 5.735657646318439

Recommendation: Focus on creating videos in the highest-engagement category, with titles like our top video!

Summary & Next Steps#

  • You learned how to analyze video metadata fields in social media datasets.
  • Metadata is critical for content discovery, ranking, and engagement.
  • Use clean, descriptive titles and think about your target category for best results.
  • Keep iteratingmost viral videos start with strong, simple metadata.
  • Like and subscribe on YouTube if you found this helpful!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.