Lesson 12 · Social Media Content Analytics
Understanding Video Metadata: Title, Tags, Category
In this lesson, we will learn how to analyze video metadata for social media and content analytics. We will focus on titles, tags, and categories using…
- CourseSocial Media Content Analytics
- Lesson12 of 41
- Video30 min
- FormatJupyter notebook · 24 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbUnderstanding Video Metadata: Title, Tags, Category#
- In this lesson, we will learn how to analyze video metadata for social media and content analytics.
- We will focus on titles, tags, and categories using real-world YouTube data.
- These elements are key for discoverability, reach, and audience engagement.
- Mastering metadata helps creators and businesses optimize their videos for better performance.
- By the end, you will be able to extract insights that guide smarter content strategies.
import pandas as pd
import numpy as np
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request
import warnings
warnings.filterwarnings('ignore')
# Setup: Load YouTube Trending Videos data
SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
def get_yt_service():
api_key = os.environ.get('YOUTUBE_API_KEY')
if api_key:
return build('youtube', 'v3', developerKey=api_key)
if Path('client_secret.json').exists():
creds = None
if Path('token_ro.pickle').exists():
with open('token_ro.pickle', 'rb') as f:
creds = pickle.load(f)
if not creds or not creds.valid:
if creds and creds.expired and creds.refresh_token:
creds.refresh(Request())
else:
flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
creds = flow.run_local_server(port=0)
with open('token_ro.pickle', 'wb') as f:
pickle.dump(creds, f)
return build('youtube', 'v3', credentials=creds)
raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
def fetch_yt_trending(max_results=200, region='US'):
youtube = get_yt_service()
records, token = [], None
while len(records) < max_results:
resp = youtube.videos().list(
part='snippet,statistics',
chart='mostPopular',
regionCode=region,
maxResults=min(50, max_results - len(records)),
pageToken=token
).execute()
for item in resp.get('items', []):
s = item['snippet']; st = item.get('statistics', {})
records.append({
'video_id': item['id'],
'trending_date': pd.Timestamp.today().date(),
'title': s.get('title', ''),
'channel_title': s.get('channelTitle', ''),
'category_id': s.get('categoryId', ''),
'views': int(st.get('viewCount', 0)),
'likes': int(st.get('likeCount', 0)),
'comment_count': int(st.get('commentCount', 0)),
})
token = resp.get('nextPageToken')
if not token: break
return pd.DataFrame(records)
try:
df = fetch_yt_trending()
print('Live trending data:', df.shape)
except Exception as e:
print(f'Falling back to synthetic: {e}')
np.random.seed(42)
n = 1000
df = pd.DataFrame({
'video_id': [f'vid{i}' for i in range(n)],
'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
'title': [f'Video Title {i}' for i in range(n)],
'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
'category_id': np.random.choice([1,2,10,22,24,28], n),
'views': np.random.randint(10000, 5000000, n),
'likes': np.random.randint(100, 200000, n),
'comment_count': np.random.randint(10, 50000, n),
})
print('Synthetic fallback:', df.shape)
print(df.head(3))
Social Media Analytics: Key Concepts and Pitfalls#
- Social media datasets track videos, posts, and user engagement.
- Metrics like views, likes, comments show what content attracts people.
- Titles, tags, and categories affect whether people discover your video.
- CTR and watch time help reveal real audience interest.
- Beginners often misinterpret engagement by ignoring reach, context, or category size.
- Always compare similar content types and watch for missing or misleading data.
# Beginner Example 1: Viewing core metadata columns
print(df[['video_id', 'title', 'category_id']].head(5))
# Beginner Example 2: Checking for missing metadata
missing_titles = df['title'].isnull().sum()
missing_categories = df['category_id'].isnull().sum()
print(f'Missing titles: {missing_titles}, Missing categories: {missing_categories}')
# Beginner Example 3: Counting videos by category
category_counts = df['category_id'].value_counts()
print('Videos per category:')
print(category_counts.head(10))
# Beginner Example 4: Find the longest and shortest video titles
df['title_length'] = df['title'].apply(len)
print('Longest title:', df.loc[df['title_length'].idxmax(), 'title'])
print('Shortest title:', df.loc[df['title_length'].idxmin(), 'title'])
# Beginner Example 5: Average engagement (likes + comments) by category
df['engagement'] = df['likes'] + df['comment_count']
avg_engagement = df.groupby('category_id')['engagement'].mean().sort_values(ascending=False)
print('Average engagement by category:')
print(avg_engagement.head(10))
# Intermediate Example 1: Top 5 most liked videos with their titles and category
top_liked = df.nlargest(5, 'likes')[['title', 'likes', 'category_id']]
print('Top 5 Most Liked Videos:')
print(top_liked)
# Intermediate Example 2: Distribution of video title lengths
title_length_stats = df['title_length'].describe()
print('Statistics for video title lengths:')
print(title_length_stats)
# Intermediate Example 3: Average views per category
cat_views = df.groupby('category_id')['views'].mean().sort_values(ascending=False)
print('Average views by category:')
print(cat_views.head(8))
# Intermediate Example 4: Engagement rate by category (engagement per view)
df['engagement_rate'] = df['engagement'] / df['views']
rate_cat = df.groupby('category_id')['engagement_rate'].mean().sort_values(ascending=False)
print('Average engagement rate by category:')
print(rate_cat.head(10))
# Intermediate Example 5: Channel with most trending videos
channel_counts = df['channel_title'].value_counts()
top_channel = channel_counts.idxmax()
top_count = channel_counts.max()
print(f'Top trending channel: {top_channel} with {top_count} trending videos')
# Intermediate Example 6: Calculate median views for videos with short titles (<20 characters)
short_titles = df[df['title_length'] < 20]
median_views_short = short_titles['views'].median()
print(f'Median views for short-titled videos: {median_views_short}')
# Advanced Example 1: Most common words in trending video titles
from collections import Counter
all_title_words = ' '.join(df['title']).lower().split()
common_words = Counter(all_title_words).most_common(10)
print('Top 10 most common words in titles:')
for word, count in common_words:
print(f'{word}: {count}')
# Advanced Example 2: Outlier detection for unusually high engagement rate
q3 = df['engagement_rate'].quantile(0.75)
iqr = df['engagement_rate'].quantile(0.75) - df['engagement_rate'].quantile(0.25)
outlier_threshold = q3 + 1.5 * iqr
outliers = df[df['engagement_rate'] > outlier_threshold]
print(f'Videos with unusually high engagement rate: {len(outliers)}')
print(outliers[['title', 'engagement_rate']].head(3))
# Advanced Example 3: Assigning text-based tags based on title keywords
def assign_tag(title):
title = title.lower()
if 'challenge' in title:
return 'challenge'
elif 'tutorial' in title or 'how to' in title:
return 'education'
elif 'review' in title or 'unboxing' in title:
return 'review'
elif 'reaction' in title:
return 'reaction'
else:
return 'other'
df['auto_tag'] = df['title'].apply(assign_tag)
print(df[['title', 'auto_tag']].head(8))
# Error Handling Example 1: Videos with zero or negative views
zero_views = df[df['views'] <= 0]
print(f'Videos with zero or negative views: {len(zero_views)}')
if not zero_views.empty:
print(zero_views.head())
# Error Handling Example 2: Check for inconsistent aggregation
category_total_likes = df.groupby('category_id')['likes'].sum().sum()
dataset_total_likes = df['likes'].sum()
if category_total_likes != dataset_total_likes:
print('Warning: Aggregation mismatch in likes!')
else:
print('Aggregation OK: Total likes match.')
# Error Handling Example 3: Handle missing category_id when grouping
if df['category_id'].isnull().any():
clean_df = df.dropna(subset=['category_id'])
else:
clean_df = df
category_mean = clean_df.groupby('category_id')['views'].mean()
print('Mean views per category (clean data):')
print(category_mean.head())
# Error Handling Example 4: Misinterpreting engagement ratios
def safe_engagement_rate(row):
if row['views'] > 0:
return row['engagement'] / row['views']
else:
return np.nan
df['safe_engagement_rate'] = df.apply(safe_engagement_rate, axis=1)
print('Example safe engagement rates:')
print(df.loc[df['views'] <= 20000, ['views', 'engagement', 'safe_engagement_rate']].head(5))
Best Practices for Video Metadata Analytics#
- Benchmark content by always using engagement rate, not just raw counts.
- Segment your audience by category to spot niche opportunities.
- Analyze trends in title, tags, and timing for growth.
- Optimize by testing different metadata approaches and measuring results.
- Be consistent in how you define key metrics over time.
# Best Practice Example 1: Benchmarking engagement for videos above and below median views
median_views = df['views'].median()
high_perf = df[df['views'] >= median_views]
low_perf = df[df['views'] < median_views]
mean_engagement_high = high_perf['engagement_rate'].mean()
mean_engagement_low = low_perf['engagement_rate'].mean()
print(f'Average engagement rate - High View Videos: {mean_engagement_high:.3f}')
print(f'Average engagement rate - Low View Videos: {mean_engagement_low:.3f}')
# Best Practice Example 2: Finding rising trends in auto-generated tags
tag_counts = df['auto_tag'].value_counts()
print('Auto-generated tag frequencies:')
print(tag_counts)
# Best Practice Example 3: Consistent metric definition - Create a clean engagement_rate column
def compute_engagement_rate(row):
if row['views'] > 0:
return row['engagement'] / row['views']
return np.nan
df['engagement_rate'] = df.apply(compute_engagement_rate, axis=1)
print('Engagement rate column updated and cleaned!')
Mini Project: Find Top-Performing Category and Video#
- Let us use everything we learned to identify:
- The category with the highest average engagement rate.
- The single most engaging video overall.
- A simple actionable recommendation for content strategy.
# Find category with highest mean engagement rate
best_cat_id = df.groupby('category_id')['engagement_rate'].mean().idxmax()
best_cat_rate = df.groupby('category_id')['engagement_rate'].mean().max()
print(f'Best-performing category (ID): {best_cat_id}')
print(f'Highest average engagement rate: {best_cat_rate:.3f}')
# Find top single video
top_video = df.loc[df['engagement_rate'].idxmax()]
print('\nTop overall video:')
print('Title:', top_video['title'])
print('Channel:', top_video['channel_title'])
print('Engagement rate:', top_video['engagement_rate'])
# High-level recommendation
print('\nRecommendation: Focus on creating videos in the highest-engagement category, with titles like our top video!')
Summary & Next Steps#
- You learned how to analyze video metadata fields in social media datasets.
- Metadata is critical for content discovery, ranking, and engagement.
- Use clean, descriptive titles and think about your target category for best results.
- Keep iteratingmost viral videos start with strong, simple metadata.
- Like and subscribe on YouTube if you found this helpful!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



