Lesson 42 · Social Media Content Analytics
Extracting Data Using the YouTube Data API
Learn how to collect trending video data from YouTube using the Data API. Understand why timely content insights help creators and brands grow their…
- CourseSocial Media Content Analytics
- Lesson42 of 41
- Video26 min
- FormatJupyter notebook · 16 code cells
What you'll learn
- Social Media Analytics Concepts for YouTube Data
- Beginner Example 1: Simple Video Counts and Sampling
- Beginner Example 2: Sorting Videos by Views
- Beginner Example 3: Basic Engagement Rate by Video
- Intermediate Example 1: Top Channels by Average Engagement Rate
- Intermediate Example 2: Trending Video Count by Category
- Intermediate Example 3: Finding Videos with High Comments/Views Ratio
- Advanced Example 1: Detecting Viral Videos Using Z-Score
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbExtracting Data Using the YouTube Data API#
Learn how to collect trending video data from YouTube using the Data API.
Understand why timely content insights help creators and brands grow their channels.
By the end, you will extract, preview, and analyze real YouTube trending metrics.
Skill: Connect analytics concepts with hands-on social media data extraction.
import pandas as pd
import numpy as np
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request
import warnings
warnings.filterwarnings('ignore')
Social Media Analytics Concepts for YouTube Data#
- The YouTube Trending dataset contains daily video metrics such as views, likes, and comments.
- Social media datasets record activities of content (e.g. videos), not just final statistics.
- Metrics like views, likes, and comments reflect both user interest and platform behavior.
- Click-through rate (CTR) and watch time are not always present in trending API data.
- Common mistake: Assuming total engagement equals quality or reach without accounting for trends and context.
# --- Data Setup: Fetching YouTube Trending Videos ---
SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
def get_yt_service():
api_key = os.environ.get('YOUTUBE_API_KEY')
if api_key:
return build('youtube', 'v3', developerKey=api_key)
if Path('client_secret.json').exists():
creds = None
if Path('token_ro.pickle').exists():
with open('token_ro.pickle', 'rb') as f:
creds = pickle.load(f)
if not creds or not creds.valid:
if creds and creds.expired and creds.refresh_token:
creds.refresh(Request())
else:
flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
creds = flow.run_local_server(port=0)
with open('token_ro.pickle', 'wb') as f:
pickle.dump(creds, f)
return build('youtube', 'v3', credentials=creds)
raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
def fetch_yt_trending(max_results=200, region='US'):
youtube = get_yt_service()
records, token = [], None
while len(records) < max_results:
resp = youtube.videos().list(
part='snippet,statistics',
chart='mostPopular',
regionCode=region,
maxResults=min(50, max_results - len(records)),
pageToken=token
).execute()
for item in resp.get('items', []):
s = item['snippet']; st = item.get('statistics', {})
records.append({
'video_id': item['id'],
'trending_date': pd.Timestamp.today().date(),
'title': s.get('title', ''),
'channel_title': s.get('channelTitle', ''),
'category_id': s.get('categoryId', ''),
'views': int(st.get('viewCount', 0)),
'likes': int(st.get('likeCount', 0)),
'comment_count': int(st.get('commentCount', 0)),
})
token = resp.get('nextPageToken')
if not token: break
return pd.DataFrame(records)
try:
df = fetch_yt_trending()
print('Live trending data:', df.shape)
except Exception as e:
print(f'Falling back to synthetic: {e}')
np.random.seed(42)
n = 1000
df = pd.DataFrame({
'video_id': [f'vid{i}' for i in range(n)],
'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
'title': [f'Video Title {i}' for i in range(n)],
'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
'category_id': np.random.choice([1,2,10,22,24,28], n),
'views': np.random.randint(10000, 5000000, n),
'likes': np.random.randint(100, 200000, n),
'comment_count': np.random.randint(10, 50000, n),
})
print('Synthetic fallback:', df.shape)
print(df.head(3))
# Previewing dataset shape and initial columns
print('Shape:', df.shape)
print('Columns:', df.columns.tolist())
Beginner Example 1: Simple Video Counts and Sampling#
- Let us calculate the number of trending videos we have in the data.
- Let us preview a few random YouTube video records.
- This helps us build intuition about data size and typical records.
# How many trending videos are in our snapshot?
n_videos = df['video_id'].nunique()
print(f'Unique trending videos: {n_videos}')
# Random sample of 5 rows for fast inspection
print(df.sample(5, random_state=42))
Beginner Example 2: Sorting Videos by Views#
- To spot popular content, sort videos by their view counts.
- Sorting lets you identify which videos are getting the most attention right now.
# Sort trending videos in descending order of view count
sorted_df = df.sort_values(by='views', ascending=False)
print(sorted_df[['title', 'channel_title', 'views']].head(7))
Beginner Example 3: Basic Engagement Rate by Video#
- Engagement rate combines likes and comments as a fraction of views.
- High engagement rates often mean deeper audience interest, not just reach.
# Calculate engagement rate for each video as (likes + comments)/views
df['engagement_rate'] = ((df['likes'] + df['comment_count']) / df['views']).round(4)
print(df[['title', 'views', 'likes', 'comment_count', 'engagement_rate']].head(5))
Intermediate Example 1: Top Channels by Average Engagement Rate#
- Let us find out which channels produce the most engaging trending videos on average.
- High channel engagement rates can signal strong community or content resonance.
# Group by channel and compute average engagement rate
ch_eng = (df.groupby('channel_title')['engagement_rate'].mean()
.sort_values(ascending=False))
print('Top channels by average engagement rate:')
print(ch_eng.head(10))
Intermediate Example 2: Trending Video Count by Category#
- Let us count how many trending videos belong to each YouTube content category.
- Category distributions reveal what topics are gaining momentum.
# Count trending videos by category ID
cat_counts = df['category_id'].value_counts()
print(cat_counts)
Intermediate Example 3: Finding Videos with High Comments/Views Ratio#
- High comments per view can signal controversy or deeply engaging topics.
- Let us list videos where viewers are especially likely to comment.
# Create a comments per view ratio and sort for the highest
df['comments_per_view'] = (df['comment_count'] / df['views']).round(4)
most_commented = df.sort_values('comments_per_view', ascending=False)
print(most_commented[['title', 'views', 'comment_count', 'comments_per_view']].head(5))
Advanced Example 1: Detecting Viral Videos Using Z-Score#
- Let us identify videos with unusually high view counts compared to all others (outliers).
- Statistically, these can be found using the Z-score method.
# Add a Z-score column based on view counts
df['views_z'] = ((df['views'] - df['views'].mean()) / df['views'].std()).round(2)
# Flagging videos with Z-score > 2.5 as viral candidates
viral_videos = df[df['views_z'] > 2.5]
print(f'Found {len(viral_videos)} viral candidate(s)')
print(viral_videos[['title', 'views', 'views_z']].head(5))
Advanced Example 2: Daily Trends and Rolling Averages#
- For ongoing trend detection, use rolling averages to smooth day-to-day swings.
- Let us aggregate total trending video views by date and apply a 7-day rolling mean.
# Aggregate total trending video views per day
daily_views = df.groupby('trending_date')['views'].sum()
# Compute a 7-day rolling average of daily total views
daily_views_smooth = daily_views.rolling(window=7, min_periods=1).mean()
print(daily_views_smooth.tail(10))
Error Handling Example 1: Detecting Missing Engagement Data#
- Sometimes API data may be missing likes or comment counts for certain videos.
- Let us check for missing or zero engagement metrics and flag them.
# Find videos with missing or zero likes/comments
missing = df[(df['likes'] == 0) | (df['comment_count'] == 0)]
print(f'Videos with zero likes or comments: {len(missing)}')
if not missing.empty:
print(missing[['title', 'likes', 'comment_count']].head(3))
Error Handling Example 2: Mistaken AggregationSum or Mean?#
- Let us show the difference between summing and averaging channel engagement rates.
- Aggregation mistakes can produce misleading ranking for creators.
# Compare sum and mean of engagement rates for channels
agg_df = df.groupby('channel_title')['engagement_rate'].agg(['sum', 'mean'])
print(agg_df.head(5))
Error Handling Example 3: Grouping by Wrong Column#
- A common mistake is grouping by a column that does not actually group what you expect.
- Let us see what happens if we group engagement by 'video_id' instead of 'channel_title'.
# Incorrect grouping exampleeach video is own group
wrong_group = df.groupby('video_id')['engagement_rate'].mean()
print(wrong_group.head(3))
Best Practices: Consistent Metric Definitions & Benchmarking#
- Always define metrics like engagement rate clearly and use them consistently.
- Use mean engagement per creator or topicnot sumfor fair benchmarking.
- Trending and viral analysis works better with rolling averages and Z-scores than with raw totals.
- Keep API authentication, data loading, and fallback plans clear and robust.
Common Analytics Patterns for Content Performance Analysis#
- Identify top-performing content by normalized engagement, not just raw numbers.
- Segment audience or content by relevant groupings (e.g., channel, category, posting day).
- Analyze both single-day spikes (viral) and longer growth patterns with rolling statistics.
- Use visualization tools to spot trends and outliers quickly.
End-to-End Example: Build a Top Trending Content List for Strategy#
- Problem: Which trending videos should a creator or brand use as benchmarks today?
- Solution: Build a list of top five videos with high views and high engagement rate.
- This narrows thousands of results to actionable insights for content improvement.
# Recommend top five videos with high views and high engagement (engagement > median)
eng_median = df['engagement_rate'].median()
top_both = df[df['engagement_rate'] > eng_median].sort_values('views', ascending=False).head(5)
print(top_both[['title', 'channel_title', 'views', 'engagement_rate']])
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



