Mathew K Analytics

Lesson 42 · Social Media Content Analytics

Extracting Data Using the YouTube Data API

Learn how to collect trending video data from YouTube using the Data API. Understand why timely content insights help creators and brands grow their…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Extracting Data Using the YouTube Data API#

  • Learn how to collect trending video data from YouTube using the Data API.

  • Understand why timely content insights help creators and brands grow their channels.

  • By the end, you will extract, preview, and analyze real YouTube trending metrics.

  • Skill: Connect analytics concepts with hands-on social media data extraction.

import pandas as pd
import numpy as np
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request
import warnings
warnings.filterwarnings('ignore')

Social Media Analytics Concepts for YouTube Data#

  • The YouTube Trending dataset contains daily video metrics such as views, likes, and comments.
  • Social media datasets record activities of content (e.g. videos), not just final statistics.
  • Metrics like views, likes, and comments reflect both user interest and platform behavior.
  • Click-through rate (CTR) and watch time are not always present in trending API data.
  • Common mistake: Assuming total engagement equals quality or reach without accounting for trends and context.
# --- Data Setup: Fetching YouTube Trending Videos ---
SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']

def get_yt_service():
    api_key = os.environ.get('YOUTUBE_API_KEY')
    if api_key:
        return build('youtube', 'v3', developerKey=api_key)
    if Path('client_secret.json').exists():
        creds = None
        if Path('token_ro.pickle').exists():
            with open('token_ro.pickle', 'rb') as f:
                creds = pickle.load(f)
        if not creds or not creds.valid:
            if creds and creds.expired and creds.refresh_token:
                creds.refresh(Request())
            else:
                flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
                creds = flow.run_local_server(port=0)
            with open('token_ro.pickle', 'wb') as f:
                pickle.dump(creds, f)
        return build('youtube', 'v3', credentials=creds)
    raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')

def fetch_yt_trending(max_results=200, region='US'):
    youtube = get_yt_service()
    records, token = [], None
    while len(records) < max_results:
        resp = youtube.videos().list(
            part='snippet,statistics',
            chart='mostPopular',
            regionCode=region,
            maxResults=min(50, max_results - len(records)),
            pageToken=token
        ).execute()
        for item in resp.get('items', []):
            s = item['snippet']; st = item.get('statistics', {})
            records.append({
                'video_id':      item['id'],
                'trending_date': pd.Timestamp.today().date(),
                'title':         s.get('title', ''),
                'channel_title': s.get('channelTitle', ''),
                'category_id':   s.get('categoryId', ''),
                'views':         int(st.get('viewCount', 0)),
                'likes':         int(st.get('likeCount', 0)),
                'comment_count': int(st.get('commentCount', 0)),
            })
        token = resp.get('nextPageToken')
        if not token: break
    return pd.DataFrame(records)

try:
    df = fetch_yt_trending()
    print('Live trending data:', df.shape)
except Exception as e:
    print(f'Falling back to synthetic: {e}')
    np.random.seed(42)
    n = 1000
    df = pd.DataFrame({
        'video_id':      [f'vid{i}' for i in range(n)],
        'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
        'title':         [f'Video Title {i}' for i in range(n)],
        'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
        'category_id':   np.random.choice([1,2,10,22,24,28], n),
        'views':         np.random.randint(10000, 5000000, n),
        'likes':         np.random.randint(100, 200000, n),
        'comment_count': np.random.randint(10, 50000, n),
    })
    print('Synthetic fallback:', df.shape)

print(df.head(3))
Live trending data: (199, 8)
      video_id trending_date  \
0  82-jTNka3uc    2026-06-02   
1  3oB9AxspVow    2026-06-02   
2  l-cyT28MFyk    2026-06-02   

                                               title     channel_title  \
0  Ariana Grande - hate that i made you love me (...  ArianaGrandeVevo   
1           The End of Oak Street | Official Trailer      Warner Bros.   
2        PLAYING HARDCORE MINECRAFT UNTIL WE BEAT IT            Jynxzi   

  category_id    views   likes  comment_count  
0          10  2065102  378053          25520  
1           1  1334321   29348           2505  
2          24   658907   14957            263  
# Previewing dataset shape and initial columns
print('Shape:', df.shape)
print('Columns:', df.columns.tolist())
Shape: (199, 8)
Columns: ['video_id', 'trending_date', 'title', 'channel_title', 'category_id', 'views', 'likes', 'comment_count']

Beginner Example 1: Simple Video Counts and Sampling#

  • Let us calculate the number of trending videos we have in the data.
  • Let us preview a few random YouTube video records.
  • This helps us build intuition about data size and typical records.
# How many trending videos are in our snapshot?
n_videos = df['video_id'].nunique()
print(f'Unique trending videos: {n_videos}')

# Random sample of 5 rows for fast inspection
print(df.sample(5, random_state=42))
Unique trending videos: 199
        video_id trending_date  \
82   Gpq9TXZxYqU    2026-06-02   
15   1RZLTuaG6_k    2026-06-02   
111  IGnQcw9ZEWs    2026-06-02   
177  G2-fsatqCuE    2026-06-02   
76   qLKKbahl2ac    2026-06-02   

                                                 title        channel_title  \
82                                 I TRIED SAKUPEN END            ZoinkLIVE   
15    Yng Lvcas & Peso Pluma - La Bebe (Remix) (Letra)        Keller Lyrics   
111  I Found A TOXIC STREAMER, So I CALLED HIS MOM ...           Techy Plus   
177  I Logged Into my FIRST Steal a Brainrot Account..            WinterSIM   
76             We SNEAK into BEAN'S Head..(Brookhaven)  Angelazz Brookhaven   

    category_id   views  likes  comment_count  
82           20  236507  10078            843  
15           10  166544    401             10  
111          20   23558    937            136  
177          20  198459   4195            659  
76           20   65757   3213            378  

Beginner Example 2: Sorting Videos by Views#

  • To spot popular content, sort videos by their view counts.
  • Sorting lets you identify which videos are getting the most attention right now.
# Sort trending videos in descending order of view count
sorted_df = df.sort_values(by='views', ascending=False)
print(sorted_df[['title', 'channel_title', 'views']].head(7))
                                                 title     channel_title  \
29                               TREASURE - ‘IF I’ M/V    TREASURE (트레저)   
106          I Went to WAR on a Hardcore Minecraft SMP            Wemmbu   
113          SIDEMEN AMONG US: HARRY POTTER CHAOS MODE       MoreSidemen   
85                   The Most SATISFYING Roblox Game..            Foltyn   
3    IShowSpeed - World Cup (Champions) [Official M...        IShowSpeed   
0    Ariana Grande - hate that i made you love me (...  ArianaGrandeVevo   
158                 JUGAMOS A LAS ESCONDIDAS EN ROBLOX         FedeGames   

       views  
29   4023274  
106  3629894  
113  2648718  
85   2134136  
3    2131908  
0    2065102  
158  2044336  

Beginner Example 3: Basic Engagement Rate by Video#

  • Engagement rate combines likes and comments as a fraction of views.
  • High engagement rates often mean deeper audience interest, not just reach.
# Calculate engagement rate for each video as (likes + comments)/views
df['engagement_rate'] = ((df['likes'] + df['comment_count']) / df['views']).round(4)
print(df[['title', 'views', 'likes', 'comment_count', 'engagement_rate']].head(5))
                                               title    views   likes  \
0  Ariana Grande - hate that i made you love me (...  2065102  378053   
1           The End of Oak Street | Official Trailer  1334321   29348   
2        PLAYING HARDCORE MINECRAFT UNTIL WE BEAT IT   658907   14957   
3  IShowSpeed - World Cup (Champions) [Official M...  2131908  422883   
4  No Peace Amongst the Stars | Warhammer 40,000 ...  1997485  146236   

   comment_count  engagement_rate  
0          25520           0.1954  
1           2505           0.0239  
2            263           0.0231  
3          39316           0.2168  
4           7188           0.0768  

Intermediate Example 1: Top Channels by Average Engagement Rate#

  • Let us find out which channels produce the most engaging trending videos on average.
  • High channel engagement rates can signal strong community or content resonance.
# Group by channel and compute average engagement rate
ch_eng = (df.groupby('channel_title')['engagement_rate'].mean()
           .sort_values(ascending=False))
print('Top channels by average engagement rate:')
print(ch_eng.head(10))
Top channels by average engagement rate:
channel_title
KOT4Q               0.3948
Ivycomb Music       0.3375
Bill McClintock     0.3173
IShowSpeed          0.2168
ArianaGrandeVevo    0.1954
AtumisManor         0.1953
Clooless Gaming     0.1860
SMTOWN              0.1846
Slushy Noobz        0.1802
Jason Schreier      0.1750
Name: engagement_rate, dtype: float64

Intermediate Example 2: Trending Video Count by Category#

  • Let us count how many trending videos belong to each YouTube content category.
  • Category distributions reveal what topics are gaining momentum.
# Count trending videos by category ID
cat_counts = df['category_id'].value_counts()
print(cat_counts)
category_id
20    144
10     29
24     14
22      5
1       4
17      1
28      1
23      1
Name: count, dtype: int64

Intermediate Example 3: Finding Videos with High Comments/Views Ratio#

  • High comments per view can signal controversy or deeply engaging topics.
  • Let us list videos where viewers are especially likely to comment.
# Create a comments per view ratio and sort for the highest
df['comments_per_view'] = (df['comment_count'] / df['views']).round(4)
most_commented = df.sort_values('comments_per_view', ascending=False)
print(most_commented[['title', 'views', 'comment_count', 'comments_per_view']].head(5))
                                                 title   views  comment_count  \
67                                 Homesick Wanderlust    7691            305   
65                         Queen Sabbath - "Rockanoid"   17934            637   
127  الرجل العنكبوت انقاذ باتمان Spider-Man Rescue ...  200553           7027   
80   MHA Voice Actors Make DEKU CONFESS in TOMODACH...  134289           3050   
18   Spider-Man Vs Hulk, Lex Warsuit First Look, Ba...   15460            333   

     comments_per_view  
67              0.0397  
65              0.0355  
127             0.0350  
80              0.0227  
18              0.0215  

Advanced Example 1: Detecting Viral Videos Using Z-Score#

  • Let us identify videos with unusually high view counts compared to all others (outliers).
  • Statistically, these can be found using the Z-score method.
# Add a Z-score column based on view counts
df['views_z'] = ((df['views'] - df['views'].mean()) / df['views'].std()).round(2)
# Flagging videos with Z-score > 2.5 as viral candidates
viral_videos = df[df['views_z'] > 2.5]
print(f'Found {len(viral_videos)} viral candidate(s)')
print(viral_videos[['title', 'views', 'views_z']].head(5))
Found 10 viral candidate(s)
                                                title    views  views_z
0   Ariana Grande - hate that i made you love me (...  2065102     2.96
3   IShowSpeed - World Cup (Champions) [Official M...  2131908     3.08
4   No Peace Amongst the Stars | Warhammer 40,000 ...  1997485     2.85
5             1000 VS 1000 Player Minecraft Civil War  2009970     2.87
29                              TREASURE - ‘IF I’ M/V  4023274     6.27

Advanced Example 2: Daily Trends and Rolling Averages#

  • For ongoing trend detection, use rolling averages to smooth day-to-day swings.
  • Let us aggregate total trending video views by date and apply a 7-day rolling mean.
# Aggregate total trending video views per day
daily_views = df.groupby('trending_date')['views'].sum()
# Compute a 7-day rolling average of daily total views
daily_views_smooth = daily_views.rolling(window=7, min_periods=1).mean()
print(daily_views_smooth.tail(10))
trending_date
2026-06-02    61969950.0
Name: views, dtype: float64

Error Handling Example 1: Detecting Missing Engagement Data#

  • Sometimes API data may be missing likes or comment counts for certain videos.
  • Let us check for missing or zero engagement metrics and flag them.
# Find videos with missing or zero likes/comments
missing = df[(df['likes'] == 0) | (df['comment_count'] == 0)]
print(f'Videos with zero likes or comments: {len(missing)}')
if not missing.empty:
    print(missing[['title', 'likes', 'comment_count']].head(3))
Videos with zero likes or comments: 5
                           title  likes  comment_count
6   MEOVV(미야오) - ‘DDI RO RI’ M/V      0           8389
26            Cold Hearted Bitch   2089              0
45                    Unbothered   5378              0

Error Handling Example 2: Mistaken AggregationSum or Mean?#

  • Let us show the difference between summing and averaging channel engagement rates.
  • Aggregation mistakes can produce misleading ranking for creators.
# Compare sum and mean of engagement rates for channels
agg_df = df.groupby('channel_title')['engagement_rate'].agg(['sum', 'mean'])
print(agg_df.head(5))
                             sum    mean
channel_title                           
3C Films                  0.1633  0.1633
AR12Gaming                0.0655  0.0655
Abnaze                    0.0288  0.0288
Affirmation Club - Topic  0.0256  0.0256
Angelazz Brookhaven       0.0546  0.0546

Error Handling Example 3: Grouping by Wrong Column#

  • A common mistake is grouping by a column that does not actually group what you expect.
  • Let us see what happens if we group engagement by 'video_id' instead of 'channel_title'.
# Incorrect grouping exampleeach video is own group
wrong_group = df.groupby('video_id')['engagement_rate'].mean()
print(wrong_group.head(3))
video_id
-8h7JN41i8M    0.0720
-E2x80YumMI    0.0541
-GQb1HUCShk    0.0607
Name: engagement_rate, dtype: float64

Best Practices: Consistent Metric Definitions & Benchmarking#

  • Always define metrics like engagement rate clearly and use them consistently.
  • Use mean engagement per creator or topicnot sumfor fair benchmarking.
  • Trending and viral analysis works better with rolling averages and Z-scores than with raw totals.
  • Keep API authentication, data loading, and fallback plans clear and robust.

Common Analytics Patterns for Content Performance Analysis#

  • Identify top-performing content by normalized engagement, not just raw numbers.
  • Segment audience or content by relevant groupings (e.g., channel, category, posting day).
  • Analyze both single-day spikes (viral) and longer growth patterns with rolling statistics.
  • Use visualization tools to spot trends and outliers quickly.

End-to-End Example: Build a Top Trending Content List for Strategy#

  • Problem: Which trending videos should a creator or brand use as benchmarks today?
  • Solution: Build a list of top five videos with high views and high engagement rate.
  • This narrows thousands of results to actionable insights for content improvement.
# Recommend top five videos with high views and high engagement (engagement > median)
eng_median = df['engagement_rate'].median()
top_both = df[df['engagement_rate'] > eng_median].sort_values('views', ascending=False).head(5)
print(top_both[['title', 'channel_title', 'views', 'engagement_rate']])
                                                 title     channel_title  \
29                               TREASURE - ‘IF I’ M/V    TREASURE (트레저)   
106          I Went to WAR on a Hardcore Minecraft SMP            Wemmbu   
3    IShowSpeed - World Cup (Champions) [Official M...        IShowSpeed   
0    Ariana Grande - hate that i made you love me (...  ArianaGrandeVevo   
5              1000 VS 1000 Player Minecraft Civil War        FlameFrags   

       views  engagement_rate  
29   4023274           0.0664  
106  3629894           0.0636  
3    2131908           0.2168  
0    2065102           0.1954  
5    2009970           0.0724  
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.