Mathew K Analytics

Lesson 61 · Social Media Content Analytics

End-to-End YouTube Analytics Project

In this lesson, we will solve a practical YouTube analytics problem using real or simulated trending video data. This project matters because understanding…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

End-to-End YouTube Analytics Project#

  • In this lesson, we will solve a practical YouTube analytics problem using real or simulated trending video data.

  • This project matters because understanding YouTube performance helps creators grow, spot trends early, and improve engagement.

  • We will explore how to identify top-performing videos, analyze audience behavior, and build actionable content insights.

  • By the end, you will be comfortable using core analytics concepts and making strategic recommendations based on data.

import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Analytics Concepts for YouTube Data#

  • Each row in a typical YouTube dataset is a trending video, including its title, channel, category, publish date, views, likes, and comments.
  • Engagement metrics like views, likes, comments, and click-through rate (CTR) measure audience interest and activity.
  • Watch time is how much total time viewers spent on the video; more watch time usually means better video quality or topic relevance.
  • New analysts often mistake high views or likes alone for success, but ratios and trends are more meaningful than raw numbers.
  • It is easy to misinterpret engagement spikes if you do not check views, likes, and comments together over time.
# Set up the dataset: YouTube Trending Videos
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request

SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']

def get_yt_service():
    api_key = os.environ.get('YOUTUBE_API_KEY')
    if api_key:
        return build('youtube', 'v3', developerKey=api_key)
    if Path('client_secret.json').exists():
        creds = None
        if Path('token_ro.pickle').exists():
            with open('token_ro.pickle', 'rb') as f:
                creds = pickle.load(f)
        if not creds or not creds.valid:
            if creds and creds.expired and creds.refresh_token:
                creds.refresh(Request())
            else:
                flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
                creds = flow.run_local_server(port=0)
            with open('token_ro.pickle', 'wb') as f:
                pickle.dump(creds, f)
        return build('youtube', 'v3', credentials=creds)
    raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')

def fetch_yt_trending(max_results=200, region='US'):
    youtube = get_yt_service()
    records, token = [], None
    while len(records) < max_results:
        resp = youtube.videos().list(
            part='snippet,statistics',
            chart='mostPopular',
            regionCode=region,
            maxResults=min(50, max_results - len(records)),
            pageToken=token
        ).execute()
        for item in resp.get('items', []):
            s = item['snippet']; st = item.get('statistics', {})
            records.append({
                'video_id':      item['id'],
                'trending_date': pd.Timestamp.today().date(),
                'title':         s.get('title', ''),
                'channel_title': s.get('channelTitle', ''),
                'category_id':   s.get('categoryId', ''),
                'views':         int(st.get('viewCount', 0)),
                'likes':         int(st.get('likeCount', 0)),
                'comment_count': int(st.get('commentCount', 0)),
            })
        token = resp.get('nextPageToken')
        if not token: break
    return pd.DataFrame(records)

try:
    df = fetch_yt_trending()
    print('Live trending data:', df.shape)
except Exception as e:
    print(f'Falling back to synthetic: {e}')
    np.random.seed(42)
    n = 1000
    df = pd.DataFrame({
        'video_id':      [f'vid{i}' for i in range(n)],
        'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
        'title':         [f'Video Title {i}' for i in range(n)],
        'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
        'category_id':   np.random.choice([1,2,10,22,24,28], n),
        'views':         np.random.randint(10000, 5000000, n),
        'likes':         np.random.randint(100, 200000, n),
        'comment_count': np.random.randint(10, 50000, n),
    })
    print('Synthetic fallback:', df.shape)

print(df.head(3))
Live trending data: (199, 8)
      video_id trending_date  \
0  82-jTNka3uc    2026-06-03   
1  3oB9AxspVow    2026-06-03   
2  Aq9EJW9XqjQ    2026-06-03   

                                               title     channel_title  \
0  Ariana Grande - hate that i made you love me (...  ArianaGrandeVevo   
1           The End of Oak Street | Official Trailer      Warner Bros.   
2               I Shouldn’t Have Moved To This Farm…            CaseOh   

  category_id    views   likes  comment_count  
0          10  4277584  537457          30194  
1           1  4374875   45933           3544  
2          20   310516   11266            888  

Example 1: What videos are trending now?#

  • It is important for content creators to know what kinds of videos are currently popular.
  • Examining the most recent trending videos helps spot emerging topics and formats.
  • We will look at the top 5 most recently trending videos.
recent_videos = df.sort_values('trending_date', ascending=False).head(5)
print('Top recent trending videos:')
print(recent_videos[['title','channel_title','category_id','views','likes','comment_count']])
Top recent trending videos:
                                                 title     channel_title  \
0    Ariana Grande - hate that i made you love me (...  ArianaGrandeVevo   
125          We Downloaded The BEST Sparking Zero Mods          DotoDoya   
127  Guess the Pro COD Player Using ONLY Gameplay (...       OpTic Dashy   
128     This Item Could Save Me | Sailing Locked (#13)           Settled   
129  Playing As WHAT THE DOG DOIN To Troll My Frien...            Daaash   

    category_id    views   likes  comment_count  
0            10  4277584  537457          30194  
125          20    66266    4399            277  
127          20    38152    1898            119  
128          20   326083   17716           1243  
129          20   178816    3230            490  

Example 2: Calculating Engagement Rate#

  • Engagement rate (likes plus comments divided by views) shows how much the audience interacts relative to view count.
  • High engagement means an active, interested audience, not just passive viewing.
  • Let us compute the engagement rate for each trending video.
df['engagement_rate'] = (df['likes'] + df['comment_count']) / df['views']
print('Engagement rate calculated for each video.')
print(df[['title','engagement_rate']].head(3))
Engagement rate calculated for each video.
                                               title  engagement_rate
0  Ariana Grande - hate that i made you love me (...         0.132704
1           The End of Oak Street | Official Trailer         0.011309
2               I Shouldn’t Have Moved To This Farm…         0.039141

Example 3: Top Performing Videos by Views#

  • Sometimes the best-performing content is simply the most watched.
  • Identifying the videos with the highest view counts helps you learn what topics or formats attract the biggest audiences.
top_views = df.sort_values('views', ascending=False).head(5)
print('Most watched trending videos:')
print(top_views[['title','channel_title','views','likes','engagement_rate']])
Most watched trending videos:
                                                title     channel_title  \
3   IShowSpeed - World Cup (Champions) [Official M...        IShowSpeed   
1            The End of Oak Street | Official Trailer      Warner Bros.   
0   Ariana Grande - hate that i made you love me (...  ArianaGrandeVevo   
9                        MEOVV(미야오) - ‘DDI RO RI’ M/V     THEBLACKLABEL   
10  Cocktail 2 Official Trailer | Shahid Kapoor, K...     Maddock Films   

      views   likes  engagement_rate  
3   4700685  709266         0.163314  
1   4374875   45933         0.011309  
0   4277584  537457         0.132704  
9   3488616       0         0.002987  
10  3081191   40111         0.013884  

Example 4: Intermediate - Most Engaging Videos#

  • Let us find the trending videos with the highest engagement rates, not just the most total views.
  • This helps discover content that viewers truly respond to with likes or comments (even if total views are less impressive).
top_engagement = df.sort_values('engagement_rate', ascending=False).head(5)
print('Top videos by engagement rate:')
print(top_engagement[['title','channel_title','views','likes','comment_count','engagement_rate']])
Top videos by engagement rate:
                                                title           channel_title  \
34  First Glimpse! The Altar Of Pergamon: Trailer ...  Jonathan Cahn Official   
42        3Quency - Girls Talk (Official Music Video)                 3Quency   
21                             Vince Staples - Cotton           Vince Staples   
63          bleood - ding dong (official music video)                  bleood   
30      Mastodon - Your Ghost Again (Official Visual)                Mastodon   

    views  likes  comment_count  engagement_rate  
34   3491    788             45         0.238614  
42  24152   5006            428         0.224992  
21  52972   9828            639         0.197595  
63  31618   5334            630         0.188627  
30  47767   7375            725         0.169573  

Example 5: Intermediate - Trending Categories Comparison#

  • Understanding which video categories trend most can help inform future content choices.
  • We will summarize average views and engagement by category.
category_summary = df.groupby('category_id').agg({'views':'mean','engagement_rate':'mean','video_id':'count'}).reset_index()
category_summary = category_summary.rename(columns={'video_id':'video_count'})
print(category_summary.sort_values('views', ascending=False))
  category_id         views  engagement_rate  video_count
0           1  2.220568e+06         0.028001            3
1          10  5.404715e+05         0.060598           24
5          24  4.207682e+05         0.060527           15
6          28  3.061750e+05         0.044772            1
3          20  2.459984e+05         0.051204          145
2          17  1.895295e+05         0.060747            2
4          22  1.008580e+05         0.081450            9

Example 6: Intermediate - Top Channels This Week#

  • We can see which channels appear most often in the trending list over the past week.
  • This shows which creators are currently dominating the trending space.
one_week_ago = df['trending_date'].max() - pd.Timedelta(days=7)
recent_week = df[df['trending_date'] >= one_week_ago]
top_channels = recent_week['channel_title'].value_counts().head(5)
print('Channels with most trending videos in the past week:')
print(top_channels)
Channels with most trending videos in the past week:
channel_title
ArianaGrandeVevo      1
Warner Bros.          1
CaseOh                1
IShowSpeed            1
Paramount Pictures    1
Name: count, dtype: int64

Example 7: Advanced - Time Series of Average Views per Day#

  • Tracking how the average number of views per trending video changes over time reveals growth, seasonality, or major spikes.
  • This helps spot viral periods or slowdowns for the platform as a whole.
daily_views = df.groupby('trending_date')['views'].mean().reset_index()
import matplotlib.pyplot as plt
plt.figure(figsize=(10,4))
plt.plot(daily_views['trending_date'], daily_views['views'], marker='o', alpha=0.6)
plt.title('Average Trending Video Views per Day')
plt.xlabel('Date')
plt.ylabel('Avg Views')
plt.tight_layout()
plt.show()
No description has been provided for this image

Example 8: Advanced - Detecting Potential Viral Videos#

  • A viral video climbs to extreme engagement or view levels much higher than normal.
  • Let us flag videos that are in the top 1% of both views and engagement rate.
top_views_thresh = df['views'].quantile(0.99)
top_engage_thresh = df['engagement_rate'].quantile(0.99)
viral_candidates = df[(df['views'] >= top_views_thresh) & (df['engagement_rate'] >= top_engage_thresh)]
print('Potential viral trending videos:')
print(viral_candidates[['title','channel_title','views','likes','engagement_rate']])
Potential viral trending videos:
Empty DataFrame
Columns: [title, channel_title, views, likes, engagement_rate]
Index: []

Error Handling: What if Engagement Numbers Are Missing?#

  • In real YouTube data, likes or comments can sometimes be missing or not public.
  • We need to make sure our calculations handle nulls gracefully.
# Simulate missing values
df_missing = df.copy()
df_missing.loc[df_missing.sample(frac=0.02, random_state=42).index, 'likes'] = np.nan
df_missing['engagement_rate'] = (df_missing['likes'].fillna(0) + df_missing['comment_count'].fillna(0)) / df_missing['views']
print('Null-safe engagement rate (likes NaNs treated as zero):')
print(df_missing[['likes','comment_count','views','engagement_rate']].head(3))
Null-safe engagement rate (likes NaNs treated as zero):
      likes  comment_count    views  engagement_rate
0  537457.0          30194  4277584         0.132704
1   45933.0           3544  4374875         0.011309
2   11266.0            888   310516         0.039141

Error Handling: Incorrect Aggregations Can Mislead#

  • If you forget to group by channel or category before averaging, you can get totally misleading insight.
  • Always double-check group-by and aggregation steps.
# WRONG: averaging engagement with no grouping
wrong_avg = df['engagement_rate'].mean()
print(f'Overall avg engagement rate (misleading): {wrong_avg:.4f}')

# CORRECT: average engagement by category
correct_avg = df.groupby('category_id')['engagement_rate'].mean()
print('Avg engagement rate per category (correct): ')
print(correct_avg)
Overall avg engagement rate (misleading): 0.0541
Avg engagement rate per category (correct): 
category_id
1     0.028001
10    0.060598
17    0.060747
20    0.051204
22    0.081450
24    0.060527
28    0.044772
Name: engagement_rate, dtype: float64

Error Handling: Misinterpreting Ratios Like CTR#

  • You must never interpret CTR, engagement rates, or other ratios without knowing their denominator and context.
  • A high CTR on very few impressions is less impactful than a high CTR on a huge audience.
few_views = df.sort_values('views').head(5)
print('Smallest audience videos with engagement rates:')
print(few_views[['title','views','engagement_rate']])
Smallest audience videos with engagement rates:
                                                 title  views  engagement_rate
34   First Glimpse! The Altar Of Pergamon: Trailer ...   3491         0.238614
31   Modern Warfare 4 Captain Price Down Teaser Tra...   4932         0.015207
161  Things Get Crazy Ahead Of Sony's State of Play...  16610         0.142264
24   Jarahn - Malasang (Audio) feat. Saii Kay x Tat...  18593         0.060937
64                    Let’s Play: Escape the Backrooms  20182         0.090179

Best Practice: Always Use Clear Metric Definitions#

  • Define each metric you report, and include the denominator (such as likes divided by views) in your analysis.
  • Share engagement and reach metrics together so creators understand both growth and loyalty.

Best Practice: Benchmark Against Peers#

  • Instead of just looking at your own channel, compare performance to similar channels or the weekly trending set.
  • This reveals whether your growth is actually above average or falling behind the market.
channel_avg = df.groupby('channel_title')['views'].mean().sort_values(ascending=False).head(5)
print('Channels with top average trending views:')
print(channel_avg)
Channels with top average trending views:
channel_title
IShowSpeed          4700685.0
Warner Bros.        4374875.0
ArianaGrandeVevo    4277584.0
THEBLACKLABEL       3488616.0
Maddock Films       3081191.0
Name: views, dtype: float64

Analytics Pattern: Rolling Averages for Smarter Trend Detection#

  • Rolling averages smooth out day-to-day noise so you see real upward or downward trends.
  • They are especially helpful for visualizing multi-week growth or dips in engagement.
daily = df.groupby('trending_date').agg({'views':'mean','engagement_rate':'mean'}).reset_index()
daily['views_rolling7'] = daily['views'].rolling(7, min_periods=1).mean()
import matplotlib.pyplot as plt
plt.figure(figsize=(10,4))
plt.plot(daily['trending_date'], daily['views_rolling7'], label='7-Day Rolling Mean', color='orange')
plt.title('7-Day Rolling Average of Trending Video Views')
plt.xlabel('Date')
plt.ylabel('Rolling Avg Views')
plt.legend()
plt.tight_layout()
plt.show()
No description has been provided for this image

Content Optimization: Find the Best Day to Post#

  • Posting on days when engagement is highest can make it easier to reach new audiences.
  • Let us see which weekday has the highest average engagement rate.
df['weekday'] = pd.to_datetime(df['trending_date']).dt.day_name()
weekday_engage = df.groupby('weekday')['engagement_rate'].mean().sort_values(ascending=False)
print('Average engagement rate by weekday:')
print(weekday_engage)
Average engagement rate by weekday:
weekday
Wednesday    0.054121
Name: engagement_rate, dtype: float64

End-to-End Mini Project: Actionable Content Strategy#

  • Your goal is to deliver a short report with three recommendations for a YouTube creator based on trending video data.
  • Steps:
    • Identify the channel with the most trending videos this month.
      • Find the top three video titles with the highest engagement rate from that channel.
        • Recommend the best day they should post new content for maximum engagement.
this_month = pd.to_datetime(df['trending_date']).dt.to_period('M').max().to_timestamp()
month_df = df[pd.to_datetime(df['trending_date']).dt.to_period('M') == this_month.to_period('M')]

# 1. Channel with most trending videos this month
top_channel = month_df['channel_title'].value_counts().index[0]
print(f'Top trending channel this month: {top_channel}')

# 2. Top three high-engagement videos from that channel
ch_videos = month_df[month_df['channel_title'] == top_channel]
top_ch_videos = ch_videos.sort_values('engagement_rate', ascending=False).head(3)
print('Most engaging video titles from that channel:')
print(top_ch_videos['title'])

# 3. Best weekday for engagement (use all their videos this month)
best_day = ch_videos.groupby('weekday')['engagement_rate'].mean().idxmax()
print(f'Recommended best day to post: {best_day}')
Top trending channel this month: ArianaGrandeVevo
Most engaging video titles from that channel:
0    Ariana Grande - hate that i made you love me (...
Name: title, dtype: object
Recommended best day to post: Wednesday
# Optional: Save a brief summary report
output_text = f"Best channel this month: {top_channel}\n" + \
              "Top engaging titles: " + ', '.join(top_ch_videos['title'].tolist()) + "\n" + \
              f"Best day to post: {best_day}\n"
with open('yt_trending_summary.txt', 'w') as f:
    f.write(output_text)
print('Wrote report to yt_trending_summary.txt')
Wrote report to yt_trending_summary.txt

YouTube Analytics Certification and Next Steps#

  • You now know how to analyze trending YouTube data, detect viral patterns, and produce actionable content strategy.
  • Keep practicing these analyses and try applying them to your own or public YouTube datasets.
  • Like and subscribe to the official channel for more social media analytics projects!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.