Mathew K Analytics

Lesson 34 · Social Media Content Analytics

Analyzing Content Categories and Niches

In this lesson, we will learn how to analyze content categories and niches using real social media data. Knowing which content categories perform best helps…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Analyzing Content Categories and Niches#

  • In this lesson, we will learn how to analyze content categories and niches using real social media data.
  • Knowing which content categories perform best helps creators and brands focus on what matters most.
  • You will explore engagement metrics, identify top-performing niches, and discover patterns that drive growth.
  • By the end, you will be able to apply these techniques to your own content or business strategy.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
np.random.seed(42)

What Are Content Categories and Niches?#

  • In social media datasets, each row usually represents a single post or video.
  • Categories often indicate broad topics, such as Education, Gaming, or News.
  • Niches are more specific sub-topics within those categories.
  • Metrics like views, likes, comments, and watch time reflect how audiences engage with content.
  • Click-through rate (CTR) shows how often users click after seeing content.
  • Beginners sometimes misinterpret engagement: high views in one category can mean little if engagement rates are low.
  • Grouping by category lets us compare how different topics perform.
# Dataset 1: Load YouTube Trending Videos
try:
    import os, pickle
    from pathlib import Path
    from googleapiclient.discovery import build
    from google_auth_oauthlib.flow import InstalledAppFlow
    from google.auth.transport.requests import Request
    SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
    def get_yt_service():
        api_key = os.environ.get('YOUTUBE_API_KEY')
        if api_key:
            return build('youtube', 'v3', developerKey=api_key)
        if Path('client_secret.json').exists():
            creds = None
            if Path('token_ro.pickle').exists():
                with open('token_ro.pickle', 'rb') as f:
                    creds = pickle.load(f)
            if not creds or not creds.valid:
                if creds and creds.expired and creds.refresh_token:
                    creds.refresh(Request())
                else:
                    flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
                    creds = flow.run_local_server(port=0)
                with open('token_ro.pickle', 'wb') as f:
                    pickle.dump(creds, f)
            return build('youtube', 'v3', credentials=creds)
        raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
    def fetch_yt_trending(max_results=200, region='US'):
        youtube = get_yt_service()
        records, token = [], None
        while len(records) < max_results:
            resp = youtube.videos().list(
                part='snippet,statistics',
                chart='mostPopular',
                regionCode=region,
                maxResults=min(50, max_results - len(records)),
                pageToken=token
            ).execute()
            for item in resp.get('items', []):
                s = item['snippet']; st = item.get('statistics', {})
                records.append({
                    'video_id':      item['id'],
                    'trending_date': pd.Timestamp.today().date(),
                    'title':         s.get('title', ''),
                    'channel_title': s.get('channelTitle', ''),
                    'category_id':   s.get('categoryId', ''),
                    'views':         int(st.get('viewCount', 0)),
                    'likes':         int(st.get('likeCount', 0)),
                    'comment_count': int(st.get('commentCount', 0)),
                })
            token = resp.get('nextPageToken')
            if not token: break
        return pd.DataFrame(records)
    df = fetch_yt_trending()
    print('Live trending data:', df.shape)
except Exception as e:
    print(f'Falling back to synthetic: {e}')
    np.random.seed(42)
    n = 1000
    df = pd.DataFrame({
        'video_id':      [f'vid{i}' for i in range(n)],
        'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
        'title':         [f'Video Title {i}' for i in range(n)],
        'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
        'category_id':   np.random.choice([1,2,10,22,24,28], n),
        'views':         np.random.randint(10000, 5000000, n),
        'likes':         np.random.randint(100, 200000, n),
        'comment_count': np.random.randint(10, 50000, n),
    })
    print('Synthetic fallback:', df.shape)
print(df.head(3))
Live trending data: (199, 8)
      video_id trending_date  \
0  aHM2RKeMaeM    2026-06-02   
1  l0hD3KBwoiA    2026-06-02   
2  Vagb9BqdX8g    2026-06-02   

                                               title     channel_title  \
0                       MEOVV(미야오) - ‘DDI RO RI’ M/V     THEBLACKLABEL   
1  No Peace Amongst the Stars | Warhammer 40,000 ...  Warhammer 40,000   
2                 We Almost Lost EVERYTHING Gambling        SMii7Yplus   

  category_id    views   likes  comment_count  
0          10   804547       0           7454  
1          20  1699841  131643           6605  
2          20  1197396   60717           1639  
# Example: Show unique category IDs
category_counts = df['category_id'].value_counts()
print('Number of videos per category ID:')
print(category_counts)
Number of videos per category ID:
category_id
20    144
10     28
24     12
22      9
17      2
23      2
28      1
1       1
Name: count, dtype: int64

Example Starter: Basic Engagement Rate by Category#

  • A simple way to compare categories is to compute the average engagement rate.
  • Engagement rate can be likes divided by views multiplied by 100.
  • Let us group by category_id and look at the average engagement rate.
df['engagement_rate'] = df['likes'] / df['views'] * 100
cat_engagement = df.groupby('category_id')['engagement_rate'].mean().sort_values(ascending=False)
print('Avg Engagement Rate by Category:')
print(cat_engagement)
Avg Engagement Rate by Category:
category_id
28    5.981826
10    5.449388
22    4.916061
17    4.875619
1     4.759036
20    4.452056
23    4.190939
24    3.373937
Name: engagement_rate, dtype: float64

Beginner Example 1: Identify Top Categories by Likes#

  • Let us find which category gets the most total likes, not just on average.
  • Summing likes by category_id reveals which topics attract the most engagement overall.
likes_per_cat = df.groupby('category_id')['likes'].sum().sort_values(ascending=False)
print('Total Likes by Category:')
print(likes_per_cat.head())
Total Likes by Category:
category_id
20    1798752
10    1076623
24     356679
22     137672
1       23623
Name: likes, dtype: int64

Beginner Example 2: Most Commented Categories#

  • Engagement can also be measured by comments.
  • Let us check which category has the most total comments.
comments_per_cat = df.groupby('category_id')['comment_count'].sum().sort_values(ascending=False)
print('Total Comments by Category:')
print(comments_per_cat.head(3))
Total Comments by Category:
category_id
20    169847
10     70814
24     46110
Name: comment_count, dtype: int64

Beginner Example 3: Most Popular Videos by Category#

  • Sometimes, a single viral video drives a category's success.
  • Let us find, for each category, the video with the most views.
idx = df.groupby('category_id')['views'].idxmax()
top_videos = df.loc[idx, ['category_id', 'video_id', 'title', 'views']]
print('Most Viewed Video in Each Category:')
print(top_videos.sort_values('views', ascending=False).head())
Most Viewed Video in Each Category:
    category_id     video_id  \
15           24  0JlMjgqduVw   
45           10  v1t4MTqdfyI   
52           20  _hyFGrcVv5U   
72           22  r4auxPbAmP4   
139           1  Eola-v_VtGM   

                                                 title    views  
15   House of the Dragon Season 3 | Official Final ...  9388583  
45   Ariana Grande - hate that i made you love me (...  5106483  
52           I Went to WAR on a Hardcore Minecraft SMP  3536006  
72           SIDEMEN AMONG US: HARRY POTTER CHAOS MODE  2475783  
139           The First Bf 109 Shot Down by a Spitfire   496382  

Intermediate Example 1: Normalize Likes by Number of Videos#

  • Some categories might have many more videos than others.
  • To compare fairly, calculate the average likes per video in each category.
avg_likes = df.groupby('category_id')['likes'].mean().sort_values(ascending=False)
print('Avg Likes per Video by Category:')
print(avg_likes.head(4))
Avg Likes per Video by Category:
category_id
10    38450.821429
24    29723.250000
1     23623.000000
22    15296.888889
Name: likes, dtype: float64

Intermediate Example 2: Engagement Distribution in a Category#

  • We can look closer at one popular category's engagement distribution.
  • Let us pick the category with the most likes and plot the distribution.
import matplotlib.pyplot as plt
top_cat = likes_per_cat.index[0]
likes_top_cat = df[df['category_id'] == top_cat]['likes']
plt.hist(likes_top_cat, bins=30, color='skyblue', edgecolor='k')
plt.title(f'Likes Distribution for Category {top_cat}')
plt.xlabel('Likes')
plt.ylabel('Frequency')
plt.show()
No description has been provided for this image

Intermediate Example 3: Compare Engagement Across Top 3 Categories#

  • Let us compare average engagement rates side-by-side for the three most liked categories.
  • Visual comparisons help spot which niche is best for building an audience.
top3 = likes_per_cat.index[:3]
rates = cat_engagement[top3]
plt.bar([str(cat) for cat in top3], rates)
plt.title('Avg Engagement Rate of Top 3 Categories')
plt.xlabel('Category ID')
plt.ylabel('Avg Engagement Rate (%)')
plt.show()
No description has been provided for this image

Advanced Example 1: Detect Outlier Videos in a Niche#

  • Outliers can signal viral content or even errors.
  • Let us use Z-score to flag videos with unusually high engagement in one category.
cat_id = top3[0]
cat_df = df[df['category_id'] == cat_id]
mean_e = cat_df['engagement_rate'].mean()
std_e = cat_df['engagement_rate'].std()
cat_df['z'] = (cat_df['engagement_rate'] - mean_e) / std_e
outliers = cat_df[cat_df['z'] > 2]
print(f'Outlier videos (Z > 2) in category {cat_id}:')
print(outliers[['video_id', 'title', 'engagement_rate', 'z']].head())
Outlier videos (Z > 2) in category 20:
        video_id                                              title  \
56   wtJn6AVmk8g  The Phan Relationship is moving too fast - Tom...   
58   oRjkDgIPyvA  THIS HOME INVASION HORROR GAME WILL NEVER LET ...   
124  iAe2eZwEd7g                      Choose Your Fate 🔮| Episode 3   
135  -yuy9rymThs                      Rec Room AMA - the final AMA!   
164  mCimRXCWiZo  Can You Beat Genshin Impact Only Using Nahida??!!   

     engagement_rate         z  
56         15.496033  3.747979  
58         10.506441  2.054668  
124        11.284692  2.318782  
135        11.785996  2.488909  
164        11.469271  2.381422  

Advanced Example 2: Time Analysis for a Category#

  • Growth over time helps understand seasonality or trends in a niche.
  • Let us plot average engagement rate in a category over time.
cat_timeseries = cat_df.groupby('trending_date')['engagement_rate'].mean()
cat_timeseries.plot(figsize=(10,4))
plt.title(f'Engagement Rate Over Time in Category {cat_id}')
plt.ylabel('Avg Engagement Rate (%)')
plt.xlabel('Date')
plt.tight_layout()
plt.show()
No description has been provided for this image

Error Handling Example 1: Missing Likes or Views#

  • Data might have missing (NaN) values which break calculations.
  • Let us deliberately add some missing values and handle them.
# Add missing values
df_missing = df.copy()
df_missing.loc[::50, 'likes'] = np.nan
df_missing.loc[::77, 'views'] = np.nan
# Try to compute engagement rate, then fill missing
df_missing['engagement_rate'] = df_missing['likes'] / df_missing['views'] * 100
missing_count = df_missing['engagement_rate'].isna().sum()
print(f'Missing engagement_rate values: {missing_count}')
df_missing['engagement_rate'] = df_missing['engagement_rate'].fillna(0)
print('Handled missing engagement_rate. Any left?:', df_missing['engagement_rate'].isna().sum())
Missing engagement_rate values: 6
Handled missing engagement_rate. Any left?: 0

Error Handling Example 2: Incorrect Metric Aggregation#

  • Beginners sometimes sum engagement rates instead of averaging them.
  • Let us demonstrate why summing is wrong.
cat = top3[1]
rates_sum = df[df['category_id']==cat]['engagement_rate'].sum()
rates_mean = df[df['category_id']==cat]['engagement_rate'].mean()
print(f'Category {cat} summed engagement rate: {rates_sum:.1f}')
print(f'Category {cat} averaged engagement rate: {rates_mean:.2f}')
Category 10 summed engagement rate: 152.6
Category 10 averaged engagement rate: 5.45

Error Handling Example 3: Incorrect Grouping#

  • If you forget to group by category, your analysis will be misleading.
  • Let us see the difference between grouped and ungrouped results.
ungrouped_like_avg = df['likes'].mean()
grouped_like_avg = df.groupby('category_id')['likes'].mean().mean()
print(f'Average likes without grouping: {ungrouped_like_avg:.1f}')
print(f'Average of like averages by category: {grouped_like_avg:.1f}')
Average likes without grouping: 17265.9
Average of like averages by category: 18170.8

Best Practice 1: Always Define Categories Clearly#

  • Make sure category IDs or names are mapped and documented for stakeholders.
  • Use lookups or mapping dictionaries if available.

Best Practice 2: Benchmark Content Performance#

  • Benchmarking means comparing metrics against averages in the same niche.
  • This prevents overestimating performance due to a single outlier.

Best Practice 3: Regularly Review Content Mix#

  • Analyze the mix of categories and check if your content strategy is oversaturating or underserving a niche.
  • This helps with channel growth and retention.
# Tiny End-to-End Problem: Best Niche for New Channel
niche_stats = cat_engagement.to_frame(name='avg_engagement')
niche_stats['video_count'] = df['category_id'].value_counts()
niche_stats['avg_likes'] = df.groupby('category_id')['likes'].mean()
best_niches = niche_stats.query('video_count > 40').sort_values(by='avg_engagement', ascending=False)
print('Recommended niche(s) for high engagement and good video volume:')
print(best_niches.head(3))
Recommended niche(s) for high engagement and good video volume:
             avg_engagement  video_count     avg_likes
category_id                                           
20                 4.452056          144  12491.333333

Key Takeaways#

  • Analyzing categories and niches guides smarter content decisions.
  • Always compare both engagement rate and total engagement.
  • Outlier detection, error handling, and benchmarking help make your insights reliable.
  • Review your category mix regularly for optimal channel growth.
  • Try analyzing your own data next!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.