Lesson 34 · Social Media Content Analytics
Analyzing Content Categories and Niches
In this lesson, we will learn how to analyze content categories and niches using real social media data. Knowing which content categories perform best helps…
- CourseSocial Media Content Analytics
- Lesson34 of 41
- Video24 min
- FormatJupyter notebook · 17 code cells
What you'll learn
- What Are Content Categories and Niches?
- Example Starter: Basic Engagement Rate by Category
- Beginner Example 1: Identify Top Categories by Likes
- Beginner Example 2: Most Commented Categories
- Beginner Example 3: Most Popular Videos by Category
- Intermediate Example 1: Normalize Likes by Number of Videos
- Intermediate Example 2: Engagement Distribution in a Category
- Intermediate Example 3: Compare Engagement Across Top 3 Categories
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbAnalyzing Content Categories and Niches#
- In this lesson, we will learn how to analyze content categories and niches using real social media data.
- Knowing which content categories perform best helps creators and brands focus on what matters most.
- You will explore engagement metrics, identify top-performing niches, and discover patterns that drive growth.
- By the end, you will be able to apply these techniques to your own content or business strategy.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
np.random.seed(42)
What Are Content Categories and Niches?#
- In social media datasets, each row usually represents a single post or video.
- Categories often indicate broad topics, such as Education, Gaming, or News.
- Niches are more specific sub-topics within those categories.
- Metrics like views, likes, comments, and watch time reflect how audiences engage with content.
- Click-through rate (CTR) shows how often users click after seeing content.
- Beginners sometimes misinterpret engagement: high views in one category can mean little if engagement rates are low.
- Grouping by category lets us compare how different topics perform.
# Dataset 1: Load YouTube Trending Videos
try:
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request
SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
def get_yt_service():
api_key = os.environ.get('YOUTUBE_API_KEY')
if api_key:
return build('youtube', 'v3', developerKey=api_key)
if Path('client_secret.json').exists():
creds = None
if Path('token_ro.pickle').exists():
with open('token_ro.pickle', 'rb') as f:
creds = pickle.load(f)
if not creds or not creds.valid:
if creds and creds.expired and creds.refresh_token:
creds.refresh(Request())
else:
flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
creds = flow.run_local_server(port=0)
with open('token_ro.pickle', 'wb') as f:
pickle.dump(creds, f)
return build('youtube', 'v3', credentials=creds)
raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
def fetch_yt_trending(max_results=200, region='US'):
youtube = get_yt_service()
records, token = [], None
while len(records) < max_results:
resp = youtube.videos().list(
part='snippet,statistics',
chart='mostPopular',
regionCode=region,
maxResults=min(50, max_results - len(records)),
pageToken=token
).execute()
for item in resp.get('items', []):
s = item['snippet']; st = item.get('statistics', {})
records.append({
'video_id': item['id'],
'trending_date': pd.Timestamp.today().date(),
'title': s.get('title', ''),
'channel_title': s.get('channelTitle', ''),
'category_id': s.get('categoryId', ''),
'views': int(st.get('viewCount', 0)),
'likes': int(st.get('likeCount', 0)),
'comment_count': int(st.get('commentCount', 0)),
})
token = resp.get('nextPageToken')
if not token: break
return pd.DataFrame(records)
df = fetch_yt_trending()
print('Live trending data:', df.shape)
except Exception as e:
print(f'Falling back to synthetic: {e}')
np.random.seed(42)
n = 1000
df = pd.DataFrame({
'video_id': [f'vid{i}' for i in range(n)],
'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
'title': [f'Video Title {i}' for i in range(n)],
'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
'category_id': np.random.choice([1,2,10,22,24,28], n),
'views': np.random.randint(10000, 5000000, n),
'likes': np.random.randint(100, 200000, n),
'comment_count': np.random.randint(10, 50000, n),
})
print('Synthetic fallback:', df.shape)
print(df.head(3))
# Example: Show unique category IDs
category_counts = df['category_id'].value_counts()
print('Number of videos per category ID:')
print(category_counts)
Example Starter: Basic Engagement Rate by Category#
- A simple way to compare categories is to compute the average engagement rate.
- Engagement rate can be likes divided by views multiplied by 100.
- Let us group by category_id and look at the average engagement rate.
df['engagement_rate'] = df['likes'] / df['views'] * 100
cat_engagement = df.groupby('category_id')['engagement_rate'].mean().sort_values(ascending=False)
print('Avg Engagement Rate by Category:')
print(cat_engagement)
Beginner Example 1: Identify Top Categories by Likes#
- Let us find which category gets the most total likes, not just on average.
- Summing likes by category_id reveals which topics attract the most engagement overall.
likes_per_cat = df.groupby('category_id')['likes'].sum().sort_values(ascending=False)
print('Total Likes by Category:')
print(likes_per_cat.head())
Beginner Example 2: Most Commented Categories#
- Engagement can also be measured by comments.
- Let us check which category has the most total comments.
comments_per_cat = df.groupby('category_id')['comment_count'].sum().sort_values(ascending=False)
print('Total Comments by Category:')
print(comments_per_cat.head(3))
Beginner Example 3: Most Popular Videos by Category#
- Sometimes, a single viral video drives a category's success.
- Let us find, for each category, the video with the most views.
idx = df.groupby('category_id')['views'].idxmax()
top_videos = df.loc[idx, ['category_id', 'video_id', 'title', 'views']]
print('Most Viewed Video in Each Category:')
print(top_videos.sort_values('views', ascending=False).head())
Intermediate Example 1: Normalize Likes by Number of Videos#
- Some categories might have many more videos than others.
- To compare fairly, calculate the average likes per video in each category.
avg_likes = df.groupby('category_id')['likes'].mean().sort_values(ascending=False)
print('Avg Likes per Video by Category:')
print(avg_likes.head(4))
Intermediate Example 2: Engagement Distribution in a Category#
- We can look closer at one popular category's engagement distribution.
- Let us pick the category with the most likes and plot the distribution.
import matplotlib.pyplot as plt
top_cat = likes_per_cat.index[0]
likes_top_cat = df[df['category_id'] == top_cat]['likes']
plt.hist(likes_top_cat, bins=30, color='skyblue', edgecolor='k')
plt.title(f'Likes Distribution for Category {top_cat}')
plt.xlabel('Likes')
plt.ylabel('Frequency')
plt.show()
Intermediate Example 3: Compare Engagement Across Top 3 Categories#
- Let us compare average engagement rates side-by-side for the three most liked categories.
- Visual comparisons help spot which niche is best for building an audience.
top3 = likes_per_cat.index[:3]
rates = cat_engagement[top3]
plt.bar([str(cat) for cat in top3], rates)
plt.title('Avg Engagement Rate of Top 3 Categories')
plt.xlabel('Category ID')
plt.ylabel('Avg Engagement Rate (%)')
plt.show()
Advanced Example 1: Detect Outlier Videos in a Niche#
- Outliers can signal viral content or even errors.
- Let us use Z-score to flag videos with unusually high engagement in one category.
cat_id = top3[0]
cat_df = df[df['category_id'] == cat_id]
mean_e = cat_df['engagement_rate'].mean()
std_e = cat_df['engagement_rate'].std()
cat_df['z'] = (cat_df['engagement_rate'] - mean_e) / std_e
outliers = cat_df[cat_df['z'] > 2]
print(f'Outlier videos (Z > 2) in category {cat_id}:')
print(outliers[['video_id', 'title', 'engagement_rate', 'z']].head())
Advanced Example 2: Time Analysis for a Category#
- Growth over time helps understand seasonality or trends in a niche.
- Let us plot average engagement rate in a category over time.
cat_timeseries = cat_df.groupby('trending_date')['engagement_rate'].mean()
cat_timeseries.plot(figsize=(10,4))
plt.title(f'Engagement Rate Over Time in Category {cat_id}')
plt.ylabel('Avg Engagement Rate (%)')
plt.xlabel('Date')
plt.tight_layout()
plt.show()
Error Handling Example 1: Missing Likes or Views#
- Data might have missing (NaN) values which break calculations.
- Let us deliberately add some missing values and handle them.
# Add missing values
df_missing = df.copy()
df_missing.loc[::50, 'likes'] = np.nan
df_missing.loc[::77, 'views'] = np.nan
# Try to compute engagement rate, then fill missing
df_missing['engagement_rate'] = df_missing['likes'] / df_missing['views'] * 100
missing_count = df_missing['engagement_rate'].isna().sum()
print(f'Missing engagement_rate values: {missing_count}')
df_missing['engagement_rate'] = df_missing['engagement_rate'].fillna(0)
print('Handled missing engagement_rate. Any left?:', df_missing['engagement_rate'].isna().sum())
Error Handling Example 2: Incorrect Metric Aggregation#
- Beginners sometimes sum engagement rates instead of averaging them.
- Let us demonstrate why summing is wrong.
cat = top3[1]
rates_sum = df[df['category_id']==cat]['engagement_rate'].sum()
rates_mean = df[df['category_id']==cat]['engagement_rate'].mean()
print(f'Category {cat} summed engagement rate: {rates_sum:.1f}')
print(f'Category {cat} averaged engagement rate: {rates_mean:.2f}')
Error Handling Example 3: Incorrect Grouping#
- If you forget to group by category, your analysis will be misleading.
- Let us see the difference between grouped and ungrouped results.
ungrouped_like_avg = df['likes'].mean()
grouped_like_avg = df.groupby('category_id')['likes'].mean().mean()
print(f'Average likes without grouping: {ungrouped_like_avg:.1f}')
print(f'Average of like averages by category: {grouped_like_avg:.1f}')
Best Practice 1: Always Define Categories Clearly#
- Make sure category IDs or names are mapped and documented for stakeholders.
- Use lookups or mapping dictionaries if available.
Best Practice 2: Benchmark Content Performance#
- Benchmarking means comparing metrics against averages in the same niche.
- This prevents overestimating performance due to a single outlier.
Best Practice 3: Regularly Review Content Mix#
- Analyze the mix of categories and check if your content strategy is oversaturating or underserving a niche.
- This helps with channel growth and retention.
# Tiny End-to-End Problem: Best Niche for New Channel
niche_stats = cat_engagement.to_frame(name='avg_engagement')
niche_stats['video_count'] = df['category_id'].value_counts()
niche_stats['avg_likes'] = df.groupby('category_id')['likes'].mean()
best_niches = niche_stats.query('video_count > 40').sort_values(by='avg_engagement', ascending=False)
print('Recommended niche(s) for high engagement and good video volume:')
print(best_niches.head(3))
Key Takeaways#
- Analyzing categories and niches guides smarter content decisions.
- Always compare both engagement rate and total engagement.
- Outlier detection, error handling, and benchmarking help make your insights reliable.
- Review your category mix regularly for optimal channel growth.
- Try analyzing your own data next!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



