Lesson 51 · Social Media Content Analytics
Introduction to Machine Learning for Social Media
This lesson explores how machine learning helps social media creators and analysts understand and improve their content. You will learn to analyze real…
- CourseSocial Media Content Analytics
- Lesson51 of 41
- Video30 min
- FormatJupyter notebook · 22 code cells
What you'll learn
- Core Concepts: Social Media Analytics and Machine Learning
- Example 1: Calculate Engagement Rate for Trending Videos
- Example 2: Find Top-Performing Trending Videos by Engagement
- Example 3: Average Engagement by Content Category
- Example 4: Identify Most Discussed Trending Videos (by Comments)
- Example 5: Trends in Views Over Time
- Common Pitfall: Missing Engagement Values
- Error Example: Misinterpreting Ratios like CTR
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIntroduction to Machine Learning for Social Media#
This lesson explores how machine learning helps social media creators and analysts understand and improve their content.
You will learn to analyze real engagement metrics, detect trends, and make smarter content decisions using Python.
By the end, you will use real social media datasets to generate insights and build data-driven content strategies.
No prior machine learning experience needed, but familiarity with Python and Jupyter is recommended.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
Core Concepts: Social Media Analytics and Machine Learning#
- Social media datasets represent posts, videos, dates, and performance metrics like views, likes, and comments.
- Key metrics include:
- Views: How many times your content was seen.
- Likes: Positive engagements showing appreciation.
- Comments: Direct feedback from viewers.
- CTR (Click-Through Rate): The percentage of people who clicked your post or video after seeing it.
Watch Time: Total minutes watched, showing actual interest.
Common mistakes include:
- Only looking at total likes, not rates or context.
- Ignoring how different content types perform.
- Mixing up ratios like CTR and engagement rate.
- Ignoring how different content types perform.
- Only looking at total likes, not rates or context.
- CTR (Click-Through Rate): The percentage of people who clicked your post or video after seeing it.
- Comments: Direct feedback from viewers.
- Likes: Positive engagements showing appreciation.
- Views: How many times your content was seen.
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request
SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
def get_yt_service():
api_key = os.environ.get('YOUTUBE_API_KEY')
if api_key:
return build('youtube', 'v3', developerKey=api_key)
if Path('client_secret.json').exists():
creds = None
if Path('token_ro.pickle').exists():
with open('token_ro.pickle', 'rb') as f:
creds = pickle.load(f)
if not creds or not creds.valid:
if creds and creds.expired and creds.refresh_token:
creds.refresh(Request())
else:
flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
creds = flow.run_local_server(port=0)
with open('token_ro.pickle', 'wb') as f:
pickle.dump(creds, f)
return build('youtube', 'v3', credentials=creds)
raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
def fetch_yt_trending(max_results=200, region='US'):
youtube = get_yt_service()
records, token = [], None
while len(records) < max_results:
resp = youtube.videos().list(
part='snippet,statistics',
chart='mostPopular',
regionCode=region,
maxResults=min(50, max_results - len(records)),
pageToken=token
).execute()
for item in resp.get('items', []):
s = item['snippet']; st = item.get('statistics', {})
records.append({
'video_id': item['id'],
'trending_date': pd.Timestamp.today().date(),
'title': s.get('title', ''),
'channel_title': s.get('channelTitle', ''),
'category_id': s.get('categoryId', ''),
'views': int(st.get('viewCount', 0)),
'likes': int(st.get('likeCount', 0)),
'comment_count': int(st.get('commentCount', 0)),
})
token = resp.get('nextPageToken')
if not token: break
return pd.DataFrame(records)
try:
df = fetch_yt_trending()
print('Live trending data:', df.shape)
except Exception as e:
print(f'Falling back to synthetic: {e}')
np.random.seed(42)
n = 1000
df = pd.DataFrame({
'video_id': [f'vid{i}' for i in range(n)],
'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
'title': [f'Video Title {i}' for i in range(n)],
'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
'category_id': np.random.choice([1,2,10,22,24,28], n),
'views': np.random.randint(10000, 5000000, n),
'likes': np.random.randint(100, 200000, n),
'comment_count': np.random.randint(10, 50000, n),
})
print('Synthetic fallback:', df.shape)
print(df.head(3))
Example 1: Calculate Engagement Rate for Trending Videos#
- Engagement rate is a key metric that tells us how much viewers interact with content.
- It is defined as: (Likes + Comments) / Views.
- A higher engagement rate often means the content resonates more with the audience.
df['engagement_rate'] = (df['likes'] + df['comment_count']) / df['views']
print(df[['title', 'views', 'likes', 'comment_count', 'engagement_rate']].head())
Example 2: Find Top-Performing Trending Videos by Engagement#
- Sometimes a video with fewer total views still has the highest engagement rate.
- We want to know which titles generate the most loyal or interested audience.
top5_engaged = df.sort_values('engagement_rate', ascending=False).head(5)
print(top5_engaged[['title', 'views', 'likes', 'comment_count', 'engagement_rate']])
Example 3: Average Engagement by Content Category#
- Different categories (like Music, Gaming, or News) might show very different engagement patterns.
- Let us check which category receives the most interaction from viewers.
cat_group = df.groupby('category_id')['engagement_rate'].mean().sort_values(ascending=False)
print(cat_group)
Example 4: Identify Most Discussed Trending Videos (by Comments)#
- Engagement is not just likessometimes comments tell us which videos spark more discussion.
- Let us sort and see which trending videos have the most comments.
top_commented = df.sort_values('comment_count', ascending=False).head(5)
print(top_commented[['title', 'views', 'likes', 'comment_count']])
Example 5: Trends in Views Over Time#
- Sometimes, it is important to look for patterns in how video views change over time.
- Let us visualize views by trending date to spot any time-based trends.
import matplotlib.pyplot as plt
df_trend = df.groupby('trending_date')['views'].mean()
plt.figure(figsize=(10,4))
plt.plot(df_trend.index, df_trend.values, marker='o')
plt.title('Average Trending Video Views over Time')
plt.xlabel('Trending Date')
plt.ylabel('Average Views')
plt.tight_layout()
plt.show()
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df_posts = pd.DataFrame({
'post_id': range(1, n_posts+1),
'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
'views': views,
'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df_posts.head(3))
df_posts['engagement_rate'] = (df_posts['likes'] + df_posts['comments'] + df_posts['shares']) / df_posts['views']
print(df_posts[['platform', 'views', 'likes', 'comments', 'shares', 'engagement_rate']].head())
eng_by_platform = df_posts.groupby('platform')['engagement_rate'].mean()
print(eng_by_platform)
post_time_group = df_posts.groupby(df_posts['date'].dt.hour)['views'].mean()
plt.figure(figsize=(8,4))
plt.bar(post_time_group.index, post_time_group.values)
plt.title('Average Views per Posting Hour')
plt.xlabel('Hour of Day')
plt.ylabel('Average Views')
plt.xticks(range(0,24))
plt.tight_layout()
plt.show()
n_videos = 300
views = np.random.randint(100, 500000, n_videos)
df_yt = pd.DataFrame({
'video_id': range(1, n_videos+1),
'publish_date': pd.date_range('2022-01-01', periods=n_videos, freq='D'),
'views': views,
'watch_time': np.random.randint(1000, 500000, n_videos),
'likes': (views * np.random.uniform(0.01, 0.08, n_videos)).astype(int),
'comments': (views * np.random.uniform(0.001, 0.02, n_videos)).astype(int),
'ctr': np.round(np.random.uniform(2, 10, n_videos), 2)
})
print(df_yt.head(3))
df_yt['avg_watch_per_view'] = df_yt['watch_time'] / df_yt['views']
print(df_yt[['views', 'watch_time', 'avg_watch_per_view']].head())
top_ctr = df_yt.sort_values('ctr', ascending=False).head(5)
print(top_ctr[['video_id', 'views', 'ctr', 'likes', 'comments']])
# Flag potentially viral posts: high engagement & sharp view spikes
threshold = df_posts['engagement_rate'].quantile(0.95)
viral = df_posts[(df_posts['engagement_rate'] > threshold) & (df_posts['views'] > df_posts['views'].quantile(0.90))]
print(viral[['platform', 'views', 'engagement_rate']])
# Week-over-week growth for YouTube video views
yt_week = df_yt.set_index('publish_date')['views'].resample('W').sum()
yt_growth = yt_week.pct_change().fillna(0)
plt.figure(figsize=(8, 4))
plt.plot(yt_week.index, yt_growth.values, marker='o')
plt.title('Week-over-Week Growth in YouTube Video Views')
plt.xlabel('Week')
plt.ylabel('Growth Rate')
plt.axhline(0, color='red', linestyle='--')
plt.tight_layout()
plt.show()
# Simple machine learning: predict engagement rate based on views, likes, comments (linear regression)
from sklearn.linear_model import LinearRegression
X = df_posts[['views', 'likes', 'comments']]
y = df_posts['engagement_rate']
model = LinearRegression()
model.fit(X, y)
predicted = model.predict(X)
print('Sample actual vs predicted engagement rates:')
print(np.round(np.c_[y[:5], predicted[:5]], 4))
# Evaluate the prediction error using mean squared error
from sklearn.metrics import mean_squared_error
mse = mean_squared_error(y, predicted)
print(f'Mean Squared Error: {mse:.5f}')
Common Pitfall: Missing Engagement Values#
- Sometimes engagement data is missing or set to zero by accident.
- Let us see how to handle missing or outlier values safely.
# Simulate missing data: set some likes/comments/shares to np.nan
df_posts_missing = df_posts.copy()
mask = np.random.rand(len(df_posts_missing)) < 0.05
df_posts_missing.loc[mask, ['likes', 'comments', 'shares']] = np.nan
missing_summary = df_posts_missing[['likes', 'comments', 'shares']].isnull().sum()
print('Missing counts by column:')
print(missing_summary)
# Filling missing values before calculating engagement rate
df_filled = df_posts_missing.fillna(0)
df_filled['engagement_rate'] = (df_filled['likes'] + df_filled['comments'] + df_filled['shares']) / df_filled['views']
print('Engagement rates after filling missing values:')
print(df_filled['engagement_rate'].head())
Error Example: Misinterpreting Ratios like CTR#
- Mixing up ratios such as CTR (click-through rate) or engagement rate can lead to bad decisions.
- Always check which value is the numerator and which is the denominator!
# Incorrect CTR calculation example (wrong denominator!)
df_yt['ctr_wrong'] = df_yt['likes'] / df_yt['views']
print('First 3 wrong CTRs:', df_yt['ctr_wrong'].head(3).values)
# Correct CTR is already provided (percentage of impressions that resulted in a click)
Best Practices: Social Media Analytics and Machine Learning#
- Benchmark content performance within types (compare your videos to similar ones).
- Segment your audience and content (by topic, time, or platform).
- Measure growth and trends over time, not just top results.
- Use data to refine your posting schedule and content format.
- Always define metrics (engagement, CTR, watch time) clearly and consistently.
End-to-End Example: Data-Driven Content Strategy#
- Let us put all the pieces together in one workflow.
- We will find your highest-engagement platform and best post type, then make a content recommendation.
- Great analysts move from data > insight > recommendation.
# 1. Platform with highest average engagement
best_platform = eng_by_platform.idxmax()
best_rate = eng_by_platform.max()
# 2. On that platform, find the most engaging post (by engagement rate)
best_row = df_posts[df_posts['platform']==best_platform].sort_values('engagement_rate', ascending=False).iloc[0]
print(f'Best Platform: {best_platform}\nAverage Engagement Rate: {best_rate:.4f}')
print('Best Post Example:')
print(best_row[['date', 'views', 'likes', 'comments', 'shares', 'engagement_rate']])
YouTube Call To Action#
- Subscribe to the channel to learn more about data-driven social media analysis!
Quick Recap and Practice Prompt#
- In this lesson, you have learned to analyze social media content data using Python and basic machine learning.
- You have explored engagement patterns, posting strategy, and error traps.
- Your practice: Try applying these steps to a new dataset (YouTube Analytics or Instagram Insights) and write down three data-driven recommendations.
- For bonus points, plot a new metric or try different models!
- Like this walkthrough? Give it a thumbs-up!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



