Lesson 41 · Social Media Content Analytics
Introduction to Social Media APIs
In this lesson, we will explore how to access and analyze social media data using APIs. You will learn what kind of insights are possible with public and…
- CourseSocial Media Content Analytics
- Lesson41 of 41
- Video29 min
- FormatJupyter notebook · 23 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIntroduction to Social Media APIs#
- In this lesson, we will explore how to access and analyze social media data using APIs.
- You will learn what kind of insights are possible with public and programmatic content data.
- Understanding APIs enables content creators and businesses to make data-driven decisions.
- By the end of this lesson, you will be able to fetch trending content, measure engagement, and spot actionable trends.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request
What are Social Media APIs and Content Datasets?#
- APIs are software interfaces that let you request structured data from platforms like YouTube or Instagram.
- Using APIs, you can programmatically download data about videos, posts, comments, and user engagement.
- Social media datasets typically include metrics such as views, likes, comments, shares, and click-through rates.
- For example, YouTube offers live trending video lists and detailed analytics via its API.
- Mistakes can occur if you do not understand the meaning behind each metric or misuse data across platforms.
# BEGINNER EXAMPLE 1: Fetch public YouTube trending videos using the API if possible, falling back to synthetic data
SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
def get_yt_service():
api_key = os.environ.get('YOUTUBE_API_KEY')
if api_key:
return build('youtube', 'v3', developerKey=api_key)
if Path('client_secret.json').exists():
creds = None
if Path('token_ro.pickle').exists():
with open('token_ro.pickle', 'rb') as f:
creds = pickle.load(f)
if not creds or not creds.valid:
if creds and creds.expired and creds.refresh_token:
creds.refresh(Request())
else:
flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
creds = flow.run_local_server(port=0)
with open('token_ro.pickle', 'wb') as f:
pickle.dump(creds, f)
return build('youtube', 'v3', credentials=creds)
raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
def fetch_yt_trending(max_results=200, region='US'):
youtube = get_yt_service()
records, token = [], None
while len(records) < max_results:
resp = youtube.videos().list(
part='snippet,statistics',
chart='mostPopular',
regionCode=region,
maxResults=min(50, max_results - len(records)),
pageToken=token
).execute()
for item in resp.get('items', []):
s = item['snippet']; st = item.get('statistics', {})
records.append({
'video_id': item['id'],
'trending_date': pd.Timestamp.today().date(),
'title': s.get('title', ''),
'channel_title': s.get('channelTitle', ''),
'category_id': s.get('categoryId', ''),
'views': int(st.get('viewCount', 0)),
'likes': int(st.get('likeCount', 0)),
'comment_count': int(st.get('commentCount', 0)),
})
token = resp.get('nextPageToken')
if not token: break
return pd.DataFrame(records)
try:
df = fetch_yt_trending()
print('Live trending data:', df.shape)
except Exception as e:
print(f'Falling back to synthetic: {e}')
np.random.seed(42)
n = 1000
df = pd.DataFrame({
'video_id': [f'vid{i}' for i in range(n)],
'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
'title': [f'Video Title {i}' for i in range(n)],
'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
'category_id': np.random.choice([1,2,10,22,24,28], n),
'views': np.random.randint(10000, 5000000, n),
'likes': np.random.randint(100, 200000, n),
'comment_count': np.random.randint(10, 50000, n),
})
print('Synthetic fallback:', df.shape)
print(df.head(3))
# BEGINNER EXAMPLE 2: Load a synthetic cross-platform content dataset
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df_content = pd.DataFrame({
'post_id': range(1, n_posts+1),
'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
'views': views,
'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df_content.shape)
print(df_content.head(3))
# BEGINNER EXAMPLE 3: Calculate a simple engagement rate per post
df_content['engagement_rate'] = ((df_content['likes'] + df_content['comments'] + df_content['shares']) / df_content['views']) * 100
df_content['engagement_rate'] = df_content['engagement_rate'].round(2)
df_content[['platform','views','likes','comments','shares','engagement_rate']].head(5)
Core Social Media Metrics: Definitions and Pitfalls#
- Views show exposure, but do not guarantee active interest.
- Likes, comments, and shares measure different types of audience engagement.
- The engagement rate helps normalize post performance across audiences.
- Watch time and click-through rate (CTR) reveal audience attention and interaction quality.
- A common mistake is comparing raw metrics across platforms without normalizing.
# BEGINNER EXAMPLE 4: Identify the top 3 posts with highest engagement rate
top_posts = df_content.sort_values('engagement_rate', ascending=False).head(3)
print('Top 3 most engaging posts:')
print(top_posts[['platform','views','likes','comments','shares','engagement_rate']])
# BEGINNER EXAMPLE 5: Compare average engagement by platform
platform_avg_engagement = df_content.groupby('platform')['engagement_rate'].mean().round(2).sort_values(ascending=False)
print('Average engagement rate by platform:')
print(platform_avg_engagement)
# BEGINNER EXAMPLE 6: Visualize the distribution of engagement rates
import matplotlib.pyplot as plt
plt.figure(figsize=(8,5))
df_content['engagement_rate'].hist(bins=25, color='skyblue')
plt.title('Distribution of Engagement Rates (%)')
plt.xlabel('Engagement Rate (%)')
plt.ylabel('Number of Posts')
plt.grid(False)
plt.tight_layout()
plt.show()
# INTERMEDIATE EXAMPLE 1: Load and examine YouTube analytics data with watch time and CTR
np.random.seed(42)
n_videos = 300
views = np.random.randint(100, 500000, n_videos)
df_analytics = pd.DataFrame({
'video_id': range(1, n_videos+1),
'publish_date': pd.date_range('2022-01-01', periods=n_videos, freq='D'),
'views': views,
'watch_time': np.random.randint(1000, 500000, n_videos),
'likes': (views * np.random.uniform(0.01, 0.08, n_videos)).astype(int),
'comments': (views * np.random.uniform(0.001, 0.02, n_videos)).astype(int),
'ctr': np.round(np.random.uniform(2, 10, n_videos), 2)
})
print(df_analytics.shape)
print(df_analytics.head(3))
# INTERMEDIATE EXAMPLE 2: Find videos with highest click-through rate (CTR)
top_ctr = df_analytics.sort_values('ctr', ascending=False).head(5)
print('Top 5 videos by click-through rate:')
print(top_ctr[['video_id','publish_date','views','ctr']])
# INTERMEDIATE EXAMPLE 3: Compare watch time by upload month
df_analytics['month'] = df_analytics['publish_date'].dt.to_period('M')
monthly_watch = df_analytics.groupby('month')['watch_time'].sum()
monthly_watch.plot(kind='bar', color='seagreen', figsize=(10,4), title='Total Watch Time by Month')
plt.xlabel('Month')
plt.ylabel('Total Watch Time (minutes)')
plt.tight_layout()
plt.show()
# INTERMEDIATE EXAMPLE 4: Detect months with above-average CTR
mean_ctr = df_analytics['ctr'].mean()
good_months = df_analytics.groupby('month')['ctr'].mean() > mean_ctr
print('Months with average CTR above dataset mean:')
print(good_months[good_months].index.tolist())
# INTERMEDIATE EXAMPLE 5: Identify videos with more than twice average engagement
avg_likes = df_analytics['likes'].mean()
above2x = df_analytics[df_analytics['likes'] > 2 * avg_likes]
print(f'Videos with >2x average likes: {len(above2x)} out of {len(df_analytics)}')
# INTERMEDIATE EXAMPLE 6: Explore correlation between CTR and watch time
corr = df_analytics[['ctr','watch_time']].corr().iloc[0,1]
print(f'Correlation between CTR and watch time: {corr:.2f}')
# ADVANCED EXAMPLE 1: Extract and rank content categories from trending videos
top_cats = df['category_id'].value_counts().head(5)
print('Most common content categories among trending videos:')
print(top_cats)
# ADVANCED EXAMPLE 2: Time series trend of engagement rate for one year
np.random.seed(42)
n_days = 365
dates = pd.date_range('2023-01-01', periods=n_days, freq='D')
views = np.random.randint(1000, 50000, n_days)
likes = (views * np.random.uniform(0.03, 0.12, n_days)).astype(int)
df_ts = pd.DataFrame({
'date': dates,
'views': views,
'likes': likes,
'engagement_rate': np.round(likes / views * 100, 2)
})
plt.figure(figsize=(12,4))
plt.plot(df_ts['date'], df_ts['engagement_rate'], color='purple')
plt.title('Engagement Rate Trend Over One Year')
plt.ylabel('Engagement Rate (%)')
plt.xlabel('Date')
plt.tight_layout()
plt.show()
# ADVANCED EXAMPLE 3: Find best day-of-week to post for highest engagement rate
df_ts['dayofweek'] = df_ts['date'].dt.day_name()
dow_avg = df_ts.groupby('dayofweek')['engagement_rate'].mean().sort_values(ascending=False)
print('Average engagement rate by day of week:')
print(dow_avg)
# ERROR HANDLING EXAMPLE 1: Handling missing engagement values
df_content_missing = df_content.copy()
np.random.seed(42)
mask = np.random.rand(len(df_content_missing)) < 0.05
df_content_missing.loc[mask, 'likes'] = np.nan
missing_count = df_content_missing['likes'].isna().sum()
print(f'There are {missing_count} posts with missing like counts.')
# ERROR HANDLING EXAMPLE 2: Incorrect aggregation - median vs. mean engagement
mean_eng = df_content['engagement_rate'].mean().round(2)
median_eng = df_content['engagement_rate'].median().round(2)
print(f'Mean engagement rate: {mean_eng}%')
print(f'Median engagement rate: {median_eng}%')
# ERROR HANDLING EXAMPLE 3: Misinterpreted CTR as percent instead of ratio
df_analytics['ctr_ratio'] = df_analytics['ctr'] / 100
err_example = df_analytics[['ctr','ctr_ratio']].head(3)
print('Demonstrating correct conversion of CTR percent to decimal ratio:')
print(err_example)
# ERROR HANDLING EXAMPLE 4: Wrong grouping logic for viral content detection
top_channels = df.groupby('channel_title')['views'].sum().sort_values(ascending=False).head(3)
print('Total views by channel (sum across videos, not per video):')
print(top_channels)
Best Practices and Common Analytics Patterns#
- Benchmark performance using consistent metrics so your comparisons are fair.
- Segment your audience to find top fans and periodic trends.
- Use timelines to detect growth (or decline) in content engagement.
- Compare across platforms only after normalizing scales and audience size.
- Always check for missing values and unusual outliers before drawing conclusions.
# TINY END-TO-END EXAMPLE: Recommend a social media content strategy
high_eng_posts = df_content[df_content['engagement_rate'] > df_content['engagement_rate'].mean()]
most_pop_platform = high_eng_posts['platform'].mode()[0]
best_time = high_eng_posts['date'].dt.hour.mode()[0]
print('Content Strategy Recommendation:')
print(f'Focus new posts on {most_pop_platform} and schedule around {best_time}:00 for higher engagement.')
print('You have completed the API-driven Social Media Analytics lesson!')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



