Lesson 1 · Social Media Content Analytics
Introduction to Social Media and Content Analytics
In this lesson, we will learn how to analyze real-world social media and content data using Python. Social media analytics help content creators and…
- CourseSocial Media Content Analytics
- Lesson1 of 41
- Video28 min
- FormatJupyter notebook · 22 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIntroduction to Social Media and Content Analytics#
- In this lesson, we will learn how to analyze real-world social media and content data using Python.
- Social media analytics help content creators and businesses understand what works, grow their audience, and optimize their strategy.
- We will use datasets from platforms like YouTube, Instagram, and TikTok to uncover content performance and engagement trends.
- You will see how to calculate engagement rates, find top-performing posts, and avoid common data interpretation mistakes.
- By the end, you will be able to generate actionable insights using real analytics methods.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
Core Social Media Analytics Concepts#
- Social media datasets contain information about content, such as posts or videos, their engagement, and audience reactions.
- Key metrics include views (how many watched), likes (positive feedback), comments (audience interaction), CTR (click-through rate), and watch time (how long people engage).
- Beginners sometimes misread engagement data by focusing only on total numbers and not on engagement relative to views, or by mixing up metrics from different categories of content.
- Correct analysis uses ratios and comparisons to understand why a piece of content performed well or poorly.
# Beginner Example 1: Load a synthetic social media content dataset
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df_content = pd.DataFrame({
'post_id': range(1, n_posts+1),
'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
'views': views,
'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df_content.head(3))
# Beginner Example 2: Calculate the overall engagement rate
df_content['engagement_rate'] = ((df_content['likes'] + df_content['comments'] + df_content['shares']) / df_content['views']) * 100
df_content['engagement_rate'] = df_content['engagement_rate'].round(2)
print(df_content[['post_id', 'platform', 'views', 'likes', 'comments', 'shares', 'engagement_rate']].head(5))
# Beginner Example 3: Find the post with the highest engagement rate
top_engaged = df_content.loc[df_content['engagement_rate'].idxmax()]
print(f"Post {top_engaged['post_id']} on {top_engaged['platform']} had the highest engagement rate: {top_engaged['engagement_rate']}%")
print(top_engaged[['post_id', 'platform', 'views', 'likes', 'comments', 'shares', 'engagement_rate']])
# Intermediate Example 1: Compare engagement rates between platforms
platform_group = df_content.groupby('platform')['engagement_rate'].mean().round(2)
print('Average engagement rate by platform:')
print(platform_group)
# Intermediate Example 2: Identify top 5 posts by total comments
top_comments = df_content.nlargest(5, 'comments')[['post_id', 'platform', 'views', 'comments', 'engagement_rate']]
print('Top 5 posts by comment count:')
print(top_comments)
# Intermediate Example 3: Calculate posting frequency over time
posts_per_week = df_content.set_index('date').resample('W').size()
print('Number of posts per week:')
print(posts_per_week.head())
# Intermediate Example 4: Load YouTube Trending Videos data (with fallback)
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request
SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
def get_yt_service():
api_key = os.environ.get('YOUTUBE_API_KEY')
if api_key:
return build('youtube', 'v3', developerKey=api_key)
if Path('client_secret.json').exists():
creds = None
if Path('token_ro.pickle').exists():
with open('token_ro.pickle', 'rb') as f:
creds = pickle.load(f)
if not creds or not creds.valid:
if creds and creds.expired and creds.refresh_token:
creds.refresh(Request())
else:
flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
creds = flow.run_local_server(port=0)
with open('token_ro.pickle', 'wb') as f:
pickle.dump(creds, f)
return build('youtube', 'v3', credentials=creds)
raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
def fetch_yt_trending(max_results=200, region='US'):
youtube = get_yt_service()
records, token = [], None
while len(records) < max_results:
resp = youtube.videos().list(
part='snippet,statistics',
chart='mostPopular',
regionCode=region,
maxResults=min(50, max_results - len(records)),
pageToken=token
).execute()
for item in resp.get('items', []):
s = item['snippet']; st = item.get('statistics', {})
records.append({
'video_id': item['id'],
'trending_date': pd.Timestamp.today().date(),
'title': s.get('title', ''),
'channel_title': s.get('channelTitle', ''),
'category_id': s.get('categoryId', ''),
'views': int(st.get('viewCount', 0)),
'likes': int(st.get('likeCount', 0)),
'comment_count': int(st.get('commentCount', 0)),
})
token = resp.get('nextPageToken')
if not token: break
return pd.DataFrame(records)
try:
df_trending = fetch_yt_trending()
print('Live trending data:', df_trending.shape)
except Exception as e:
print(f'Falling back to synthetic: {e}')
np.random.seed(42)
n = 1000
df_trending = pd.DataFrame({
'video_id': [f'vid{i}' for i in range(n)],
'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
'title': [f'Video Title {i}' for i in range(n)],
'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
'category_id': np.random.choice([1,2,10,22,24,28], n),
'views': np.random.randint(10000, 5000000, n),
'likes': np.random.randint(100, 200000, n),
'comment_count': np.random.randint(10, 50000, n),
})
print('Synthetic fallback:', df_trending.shape)
print(df_trending.head(3))
# Intermediate Example 5: Calculate like rate and comment rate on trending videos
df_trending['like_rate'] = (df_trending['likes'] / df_trending['views'] * 100).round(2)
df_trending['comment_rate'] = (df_trending['comment_count'] / df_trending['views'] * 100).round(2)
print(df_trending[['title', 'views', 'likes', 'comment_count', 'like_rate', 'comment_rate']].head(5))
# Intermediate Example 6: Find top 3 trending videos by like rate
top_like_rate = df_trending.nlargest(3, 'like_rate')[['title', 'channel_title', 'views', 'likes', 'like_rate']]
print('Top 3 trending videos by like rate:')
print(top_like_rate)
# Advanced Example 1: Analyze content performance across categories
cat_group = df_trending.groupby('category_id')['like_rate'].mean().round(2)
print('Average like rate by content category:')
print(cat_group)
# Advanced Example 2: Detect possible viral videos by engagement outliers
viral_threshold = df_trending['engagement_rate'] = ((df_trending['likes'] + df_trending['comment_count']) / df_trending['views']) * 100
viral_videos = df_trending[df_trending['engagement_rate'] > df_trending['engagement_rate'].quantile(0.995)]
print(f'There are {len(viral_videos)} potential viral videos detected as engagement outliers.')
print(viral_videos[['title', 'views', 'likes', 'comment_count', 'engagement_rate']].head())
# Advanced Example 3: Time-series analysis of social media engagement trends
n_days = 365
dates = pd.date_range('2023-01-01', periods=n_days, freq='D')
views_series = np.random.randint(1000, 50000, n_days)
likes_series = (views_series * np.random.uniform(0.03, 0.12, n_days)).astype(int)
df_timeseries = pd.DataFrame({
'date': dates,
'views': views_series,
'likes': likes_series,
'engagement_rate': np.round(likes_series / views_series * 100, 2)
})
trend = df_timeseries['engagement_rate'].rolling(window=30).mean()
print('30-day average engagement rate trend:')
print(trend.tail(10))
# Advanced Example 4: Segmenting posts by engagement quartile
quartiles = pd.qcut(df_content['engagement_rate'], 4, labels=['Low', 'Medium', 'High', 'Top'])
seg_counts = quartiles.value_counts().sort_index()
print('Distribution of posts by engagement quartile:')
print(seg_counts)
# Error Example 1: Handle missing likes or comments
df_content_missing = df_content.copy()
df_content_missing.loc[0, 'likes'] = np.nan # Force a missing value
missing_likes = df_content_missing['likes'].isna().sum()
print(f'There are {missing_likes} posts with missing like values.')
df_content_missing['likes'] = df_content_missing['likes'].fillna(0)
print('After fill, missing likes:', df_content_missing["likes"].isna().sum())
# Error Example 2: Misinterpreting ratios (avoid division by zero)
df_buggy = df_content.copy()
df_buggy.loc[1, 'views'] = 0 # Simulate a data error
try:
df_buggy['like_rate'] = (df_buggy['likes'] / df_buggy['views']) * 100
except Exception as e:
print('Error encountered:', e)
df_buggy['like_rate'] = np.where(df_buggy['views'] == 0, 0, (df_buggy['likes'] / df_buggy['views']) * 100)
print('Like rate with division by zero handled:', df_buggy[['post_id', 'views', 'likes', 'like_rate']].head(3))
# Error Example 3: Incorrect category grouping
wrong_group = df_trending.groupby('trending_date')['like_rate'].mean().head()
correct_group = df_trending.groupby('category_id')['like_rate'].mean().head()
print('Grouping by date (incorrect for category analysis):')
print(wrong_group)
print('\nGrouping by category (correct for category analysis):')
print(correct_group)
# Best Practice 1: Benchmark your content performance
mean_engagement = df_content['engagement_rate'].mean().round(2)
print(f'The average engagement rate across all content is {mean_engagement}%.')
if mean_engagement < 2:
print('Your engagement is below typical benchmarks. Look for improvement areas!')
elif mean_engagement > 8:
print('Congrats! Your content is performing above average.')
else:
print('You are near the average engagement rate. Aim for Top Quartile!')
# Best Practice 2: Consistent metric definitions and usage
engagement_cols = [c for c in df_content.columns if 'engagement_rate' in c or 'like_rate' in c]
print('Engagement-related columns:', engagement_cols)
# Best Practice 3: Segment audience by platform and analyze engagement
aud_seg = df_content.groupby('platform')['engagement_rate'].describe()[['mean','std','min','max']]
print('Audience engagement statistics by platform:')
print(aud_seg)
# End-to-End Analytics Problem: Recommend a content strategy
platform_perf = df_content.groupby('platform').agg({'views': 'mean', 'likes': 'mean', 'engagement_rate': 'mean'}).round(2)
best_platform = platform_perf['engagement_rate'].idxmax()
suggestion = f'Recommend focusing more content on {best_platform}, which currently has the highest engagement rate.'
print(platform_perf)
print(suggestion)
YouTube Call to Action#
- If you enjoyed these analytics walkthroughs, check out our channel for more step-by-step social data tutorials!
# Save engagement summary for further exploration
platform_perf.to_csv('platform_engagement_summary.csv', index=True)
print('Saved platform engagement summary as CSV.')
Congratulations! You have completed the introduction to social media and content analytics.#
- Practice daily on real content data. Track your growth and adjust strategies using what you learned.
- Subscribe for more hands-on analytics lessons.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



