Lesson 8 · Social Media Content Analytics
Numerical Analysis with NumPy for Engagement Data
Analyze social media engagement (likes, comments, views) using Python and NumPy. See how numerical methods help content creators and teams understand…
- CourseSocial Media Content Analytics
- Lesson8 of 41
- Video21 min
- FormatJupyter notebook · 25 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbNumerical Analysis with NumPy for Engagement Data#
- Analyze social media engagement (likes, comments, views) using Python and NumPy.
- See how numerical methods help content creators and teams understand trends.
- Learn to measure performance, find outliers, and optimize social content.
- Produce actionable insights to grow reach and engagement.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
Social Media Analytics Concepts#
- Social media datasets track videos or posts and their engagement statistics.
- Common metrics: views, likes, comments, shares, click-through rate, watch time.
- Higher views do not always mean better engagement. Percentages matter.
- NumPy makes it easy to calculate averages, ratios, growth, and identify outliers.
- Beginners often confuse absolute counts with rates, or group by the wrong column.
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df = pd.DataFrame({
'post_id': range(1, n_posts+1),
'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
'views': views,
'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df.shape)
print(df.head(3))
print('Average number of views:', np.mean(df['views']))
print('Maximum views on a single post:', np.max(df['views']))
print('Minimum views on a single post:', np.min(df['views']))
print('Standard deviation of likes:', np.std(df['likes']))
engagement_rate = df['likes'] / df['views'] * 100
print('Average engagement rate (%):', np.mean(engagement_rate))
platform_avg_likes = df.groupby('platform')['likes'].mean()
print(platform_avg_likes)
most_liked = df.iloc[np.argmax(df['likes'])]
print('Post with most likes:')
print(most_liked[['platform', 'likes', 'views', 'comments']])
high_engagement = df[engagement_rate > np.percentile(engagement_rate, 90)]
print('Number of posts with top 10% engagement rate:', high_engagement.shape[0])
posts_recent = df[df['date'] >= '2023-05-01']
recent_avg_likes = np.mean(posts_recent['likes'])
print('Average likes for posts since May 2023:', recent_avg_likes)
daily_likes = df.groupby(df['date'].dt.date)['likes'].sum()
print('Likes per day (first 5 days):')
print(daily_likes.head())
corr_val = np.corrcoef(df['shares'], df['comments'])[0,1]
print('Correlation between shares and comments:', round(corr_val, 3))
threshold = np.percentile(df['views'], 98)
viral_posts = df[df['views'] > threshold]
print('Viral posts (top 2% by views):', viral_posts.shape[0])
from scipy.stats import zscore
like_zscores = zscore(df['likes'])
outlier_like_posts = df[np.abs(like_zscores) > 3]
print('Posts with like count outliers:', outlier_like_posts.shape[0])
pivot = pd.pivot_table(df, values='likes', index='platform', columns=df['date'].dt.month, aggfunc=np.mean)
print('Monthly average likes per platform:')
print(pivot.head())
growth = daily_likes.pct_change().replace([np.inf, -np.inf], np.nan).fillna(0)
avg_growth = np.mean(growth[growth != 0]) * 100
print('Average daily like growth rate (%):', round(avg_growth, 2))
df['score'] = (df['likes']/np.max(df['likes']))*0.5 + (df['shares']/np.max(df['shares']))*0.3 + (df['comments']/np.max(df['comments']))*0.2
top_score_posts = df.sort_values('score', ascending=False).head(5)
print('Top 5 posts by engagement score:')
print(top_score_posts[['platform', 'views', 'likes', 'comments', 'shares', 'score']])
# Simulate missing likes data for 10 random posts
missing_idx = np.random.choice(df.index, 10, replace=False)
df.loc[missing_idx, 'likes'] = np.nan
print('Posts with missing likes:', df['likes'].isna().sum())
# Calculate engagement ignoring missing values
engagement_no_na = (df['likes'] / df['views'] * 100).dropna()
print('Mean engagement rate without missing:', np.mean(engagement_no_na))
# Incorrect aggregation: summing instead of averaging
wrong_summary = df.groupby('platform')['likes'].sum()
print('Wrong: Total likes per platform, not average')
print(wrong_summary)
# Misinterpreting ratios: wrong engagement formula
wrong_engagement = df['comments'] / df['likes']
print('First 3 (wrong) engagement ratios:', wrong_engagement.head(3).round(2))
# Wrong grouping: grouping by date and not platform
wrong_group = df.groupby('date')['likes'].mean().head()
print('First 5 averages grouped by date only:')
print(wrong_group)
Best Practices for Social Media Analytics#
- Use average (not total) engagement for fair platform comparisons.
- Segment your audience by platform, date, or content type to spot patterns.
- Track growth rates and seasonal trends using NumPy percentiles and difference calculations.
- Benchmark against your own average and industry norms, not just competitors.
- Always define your metric formulas and stick with them throughout.
- Optimize by focusing on top-scoring content and learning from past outliers.
# End-to-end: Identify which platform is best for engagement
platform_engage = df.groupby('platform').apply(lambda x: (x['likes']/x['views']*100).mean())
top_platform = platform_engage.idxmax()
print('Best platform for engagement rate:', top_platform)
print('Rates by platform:')
print(platform_engage.round(2))
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



