Mathew K Analytics

Lesson 8 · Social Media Content Analytics

Numerical Analysis with NumPy for Engagement Data

Analyze social media engagement (likes, comments, views) using Python and NumPy. See how numerical methods help content creators and teams understand…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Numerical Analysis with NumPy for Engagement Data#

  • Analyze social media engagement (likes, comments, views) using Python and NumPy.
  • See how numerical methods help content creators and teams understand trends.
  • Learn to measure performance, find outliers, and optimize social content.
  • Produce actionable insights to grow reach and engagement.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Social Media Analytics Concepts#

  • Social media datasets track videos or posts and their engagement statistics.
  • Common metrics: views, likes, comments, shares, click-through rate, watch time.
  • Higher views do not always mean better engagement. Percentages matter.
  • NumPy makes it easy to calculate averages, ratios, growth, and identify outliers.
  • Beginners often confuse absolute counts with rates, or group by the wrong column.
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df = pd.DataFrame({
    'post_id': range(1, n_posts+1),
    'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
    'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
    'views': views,
    'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
    'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
    'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df.shape)
print(df.head(3))
(500, 7)
   post_id   platform                date  views  likes  comments  shares
0        1  Instagram 2023-01-01 00:00:00  15895   2356       116     253
1        2  Instagram 2023-01-01 06:00:00    960     94        16      26
2        3  Instagram 2023-01-01 12:00:00  76920   3910      2879    1185
print('Average number of views:', np.mean(df['views']))
Average number of views: 50730.112
print('Maximum views on a single post:', np.max(df['views']))
Maximum views on a single post: 99813
print('Minimum views on a single post:', np.min(df['views']))
Minimum views on a single post: 306
print('Standard deviation of likes:', np.std(df['likes']))
Standard deviation of likes: 3379.3268679688267
engagement_rate = df['likes'] / df['views'] * 100
print('Average engagement rate (%):', np.mean(engagement_rate))
Average engagement rate (%): 8.51350662452298
platform_avg_likes = df.groupby('platform')['likes'].mean()
print(platform_avg_likes)
platform
Instagram    4702.827160
TikTok       4491.143791
YouTube      4020.610811
Name: likes, dtype: float64
most_liked = df.iloc[np.argmax(df['likes'])]
print('Post with most likes:')
print(most_liked[['platform', 'likes', 'views', 'comments']])
Post with most likes:
platform    TikTok
likes        13839
views        98906
comments      3522
Name: 138, dtype: object
high_engagement = df[engagement_rate > np.percentile(engagement_rate, 90)]
print('Number of posts with top 10% engagement rate:', high_engagement.shape[0])
Number of posts with top 10% engagement rate: 50
posts_recent = df[df['date'] >= '2023-05-01']
recent_avg_likes = np.mean(posts_recent['likes'])
print('Average likes for posts since May 2023:', recent_avg_likes)
Average likes for posts since May 2023: 5323.6
daily_likes = df.groupby(df['date'].dt.date)['likes'].sum()
print('Likes per day (first 5 days):')
print(daily_likes.head())
Likes per day (first 5 days):
date
2023-01-01     8187
2023-01-02     9940
2023-01-03     8545
2023-01-04    12893
2023-01-05    23719
Name: likes, dtype: int64
corr_val = np.corrcoef(df['shares'], df['comments'])[0,1]
print('Correlation between shares and comments:', round(corr_val, 3))
Correlation between shares and comments: 0.491
threshold = np.percentile(df['views'], 98)
viral_posts = df[df['views'] > threshold]
print('Viral posts (top 2% by views):', viral_posts.shape[0])
Viral posts (top 2% by views): 10
from scipy.stats import zscore
like_zscores = zscore(df['likes'])
outlier_like_posts = df[np.abs(like_zscores) > 3]
print('Posts with like count outliers:', outlier_like_posts.shape[0])
Posts with like count outliers: 0
pivot = pd.pivot_table(df, values='likes', index='platform', columns=df['date'].dt.month, aggfunc=np.mean)
print('Monthly average likes per platform:')
print(pivot.head())
Monthly average likes per platform:
date                 1            2            3            4            5
platform                                                                  
Instagram  4227.416667  5853.289474  4257.682927  4391.538462  5176.375000
TikTok     4193.916667  3741.393939  4463.333333  5414.658537  4288.142857
YouTube    4091.557692  4209.073171  3785.404255  3638.050000  7008.800000
growth = daily_likes.pct_change().replace([np.inf, -np.inf], np.nan).fillna(0)
avg_growth = np.mean(growth[growth != 0]) * 100
print('Average daily like growth rate (%):', round(avg_growth, 2))
Average daily like growth rate (%): 14.51
df['score'] = (df['likes']/np.max(df['likes']))*0.5 + (df['shares']/np.max(df['shares']))*0.3 + (df['comments']/np.max(df['comments']))*0.2
top_score_posts = df.sort_values('score', ascending=False).head(5)
print('Top 5 posts by engagement score:')
print(top_score_posts[['platform', 'views', 'likes', 'comments', 'shares', 'score']])
Top 5 posts by engagement score:
      platform  views  likes  comments  shares     score
392     TikTok  99813  12361      4424    2837  0.939367
192  Instagram  97604  13347      3566    2493  0.901229
276  Instagram  96701  13427      3435    2410  0.889634
74   Instagram  99399  13323      1894    2655  0.844639
422    YouTube  93948  12615      2380    2571  0.831353
# Simulate missing likes data for 10 random posts
missing_idx = np.random.choice(df.index, 10, replace=False)
df.loc[missing_idx, 'likes'] = np.nan
print('Posts with missing likes:', df['likes'].isna().sum())
Posts with missing likes: 10
# Calculate engagement ignoring missing values
engagement_no_na = (df['likes'] / df['views'] * 100).dropna()
print('Mean engagement rate without missing:', np.mean(engagement_no_na))
Mean engagement rate without missing: 8.512296715263487
# Incorrect aggregation: summing instead of averaging
wrong_summary = df.groupby('platform')['likes'].sum()
print('Wrong: Total likes per platform, not average')
print(wrong_summary)
Wrong: Total likes per platform, not average
platform
Instagram    746391.0
TikTok       674158.0
YouTube      736259.0
Name: likes, dtype: float64
# Misinterpreting ratios: wrong engagement formula
wrong_engagement = df['comments'] / df['likes']
print('First 3 (wrong) engagement ratios:', wrong_engagement.head(3).round(2))
First 3 (wrong) engagement ratios: 0    0.05
1    0.17
2    0.74
dtype: float64
# Wrong grouping: grouping by date and not platform
wrong_group = df.groupby('date')['likes'].mean().head()
print('First 5 averages grouped by date only:')
print(wrong_group)
First 5 averages grouped by date only:
date
2023-01-01 00:00:00    2356.0
2023-01-01 06:00:00      94.0
2023-01-01 12:00:00    3910.0
2023-01-01 18:00:00    1827.0
2023-01-02 00:00:00     253.0
Name: likes, dtype: float64

Best Practices for Social Media Analytics#

  • Use average (not total) engagement for fair platform comparisons.
  • Segment your audience by platform, date, or content type to spot patterns.
  • Track growth rates and seasonal trends using NumPy percentiles and difference calculations.
  • Benchmark against your own average and industry norms, not just competitors.
  • Always define your metric formulas and stick with them throughout.
  • Optimize by focusing on top-scoring content and learning from past outliers.
# End-to-end: Identify which platform is best for engagement
platform_engage = df.groupby('platform').apply(lambda x: (x['likes']/x['views']*100).mean())
top_platform = platform_engage.idxmax()
print('Best platform for engagement rate:', top_platform)
print('Rates by platform:')
print(platform_engage.round(2))
Best platform for engagement rate: Instagram
Rates by platform:
platform
Instagram    9.09
TikTok       8.58
YouTube      7.97
dtype: float64
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.