Mathew K Analytics

Lesson 27 · Social Media Content Analytics

Analyzing Audience Behavior and Engagement Patterns

In this lesson, we will explore how to analyze social media audiences by examining their behavior and engagement with content. Understanding engagement…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Analyzing Audience Behavior and Engagement Patterns#

  • In this lesson, we will explore how to analyze social media audiences by examining their behavior and engagement with content.
  • Understanding engagement patterns helps creators and businesses learn what works, connect with viewers, and adapt their strategies for growth.
  • We will use real content datasets to calculate engagement rates, spot trends, and build recommendations for content strategies.
  • By the end, you will be able to spot top-performing content, detect when audiences are most active, and suggest improvements based on data.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Understanding Social Media Analytics Data#

  • Social media datasets can represent videos, posts, comment streams, or time series of views and likes.
  • Common metrics include views (how many saw the content), likes (positive feedback), comments (interactions), watch time (min total watched), and CTR (click-through rate).
  • It is easy to misinterpret raw engagement numbershigher views do not always mean better engagement.
  • Ratios like engagement rate and CTR help compare differently performing content fairly.
# Load sample multimodal social media analytics data (posts, engagement, time series)
np.random.seed(42)

# Content dataset: Social Media Content (multiple platforms, posts)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
sm_content_df = pd.DataFrame({
    'post_id': range(1, n_posts+1),
    'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
    'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
    'views': views,
    'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
    'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
    'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(sm_content_df.head(3))
   post_id   platform                date  views  likes  comments  shares
0        1  Instagram 2023-01-01 00:00:00  15895   2356       116     253
1        2  Instagram 2023-01-01 06:00:00    960     94        16      26
2        3  Instagram 2023-01-01 12:00:00  76920   3910      2879    1185
# Load and preview social_media_time_series (daily audience trends)
n_days = 365
dates = pd.date_range('2023-01-01', periods=n_days, freq='D')
views = np.random.randint(1000, 50000, n_days)
likes = (views * np.random.uniform(0.03, 0.12, n_days)).astype(int)
time_series_df = pd.DataFrame({
    'date': dates,
    'views': views,
    'likes': likes,
    'engagement_rate': np.round(likes / views * 100, 2)
})
print(time_series_df.head(3))
        date  views  likes  engagement_rate
0 2023-01-01  12439   1040             8.36
1 2023-01-02  42855   1676             3.91
2 2023-01-03  17021   1808            10.62

Beginner Example 1: Calculating Engagement Rate for Posts#

  • Engagement rate is likes plus comments plus shares, divided by views, then multiplied by 100 to get a percentage.
  • It shows how actively viewers interact with a post, relative to its reach.
  • A higher engagement rate means your audience is interacting more, even if views are not the highest.
# Calculate engagement rate column
sm_content_df['engagement_rate'] = (sm_content_df['likes'] + sm_content_df['comments'] + sm_content_df['shares']) / sm_content_df['views'] * 100
print(sm_content_df[['post_id', 'views', 'likes', 'comments', 'shares', 'engagement_rate']].head(5))
   post_id  views  likes  comments  shares  engagement_rate
0        1  15895   2356       116     253        17.143756
1        2    960     94        16      26        14.166667
2        3  76920   3910      2879    1185        10.366615
3        4  54986   1827       488    1637         7.187284
4        5   6365    253       261     163        10.636292

Beginner Example 2: Comparing Engagement Across Platforms#

  • It is not fair to compare posts from TikTok and YouTube only by total viewstheir audiences behave differently.
  • Instead, compare posts by average engagement rate per platform.
  • This helps discover where an audience is most active or loyal.
# Calculate average engagement rate by platform
platform_means = sm_content_df.groupby('platform')['engagement_rate'].mean().sort_values(ascending=False)
print(platform_means)
platform
Instagram    13.044316
TikTok       12.690745
YouTube      12.163010
Name: engagement_rate, dtype: float64

Beginner Example 3: Identifying Top-Performing Posts#

  • High total engagement is not always the same as high engagement rate.
  • Sorting by engagement rate highlights posts that inspire the most audience action, regardless of their reach.
  • Top-performing content can inform what to create more of.
# Find top 5 posts by engagement rate
top5_engaged = sm_content_df.sort_values('engagement_rate', ascending=False).head(5)
print(top5_engaged[['post_id', 'platform', 'views', 'engagement_rate']])
     post_id   platform  views  engagement_rate
351      352  Instagram  74643        21.781011
10        11  Instagram  16123        21.205731
383      384     TikTok  36731        20.925104
423      424     TikTok  53021        20.882292
87        88     TikTok  82898        20.698931

Intermediate Example 1: Audience Behavior Over Time#

  • Analyzing how engagement changes over months helps spot trendslike content that performs better in certain seasons.
  • These trends can guide when to post or what content themes to use.
# Calculate monthly average engagement rate
sm_content_df['month'] = sm_content_df['date'].dt.to_period('M')
monthly_trends = sm_content_df.groupby('month')['engagement_rate'].mean()
print(monthly_trends.head())
month
2023-01    12.434166
2023-02    12.373790
2023-03    12.416031
2023-04    12.917960
2023-05    14.378799
Freq: M, Name: engagement_rate, dtype: float64
# Plot engagement rate over time (line chart)
import matplotlib.pyplot as plt
monthly_trends.plot(title='Average Monthly Engagement Rate (%)', marker='o')
plt.ylabel('Engagement Rate (%)')
plt.xlabel('Month')
plt.tight_layout()
plt.show()
No description has been provided for this image

Intermediate Example 2: Detecting Audience Growth Trends#

  • Analyzing growth in total views and engagement over time can signal rising audience interest.
  • A healthy growth pattern suggests a growing and potentially loyal following.
# Calculate audience growth: monthly total views
monthly_views = sm_content_df.groupby('month')['views'].sum()
monthly_likes = sm_content_df.groupby('month')['likes'].sum()
plt.plot(monthly_views.index.astype(str), monthly_views.values, label='Total Views')
plt.plot(monthly_likes.index.astype(str), monthly_likes.values, label='Total Likes')
plt.title('Monthly Audience Growth')
plt.xlabel('Month')
plt.ylabel('Count')
plt.legend()
plt.tight_layout()
plt.show()
No description has been provided for this image

Intermediate Example 3: Spotting Viral Content Patterns#

  • Viral posts often have exceptionally high engagement rates compared to past content.
  • Detecting spikes in engagement helps you recognize what resonates unexpectedly well.
# Define viral threshold: engagement rate > 2 std deviations above mean
mean_rate = sm_content_df['engagement_rate'].mean()
std_rate = sm_content_df['engagement_rate'].std()
viral_threshold = mean_rate + 2 * std_rate
viral_posts = sm_content_df[sm_content_df['engagement_rate'] > viral_threshold]
print('Viral Posts Found:', len(viral_posts))
print(viral_posts[['post_id', 'platform', 'views', 'engagement_rate']])
Viral Posts Found: 4
     post_id   platform  views  engagement_rate
10        11  Instagram  16123        21.205731
351      352  Instagram  74643        21.781011
383      384     TikTok  36731        20.925104
423      424     TikTok  53021        20.882292

Advanced Example 1: Segmenting Audiences by Platform and Month#

  • Breaking down engagement by both platform and time reveals if certain platforms are growing or losing attention.
  • This dual segmentation is useful for advanced targeting or scheduling strategies.
# Platform and time-based audience segment analysis
segment_trends = sm_content_df.groupby(['platform', 'month'])['engagement_rate'].mean().unstack('platform')
segment_trends.plot(figsize=(10,5), title='Engagement Rate by Platform and Month', marker='o')
plt.ylabel('Engagement Rate (%)')
plt.xlabel('Month')
plt.tight_layout()
plt.show()
No description has been provided for this image

Advanced Example 2: Predicting Engagement Spikes with Rolling Averages#

  • Rolling averages smooth out short-term jumps to reveal when engagement is rising or declining.
  • This lets you spot real momentum, not just random spikes.
# Use 14-day rolling average for engagement rate
time_series_df['rolling_engagement'] = time_series_df['engagement_rate'].rolling(window=14, min_periods=1).mean()
plt.plot(time_series_df['date'], time_series_df['engagement_rate'], color='lightgray', label='Daily Engagement Rate')
plt.plot(time_series_df['date'], time_series_df['rolling_engagement'], color='red', label='14-day Rolling Avg')
plt.title('Engagement Rate Trend with Rolling Average')
plt.xlabel('Date')
plt.ylabel('Engagement Rate (%)')
plt.legend()
plt.tight_layout()
plt.show()
No description has been provided for this image

Error Handling Example 1: Dealing with Missing Engagement Data#

  • Sometimes, engagement data is missing due to platform errors or data collection gaps.
  • Detect missing data before analysis so your results are not distorted.
# Randomly introduce missing values for demonstration
sm_content_df.loc[sm_content_df.sample(frac=0.01).index, 'likes'] = np.nan
missing_counts = sm_content_df.isnull().sum()
print('Number of missing values in each column:')
print(missing_counts)
Number of missing values in each column:
post_id            0
platform           0
date               0
views              0
likes              5
comments           0
shares             0
engagement_rate    0
month              0
dtype: int64
# Fill or drop missing values to clean data for later calculations
sm_content_df['likes'] = sm_content_df['likes'].fillna(0)
print(sm_content_df['likes'].isnull().sum())
0

Error Handling Example 2: Detecting Incorrect Aggregation#

  • Summing engagement metrics without grouping can produce inflated and meaningless totals.
  • Always check if you need to group by platform, date, or post before aggregating.
# Bad practice: Aggregate engagement without grouping
bad_sum = sm_content_df['likes'].sum()
print('Sum of all likes (across all posts):', bad_sum)
Sum of all likes (across all posts): 2166980.0
# Correct way: Group by platform before aggregating
correct_sum = sm_content_df.groupby('platform')['likes'].sum()
print('Total likes by platform:')
print(correct_sum)
Total likes by platform:
platform
Instagram    752727.0
TikTok       678006.0
YouTube      736247.0
Name: likes, dtype: float64

Error Handling Example 3: Interpreting Engagement Ratios Consistently#

  • Engagement ratios like CTR or engagement rate should have clear, consistent definitions.
  • Always check formulas to avoid accidental division by zero or changing interpretations.
# Consistency check: Engagement rate when views are zero
test_df = sm_content_df.copy()
test_df.loc[0, 'views'] = 0   # Set up a zero-views example
try:
    test_df['test_engagement'] = (test_df['likes'] + test_df['comments'] + test_df['shares']) / test_df['views'] * 100
    print(test_df.loc[0, 'test_engagement'])
except ZeroDivisionError:
    print('Error: Division by zero occurred!')
inf
# Safe division: Replace zero views with np.nan before division
test_df['safe_engagement'] = np.where(test_df['views'] == 0, np.nan,
    (test_df['likes'] + test_df['comments'] + test_df['shares']) / test_df['views'] * 100)
print(test_df[['views', 'safe_engagement']].head(3))
   views  safe_engagement
0      0              NaN
1    960        14.166667
2  76920        10.366615

Best Practices in Audience Analytics#

  • Always benchmark content against a moving average, not just single-day spikes.
  • Segment audiences (platform, time, content type) to spot growth or risk areas.
  • Define engagement and viral thresholds before, not after, looking at your results.
  • Use safe calculations and handle missing data proactively.
  • Take time to review what worked for top contentsuccess is rarely random.

Common Analytics Patterns for Content Strategy#

  • Track engagement and growth over time, and compare to past periods.
  • Identify which audiences, platforms, or content types are most active.
  • Spot high- and low-performing content to adjust future strategies.
  • Use data-driven recommendations for best posting times and themes.
# Tiny End-to-End Problem: Recommend a Content Strategy
# Step 1: Identify platform with highest median engagement rate
platform_median = sm_content_df.groupby('platform')['engagement_rate'].median().sort_values(ascending=False)
best_platform = platform_median.index[0]
print('Highest-median platform:', best_platform)

# Step 2: Find the best posting month for that platform
best_month = sm_content_df[sm_content_df['platform'] == best_platform].groupby('month')['engagement_rate'].mean().idxmax()
print('Best month to post on', best_platform, 'is', best_month)

# Step 3: Find sample posts to study for future inspiration
sample_posts = sm_content_df[(sm_content_df['platform'] == best_platform) & (sm_content_df['month'] == best_month)].sort_values('engagement_rate', ascending=False).head(3)
print('Top example posts:')
print(sample_posts[['post_id', 'views', 'engagement_rate']])
Highest-median platform: TikTok
Best month to post on TikTok is 2023-04
Top example posts:
     post_id  views  engagement_rate
383      384  36731        20.925104
423      424  53021        20.882292
429      430  56761        20.094783

Lesson Complete!#

  • You have learned to analyze engagement and audience patterns with real social media analytics data.
  • Use these techniques to guide your next content decisionand keep practicing with your own numbers.
  • If you enjoyed this lesson, subscribe to our YouTube for more data-driven guides!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.