Mathew K Analytics

Lesson 24 · Social Media Content Analytics

Visualizing Content Performance

In this lesson, we will learn how to analyze and visualize social media content performance using Python. Understanding content performance helps creators…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Visualizing Content Performance#

  • In this lesson, we will learn how to analyze and visualize social media content performance using Python.
  • Understanding content performance helps creators and businesses make strategic decisions.
  • We will explore metrics like views, likes, comments, and engagement rates across several platforms.
  • By the end, you will be able to identify top-performing posts and trends to optimize your content strategy.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')

What Is Social Media Content Performance?#

  • Social media datasets represent posts, videos, or campaigns with columns like views, likes, and comments.
  • Metrics such as views show how many people saw the content, while likes and comments measure engagement.
  • Click-through rate (CTR) shows what percent of viewers clicked a link or interacted further.
  • Watch time indicates how long people spent watching a video.
  • Beginners often mistake high views for high engagement it is important to compare multiple metrics together.
  • Misreading ratios or forgetting to consider the content type can lead to wrong decisions.
# Load the Social Media Content Dataset
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df = pd.DataFrame({
    'post_id': range(1, n_posts+1),
    'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
    'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
    'views': views,
    'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
    'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
    'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df.shape)
print(df.head(3))
(500, 7)
   post_id   platform                date  views  likes  comments  shares
0        1  Instagram 2023-01-01 00:00:00  15895   2356       116     253
1        2  Instagram 2023-01-01 06:00:00    960     94        16      26
2        3  Instagram 2023-01-01 12:00:00  76920   3910      2879    1185

Beginner Example 1: Plot Views Over Time#

  • A line plot is a simple way to see how total views changed over time.
  • This lets you spot content growth or identify peak periods of activity.
  • We can use groupby to sum daily views.
# Group by date and plot views
daily_views = df.groupby(df['date'].dt.date)['views'].sum()
plt.figure(figsize=(10,4))
plt.plot(daily_views.index, daily_views.values, marker='o')
plt.title('Total Views per Day')
plt.xlabel('Date')
plt.ylabel('Views')
plt.tight_layout()
plt.show()
No description has been provided for this image

Beginner Example 2: Bar Chart of Platform Post Counts#

  • Bar charts help compare how much content each platform has.
  • Knowing which platform you publish on the most can guide your strategy.
  • Let us see which platform is most active in your dataset.
# Count number of posts per platform
platform_counts = df['platform'].value_counts()
sns.barplot(x=platform_counts.index, y=platform_counts.values, palette='deep')
plt.title('Number of Posts per Platform')
plt.xlabel('Platform')
plt.ylabel('Number of Posts')
plt.show()
No description has been provided for this image

Beginner Example 3: Average Engagement Rate by Platform#

  • Engagement rate tells you how effectively your content is engaging viewers.
  • It is commonly calculated as (likes + comments + shares) divided by views, times 100.
  • We will calculate the average engagement rate per platform.
# Calculate engagement rate for each post and take platform averages
df['engagement_rate'] = (df['likes'] + df['comments'] + df['shares']) / df['views'] * 100
avg_engagement = df.groupby('platform')['engagement_rate'].mean()
avg_engagement.plot(kind='bar', color=['#c4302b','#e1306c','#69c9d0'])
plt.title('Average Engagement Rate by Platform (%)')
plt.ylabel('Engagement Rate (%)')
plt.xlabel('Platform')
plt.show()
No description has been provided for this image

Intermediate Example 1: Distribution of Views per Post#

  • Visualization of the distribution can show if your content often goes viral or mostly gets moderate attention.
  • A histogram or KDE plot reveals if you have a few outliers with very high views.
  • Let us plot the distribution of view counts.
# Histogram and KDE for view distributions
plt.figure(figsize=(8,4))
sns.histplot(df['views'], bins=30, kde=True, color='skyblue', edgecolor='white')
plt.title('Distribution of Views per Post')
plt.xlabel('Views')
plt.ylabel('Number of Posts')
plt.tight_layout()
plt.show()
No description has been provided for this image

Intermediate Example 2: Correlation Heatmap#

  • Correlation heatmaps show the relationships between numeric metrics in your content data.
  • For example, are likes and shares always linked, or does more watch time mean more comments?
  • Let us compute and plot a correlation heatmap.
# Compute correlations and plot heatmap
corr = df[['views','likes','comments','shares','engagement_rate']].corr()
plt.figure(figsize=(7,6))
sns.heatmap(corr, annot=True, cmap='YlGnBu', fmt='.2f')
plt.title('Correlation Between Content Metrics')
plt.show()
No description has been provided for this image

Intermediate Example 3: Top 10 Posts by Engagement Rate#

  • Finding top content helps focus effort on what performs best.
  • Sort posts by engagement rate, and show their key info.
  • This helps you analyze the qualities of your strongest content.
# Get top 10 posts by engagement rate
top_posts = df.sort_values('engagement_rate', ascending=False).head(10)
display_cols = ['post_id','platform','date','views','likes','comments','shares','engagement_rate']
print(top_posts[display_cols])
     post_id   platform                date  views  likes  comments  shares  \
351      352  Instagram 2023-03-29 18:00:00  74643  10653      3569    2036   
10        11  Instagram 2023-01-03 12:00:00  16123   2202       791     426   
383      384     TikTok 2023-04-06 18:00:00  36731   4965      1769     952   
423      424     TikTok 2023-04-16 18:00:00  53021   7572      2619     881   
87        88     TikTok 2023-01-22 18:00:00  82898  11426      3823    1910   
310      311     TikTok 2023-03-19 12:00:00  77605  11003      2885    1949   
429      430     TikTok 2023-04-18 06:00:00  56761   8091      1982    1333   
56        57  Instagram 2023-01-15 00:00:00  36020   5239      1237     762   
182      183  Instagram 2023-02-15 12:00:00  61473   8909      2251    1154   
335      336     TikTok 2023-03-25 18:00:00  47433   6605      1487    1363   

     engagement_rate  
351        21.781011  
10         21.205731  
383        20.925104  
423        20.882292  
87         20.698931  
310        20.407190  
429        20.094783  
56         20.094392  
182        20.031559  
335        19.933380  

Advanced Example 1: Content Performance Heatmap by Platform and Day#

  • Some platforms may perform better on certain days of the week.
  • A heatmap helps visualize posting patterns and content success by day and platform.
  • Let us analyze average engagement rates for each combination.
# Add weekday column and compute pivot table
df['weekday'] = df['date'].dt.day_name()
pivot = df.pivot_table(index='weekday', columns='platform', values='engagement_rate', aggfunc='mean')
weekday_order = ['Monday','Tuesday','Wednesday','Thursday','Friday','Saturday','Sunday']
pivot = pivot.reindex(weekday_order)
plt.figure(figsize=(7,5))
sns.heatmap(pivot, annot=True, cmap='coolwarm', fmt='.1f')
plt.title('Avg Engagement Rate by Platform and Weekday')
plt.xlabel('Platform')
plt.ylabel('Weekday')
plt.show()
No description has been provided for this image

Advanced Example 2: Analyzing Viewer Growth Trends#

  • A rolling average smooths out spikes and reveals the direction of audience growth.
  • Plot a rolling sum of views to see long-term patterns.
  • This helps when planning for seasonal shifts or campaign timing.
# Rolling 7-day views trend
daily_views = df.groupby(df['date'].dt.date)['views'].sum()
rolling_views = daily_views.rolling(window=7, min_periods=1).mean()
plt.figure(figsize=(10,4))
plt.plot(daily_views.index, daily_views, label='Daily Views', alpha=0.4)
plt.plot(rolling_views.index, rolling_views, color='red', label='7-Day Rolling Average')
plt.title('Daily Views and 7-Day Trend')
plt.xlabel('Date')
plt.ylabel('Views')
plt.legend()
plt.tight_layout()
plt.show()
No description has been provided for this image

Error Handling Example 1: Handling Missing Engagement Data#

  • Sometimes platforms fail to return all metrics values could be missing or blank.
  • If you calculate engagement with NA values, you will get errors or wrong results.
  • Let us simulate missing data and handle it nicely.
# Simulate missing likes and fill with zeros before engaging
df_missing = df.copy()
missing_idx = np.random.choice(df.index, size=20, replace=False)
df_missing.loc[missing_idx, 'likes'] = np.nan
df_missing['likes_filled'] = df_missing['likes'].fillna(0)
# Compute engagement rate safely with the filled values
df_missing['engagement_rate_fixed'] = (df_missing['likes_filled'] + df_missing['comments'] + df_missing['shares']) / df_missing['views'] * 100
print(df_missing[['likes','likes_filled','engagement_rate_fixed']].head(10))
    likes  likes_filled  engagement_rate_fixed
0  2356.0        2356.0              17.143756
1    94.0          94.0              14.166667
2  3910.0        3910.0              10.366615
3  1827.0        1827.0               7.187284
4   253.0         253.0              10.636292
5  4287.0        4287.0              10.078074
6  1524.0        1524.0               9.468011
7  3876.0        3876.0               4.993265
8     NaN           0.0               3.974588
9  2567.0        2567.0               8.578102

Error Handling Example 2: Incorrect Metric Aggregation#

  • It is easy to accidentally sum engagement rate percentages, which gives meaningless totals.
  • Always aggregate raw counts first, then compute engagement rates.
  • Let us show a wrong way and the correct way.
# Sum engagement rates incorrectly vs correctly
sum_wrong = df[df['platform']=='YouTube']['engagement_rate'].sum()
print(f"Wrong sum of engagement rates: {sum_wrong:.2f}")
sum_right = (df[df['platform']=='YouTube'][['likes','comments','shares']].sum().sum() / df[df['platform']=='YouTube']['views'].sum()) * 100
print(f"Correct overall engagement rate: {sum_right:.2f}%")
Wrong sum of engagement rates: 2250.16
Correct overall engagement rate: 12.36%

Error Handling Example 3: Misinterpreting Ratio Metrics#

  • Beginners often confuse raw counts with rates or ratios.
  • For example, having more views does not always mean higher CTR or engagement.
  • Always compare ratios within similar contexts.
# Compare the most viewed post to its engagement rate
max_view_idx = df['views'].idxmax()
most_viewed_post = df.loc[max_view_idx]
print(f"Post ID: {most_viewed_post['post_id']}")
print(f"Platform: {most_viewed_post['platform']}")
print(f"Views: {most_viewed_post['views']}")
print(f"Engagement Rate: {most_viewed_post['engagement_rate']:.2f}%")
Post ID: 393
Platform: TikTok
Views: 99813
Engagement Rate: 19.66%

Best Practices: Consistent Metric Definitions#

  • Always define metrics like "engagement rate" the same way each time.
  • Use raw counts to calculate rates after grouping or filtering.
  • Visualize your metrics together to spot trends, do not rely on one chart.
  • Segment by platform or posting time to reveal actionable differences.
  • Check for missing or outlier data before making recommendations.
  • Iterate your plots as your strategy evolves.
# Save a summary heatmap of engagement rates to file
fig, ax = plt.subplots(figsize=(6,4))
sns.heatmap(pivot, annot=True, cmap='coolwarm', fmt='.1f', ax=ax)
plt.title('Avg Engagement Rate by Platform and Weekday')
plt.xlabel('Platform')
plt.ylabel('Weekday')
fig.tight_layout()
fig.savefig('engagement_heatmap.png')
plt.close(fig)

Social Media Analytics Workflow: End-to-End Example#

  • Let us solve a simple real-world problem from start to finish.
  • You want to recommend a posting schedule and best platform for future content.
  • We will: 1) identify when and where content performs best, 2) summarize key stats, 3) suggest a data-driven improvement.
# Which platform and weekday combination has the highest avg engagement?
best = pivot.stack().idxmax()
best_value = pivot.stack().max()
print(f"Best time/platform: {best[1]} on {best[0]} with {best_value:.2f}% average engagement rate.")
Best time/platform: Instagram on Saturday with 13.97% average engagement rate.
# Save summary as text recommendation
adv_txt = f"Recommendation: For highest engagement, post on {best[1]} every {best[0]}. Max observed average engagement rate is {best_value:.2f}%."
with open('content_strategy.txt','w') as f:
    f.write(adv_txt)
print(adv_txt)
Recommendation: For highest engagement, post on Instagram every Saturday. Max observed average engagement rate is 13.97%.

Next Steps: Extend Your Analysis#

  • Try exploring more datasets, like YouTube Trending Videos or Analytics.
  • Integrate sentiment analysis from comments for deeper insights.
  • Compare your performance to industry benchmarks.
  • Refine your schedule, content mix, or format based on the visualized insights.
  • For full YouTube analysis workflow, search for "Data School social media" on YouTube.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.