Lesson 24 · Social Media Content Analytics
Visualizing Content Performance
In this lesson, we will learn how to analyze and visualize social media content performance using Python. Understanding content performance helps creators…
- CourseSocial Media Content Analytics
- Lesson24 of 41
- Video24 min
- FormatJupyter notebook · 16 code cells
What you'll learn
- What Is Social Media Content Performance?
- Beginner Example 1: Plot Views Over Time
- Beginner Example 2: Bar Chart of Platform Post Counts
- Beginner Example 3: Average Engagement Rate by Platform
- Intermediate Example 1: Distribution of Views per Post
- Intermediate Example 2: Correlation Heatmap
- Intermediate Example 3: Top 10 Posts by Engagement Rate
- Advanced Example 1: Content Performance Heatmap by Platform and Day
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbVisualizing Content Performance#
- In this lesson, we will learn how to analyze and visualize social media content performance using Python.
- Understanding content performance helps creators and businesses make strategic decisions.
- We will explore metrics like views, likes, comments, and engagement rates across several platforms.
- By the end, you will be able to identify top-performing posts and trends to optimize your content strategy.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')
What Is Social Media Content Performance?#
- Social media datasets represent posts, videos, or campaigns with columns like views, likes, and comments.
- Metrics such as views show how many people saw the content, while likes and comments measure engagement.
- Click-through rate (CTR) shows what percent of viewers clicked a link or interacted further.
- Watch time indicates how long people spent watching a video.
- Beginners often mistake high views for high engagement it is important to compare multiple metrics together.
- Misreading ratios or forgetting to consider the content type can lead to wrong decisions.
# Load the Social Media Content Dataset
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df = pd.DataFrame({
'post_id': range(1, n_posts+1),
'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
'views': views,
'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df.shape)
print(df.head(3))
Beginner Example 1: Plot Views Over Time#
- A line plot is a simple way to see how total views changed over time.
- This lets you spot content growth or identify peak periods of activity.
- We can use groupby to sum daily views.
# Group by date and plot views
daily_views = df.groupby(df['date'].dt.date)['views'].sum()
plt.figure(figsize=(10,4))
plt.plot(daily_views.index, daily_views.values, marker='o')
plt.title('Total Views per Day')
plt.xlabel('Date')
plt.ylabel('Views')
plt.tight_layout()
plt.show()
Beginner Example 2: Bar Chart of Platform Post Counts#
- Bar charts help compare how much content each platform has.
- Knowing which platform you publish on the most can guide your strategy.
- Let us see which platform is most active in your dataset.
# Count number of posts per platform
platform_counts = df['platform'].value_counts()
sns.barplot(x=platform_counts.index, y=platform_counts.values, palette='deep')
plt.title('Number of Posts per Platform')
plt.xlabel('Platform')
plt.ylabel('Number of Posts')
plt.show()
Beginner Example 3: Average Engagement Rate by Platform#
- Engagement rate tells you how effectively your content is engaging viewers.
- It is commonly calculated as (likes + comments + shares) divided by views, times 100.
- We will calculate the average engagement rate per platform.
# Calculate engagement rate for each post and take platform averages
df['engagement_rate'] = (df['likes'] + df['comments'] + df['shares']) / df['views'] * 100
avg_engagement = df.groupby('platform')['engagement_rate'].mean()
avg_engagement.plot(kind='bar', color=['#c4302b','#e1306c','#69c9d0'])
plt.title('Average Engagement Rate by Platform (%)')
plt.ylabel('Engagement Rate (%)')
plt.xlabel('Platform')
plt.show()
Intermediate Example 1: Distribution of Views per Post#
- Visualization of the distribution can show if your content often goes viral or mostly gets moderate attention.
- A histogram or KDE plot reveals if you have a few outliers with very high views.
- Let us plot the distribution of view counts.
# Histogram and KDE for view distributions
plt.figure(figsize=(8,4))
sns.histplot(df['views'], bins=30, kde=True, color='skyblue', edgecolor='white')
plt.title('Distribution of Views per Post')
plt.xlabel('Views')
plt.ylabel('Number of Posts')
plt.tight_layout()
plt.show()
Intermediate Example 2: Correlation Heatmap#
- Correlation heatmaps show the relationships between numeric metrics in your content data.
- For example, are likes and shares always linked, or does more watch time mean more comments?
- Let us compute and plot a correlation heatmap.
# Compute correlations and plot heatmap
corr = df[['views','likes','comments','shares','engagement_rate']].corr()
plt.figure(figsize=(7,6))
sns.heatmap(corr, annot=True, cmap='YlGnBu', fmt='.2f')
plt.title('Correlation Between Content Metrics')
plt.show()
Intermediate Example 3: Top 10 Posts by Engagement Rate#
- Finding top content helps focus effort on what performs best.
- Sort posts by engagement rate, and show their key info.
- This helps you analyze the qualities of your strongest content.
# Get top 10 posts by engagement rate
top_posts = df.sort_values('engagement_rate', ascending=False).head(10)
display_cols = ['post_id','platform','date','views','likes','comments','shares','engagement_rate']
print(top_posts[display_cols])
Advanced Example 1: Content Performance Heatmap by Platform and Day#
- Some platforms may perform better on certain days of the week.
- A heatmap helps visualize posting patterns and content success by day and platform.
- Let us analyze average engagement rates for each combination.
# Add weekday column and compute pivot table
df['weekday'] = df['date'].dt.day_name()
pivot = df.pivot_table(index='weekday', columns='platform', values='engagement_rate', aggfunc='mean')
weekday_order = ['Monday','Tuesday','Wednesday','Thursday','Friday','Saturday','Sunday']
pivot = pivot.reindex(weekday_order)
plt.figure(figsize=(7,5))
sns.heatmap(pivot, annot=True, cmap='coolwarm', fmt='.1f')
plt.title('Avg Engagement Rate by Platform and Weekday')
plt.xlabel('Platform')
plt.ylabel('Weekday')
plt.show()
Advanced Example 2: Analyzing Viewer Growth Trends#
- A rolling average smooths out spikes and reveals the direction of audience growth.
- Plot a rolling sum of views to see long-term patterns.
- This helps when planning for seasonal shifts or campaign timing.
# Rolling 7-day views trend
daily_views = df.groupby(df['date'].dt.date)['views'].sum()
rolling_views = daily_views.rolling(window=7, min_periods=1).mean()
plt.figure(figsize=(10,4))
plt.plot(daily_views.index, daily_views, label='Daily Views', alpha=0.4)
plt.plot(rolling_views.index, rolling_views, color='red', label='7-Day Rolling Average')
plt.title('Daily Views and 7-Day Trend')
plt.xlabel('Date')
plt.ylabel('Views')
plt.legend()
plt.tight_layout()
plt.show()
Error Handling Example 1: Handling Missing Engagement Data#
- Sometimes platforms fail to return all metrics values could be missing or blank.
- If you calculate engagement with NA values, you will get errors or wrong results.
- Let us simulate missing data and handle it nicely.
# Simulate missing likes and fill with zeros before engaging
df_missing = df.copy()
missing_idx = np.random.choice(df.index, size=20, replace=False)
df_missing.loc[missing_idx, 'likes'] = np.nan
df_missing['likes_filled'] = df_missing['likes'].fillna(0)
# Compute engagement rate safely with the filled values
df_missing['engagement_rate_fixed'] = (df_missing['likes_filled'] + df_missing['comments'] + df_missing['shares']) / df_missing['views'] * 100
print(df_missing[['likes','likes_filled','engagement_rate_fixed']].head(10))
Error Handling Example 2: Incorrect Metric Aggregation#
- It is easy to accidentally sum engagement rate percentages, which gives meaningless totals.
- Always aggregate raw counts first, then compute engagement rates.
- Let us show a wrong way and the correct way.
# Sum engagement rates incorrectly vs correctly
sum_wrong = df[df['platform']=='YouTube']['engagement_rate'].sum()
print(f"Wrong sum of engagement rates: {sum_wrong:.2f}")
sum_right = (df[df['platform']=='YouTube'][['likes','comments','shares']].sum().sum() / df[df['platform']=='YouTube']['views'].sum()) * 100
print(f"Correct overall engagement rate: {sum_right:.2f}%")
Error Handling Example 3: Misinterpreting Ratio Metrics#
- Beginners often confuse raw counts with rates or ratios.
- For example, having more views does not always mean higher CTR or engagement.
- Always compare ratios within similar contexts.
# Compare the most viewed post to its engagement rate
max_view_idx = df['views'].idxmax()
most_viewed_post = df.loc[max_view_idx]
print(f"Post ID: {most_viewed_post['post_id']}")
print(f"Platform: {most_viewed_post['platform']}")
print(f"Views: {most_viewed_post['views']}")
print(f"Engagement Rate: {most_viewed_post['engagement_rate']:.2f}%")
Best Practices: Consistent Metric Definitions#
- Always define metrics like "engagement rate" the same way each time.
- Use raw counts to calculate rates after grouping or filtering.
- Visualize your metrics together to spot trends, do not rely on one chart.
- Segment by platform or posting time to reveal actionable differences.
- Check for missing or outlier data before making recommendations.
- Iterate your plots as your strategy evolves.
# Save a summary heatmap of engagement rates to file
fig, ax = plt.subplots(figsize=(6,4))
sns.heatmap(pivot, annot=True, cmap='coolwarm', fmt='.1f', ax=ax)
plt.title('Avg Engagement Rate by Platform and Weekday')
plt.xlabel('Platform')
plt.ylabel('Weekday')
fig.tight_layout()
fig.savefig('engagement_heatmap.png')
plt.close(fig)
Social Media Analytics Workflow: End-to-End Example#
- Let us solve a simple real-world problem from start to finish.
- You want to recommend a posting schedule and best platform for future content.
- We will: 1) identify when and where content performs best, 2) summarize key stats, 3) suggest a data-driven improvement.
# Which platform and weekday combination has the highest avg engagement?
best = pivot.stack().idxmax()
best_value = pivot.stack().max()
print(f"Best time/platform: {best[1]} on {best[0]} with {best_value:.2f}% average engagement rate.")
# Save summary as text recommendation
adv_txt = f"Recommendation: For highest engagement, post on {best[1]} every {best[0]}. Max observed average engagement rate is {best_value:.2f}%."
with open('content_strategy.txt','w') as f:
f.write(adv_txt)
print(adv_txt)
Next Steps: Extend Your Analysis#
- Try exploring more datasets, like YouTube Trending Videos or Analytics.
- Integrate sentiment analysis from comments for deeper insights.
- Compare your performance to industry benchmarks.
- Refine your schedule, content mix, or format based on the visualized insights.
- For full YouTube analysis workflow, search for "Data School social media" on YouTube.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



