Lesson 54 · Social Media Content Analytics
Feature Importance for Engagement Prediction
We will learn how to predict engagement on social media content using Python. This lesson shows how to identify which features influence engagement the…
- CourseSocial Media Content Analytics
- Lesson54 of 41
- Video21 min
- FormatJupyter notebook · 15 code cells
What you'll learn
- Key Analytics Concepts Used in This Lesson
- Beginner Example 1: Calculating Engagement Rate
- Beginner Example 2: Sorting for Top Engagement
- Beginner Example 3: Visualizing Engagement Rate by Platform
- Intermediate Example 1: Correlation Matrix for Engagement Features
- Intermediate Example 2: Encoding Platform for Modeling
- Intermediate Example 3: Simple Engagement Prediction Model
- Advanced Example 1: Random Forest Feature Importance
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbFeature Importance for Engagement Prediction#
- We will learn how to predict engagement on social media content using Python.
- This lesson shows how to identify which features influence engagement the most.
- Content creators and businesses need to know what makes content successful.
- We will use hands-on code to explore, model, and explain real content data.
- You will see how feature importance techniques can guide content strategy.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.ensemble import RandomForestRegressor
from sklearn.tree import DecisionTreeRegressor
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error, r2_score
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')
Key Analytics Concepts Used in This Lesson#
- Social media datasets represent posts, videos, and user interactions.
- Engagement metrics include views, likes, comments, shares, watch time, and CTR.
- Views measure exposure, while likes and comments reflect user reactions.
- Click-through rate (CTR) shows how often users act on content.
- Watch time tracks attention span and engagement depth.
- Beginners may confuse cause and effect: high views can drive likes, but also depend on other factors.
- It is important not to assume correlation means causation in engagement analytics.
# Load a social media dataset simulating multichannel post engagement
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df = pd.DataFrame({
'post_id': range(1, n_posts+1),
'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
'views': views,
'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df.shape)
print(df.head(3))
Beginner Example 1: Calculating Engagement Rate#
- Engagement rate helps compare posts of very different sizes.
- It is most often defined as (likes + comments + shares) divided by views, times 100.
- Higher engagement rates suggest content is resonating with audiences.
df['engagement_rate'] = (df['likes'] + df['comments'] + df['shares']) / df['views'] * 100
print(df[['platform', 'views', 'likes', 'comments', 'shares', 'engagement_rate']].head())
Beginner Example 2: Sorting for Top Engagement#
- You can find which posts performed best for engagement rate.
- Sorting metrics helps surface which content is actually working.
top5 = df.sort_values('engagement_rate', ascending=False).head(5)
print(top5[['post_id', 'platform', 'views', 'engagement_rate']])
Beginner Example 3: Visualizing Engagement Rate by Platform#
- Social platform type can affect overall engagement performance.
- Visuals help spot patterns that would be missed in tables.
plt.figure(figsize=(8,5))
sns.boxplot(x='platform', y='engagement_rate', data=df)
plt.title('Engagement Rate Distribution by Platform')
plt.ylabel('Engagement Rate (%)')
plt.xlabel('Platform')
plt.show()
Intermediate Example 1: Correlation Matrix for Engagement Features#
- Correlation matrices show which features move together.
- This is a first step before building feature importance models.
- Some features may be highly correlated and not add extra value for prediction.
corr = df[['views', 'likes', 'comments', 'shares', 'engagement_rate']].corr()
plt.figure(figsize=(6,5))
sns.heatmap(corr, annot=True, cmap='coolwarm')
plt.title('Correlation Matrix of Engagement Metrics')
plt.show()
Intermediate Example 2: Encoding Platform for Modeling#
- Machine learning models need purely numeric features.
- We must convert text columns like 'platform' into numerical codes.
- Skipping this step is a common beginner error.
df['platform_code'] = df['platform'].map({'YouTube':0, 'Instagram':1, 'TikTok':2})
print(df[['platform', 'platform_code']].head())
Intermediate Example 3: Simple Engagement Prediction Model#
- We will predict engagement rate using a Decision Tree.
- This will let us see which features start to become important.
- Decision trees are easy to interpret for beginners.
features = ['views', 'likes', 'comments', 'shares', 'platform_code']
X = df[features]
y = df['engagement_rate']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
tree = DecisionTreeRegressor(random_state=42, max_depth=4)
tree.fit(X_train, y_train)
y_pred = tree.predict(X_test)
mse = mean_squared_error(y_test, y_pred)
r2 = r2_score(y_test, y_pred)
print('MSE:', round(mse, 2), '| R2:', round(r2, 3))
Advanced Example 1: Random Forest Feature Importance#
- Random Forest ensembles can better reveal patterns in real social media data.
- Feature importances show which factors have the biggest influence on engagement.
- This is key for building actionable strategies.
rf = RandomForestRegressor(random_state=42, n_estimators=100)
rf.fit(X_train, y_train)
importances = rf.feature_importances_
forest_importance = pd.Series(importances, index=features)
forest_importance.sort_values(ascending=True).plot(kind='barh', color='navy')
plt.title('Random Forest Feature Importance')
plt.xlabel('Importance Score')
plt.show()
Advanced Example 2: Partial Dependence Plot for Top Feature#
- Partial dependence plots show how engagement changes as one feature varies.
- This helps interpret whether more views, comments, or shares really improve engagement.
- We will plot the relationship between the most important feature and predicted engagement rate.
from sklearn.inspection import PartialDependenceDisplay
top_feature = forest_importance.idxmax()
PartialDependenceDisplay.from_estimator(rf, X_test, [top_feature],
feature_names=features,
grid_resolution=30)
plt.title(f'Partial Dependence: {top_feature} vs Engagement Rate')
plt.show()
Error Handling Example: Missing Engagement Values#
- Real data often has missing or zero engagement numbers.
- Failing to handle missing data can break models and give misleading results.
- Always check for missing or zero values before modeling.
df_missing = df.copy()
df_missing.loc[df_missing.sample(frac=0.05, random_state=42).index, 'comments'] = np.nan
print('Missing comments:', df_missing["comments"].isna().sum())
df_missing['comments'] = df_missing['comments'].fillna(0)
print('After fillna, missing comments:', df_missing['comments'].isna().sum())
Error Handling Example: Incorrect Aggregation#
- Aggregating engagement the wrong way can distort content rankings.
- Always double-check groupby columns and aggregations when benchmarking.
# Wrong: sums engagement across all posts, losing date info
agg_wrong = df.groupby('platform').sum(numeric_only=True)['engagement_rate']
print('Aggregation error example:\n', agg_wrong)
# Correct: calculates average engagement per post for each platform
agg_correct = df.groupby('platform')['engagement_rate'].mean()
print('Correct aggregation:\n', agg_correct)
Best Practices: Content Benchmarking and Trends#
- Benchmarking helps compare your posts to platform and industry averages.
- Audience segmentation splits results by platform, topic, or audience type.
- Tracking engagement trends helps spot growth, stagnation, or decline.
- Define metrics consistently so your results can be trusted.
trend = df.groupby(df['date'].dt.to_period('M'))['engagement_rate'].mean()
trend.plot(marker='o', figsize=(9,4))
plt.title('Monthly Engagement Rate Trend')
plt.ylabel('Avg Engagement Rate (%)')
plt.xlabel('Month')
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
End-to-End Example: Social Media Analytics to Actionable Strategy#
- Let us bring all these steps together: from raw data, to benchmark, to feature importance, to a content strategy tip.
- We will find which factor most influences engagement and suggest an improvement.
mean_engagement = df['engagement_rate'].mean()
topfactor = forest_importance.idxmax()
tip = f'Focus on improving {topfactor}: it has the biggest impact on engagement! Your average engagement rate is {mean_engagement:.2f}%.'
print(tip)
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



