Mathew K Analytics

Lesson 54 · Social Media Content Analytics

Feature Importance for Engagement Prediction

We will learn how to predict engagement on social media content using Python. This lesson shows how to identify which features influence engagement the…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Feature Importance for Engagement Prediction#

  • We will learn how to predict engagement on social media content using Python.
  • This lesson shows how to identify which features influence engagement the most.
  • Content creators and businesses need to know what makes content successful.
  • We will use hands-on code to explore, model, and explain real content data.
  • You will see how feature importance techniques can guide content strategy.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.ensemble import RandomForestRegressor
from sklearn.tree import DecisionTreeRegressor
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error, r2_score
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')

Key Analytics Concepts Used in This Lesson#

  • Social media datasets represent posts, videos, and user interactions.
  • Engagement metrics include views, likes, comments, shares, watch time, and CTR.
  • Views measure exposure, while likes and comments reflect user reactions.
  • Click-through rate (CTR) shows how often users act on content.
  • Watch time tracks attention span and engagement depth.
  • Beginners may confuse cause and effect: high views can drive likes, but also depend on other factors.
  • It is important not to assume correlation means causation in engagement analytics.
# Load a social media dataset simulating multichannel post engagement
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df = pd.DataFrame({
    'post_id': range(1, n_posts+1),
    'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
    'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
    'views': views,
    'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
    'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
    'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df.shape)
print(df.head(3))
(500, 7)
   post_id   platform                date  views  likes  comments  shares
0        1  Instagram 2023-01-01 00:00:00  15895   2356       116     253
1        2  Instagram 2023-01-01 06:00:00    960     94        16      26
2        3  Instagram 2023-01-01 12:00:00  76920   3910      2879    1185

Beginner Example 1: Calculating Engagement Rate#

  • Engagement rate helps compare posts of very different sizes.
  • It is most often defined as (likes + comments + shares) divided by views, times 100.
  • Higher engagement rates suggest content is resonating with audiences.
df['engagement_rate'] = (df['likes'] + df['comments'] + df['shares']) / df['views'] * 100
print(df[['platform', 'views', 'likes', 'comments', 'shares', 'engagement_rate']].head())
    platform  views  likes  comments  shares  engagement_rate
0  Instagram  15895   2356       116     253        17.143756
1  Instagram    960     94        16      26        14.166667
2  Instagram  76920   3910      2879    1185        10.366615
3    YouTube  54986   1827       488    1637         7.187284
4  Instagram   6365    253       261     163        10.636292

Beginner Example 2: Sorting for Top Engagement#

  • You can find which posts performed best for engagement rate.
  • Sorting metrics helps surface which content is actually working.
top5 = df.sort_values('engagement_rate', ascending=False).head(5)
print(top5[['post_id', 'platform', 'views', 'engagement_rate']])
     post_id   platform  views  engagement_rate
351      352  Instagram  74643        21.781011
10        11  Instagram  16123        21.205731
383      384     TikTok  36731        20.925104
423      424     TikTok  53021        20.882292
87        88     TikTok  82898        20.698931

Beginner Example 3: Visualizing Engagement Rate by Platform#

  • Social platform type can affect overall engagement performance.
  • Visuals help spot patterns that would be missed in tables.
plt.figure(figsize=(8,5))
sns.boxplot(x='platform', y='engagement_rate', data=df)
plt.title('Engagement Rate Distribution by Platform')
plt.ylabel('Engagement Rate (%)')
plt.xlabel('Platform')
plt.show()
No description has been provided for this image

Intermediate Example 1: Correlation Matrix for Engagement Features#

  • Correlation matrices show which features move together.
  • This is a first step before building feature importance models.
  • Some features may be highly correlated and not add extra value for prediction.
corr = df[['views', 'likes', 'comments', 'shares', 'engagement_rate']].corr()
plt.figure(figsize=(6,5))
sns.heatmap(corr, annot=True, cmap='coolwarm')
plt.title('Correlation Matrix of Engagement Metrics')
plt.show()
No description has been provided for this image

Intermediate Example 2: Encoding Platform for Modeling#

  • Machine learning models need purely numeric features.
  • We must convert text columns like 'platform' into numerical codes.
  • Skipping this step is a common beginner error.
df['platform_code'] = df['platform'].map({'YouTube':0, 'Instagram':1, 'TikTok':2})
print(df[['platform', 'platform_code']].head())
    platform  platform_code
0  Instagram              1
1  Instagram              1
2  Instagram              1
3    YouTube              0
4  Instagram              1

Intermediate Example 3: Simple Engagement Prediction Model#

  • We will predict engagement rate using a Decision Tree.
  • This will let us see which features start to become important.
  • Decision trees are easy to interpret for beginners.
features = ['views', 'likes', 'comments', 'shares', 'platform_code']
X = df[features]
y = df['engagement_rate']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
tree = DecisionTreeRegressor(random_state=42, max_depth=4)
tree.fit(X_train, y_train)
y_pred = tree.predict(X_test)
mse = mean_squared_error(y_test, y_pred)
r2 = r2_score(y_test, y_pred)
print('MSE:', round(mse, 2), '| R2:', round(r2, 3))
MSE: 6.87 | R2: 0.553

Advanced Example 1: Random Forest Feature Importance#

  • Random Forest ensembles can better reveal patterns in real social media data.
  • Feature importances show which factors have the biggest influence on engagement.
  • This is key for building actionable strategies.
rf = RandomForestRegressor(random_state=42, n_estimators=100)
rf.fit(X_train, y_train)
importances = rf.feature_importances_
forest_importance = pd.Series(importances, index=features)
forest_importance.sort_values(ascending=True).plot(kind='barh', color='navy')
plt.title('Random Forest Feature Importance')
plt.xlabel('Importance Score')
plt.show()
No description has been provided for this image

Advanced Example 2: Partial Dependence Plot for Top Feature#

  • Partial dependence plots show how engagement changes as one feature varies.
  • This helps interpret whether more views, comments, or shares really improve engagement.
  • We will plot the relationship between the most important feature and predicted engagement rate.
from sklearn.inspection import PartialDependenceDisplay
top_feature = forest_importance.idxmax()
PartialDependenceDisplay.from_estimator(rf, X_test, [top_feature],
                                       feature_names=features,
                                       grid_resolution=30)
plt.title(f'Partial Dependence: {top_feature} vs Engagement Rate')
plt.show()
No description has been provided for this image

Error Handling Example: Missing Engagement Values#

  • Real data often has missing or zero engagement numbers.
  • Failing to handle missing data can break models and give misleading results.
  • Always check for missing or zero values before modeling.
df_missing = df.copy()
df_missing.loc[df_missing.sample(frac=0.05, random_state=42).index, 'comments'] = np.nan
print('Missing comments:', df_missing["comments"].isna().sum())
df_missing['comments'] = df_missing['comments'].fillna(0)
print('After fillna, missing comments:', df_missing['comments'].isna().sum())
Missing comments: 25
After fillna, missing comments: 0

Error Handling Example: Incorrect Aggregation#

  • Aggregating engagement the wrong way can distort content rankings.
  • Always double-check groupby columns and aggregations when benchmarking.
# Wrong: sums engagement across all posts, losing date info
agg_wrong = df.groupby('platform').sum(numeric_only=True)['engagement_rate']
print('Aggregation error example:\n', agg_wrong)
# Correct: calculates average engagement per post for each platform
agg_correct = df.groupby('platform')['engagement_rate'].mean()
print('Correct aggregation:\n', agg_correct)
Aggregation error example:
 platform
Instagram    2113.179174
TikTok       1941.683993
YouTube      2250.156922
Name: engagement_rate, dtype: float64
Correct aggregation:
 platform
Instagram    13.044316
TikTok       12.690745
YouTube      12.163010
Name: engagement_rate, dtype: float64

Best Practices: Content Benchmarking and Trends#

  • Benchmarking helps compare your posts to platform and industry averages.
  • Audience segmentation splits results by platform, topic, or audience type.
  • Tracking engagement trends helps spot growth, stagnation, or decline.
  • Define metrics consistently so your results can be trusted.
trend = df.groupby(df['date'].dt.to_period('M'))['engagement_rate'].mean()
trend.plot(marker='o', figsize=(9,4))
plt.title('Monthly Engagement Rate Trend')
plt.ylabel('Avg Engagement Rate (%)')
plt.xlabel('Month')
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
No description has been provided for this image

End-to-End Example: Social Media Analytics to Actionable Strategy#

  • Let us bring all these steps together: from raw data, to benchmark, to feature importance, to a content strategy tip.
  • We will find which factor most influences engagement and suggest an improvement.
mean_engagement = df['engagement_rate'].mean()
topfactor = forest_importance.idxmax()
tip = f'Focus on improving {topfactor}: it has the biggest impact on engagement! Your average engagement rate is {mean_engagement:.2f}%.'
print(tip)
Focus on improving likes: it has the biggest impact on engagement! Your average engagement rate is 12.61%.
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.