Mathew K Analytics

Lesson 19 · Social Media Content Analytics

Encoding Categorical Variables in Social Media Data

Many real-world social media datasets include text or category columns, such as platform type or channel name. Categorical variables are text or ID values…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Encoding Categorical Variables in Social Media Data#

  • Many real-world social media datasets include text or category columns, such as platform type or channel name.
  • Categorical variables are text or ID values that must be converted before you analyze or model your data.
  • For social media and content analytics, correctly encoding these variables allows you to segment, compare, and benchmark performance across different audience or creator groups.
  • In this lesson, you will learn why encoding is important, how to safely transform social media category data, and the impact on your analytics.
  • You will see hands-on steps for encoding platforms, categories, and channels in real datasets, with real engagement analysis examples.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Understanding Social Media Content Analytics Data#

  • Social media datasets often record content like videos, posts, and comments, with columns for views, likes, shares, and other metrics.
  • Categorical variables can include things like platform (YouTube, Instagram), category (Music, Sports), or specific channel names.
  • Engagement metrics reflect user response but may be influenced by content type or audience.
  • Beginners often forget to encode or mishandle category columns, causing misleading groupings or failed modeling.
  • If you analyze engagement by category, but categories remain as text, you may get errors or miss crucial trends.
# Load a social media posts dataset with categorical variables
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df = pd.DataFrame({
    'post_id': range(1, n_posts+1),
    'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
    'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
    'views': views,
    'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
    'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
    'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df.head(3))
   post_id   platform                date  views  likes  comments  shares
0        1  Instagram 2023-01-01 00:00:00  15895   2356       116     253
1        2  Instagram 2023-01-01 06:00:00    960     94        16      26
2        3  Instagram 2023-01-01 12:00:00  76920   3910      2879    1185

What Are Categorical Variables?#

  • A categorical variable holds text or codes representing groups, types, or classes.
  • Examples in social media: platform (YouTube, Instagram), category (Music, Gaming), region, device type, or user segment.
  • Categorical columns allow you to compare engagement or behavior between different groups.
  • They must be encoded (converted to numbers or binary variables) for most analytics, plotting, or modeling.
# Check the types of columns in our dataset
print(df.dtypes)
post_id              int64
platform            object
date        datetime64[ns]
views                int32
likes                int64
comments             int64
shares               int64
dtype: object
# Beginner: Count the number of unique categories in 'platform'
platform_counts = df['platform'].value_counts()
print(platform_counts)
platform
YouTube      185
Instagram    162
TikTok       153
Name: count, dtype: int64
# Beginner: Calculate average engagement rate per platform without encoding
df['engagement_rate'] = (df['likes'] + df['comments'] + df['shares']) / df['views'] * 100
print(df.groupby('platform')['engagement_rate'].mean())
platform
Instagram    13.044316
TikTok       12.690745
YouTube      12.163010
Name: engagement_rate, dtype: float64
# Beginner: Encode platform using pandas 'category' dtype
df['platform_cat'] = df['platform'].astype('category')
print(df[['platform','platform_cat']].head())
    platform platform_cat
0  Instagram    Instagram
1  Instagram    Instagram
2  Instagram    Instagram
3    YouTube      YouTube
4  Instagram    Instagram
# Beginner: Encode platform using label encoding (integers)
df['platform_le'] = df['platform_cat'].cat.codes
print(df[['platform','platform_le']].head(6))
    platform  platform_le
0  Instagram            0
1  Instagram            0
2  Instagram            0
3    YouTube            2
4  Instagram            0
5  Instagram            0
# Beginner: One-hot encode platform for use in machine learning
platform_dummies = pd.get_dummies(df['platform'], prefix='platform')
print(platform_dummies.head())
   platform_Instagram  platform_TikTok  platform_YouTube
0                True            False             False
1                True            False             False
2                True            False             False
3               False            False              True
4                True            False             False
# Intermediate: Join one-hot columns to original DataFrame
df = pd.concat([df, platform_dummies], axis=1)
print(df.head(3))
   post_id   platform                date  views  likes  comments  shares  \
0        1  Instagram 2023-01-01 00:00:00  15895   2356       116     253   
1        2  Instagram 2023-01-01 06:00:00    960     94        16      26   
2        3  Instagram 2023-01-01 12:00:00  76920   3910      2879    1185   

   engagement_rate platform_cat  platform_le  platform_Instagram  \
0        17.143756    Instagram            0                True   
1        14.166667    Instagram            0                True   
2        10.366615    Instagram            0                True   

   platform_TikTok  platform_YouTube  
0            False             False  
1            False             False  
2            False             False  
# Intermediate: Calculate engagement rates by one-hot-encoded platform
avg_engagement_onehot = df.groupby(['platform_YouTube','platform_Instagram','platform_TikTok'])['engagement_rate'].mean()
print(avg_engagement_onehot)
platform_YouTube  platform_Instagram  platform_TikTok
False             False               True               12.690745
                  True                False              13.044316
True              False               False              12.163010
Name: engagement_rate, dtype: float64

Categorical Encoding for Other Social Media Columns#

  • You can encode content category, channel name, or region with the same approach.
  • For categories with many unique values (such as channel titles), use grouping/aggregation or target encoding.
  • Binary encoding (one-hot) is safest for non-ordinal categories with a small number of groups.
  • Ordinal encoding applies only when a true ranking exists.
# Intermediate: Add a fake content category and encode it
df['category'] = np.random.choice(['Education','Entertainment','Lifestyle'], size=df.shape[0])
df['category_cat'] = df['category'].astype('category')
df['category_le'] = df['category_cat'].cat.codes
print(df[['category','category_le']].head())
        category  category_le
0  Entertainment            1
1      Education            0
2      Education            0
3      Lifestyle            2
4  Entertainment            1
# Intermediate: One-hot encode the category column
category_dummies = pd.get_dummies(df['category'], prefix='category')
df = pd.concat([df, category_dummies], axis=1)
print(df.head(3))
   post_id   platform                date  views  likes  comments  shares  \
0        1  Instagram 2023-01-01 00:00:00  15895   2356       116     253   
1        2  Instagram 2023-01-01 06:00:00    960     94        16      26   
2        3  Instagram 2023-01-01 12:00:00  76920   3910      2879    1185   

   engagement_rate platform_cat  platform_le  platform_Instagram  \
0        17.143756    Instagram            0                True   
1        14.166667    Instagram            0                True   
2        10.366615    Instagram            0                True   

   platform_TikTok  platform_YouTube       category   category_cat  \
0            False             False  Entertainment  Entertainment   
1            False             False      Education      Education   
2            False             False      Education      Education   

   category_le  category_Education  category_Entertainment  category_Lifestyle  
0            1               False                    True               False  
1            0                True                   False               False  
2            0                True                   False               False  
# Intermediate: Pivot engagement rates by platform and category
engagement_pivot = df.pivot_table(index='platform', columns='category', values='engagement_rate', aggfunc='mean')
print(engagement_pivot)
category   Education  Entertainment  Lifestyle
platform                                      
Instagram  13.329860      13.059232  12.737610
TikTok     13.104670      13.062484  12.026173
YouTube    12.274175      12.288705  11.905068
# Advanced: Simulate channel_title as high-cardinality categorical variable
df['channel_title'] = np.random.choice([f'Channel_{i}' for i in range(1,61)], size=df.shape[0])
print(df['channel_title'].value_counts().head(3))
channel_title
Channel_13    15
Channel_47    14
Channel_17    14
Name: count, dtype: int64
# Advanced: Group by channel_title and compute mean engagement
top_channels = df.groupby('channel_title')['engagement_rate'].mean().sort_values(ascending=False).head(5)
print(top_channels)
channel_title
Channel_48    16.057606
Channel_55    15.526119
Channel_54    15.337023
Channel_39    14.970264
Channel_2     14.931884
Name: engagement_rate, dtype: float64
# Advanced: Target encoding example for channel_title (mean encoding)
channel_mean_engagement = df.groupby('channel_title')['engagement_rate'].transform('mean')
df['channel_mean_encoded'] = channel_mean_engagement
print(df[['channel_title','channel_mean_encoded']].head(3))
  channel_title  channel_mean_encoded
0    Channel_27             13.832927
1    Channel_11             13.079897
2    Channel_57             11.462865

Error Handling and Debugging in Categorical Encoding#

  • Missing category values can cause errors when encoding or grouping.
  • Different spelling or case (such as YouTube vs youtube) creates accidental new categories.
  • Including rare categories in one-hot encoding can create noisy, mostly-zero columns.
  • Always check for NaNs and standardize category labels before encoding.
  • Double-check column data types after encoding to avoid silent bugs.
# Debugging: Check for missing values in category columns
print(df[['platform','category','channel_title']].isnull().sum())
platform         0
category         0
channel_title    0
dtype: int64
# Debugging: Introduce and handle a category spelling error
df.loc[df.index[0], 'platform'] = 'youtube'  # lowercase typo
df['platform'] = df['platform'].str.title()
print(df['platform'].unique())
['Youtube' 'Instagram' 'Tiktok']
# Debugging: Handle missing categories by filling with 'Unknown'
df.loc[5, 'category'] = np.nan
df['category_filled'] = df['category'].fillna('Unknown')
print(df[['category','category_filled']].head(7))
        category category_filled
0  Entertainment   Entertainment
1      Education       Education
2      Education       Education
3      Lifestyle       Lifestyle
4  Entertainment   Entertainment
5            NaN         Unknown
6      Lifestyle       Lifestyle
# Debugging: Check types after encoding
print(df[['platform_le','category_le','channel_mean_encoded']].dtypes)
platform_le                int8
category_le                int8
channel_mean_encoded    float64
dtype: object

Best Practices and Analytics Patterns with Encoded Categories#

  • Use category encoding to benchmark content performance by platform, category, or channel.
  • Grouping and pivoting on category columns lets you spot patterns and trends.
  • For audience segmentation, encode user location or device type to find key segments.
  • Prefer binary or one-hot for few categories, target encoding or aggregation for many.
  • Always document what each encoded column represents to avoid model mistakes.
# Benchmarking: Find top-performing content category overall
best_category = df.groupby('category_filled')['engagement_rate'].mean().idxmax()
print('Top performing content category:', best_category)
Top performing content category: Education
# Segmentation: Compare engagement for a single channel across platforms
focus_channel = df['channel_title'].iloc[0]
channel_df = df[df['channel_title']==focus_channel]
seg_result = channel_df.groupby('platform')['engagement_rate'].mean()
print('Engagement by platform for', focus_channel)
print(seg_result)
Engagement by platform for Channel_27
platform
Instagram    11.802925
Youtube      15.862929
Name: engagement_rate, dtype: float64

Tiny End-to-End Analytics Problem: Recommend a Content Strategy#

  • Task: From all posts, identify which platform and category together yield the best engagement rates.
  • Your strategy: For new content, prioritize posting that matches this top segment.
  • Encoding the categories ensures the results are reliable and repeatable.
# Solution: Find (platform, category) combo with highest mean engagement
combo = df.groupby(['platform','category_filled'])['engagement_rate'].mean().sort_values(ascending=False).index[0]
print('Best combination:', combo)
Best combination: ('Instagram', 'Education')
# Save the processed DataFrame for further analysis or reporting
df.to_csv('encoded_social_media.csv', index=False)

YouTube Quick Tip: Keep Learning and Try Encoding on Real Data!#

  • Watch videos about pandas encoding or scikit-learn preprocessing for more examples.
  • Practice by encoding your own channel or post dataset, then try grouping and segmentation.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.