Mathew K Analytics

Lesson 7 · Social Media Content Analytics

Working with Lists, Dictionaries, and Functions for Social Media Content Analytics

In this lesson, we will solve real-world social media analytics problems using Python lists, dictionaries, and functions. The focus will be on analyzing…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Working with Lists, Dictionaries, and Functions for Social Media Content Analytics#

  • In this lesson, we will solve real-world social media analytics problems using Python lists, dictionaries, and functions.
  • The focus will be on analyzing content performance using post-level data and extracting actionable insights.
  • These techniques matter because they help creators and businesses understand what content works and how to improve engagement.
  • You will learn to compute and interpret metrics, detect errors, and make data-driven recommendations for social media strategy.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Core Concepts: Social Media Content Analytics Data#

  • Social media analytics datasets represent digital content such as videos or posts, each row showing key metrics and features.
  • Key engagement metrics include:
    • Views: How many times users watched or saw the content
    • Likes: Number of positive reactions from users
    • Comments: User feedback, questions, or discussions below the post
    • CTR (Click-through Rate): Percentage of users who clicked after seeing content
    • Watch Time: How long users spent watching a video
  • Engagement data is noisysometimes values are missing or have outliers.
  • Beginners often forget to normalize metrics (per post, per user, per period), leading to misleading results.
# Load a public synthetic social media content dataset
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df = pd.DataFrame({
    'post_id': range(1, n_posts+1),
    'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
    'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
    'views': views,
    'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
    'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
    'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df.shape)
print(df.head(3))
(500, 7)
   post_id   platform                date  views  likes  comments  shares
0        1  Instagram 2023-01-01 00:00:00  15895   2356       116     253
1        2  Instagram 2023-01-01 06:00:00    960     94        16      26
2        3  Instagram 2023-01-01 12:00:00  76920   3910      2879    1185

Beginner Example 1: Calculating Engagement Rate for Each Post#

  • Engagement rate helps compare how actively users interact with content, regardless of post popularity.
  • It is usually computed by dividing likes by views and multiplying by 100.
def calculate_engagement_rate(row):
    return (row['likes'] / row['views'] * 100) if row['views'] > 0 else np.nan

df['engagement_rate'] = df.apply(calculate_engagement_rate, axis=1)
print(df[['post_id', 'platform', 'views', 'likes', 'engagement_rate']].head(3))
   post_id   platform  views  likes  engagement_rate
0        1  Instagram  15895   2356        14.822271
1        2  Instagram    960     94         9.791667
2        3  Instagram  76920   3910         5.083203

Beginner Example 2: Using Lists to Track Top Post IDs#

  • Lists are useful for storing and reusing collections of values such as top-performing post IDs.
  • Here, we will start by making a list of post IDs whose engagement rate is above 10%.
high_engagement = df[df['engagement_rate'] > 10]['post_id'].tolist()
print(f'Number of posts with >10% engagement: {len(high_engagement)}')
print('Sample top post IDs:', high_engagement[:5])
Number of posts with >10% engagement: 191
Sample top post IDs: [1, 11, 15, 18, 19]

Beginner Example 3: Using Dictionaries to Summarize Post Counts by Platform#

  • Dictionaries in Python help associate keys (like platform names) with summary values (like post counts).
  • We will create a summary that shows how many posts are on each platform.
platform_counts = {}
for platform in df['platform'].unique():
    count = df[df['platform'] == platform].shape[0]
    platform_counts[platform] = count
print('Number of posts per platform:', platform_counts)
Number of posts per platform: {'Instagram': 162, 'YouTube': 185, 'TikTok': 153}

Intermediate Example 1: Defining a Function to Find Top N Posts#

  • Functions help us reuse code for custom analytics tasks.
  • We will write a function that returns the top N posts by engagement rate, including post IDs and platforms.
def get_top_posts(df, n=5):
    top = df.sort_values('engagement_rate', ascending=False).head(n)
    return [{'post_id': row['post_id'], 'platform': row['platform'], 'engagement_rate': row['engagement_rate']} for _, row in top.iterrows()]

top_5 = get_top_posts(df, 5)
print('Top 5 posts by engagement rate:')
for post in top_5:
    print(post)
Top 5 posts by engagement rate:
{'post_id': 394, 'platform': 'TikTok', 'engagement_rate': 14.972183434182238}
{'post_id': 187, 'platform': 'TikTok', 'engagement_rate': 14.957175684144557}
{'post_id': 444, 'platform': 'TikTok', 'engagement_rate': 14.93320522532843}
{'post_id': 428, 'platform': 'Instagram', 'engagement_rate': 14.884489041268017}
{'post_id': 45, 'platform': 'YouTube', 'engagement_rate': 14.868497820690285}

Intermediate Example 2: Grouping Content by Platform and Calculating Mean Engagement#

  • Grouping is a key analytics tool: it lets us compare performance of content across different platforms.
  • We will calculate the average engagement rate for each platform using groupby and dictionaries for clear output.
platform_means = df.groupby('platform')['engagement_rate'].mean().to_dict()
for platform, mean_eng in platform_means.items():
    print(f'Average engagement rate on {platform}: {mean_eng:.2f}%')
Average engagement rate on Instagram: 8.99%
Average engagement rate on TikTok: 8.63%
Average engagement rate on YouTube: 8.00%

Intermediate Example 3: Creating a Summary Table with Engagement Percentiles#

  • Beyond averages, percentiles reveal what separates normal from best-performing content.
  • We will compute the 25th, 50th (median), and 75th percentile engagement rates for each platform using dictionaries.
percentiles = [0.25, 0.5, 0.75]
summary = {}
for platform in df['platform'].unique():
    rates = df[df['platform'] == platform]['engagement_rate']
    summary[platform] = {p: np.percentile(rates, p*100) for p in percentiles}
print('Engagement rate percentiles by platform:')
for plat, pctiles in summary.items():
    print(plat, {f'{int(p*100)}th': round(val,2) for p, val in pctiles.items()})
Engagement rate percentiles by platform:
Instagram {'25th': np.float64(5.26), '50th': np.float64(9.02), '75th': np.float64(12.93)}
YouTube {'25th': np.float64(4.7), '50th': np.float64(7.86), '75th': np.float64(11.02)}
TikTok {'25th': np.float64(5.1), '50th': np.float64(8.98), '75th': np.float64(11.88)}

Advanced Example 1: Detecting Viral Posts Using a Function and List Comprehensions#

  • Viral posts can be identified by having engagement rates far above the norm for their platform.
  • Let us write a function that returns all posts above the 90th percentile for their platform.
def find_viral_posts(df):
    viral_list = []
    platforms = df['platform'].unique()
    for plat in platforms:
        threshold = np.percentile(df[df['platform'] == plat]['engagement_rate'], 90)
        viral = df[(df['platform'] == plat) & (df['engagement_rate'] > threshold)]
        viral_list.extend(viral['post_id'].tolist())
    return viral_list

viral_post_ids = find_viral_posts(df)
print(f'Viral post IDs (above 90th percentile for each platform): {viral_post_ids[:8]}...')
Viral post IDs (above 90th percentile for each platform): [1, 15, 18, 55, 57, 60, 129, 147]...

Advanced Example 2: Building a Custom Analytics Report as a List of Dictionaries#

  • Sometimes we want to build a detailed analytics report, not just raw numbers or lists.
  • We will collect summary statistics for each platform using a function, storing the outputs in a list of dictionaries.
def platform_summary(df):
    reports = []
    for plat in df['platform'].unique():
        subset = df[df['platform'] == plat]
        entry = {
            'platform': plat,
            'n_posts': subset.shape[0],
            'avg_engagement_rate': round(subset['engagement_rate'].mean(),2),
            'top_post_id': subset.sort_values('engagement_rate', ascending=False).iloc[0]['post_id']
        }
        reports.append(entry)
    return reports

report = platform_summary(df)
for item in report:
    print(item)
{'platform': 'Instagram', 'n_posts': 162, 'avg_engagement_rate': np.float64(8.99), 'top_post_id': np.int64(428)}
{'platform': 'YouTube', 'n_posts': 185, 'avg_engagement_rate': np.float64(8.0), 'top_post_id': np.int64(45)}
{'platform': 'TikTok', 'n_posts': 153, 'avg_engagement_rate': np.float64(8.63), 'top_post_id': np.int64(394)}

Error Handling Example 1: Handling Missing Engagement Values#

  • In real datasets, some engagement numbers (likes, comments, etc.) may be missing or invalid.
  • We will introduce missing data, then safely handle it with Python dictionaries and functions.
# Simulate some missing like counts
df_missing = df.copy()
df_missing.loc[np.random.choice(df_missing.index, 10, replace=False), 'likes'] = np.nan

def safe_engagement(row):
    if pd.isna(row['likes']) or row['views'] == 0:
        return np.nan
    return row['likes'] / row['views'] * 100

df_missing['safe_engagement_rate'] = df_missing.apply(safe_engagement, axis=1)
print(df_missing[['likes', 'views', 'safe_engagement_rate']].head(10))
    likes  views  safe_engagement_rate
0  2356.0  15895             14.822271
1    94.0    960              9.791667
2  3910.0  76920              5.083203
3  1827.0  54986              3.322664
4   253.0   6365              3.974863
5  4287.0  82486              5.197246
6  1524.0  37294              4.086448
7  3876.0  87598              4.424759
8  2523.0  44231              5.704144
9  2567.0  60363              4.252605

Error Handling Example 2: Detecting Invalid Category Aggregations#

  • Aggregating metrics without grouping can create misleading results (for example, all platforms combined).
  • We will show the common mistake, then fix it using groupby in a function.
# Mistake: calculate mean engagement rate on ALL posts regardless of platform
mean_all = df['engagement_rate'].mean()
print(f'Average engagement rate across all platforms: {mean_all:.2f}%')

# Fix: calculate mean for each platform with a function
def group_mean_engagement(df):
    return df.groupby('platform')['engagement_rate'].mean().to_dict()

platform_means = group_mean_engagement(df)
print('Proper per-platform means:', platform_means)
Average engagement rate across all platforms: 8.51%
Proper per-platform means: {'Instagram': 8.990956565523035, 'TikTok': 8.629246963217597, 'YouTube': 7.99969493661874}

Error Handling Example 3: Interpreting Engagement Ratios with Few Views#

  • Very high engagement rates are sometimes due to very low view counts, providing misleading insights.
  • We will create a filter to flag suspicious engagement rates resulting from few views.
def suspicious_high_engagement(row):
    return row['engagement_rate'] > 30 and row['views'] < 200

suspicious = df[df.apply(suspicious_high_engagement, axis=1)]
print(f'Number of suspiciously high engagement rate posts: {len(suspicious)}')
print(suspicious[['post_id', 'views', 'likes', 'engagement_rate']].head())
Number of suspiciously high engagement rate posts: 0
Empty DataFrame
Columns: [post_id, views, likes, engagement_rate]
Index: []

Best Practices: Reliable Content Analytics Patterns#

  • Always check for missing or outlier values before analyzing.
  • Aggregate metrics by relevant groupings (such as platform or content type).
  • Define and use consistent, well-documented metric calculations.
  • Use lists and dictionaries to organize summaries and results for reports or further analysis.
  • Design functions to automate repetitive analytics logicthis avoids mistakes.
# Example: Consistently defining engagement rate as likes divided by views times 100
def consistent_engagement(df, like_col='likes', view_col='views'):
    return (df[like_col] / df[view_col] * 100).round(2)

df['engagement_rate2'] = consistent_engagement(df)
print(df[['likes', 'views', 'engagement_rate2']].head(3))
   likes  views  engagement_rate2
0   2356  15895             14.82
1     94    960              9.79
2   3910  76920              5.08

Applying It All: An End-to-End Mini Analytics Project#

  • We will take raw post-level data, compute engagement rates, identify top performers, and recommend a posting strategy.
  • Steps:
    1. Calculate engagement rates.
    1. Identify the platform's top 5 posts.
    1. Suggest the best platform for future posts based on average engagement.
# Step 1: Engagement rates are already calculated in 'engagement_rate'
# Step 2: Identify top 5 posts overall
top5_end_to_end = df.sort_values('engagement_rate', ascending=False).head(5)
print('Top 5 posts by engagement rate:')
print(top5_end_to_end[['post_id','platform','engagement_rate','likes','views']])

# Step 3: Suggest the platform with the highest average engagement
avg_engagements = df.groupby('platform')['engagement_rate'].mean()
best_platform = avg_engagements.idxmax()
print(f'Best platform for posting, based on average engagement rate: {best_platform}')
Top 5 posts by engagement rate:
     post_id   platform  engagement_rate  likes  views
393      394     TikTok        14.972183   6755  45117
186      187     TikTok        14.957176   1432   9574
443      444     TikTok        14.933205   8082  54121
427      428  Instagram        14.884489   8292  55709
44        45    YouTube        14.868498  12997  87413
Best platform for posting, based on average engagement rate: Instagram
# YouTube CTA: Find more tutorials at YouTube.com!
print('Subscribe for more analytics walkthroughs on our channel!')
Subscribe for more analytics walkthroughs on our channel!
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.