Mathew K Analytics

Lesson 26 · Social Media Content Analytics

Introduction to Audience Analytics

In this lesson, we will explore how to analyze social media audiences using real public datasets. Audience analytics helps creators and businesses…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Introduction to Audience Analytics#

  • In this lesson, we will explore how to analyze social media audiences using real public datasets.
  • Audience analytics helps creators and businesses understand who is engaging with their content and how.
  • We will solve practical problems like identifying top-performing posts, measuring engagement, and detecting trends.
  • By the end, you will be able to extract actionable insights to optimize content strategy.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Social Media Audience Analytics Concepts#

  • Social media datasets typically include posts or videos and audience engagement metrics.
  • Metrics like views, likes, comments, shares, watch time, and CTR measure how people interact with content.
  • It is important to interpret these numbers in context; for example, high views may not mean high engagement.
  • Beginners often mistake total numbers for performance instead of using ratios and trends.
  • Audience analytics is about understanding behavior, not just counting numbers.
# Beginner Example 1: Load a sample social media content dataset
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df = pd.DataFrame({
    'post_id': range(1, n_posts+1),
    'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
    'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
    'views': views,
    'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
    'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
    'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df.head(3))
   post_id   platform                date  views  likes  comments  shares
0        1  Instagram 2023-01-01 00:00:00  15895   2356       116     253
1        2  Instagram 2023-01-01 06:00:00    960     94        16      26
2        3  Instagram 2023-01-01 12:00:00  76920   3910      2879    1185
# Beginner Example 2: Calculate basic engagement rate
df['engagement_rate'] = ((df['likes'] + df['comments'] + df['shares']) / df['views']) * 100
print(df[['post_id', 'platform', 'views', 'likes', 'comments', 'shares', 'engagement_rate']].head(3))
   post_id   platform  views  likes  comments  shares  engagement_rate
0        1  Instagram  15895   2356       116     253        17.143756
1        2  Instagram    960     94        16      26        14.166667
2        3  Instagram  76920   3910      2879    1185        10.366615
# Beginner Example 3: Find most viewed post
top_post = df.sort_values('views', ascending=False).head(1)
print(top_post[['post_id', 'platform', 'views', 'likes', 'comments', 'shares']])
     post_id platform  views  likes  comments  shares
392      393   TikTok  99813  12361      4424    2837
# Intermediate Example 1: Group by platform and calculate average engagement
platform_group = df.groupby('platform')['engagement_rate'].mean().reset_index()
print(platform_group)
    platform  engagement_rate
0  Instagram        13.044316
1     TikTok        12.690745
2    YouTube        12.163010
# Intermediate Example 2: Identify top 5% highest engagement posts
threshold = df['engagement_rate'].quantile(0.95)
top_engaged = df[df['engagement_rate'] >= threshold]
print(top_engaged[['post_id', 'platform', 'engagement_rate']].head(5))
    post_id   platform  engagement_rate
10       11  Instagram        21.205731
44       45    YouTube        19.149326
56       57  Instagram        20.094392
59       60  Instagram        19.468518
87       88     TikTok        20.698931
# Intermediate Example 3: Analyze daily engagement trends
daily = df.groupby(df['date'].dt.date)['engagement_rate'].mean().reset_index()
print(daily.head(5))
         date  engagement_rate
0  2023-01-01        12.216080
1  2023-01-02         8.793910
2  2023-01-03        12.079164
3  2023-01-04        12.890983
4  2023-01-05        16.177718
# Intermediate Example 4: Find posts with zero comments (low discussion)
no_comment_posts = df[df['comments'] == 0]
print(no_comment_posts[['post_id', 'platform', 'views', 'likes', 'comments']].head())
Empty DataFrame
Columns: [post_id, platform, views, likes, comments]
Index: []
# Intermediate Example 5: Most shared post per platform
idx = df.groupby('platform')['shares'].idxmax()
most_shared = df.loc[idx][['platform', 'post_id', 'shares', 'engagement_rate']]
print(most_shared)
      platform  post_id  shares  engagement_rate
57   Instagram       58    2721         9.874467
392     TikTok      393    2837        19.658762
166    YouTube      167    2766        13.387590
# Advanced Example 1: Simulated YouTube trending data load (fallback to synthetic)
try:
    from googleapiclient.discovery import build
    from google_auth_oauthlib.flow import InstalledAppFlow
    from google.auth.transport.requests import Request
    import os, pickle
    from pathlib import Path
    SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
    def get_yt_service():
        api_key = os.environ.get('YOUTUBE_API_KEY')
        if api_key:
            return build('youtube', 'v3', developerKey=api_key)
        if Path('client_secret.json').exists():
            creds = None
            if Path('token_ro.pickle').exists():
                with open('token_ro.pickle', 'rb') as f:
                    creds = pickle.load(f)
            if not creds or not creds.valid:
                if creds and creds.expired and creds.refresh_token:
                    creds.refresh(Request())
                else:
                    flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
                    creds = flow.run_local_server(port=0)
                with open('token_ro.pickle', 'wb') as f:
                    pickle.dump(creds, f)
            return build('youtube', 'v3', credentials=creds)
        raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
    def fetch_yt_trending(max_results=200, region='US'):
        youtube = get_yt_service()
        records, token = [], None
        while len(records) < max_results:
            resp = youtube.videos().list(
                part='snippet,statistics',
                chart='mostPopular',
                regionCode=region,
                maxResults=min(50, max_results - len(records)),
                pageToken=token
            ).execute()
            for item in resp.get('items', []):
                s = item['snippet']; st = item.get('statistics', {})
                records.append({
                    'video_id':      item['id'],
                    'trending_date': pd.Timestamp.today().date(),
                    'title':         s.get('title', ''),
                    'channel_title': s.get('channelTitle', ''),
                    'category_id':   s.get('categoryId', ''),
                    'views':         int(st.get('viewCount', 0)),
                    'likes':         int(st.get('likeCount', 0)),
                    'comment_count': int(st.get('commentCount', 0)),
                })
            token = resp.get('nextPageToken')
            if not token: break
        return pd.DataFrame(records)
    df_yt = fetch_yt_trending()
    print('Live trending data:', df_yt.shape)
except Exception as e:
    print(f'Falling back to synthetic: {e}')
    np.random.seed(42)
    n = 1000
    df_yt = pd.DataFrame({
        'video_id':      [f'vid{i}' for i in range(n)],
        'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
        'title':         [f'Video Title {i}' for i in range(n)],
        'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
        'category_id':   np.random.choice([1,2,10,22,24,28], n),
        'views':         np.random.randint(10000, 5000000, n),
        'likes':         np.random.randint(100, 200000, n),
        'comment_count': np.random.randint(10, 50000, n),
    })
    print('Synthetic fallback:', df_yt.shape)
print(df_yt.head(3))
Live trending data: (199, 8)
      video_id trending_date  \
0  J7vL5Gb-fEw    2026-06-01   
1  l0hD3KBwoiA    2026-06-01   
2  Vagb9BqdX8g    2026-06-01   

                                               title     channel_title  \
0              ONE BLO MI - DA VIBE | Official Audio           Da Vibe   
1  No Peace Amongst the Stars | Warhammer 40,000 ...  Warhammer 40,000   
2                 We Almost Lost EVERYTHING Gambling        SMii7Yplus   

  category_id    views   likes  comment_count  
0          22   106276    2128             55  
1          20  1396154  114682           5971  
2          20  1007475   52361           1503  
# Advanced Example 2: Calculate like-to-view and comment-to-view ratios for trending videos
df_yt['like_rate'] = (df_yt['likes'] / df_yt['views']) * 100
df_yt['comment_rate'] = (df_yt['comment_count'] / df_yt['views']) * 100
print(df_yt[['video_id', 'views', 'likes', 'comment_count', 'like_rate', 'comment_rate']].head(5))
      video_id    views   likes  comment_count  like_rate  comment_rate
0  J7vL5Gb-fEw   106276    2128             55   2.002334      0.051752
1  l0hD3KBwoiA  1396154  114682           5971   8.214137      0.427675
2  Vagb9BqdX8g  1007475   52361           1503   5.197251      0.149185
3  _9wkngCZ80I    14644     939             91   6.412182      0.621415
4  28yybXEERCg   209210   13025           1107   6.225802      0.529133
# Advanced Example 3: Identify channels with consistently high engagement
engagement_mean = df_yt.groupby('channel_title')[['like_rate', 'comment_rate']].mean().reset_index()
top_channels = engagement_mean.sort_values('like_rate', ascending=False).head(3)
print(top_channels)
      channel_title  like_rate  comment_rate
132          SMTOWN  28.177494      3.028701
152  TREASURE (트레저)  25.505448      4.020588
83          Koogs46  20.656680      0.435050
# Advanced Example 4: Detect possible viral videos using statistical outliers
viral_threshold = df_yt['views'].quantile(0.99)
viral_videos = df_yt[df_yt['views'] >= viral_threshold]
print(viral_videos[['video_id', 'title', 'views', 'likes']].head())
       video_id                                              title    views  \
12  0JlMjgqduVw  House of the Dragon Season 3 | Official Final ...  8531193   
28  v1t4MTqdfyI  Ariana Grande - hate that i made you love me (...  4830843   

     likes  
12   74959  
28  372267  
# Error Handling 1: Check for missing engagement values
missing_any = df.isnull().sum()
print(missing_any)
post_id            0
platform           0
date               0
views              0
likes              0
comments           0
shares             0
engagement_rate    0
dtype: int64
# Error Handling 2: Aggregate metrics incorrectly (demonstration of a common mistake)
wrong_group = df.groupby('platform')[['likes', 'comments', 'shares']].sum() / df.groupby('platform')[['views']].sum() * 100
print(wrong_group)
           comments  likes  shares  views
platform                                 
Instagram       NaN    NaN     NaN    NaN
TikTok          NaN    NaN     NaN    NaN
YouTube         NaN    NaN     NaN    NaN
# Error Handling 3: Misinterpreting click-through rate (CTR) with made-up data
yt_analytics = pd.DataFrame({
    'video_id': range(1, 6),
    'views': [10000, 20000, 5000, 35000, 8000],
    'ctr': [2.5, 15.0, 4.2, 9.0, 6.5],   # Not all are realistic values
})
print(yt_analytics)
   video_id  views   ctr
0         1  10000   2.5
1         2  20000  15.0
2         3   5000   4.2
3         4  35000   9.0
4         5   8000   6.5
# Debugging 1: Identify wrong groupby columns for content category
if 'category_id' in df_yt.columns:
    cat_agg = df_yt.groupby('category_id')['views'].mean().reset_index()
    print(cat_agg.head())
else:
    print('category_id column not found!')
  category_id          views
0           1  422988.000000
1          10  530627.480000
2          17  146731.000000
3          20  260187.790541
4          22  374770.750000

Best Practices for Audience Analytics#

  • Benchmark content using average engagement rates, not just totals.
  • Segment audience behavior by platform, category, or channel.
  • Look for trends in engagement over time and experiment with posting schedules.
  • Use ratios (like likes per view) rather than raw counts for comparing posts.
  • Optimize content by learning from both top- and low-performing posts.
# Common analytics pattern: Benchmark each post against platform average
platform_avgs = df.groupby('platform')['engagement_rate'].transform('mean')
df['above_platform_avg'] = df['engagement_rate'] > platform_avgs
print(df[['post_id', 'platform', 'engagement_rate', 'above_platform_avg']].head(5))
   post_id   platform  engagement_rate  above_platform_avg
0        1  Instagram        17.143756                True
1        2  Instagram        14.166667                True
2        3  Instagram        10.366615               False
3        4    YouTube         7.187284               False
4        5  Instagram        10.636292               False
# Common analytics pattern: Segment by posting time of day
df['hour'] = df['date'].dt.hour
hourly_engagement = df.groupby('hour')['engagement_rate'].mean().reset_index()
print(hourly_engagement.sort_values('engagement_rate', ascending=False).head())
   hour  engagement_rate
2    12        13.243221
0     0        12.631587
3    18        12.589461
1     6        11.975891

End-to-End Audience Analytics Example#

  • Let us solve a mini project: Identify the top three content strategy recommendations for maximizing engagement.
  • We will use all our previous patterns and metrics.
  • We start by finding the top performers, top times, and platforms.
# Step 1: Top 3 posts by engagement rate
top3_posts = df.sort_values('engagement_rate', ascending=False).head(3)
print(top3_posts[['post_id', 'platform', 'date', 'engagement_rate']])
     post_id   platform                date  engagement_rate
351      352  Instagram 2023-03-29 18:00:00        21.781011
10        11  Instagram 2023-01-03 12:00:00        21.205731
383      384     TikTok 2023-04-06 18:00:00        20.925104
# Step 2: Best posting time window by engagement
best_hour = hourly_engagement.sort_values('engagement_rate', ascending=False).head(1)
print(best_hour)
   hour  engagement_rate
2    12        13.243221
# Step 3: Platform with highest average engagement
best_platform_row = platform_group.sort_values('engagement_rate', ascending=False).head(1)
platform_name = best_platform_row.iloc[0]['platform']
print(f'Platform with highest engagement rate: {platform_name}')
Platform with highest engagement rate: Instagram

Content Strategy Recommendations#

    1. Focus more posts on the highest-engagement platform identified.
    1. Schedule content during the hour shown to boost engagement.
    1. Replicate top-performing post themes or formats as shown in the data.
  • Audience analytics turns raw numbers into strategy for creators and brands.
  • Try these steps with your real business or creator data for deeper insight.

Want to learn more?#

  • Subscribe on YouTube for more data-driven content strategy tutorials.
  • Share your insights or ask questions in the comments below!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.