Mathew K Analytics

Lesson 1 · Social Media Content Analytics

Introduction to Social Media and Content Analytics

In this lesson, we will learn how to analyze real-world social media and content data using Python. Social media analytics help content creators and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Introduction to Social Media and Content Analytics#

  • In this lesson, we will learn how to analyze real-world social media and content data using Python.
  • Social media analytics help content creators and businesses understand what works, grow their audience, and optimize their strategy.
  • We will use datasets from platforms like YouTube, Instagram, and TikTok to uncover content performance and engagement trends.
  • You will see how to calculate engagement rates, find top-performing posts, and avoid common data interpretation mistakes.
  • By the end, you will be able to generate actionable insights using real analytics methods.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Core Social Media Analytics Concepts#

  • Social media datasets contain information about content, such as posts or videos, their engagement, and audience reactions.
  • Key metrics include views (how many watched), likes (positive feedback), comments (audience interaction), CTR (click-through rate), and watch time (how long people engage).
  • Beginners sometimes misread engagement data by focusing only on total numbers and not on engagement relative to views, or by mixing up metrics from different categories of content.
  • Correct analysis uses ratios and comparisons to understand why a piece of content performed well or poorly.
# Beginner Example 1: Load a synthetic social media content dataset
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df_content = pd.DataFrame({
    'post_id': range(1, n_posts+1),
    'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
    'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
    'views': views,
    'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
    'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
    'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df_content.head(3))
   post_id   platform                date  views  likes  comments  shares
0        1  Instagram 2023-01-01 00:00:00  15895   2356       116     253
1        2  Instagram 2023-01-01 06:00:00    960     94        16      26
2        3  Instagram 2023-01-01 12:00:00  76920   3910      2879    1185
# Beginner Example 2: Calculate the overall engagement rate
df_content['engagement_rate'] = ((df_content['likes'] + df_content['comments'] + df_content['shares']) / df_content['views']) * 100
df_content['engagement_rate'] = df_content['engagement_rate'].round(2)
print(df_content[['post_id', 'platform', 'views', 'likes', 'comments', 'shares', 'engagement_rate']].head(5))
   post_id   platform  views  likes  comments  shares  engagement_rate
0        1  Instagram  15895   2356       116     253            17.14
1        2  Instagram    960     94        16      26            14.17
2        3  Instagram  76920   3910      2879    1185            10.37
3        4    YouTube  54986   1827       488    1637             7.19
4        5  Instagram   6365    253       261     163            10.64
# Beginner Example 3: Find the post with the highest engagement rate
top_engaged = df_content.loc[df_content['engagement_rate'].idxmax()]
print(f"Post {top_engaged['post_id']} on {top_engaged['platform']} had the highest engagement rate: {top_engaged['engagement_rate']}%")
print(top_engaged[['post_id', 'platform', 'views', 'likes', 'comments', 'shares', 'engagement_rate']])
Post 352 on Instagram had the highest engagement rate: 21.78%
post_id                  352
platform           Instagram
views                  74643
likes                  10653
comments                3569
shares                  2036
engagement_rate        21.78
Name: 351, dtype: object
# Intermediate Example 1: Compare engagement rates between platforms
platform_group = df_content.groupby('platform')['engagement_rate'].mean().round(2)
print('Average engagement rate by platform:')
print(platform_group)
Average engagement rate by platform:
platform
Instagram    13.04
TikTok       12.69
YouTube      12.16
Name: engagement_rate, dtype: float64
# Intermediate Example 2: Identify top 5 posts by total comments
top_comments = df_content.nlargest(5, 'comments')[['post_id', 'platform', 'views', 'comments', 'engagement_rate']]
print('Top 5 posts by comment count:')
print(top_comments)
Top 5 posts by comment count:
     post_id   platform  views  comments  engagement_rate
77        78     TikTok  94763      4590            13.72
392      393     TikTok  99813      4424            19.66
494      495    YouTube  93568      4389            17.27
140      141  Instagram  91082      4269            10.91
455      456  Instagram  83713      4161             9.53
# Intermediate Example 3: Calculate posting frequency over time
posts_per_week = df_content.set_index('date').resample('W').size()
print('Number of posts per week:')
print(posts_per_week.head())
Number of posts per week:
date
2023-01-01     4
2023-01-08    28
2023-01-15    28
2023-01-22    28
2023-01-29    28
Freq: W-SUN, dtype: int64
# Intermediate Example 4: Load YouTube Trending Videos data (with fallback)
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request

SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']

def get_yt_service():
    api_key = os.environ.get('YOUTUBE_API_KEY')
    if api_key:
        return build('youtube', 'v3', developerKey=api_key)
    if Path('client_secret.json').exists():
        creds = None
        if Path('token_ro.pickle').exists():
            with open('token_ro.pickle', 'rb') as f:
                creds = pickle.load(f)
        if not creds or not creds.valid:
            if creds and creds.expired and creds.refresh_token:
                creds.refresh(Request())
            else:
                flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
                creds = flow.run_local_server(port=0)
            with open('token_ro.pickle', 'wb') as f:
                pickle.dump(creds, f)
        return build('youtube', 'v3', credentials=creds)
    raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')

def fetch_yt_trending(max_results=200, region='US'):
    youtube = get_yt_service()
    records, token = [], None
    while len(records) < max_results:
        resp = youtube.videos().list(
            part='snippet,statistics',
            chart='mostPopular',
            regionCode=region,
            maxResults=min(50, max_results - len(records)),
            pageToken=token
        ).execute()
        for item in resp.get('items', []):
            s = item['snippet']; st = item.get('statistics', {})
            records.append({
                'video_id':      item['id'],
                'trending_date': pd.Timestamp.today().date(),
                'title':         s.get('title', ''),
                'channel_title': s.get('channelTitle', ''),
                'category_id':   s.get('categoryId', ''),
                'views':         int(st.get('viewCount', 0)),
                'likes':         int(st.get('likeCount', 0)),
                'comment_count': int(st.get('commentCount', 0)),
            })
        token = resp.get('nextPageToken')
        if not token: break
    return pd.DataFrame(records)

try:
    df_trending = fetch_yt_trending()
    print('Live trending data:', df_trending.shape)
except Exception as e:
    print(f'Falling back to synthetic: {e}')
    np.random.seed(42)
    n = 1000
    df_trending = pd.DataFrame({
        'video_id':      [f'vid{i}' for i in range(n)],
        'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
        'title':         [f'Video Title {i}' for i in range(n)],
        'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
        'category_id':   np.random.choice([1,2,10,22,24,28], n),
        'views':         np.random.randint(10000, 5000000, n),
        'likes':         np.random.randint(100, 200000, n),
        'comment_count': np.random.randint(10, 50000, n),
    })
    print('Synthetic fallback:', df_trending.shape)

print(df_trending.head(3))
Falling back to synthetic: <HttpError 403 when requesting https://youtube.googleapis.com/youtube/v3/videos?part=snippet%2Cstatistics&chart=mostPopular&regionCode=US&maxResults=50&key=YOUR_GOOGLE_API_KEY&alt=json returned "The request cannot be completed because you have exceeded your <a href="/youtube/v3/getting-started#quota">quota</a>.". Details: "[{'message': 'The request cannot be completed because you have exceeded your <a href="/youtube/v3/getting-started#quota">quota</a>.', 'domain': 'youtube.quota', 'reason': 'quotaExceeded'}]">
Synthetic fallback: (1000, 8)
  video_id trending_date          title channel_title  category_id    views  \
0     vid0    2023-01-01  Video Title 0      ChannelC           10  3958242   
1     vid1    2023-01-02  Video Title 1      ChannelA           10  1227060   
2     vid2    2023-01-03  Video Title 2      ChannelC           24  1041519   

    likes  comment_count  
0  144402          18611  
1  160181           4977  
2   33966          18763  
# Intermediate Example 5: Calculate like rate and comment rate on trending videos
df_trending['like_rate'] = (df_trending['likes'] / df_trending['views'] * 100).round(2)
df_trending['comment_rate'] = (df_trending['comment_count'] / df_trending['views'] * 100).round(2)
print(df_trending[['title', 'views', 'likes', 'comment_count', 'like_rate', 'comment_rate']].head(5))
           title    views   likes  comment_count  like_rate  comment_rate
0  Video Title 0  3958242  144402          18611       3.65          0.47
1  Video Title 1  1227060  160181           4977      13.05          0.41
2  Video Title 2  1041519   33966          18763       3.26          1.80
3  Video Title 3   799352  105500          20032      13.20          2.51
4  Video Title 4  4688714   84346          43473       1.80          0.93
# Intermediate Example 6: Find top 3 trending videos by like rate
top_like_rate = df_trending.nlargest(3, 'like_rate')[['title', 'channel_title', 'views', 'likes', 'like_rate']]
print('Top 3 trending videos by like rate:')
print(top_like_rate)
Top 3 trending videos by like rate:
               title channel_title  views   likes  like_rate
675  Video Title 675      ChannelC  39725  199875     503.15
821  Video Title 821      ChannelA  29675   93674     315.67
155  Video Title 155      ChannelC  38625  110255     285.45
# Advanced Example 1: Analyze content performance across categories
cat_group = df_trending.groupby('category_id')['like_rate'].mean().round(2)
print('Average like rate by content category:')
print(cat_group)
Average like rate by content category:
category_id
1      8.60
2      9.23
10    11.33
22     7.21
24    13.17
28     8.15
Name: like_rate, dtype: float64
# Advanced Example 2: Detect possible viral videos by engagement outliers
viral_threshold = df_trending['engagement_rate'] = ((df_trending['likes'] + df_trending['comment_count']) / df_trending['views']) * 100
viral_videos = df_trending[df_trending['engagement_rate'] > df_trending['engagement_rate'].quantile(0.995)]
print(f'There are {len(viral_videos)} potential viral videos detected as engagement outliers.')
print(viral_videos[['title', 'views', 'likes', 'comment_count', 'engagement_rate']].head())
There are 5 potential viral videos detected as engagement outliers.
               title  views   likes  comment_count  engagement_rate
71    Video Title 71  71367  154555          43008       276.826825
155  Video Title 155  38625  110255          23746       346.928155
301  Video Title 301  59104  100032          48702       251.647943
675  Video Title 675  39725  199875          27974       573.565765
821  Video Title 821  29675   93674           1639       321.189553
# Advanced Example 3: Time-series analysis of social media engagement trends
n_days = 365
dates = pd.date_range('2023-01-01', periods=n_days, freq='D')
views_series = np.random.randint(1000, 50000, n_days)
likes_series = (views_series * np.random.uniform(0.03, 0.12, n_days)).astype(int)
df_timeseries = pd.DataFrame({
    'date': dates,
    'views': views_series,
    'likes': likes_series,
    'engagement_rate': np.round(likes_series / views_series * 100, 2)
})
trend = df_timeseries['engagement_rate'].rolling(window=30).mean()
print('30-day average engagement rate trend:')
print(trend.tail(10))
30-day average engagement rate trend:
355    8.315667
356    8.202667
357    8.184667
358    8.389667
359    8.167667
360    8.380667
361    8.115000
362    7.859000
363    7.914667
364    7.919000
Name: engagement_rate, dtype: float64
# Advanced Example 4: Segmenting posts by engagement quartile
quartiles = pd.qcut(df_content['engagement_rate'], 4, labels=['Low', 'Medium', 'High', 'Top'])
seg_counts = quartiles.value_counts().sort_index()
print('Distribution of posts by engagement quartile:')
print(seg_counts)
Distribution of posts by engagement quartile:
engagement_rate
Low       125
Medium    125
High      125
Top       125
Name: count, dtype: int64
# Error Example 1: Handle missing likes or comments
df_content_missing = df_content.copy()
df_content_missing.loc[0, 'likes'] = np.nan  # Force a missing value
missing_likes = df_content_missing['likes'].isna().sum()
print(f'There are {missing_likes} posts with missing like values.')
df_content_missing['likes'] = df_content_missing['likes'].fillna(0)
print('After fill, missing likes:', df_content_missing["likes"].isna().sum())
There are 1 posts with missing like values.
After fill, missing likes: 0
# Error Example 2: Misinterpreting ratios (avoid division by zero)
df_buggy = df_content.copy()
df_buggy.loc[1, 'views'] = 0  # Simulate a data error
try:
    df_buggy['like_rate'] = (df_buggy['likes'] / df_buggy['views']) * 100
except Exception as e:
    print('Error encountered:', e)
df_buggy['like_rate'] = np.where(df_buggy['views'] == 0, 0, (df_buggy['likes'] / df_buggy['views']) * 100)
print('Like rate with division by zero handled:', df_buggy[['post_id', 'views', 'likes', 'like_rate']].head(3))
Like rate with division by zero handled:    post_id  views  likes  like_rate
0        1  15895   2356  14.822271
1        2      0     94   0.000000
2        3  76920   3910   5.083203
# Error Example 3: Incorrect category grouping
wrong_group = df_trending.groupby('trending_date')['like_rate'].mean().head()
correct_group = df_trending.groupby('category_id')['like_rate'].mean().head()
print('Grouping by date (incorrect for category analysis):')
print(wrong_group)
print('\nGrouping by category (correct for category analysis):')
print(correct_group)
Grouping by date (incorrect for category analysis):
trending_date
2023-01-01     3.65
2023-01-02    13.05
2023-01-03     3.26
2023-01-04    13.20
2023-01-05     1.80
Name: like_rate, dtype: float64

Grouping by category (correct for category analysis):
category_id
1      8.601412
2      9.228443
10    11.326879
22     7.207386
24    13.169401
Name: like_rate, dtype: float64
# Best Practice 1: Benchmark your content performance
mean_engagement = df_content['engagement_rate'].mean().round(2)
print(f'The average engagement rate across all content is {mean_engagement}%.')
if mean_engagement < 2:
    print('Your engagement is below typical benchmarks. Look for improvement areas!')
elif mean_engagement > 8:
    print('Congrats! Your content is performing above average.')
else:
    print('You are near the average engagement rate. Aim for Top Quartile!')
The average engagement rate across all content is 12.61%.
Congrats! Your content is performing above average.
# Best Practice 2: Consistent metric definitions and usage
engagement_cols = [c for c in df_content.columns if 'engagement_rate' in c or 'like_rate' in c]
print('Engagement-related columns:', engagement_cols)
Engagement-related columns: ['engagement_rate']
# Best Practice 3: Segment audience by platform and analyze engagement
aud_seg = df_content.groupby('platform')['engagement_rate'].describe()[['mean','std','min','max']]
print('Audience engagement statistics by platform:')
print(aud_seg)
Audience engagement statistics by platform:
                mean       std   min    max
platform                                   
Instagram  13.044444  4.186091  3.20  21.78
TikTok     12.690850  4.192301  3.37  20.93
YouTube    12.163027  3.995700  3.08  19.74
# End-to-End Analytics Problem: Recommend a content strategy
platform_perf = df_content.groupby('platform').agg({'views': 'mean', 'likes': 'mean', 'engagement_rate': 'mean'}).round(2)
best_platform = platform_perf['engagement_rate'].idxmax()
suggestion = f'Recommend focusing more content on {best_platform}, which currently has the highest engagement rate.'
print(platform_perf)
print(suggestion)
              views    likes  engagement_rate
platform                                     
Instagram  52754.95  4702.83            13.04
TikTok     50206.41  4491.14            12.69
YouTube    49390.12  4020.61            12.16
Recommend focusing more content on Instagram, which currently has the highest engagement rate.

YouTube Call to Action#

  • If you enjoyed these analytics walkthroughs, check out our channel for more step-by-step social data tutorials!
# Save engagement summary for further exploration
platform_perf.to_csv('platform_engagement_summary.csv', index=True)
print('Saved platform engagement summary as CSV.')
Saved platform engagement summary as CSV.

Congratulations! You have completed the introduction to social media and content analytics.#

  • Practice daily on real content data. Track your growth and adjust strategies using what you learned.
  • Subscribe for more hands-on analytics lessons.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.