Mathew K Analytics

Lesson 51 · Social Media Content Analytics

Introduction to Machine Learning for Social Media

This lesson explores how machine learning helps social media creators and analysts understand and improve their content. You will learn to analyze real…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Introduction to Machine Learning for Social Media#

  • This lesson explores how machine learning helps social media creators and analysts understand and improve their content.

  • You will learn to analyze real engagement metrics, detect trends, and make smarter content decisions using Python.

  • By the end, you will use real social media datasets to generate insights and build data-driven content strategies.

  • No prior machine learning experience needed, but familiarity with Python and Jupyter is recommended.

import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Core Concepts: Social Media Analytics and Machine Learning#

  • Social media datasets represent posts, videos, dates, and performance metrics like views, likes, and comments.
  • Key metrics include:
    • Views: How many times your content was seen.
      • Likes: Positive engagements showing appreciation.
        • Comments: Direct feedback from viewers.
          • CTR (Click-Through Rate): The percentage of people who clicked your post or video after seeing it.
            • Watch Time: Total minutes watched, showing actual interest.

            • Common mistakes include:

              • Only looking at total likes, not rates or context.
                • Ignoring how different content types perform.
                  • Mixing up ratios like CTR and engagement rate.
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request

SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']

def get_yt_service():
    api_key = os.environ.get('YOUTUBE_API_KEY')
    if api_key:
        return build('youtube', 'v3', developerKey=api_key)
    if Path('client_secret.json').exists():
        creds = None
        if Path('token_ro.pickle').exists():
            with open('token_ro.pickle', 'rb') as f:
                creds = pickle.load(f)
        if not creds or not creds.valid:
            if creds and creds.expired and creds.refresh_token:
                creds.refresh(Request())
            else:
                flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
                creds = flow.run_local_server(port=0)
            with open('token_ro.pickle', 'wb') as f:
                pickle.dump(creds, f)
        return build('youtube', 'v3', credentials=creds)
    raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')

def fetch_yt_trending(max_results=200, region='US'):
    youtube = get_yt_service()
    records, token = [], None
    while len(records) < max_results:
        resp = youtube.videos().list(
            part='snippet,statistics',
            chart='mostPopular',
            regionCode=region,
            maxResults=min(50, max_results - len(records)),
            pageToken=token
        ).execute()
        for item in resp.get('items', []):
            s = item['snippet']; st = item.get('statistics', {})
            records.append({
                'video_id':      item['id'],
                'trending_date': pd.Timestamp.today().date(),
                'title':         s.get('title', ''),
                'channel_title': s.get('channelTitle', ''),
                'category_id':   s.get('categoryId', ''),
                'views':         int(st.get('viewCount', 0)),
                'likes':         int(st.get('likeCount', 0)),
                'comment_count': int(st.get('commentCount', 0)),
            })
        token = resp.get('nextPageToken')
        if not token: break
    return pd.DataFrame(records)

try:
    df = fetch_yt_trending()
    print('Live trending data:', df.shape)
except Exception as e:
    print(f'Falling back to synthetic: {e}')
    np.random.seed(42)
    n = 1000
    df = pd.DataFrame({
        'video_id':      [f'vid{i}' for i in range(n)],
        'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
        'title':         [f'Video Title {i}' for i in range(n)],
        'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
        'category_id':   np.random.choice([1,2,10,22,24,28], n),
        'views':         np.random.randint(10000, 5000000, n),
        'likes':         np.random.randint(100, 200000, n),
        'comment_count': np.random.randint(10, 50000, n),
    })
    print('Synthetic fallback:', df.shape)

print(df.head(3))
Live trending data: (199, 8)
      video_id trending_date  \
0  82-jTNka3uc    2026-06-02   
1  3oB9AxspVow    2026-06-02   
2  l-cyT28MFyk    2026-06-02   

                                               title     channel_title  \
0  Ariana Grande - hate that i made you love me (...  ArianaGrandeVevo   
1           The End of Oak Street | Official Trailer      Warner Bros.   
2        PLAYING HARDCORE MINECRAFT UNTIL WE BEAT IT            Jynxzi   

  category_id    views   likes  comment_count  
0          10  3169742  467032          28197  
1           1  3029814   37920           3036  
2          24   763560   16663            387  

Example 1: Calculate Engagement Rate for Trending Videos#

  • Engagement rate is a key metric that tells us how much viewers interact with content.
  • It is defined as: (Likes + Comments) / Views.
  • A higher engagement rate often means the content resonates more with the audience.
df['engagement_rate'] = (df['likes'] + df['comment_count']) / df['views']
print(df[['title', 'views', 'likes', 'comment_count', 'engagement_rate']].head())
                                               title    views   likes  \
0  Ariana Grande - hate that i made you love me (...  3169742  467032   
1           The End of Oak Street | Official Trailer  3029814   37920   
2        PLAYING HARDCORE MINECRAFT UNTIL WE BEAT IT   763560   16663   
3  IShowSpeed - World Cup (Champions) [Official M...  3209162  551706   
4  No Peace Amongst the Stars | Warhammer 40,000 ...  2298282  156952   

   comment_count  engagement_rate  
0          28197         0.156236  
1           3036         0.013518  
2            387         0.022330  
3          47856         0.186828  
4           7601         0.071598  

Example 2: Find Top-Performing Trending Videos by Engagement#

  • Sometimes a video with fewer total views still has the highest engagement rate.
  • We want to know which titles generate the most loyal or interested audience.
top5_engaged = df.sort_values('engagement_rate', ascending=False).head(5)
print(top5_engaged[['title', 'views', 'likes', 'comment_count', 'engagement_rate']])
                                                title    views   likes  \
55      Mastodon - Your Ghost Again (Official Visual)     5842    2234   
33                             Vince Staples - Cotton    14747    4238   
53          bleood - ding dong (official music video)    19988    3954   
20  I Rebuilt The Chicago Bulls For 10 Years (Year 1)   109560   22024   
3   IShowSpeed - World Cup (Champions) [Official M...  3209162  551706   

    comment_count  engagement_rate  
55            291         0.432215  
33            321         0.309148  
53            532         0.224435  
20            439         0.205029  
3           47856         0.186828  

Example 3: Average Engagement by Content Category#

  • Different categories (like Music, Gaming, or News) might show very different engagement patterns.
  • Let us check which category receives the most interaction from viewers.
cat_group = df.groupby('category_id')['engagement_rate'].mean().sort_values(ascending=False)
print(cat_group)
category_id
22    0.084419
10    0.079922
17    0.064187
20    0.056363
24    0.055366
28    0.048305
1     0.038529
19    0.038056
Name: engagement_rate, dtype: float64

Example 4: Identify Most Discussed Trending Videos (by Comments)#

  • Engagement is not just likessometimes comments tell us which videos spark more discussion.
  • Let us sort and see which trending videos have the most comments.
top_commented = df.sort_values('comment_count', ascending=False).head(5)
print(top_commented[['title', 'views', 'likes', 'comment_count']])
                                                 title    views   likes  \
41                               TREASURE - ‘IF I’ M/V  5536570  253481   
3    IShowSpeed - World Cup (Champions) [Official M...  3209162  551706   
197          I Went to WAR on a Hardcore Minecraft SMP  3712832  198866   
0    Ariana Grande - hate that i made you love me (...  3169742  467032   
8              1000 VS 1000 Player Minecraft Civil War  2330558  128251   

     comment_count  
41           49523  
3            47856  
197          33770  
0            28197  
8            24928  

Example 5: Trends in Views Over Time#

  • Sometimes, it is important to look for patterns in how video views change over time.
  • Let us visualize views by trending date to spot any time-based trends.
import matplotlib.pyplot as plt
df_trend = df.groupby('trending_date')['views'].mean()
plt.figure(figsize=(10,4))
plt.plot(df_trend.index, df_trend.values, marker='o')
plt.title('Average Trending Video Views over Time')
plt.xlabel('Trending Date')
plt.ylabel('Average Views')
plt.tight_layout()
plt.show()
No description has been provided for this image
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df_posts = pd.DataFrame({
    'post_id': range(1, n_posts+1),
    'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
    'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
    'views': views,
    'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
    'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
    'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df_posts.head(3))
   post_id   platform                date  views  likes  comments  shares
0        1  Instagram 2023-01-01 00:00:00  15895   2356       116     253
1        2  Instagram 2023-01-01 06:00:00    960     94        16      26
2        3  Instagram 2023-01-01 12:00:00  76920   3910      2879    1185
df_posts['engagement_rate'] = (df_posts['likes'] + df_posts['comments'] + df_posts['shares']) / df_posts['views']
print(df_posts[['platform', 'views', 'likes', 'comments', 'shares', 'engagement_rate']].head())
    platform  views  likes  comments  shares  engagement_rate
0  Instagram  15895   2356       116     253         0.171438
1  Instagram    960     94        16      26         0.141667
2  Instagram  76920   3910      2879    1185         0.103666
3    YouTube  54986   1827       488    1637         0.071873
4  Instagram   6365    253       261     163         0.106363
eng_by_platform = df_posts.groupby('platform')['engagement_rate'].mean()
print(eng_by_platform)
platform
Instagram    0.130443
TikTok       0.126907
YouTube      0.121630
Name: engagement_rate, dtype: float64
post_time_group = df_posts.groupby(df_posts['date'].dt.hour)['views'].mean()
plt.figure(figsize=(8,4))
plt.bar(post_time_group.index, post_time_group.values)
plt.title('Average Views per Posting Hour')
plt.xlabel('Hour of Day')
plt.ylabel('Average Views')
plt.xticks(range(0,24))
plt.tight_layout()
plt.show()
No description has been provided for this image
n_videos = 300
views = np.random.randint(100, 500000, n_videos)
df_yt = pd.DataFrame({
    'video_id': range(1, n_videos+1),
    'publish_date': pd.date_range('2022-01-01', periods=n_videos, freq='D'),
    'views': views,
    'watch_time': np.random.randint(1000, 500000, n_videos),
    'likes': (views * np.random.uniform(0.01, 0.08, n_videos)).astype(int),
    'comments': (views * np.random.uniform(0.001, 0.02, n_videos)).astype(int),
    'ctr': np.round(np.random.uniform(2, 10, n_videos), 2)
})
print(df_yt.head(3))
   video_id publish_date   views  watch_time  likes  comments   ctr
0         1   2022-01-01  339219        9210  22163      1407  2.85
1         2   2022-01-02  435171       45363   8153      4119  3.24
2         3   2022-01-03  343801      243247  22551      1710  9.56
df_yt['avg_watch_per_view'] = df_yt['watch_time'] / df_yt['views']
print(df_yt[['views', 'watch_time', 'avg_watch_per_view']].head())
    views  watch_time  avg_watch_per_view
0  339219        9210            0.027151
1  435171       45363            0.104242
2  343801      243247            0.707523
3  137588       79279            0.576206
4   72616      192420            2.649829
top_ctr = df_yt.sort_values('ctr', ascending=False).head(5)
print(top_ctr[['video_id', 'views', 'ctr', 'likes', 'comments']])
     video_id   views   ctr  likes  comments
224       225   40297  9.98   3156       650
19         20  141659  9.82   7140       338
217       218  279813  9.74  21670       980
160       161  422481  9.68  26731      5108
14         15  129395  9.66   3134      2541
# Flag potentially viral posts: high engagement & sharp view spikes
threshold = df_posts['engagement_rate'].quantile(0.95)
viral = df_posts[(df_posts['engagement_rate'] > threshold) & (df_posts['views'] > df_posts['views'].quantile(0.90))]
print(viral[['platform', 'views', 'engagement_rate']])
      platform  views  engagement_rate
192  Instagram  97604         0.198824
276  Instagram  96701         0.199295
392     TikTok  99813         0.196588
# Week-over-week growth for YouTube video views
yt_week = df_yt.set_index('publish_date')['views'].resample('W').sum()
yt_growth = yt_week.pct_change().fillna(0)
plt.figure(figsize=(8, 4))
plt.plot(yt_week.index, yt_growth.values, marker='o')
plt.title('Week-over-Week Growth in YouTube Video Views')
plt.xlabel('Week')
plt.ylabel('Growth Rate')
plt.axhline(0, color='red', linestyle='--')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Simple machine learning: predict engagement rate based on views, likes, comments (linear regression)
from sklearn.linear_model import LinearRegression
X = df_posts[['views', 'likes', 'comments']]
y = df_posts['engagement_rate']
model = LinearRegression()
model.fit(X, y)
predicted = model.predict(X)
print('Sample actual vs predicted engagement rates:')
print(np.round(np.c_[y[:5], predicted[:5]], 4))
Sample actual vs predicted engagement rates:
[[0.1714 0.1342]
 [0.1417 0.1232]
 [0.1037 0.1002]
 [0.0719 0.0677]
 [0.1064 0.1205]]
# Evaluate the prediction error using mean squared error
from sklearn.metrics import mean_squared_error
mse = mean_squared_error(y, predicted)
print(f'Mean Squared Error: {mse:.5f}')
Mean Squared Error: 0.00049

Common Pitfall: Missing Engagement Values#

  • Sometimes engagement data is missing or set to zero by accident.
  • Let us see how to handle missing or outlier values safely.
# Simulate missing data: set some likes/comments/shares to np.nan
df_posts_missing = df_posts.copy()
mask = np.random.rand(len(df_posts_missing)) < 0.05
df_posts_missing.loc[mask, ['likes', 'comments', 'shares']] = np.nan
missing_summary = df_posts_missing[['likes', 'comments', 'shares']].isnull().sum()
print('Missing counts by column:')
print(missing_summary)
Missing counts by column:
likes       24
comments    24
shares      24
dtype: int64
# Filling missing values before calculating engagement rate
df_filled = df_posts_missing.fillna(0)
df_filled['engagement_rate'] = (df_filled['likes'] + df_filled['comments'] + df_filled['shares']) / df_filled['views']
print('Engagement rates after filling missing values:')
print(df_filled['engagement_rate'].head())
Engagement rates after filling missing values:
0    0.171438
1    0.141667
2    0.103666
3    0.071873
4    0.106363
Name: engagement_rate, dtype: float64

Error Example: Misinterpreting Ratios like CTR#

  • Mixing up ratios such as CTR (click-through rate) or engagement rate can lead to bad decisions.
  • Always check which value is the numerator and which is the denominator!
# Incorrect CTR calculation example (wrong denominator!)
df_yt['ctr_wrong'] = df_yt['likes'] / df_yt['views']
print('First 3 wrong CTRs:', df_yt['ctr_wrong'].head(3).values)

# Correct CTR is already provided (percentage of impressions that resulted in a click)
First 3 wrong CTRs: [0.06533537 0.01873516 0.06559318]

Best Practices: Social Media Analytics and Machine Learning#

  • Benchmark content performance within types (compare your videos to similar ones).
  • Segment your audience and content (by topic, time, or platform).
  • Measure growth and trends over time, not just top results.
  • Use data to refine your posting schedule and content format.
  • Always define metrics (engagement, CTR, watch time) clearly and consistently.

End-to-End Example: Data-Driven Content Strategy#

  • Let us put all the pieces together in one workflow.
  • We will find your highest-engagement platform and best post type, then make a content recommendation.
  • Great analysts move from data > insight > recommendation.
# 1. Platform with highest average engagement
best_platform = eng_by_platform.idxmax()
best_rate = eng_by_platform.max()
# 2. On that platform, find the most engaging post (by engagement rate)
best_row = df_posts[df_posts['platform']==best_platform].sort_values('engagement_rate', ascending=False).iloc[0]
print(f'Best Platform: {best_platform}\nAverage Engagement Rate: {best_rate:.4f}')
print('Best Post Example:')
print(best_row[['date', 'views', 'likes', 'comments', 'shares', 'engagement_rate']])
Best Platform: Instagram
Average Engagement Rate: 0.1304
Best Post Example:
date               2023-03-29 18:00:00
views                            74643
likes                            10653
comments                          3569
shares                            2036
engagement_rate                0.21781
Name: 351, dtype: object

YouTube Call To Action#

  • Subscribe to the channel to learn more about data-driven social media analysis!

Quick Recap and Practice Prompt#

  • In this lesson, you have learned to analyze social media content data using Python and basic machine learning.
  • You have explored engagement patterns, posting strategy, and error traps.
  • Your practice: Try applying these steps to a new dataset (YouTube Analytics or Instagram Insights) and write down three data-driven recommendations.
  • For bonus points, plot a new metric or try different models!
  • Like this walkthrough? Give it a thumbs-up!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.