Mathew K Analytics

Lesson 41 · Social Media Content Analytics

Introduction to Social Media APIs

In this lesson, we will explore how to access and analyze social media data using APIs. You will learn what kind of insights are possible with public and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Introduction to Social Media APIs#

  • In this lesson, we will explore how to access and analyze social media data using APIs.
  • You will learn what kind of insights are possible with public and programmatic content data.
  • Understanding APIs enables content creators and businesses to make data-driven decisions.
  • By the end of this lesson, you will be able to fetch trending content, measure engagement, and spot actionable trends.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
import os, pickle
from pathlib import Path
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request

What are Social Media APIs and Content Datasets?#

  • APIs are software interfaces that let you request structured data from platforms like YouTube or Instagram.
  • Using APIs, you can programmatically download data about videos, posts, comments, and user engagement.
  • Social media datasets typically include metrics such as views, likes, comments, shares, and click-through rates.
  • For example, YouTube offers live trending video lists and detailed analytics via its API.
  • Mistakes can occur if you do not understand the meaning behind each metric or misuse data across platforms.
# BEGINNER EXAMPLE 1: Fetch public YouTube trending videos using the API if possible, falling back to synthetic data
SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
def get_yt_service():
    api_key = os.environ.get('YOUTUBE_API_KEY')
    if api_key:
        return build('youtube', 'v3', developerKey=api_key)
    if Path('client_secret.json').exists():
        creds = None
        if Path('token_ro.pickle').exists():
            with open('token_ro.pickle', 'rb') as f:
                creds = pickle.load(f)
        if not creds or not creds.valid:
            if creds and creds.expired and creds.refresh_token:
                creds.refresh(Request())
            else:
                flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
                creds = flow.run_local_server(port=0)
            with open('token_ro.pickle', 'wb') as f:
                pickle.dump(creds, f)
        return build('youtube', 'v3', credentials=creds)
    raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
def fetch_yt_trending(max_results=200, region='US'):
    youtube = get_yt_service()
    records, token = [], None
    while len(records) < max_results:
        resp = youtube.videos().list(
            part='snippet,statistics',
            chart='mostPopular',
            regionCode=region,
            maxResults=min(50, max_results - len(records)),
            pageToken=token
        ).execute()
        for item in resp.get('items', []):
            s = item['snippet']; st = item.get('statistics', {})
            records.append({
                'video_id':      item['id'],
                'trending_date': pd.Timestamp.today().date(),
                'title':         s.get('title', ''),
                'channel_title': s.get('channelTitle', ''),
                'category_id':   s.get('categoryId', ''),
                'views':         int(st.get('viewCount', 0)),
                'likes':         int(st.get('likeCount', 0)),
                'comment_count': int(st.get('commentCount', 0)),
            })
        token = resp.get('nextPageToken')
        if not token: break
    return pd.DataFrame(records)
try:
    df = fetch_yt_trending()
    print('Live trending data:', df.shape)
except Exception as e:
    print(f'Falling back to synthetic: {e}')
    np.random.seed(42)
    n = 1000
    df = pd.DataFrame({
        'video_id':      [f'vid{i}' for i in range(n)],
        'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
        'title':         [f'Video Title {i}' for i in range(n)],
        'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
        'category_id':   np.random.choice([1,2,10,22,24,28], n),
        'views':         np.random.randint(10000, 5000000, n),
        'likes':         np.random.randint(100, 200000, n),
        'comment_count': np.random.randint(10, 50000, n),
    })
    print('Synthetic fallback:', df.shape)
print(df.head(3))
Live trending data: (199, 8)
      video_id trending_date  \
0  82-jTNka3uc    2026-06-02   
1  3oB9AxspVow    2026-06-02   
2  l-cyT28MFyk    2026-06-02   

                                               title     channel_title  \
0  Ariana Grande - hate that i made you love me (...  ArianaGrandeVevo   
1           The End of Oak Street | Official Trailer      Warner Bros.   
2        PLAYING HARDCORE MINECRAFT UNTIL WE BEAT IT            Jynxzi   

  category_id    views   likes  comment_count  
0          10  1917025  365018          25091  
1           1  1129769   28191           2410  
2          24   581580   14533            231  
# BEGINNER EXAMPLE 2: Load a synthetic cross-platform content dataset
np.random.seed(42)
n_posts = 500
views = np.random.randint(100, 100000, n_posts)
df_content = pd.DataFrame({
    'post_id': range(1, n_posts+1),
    'platform': np.random.choice(['YouTube','Instagram','TikTok'], n_posts),
    'date': pd.date_range('2023-01-01', periods=n_posts, freq='6h'),
    'views': views,
    'likes': (views * np.random.uniform(0.02, 0.15, n_posts)).astype(int),
    'comments': (views * np.random.uniform(0.001, 0.05, n_posts)).astype(int),
    'shares': (views * np.random.uniform(0.001, 0.03, n_posts)).astype(int)
})
print(df_content.shape)
print(df_content.head(3))
(500, 7)
   post_id   platform                date  views  likes  comments  shares
0        1  Instagram 2023-01-01 00:00:00  15895   2356       116     253
1        2  Instagram 2023-01-01 06:00:00    960     94        16      26
2        3  Instagram 2023-01-01 12:00:00  76920   3910      2879    1185
# BEGINNER EXAMPLE 3: Calculate a simple engagement rate per post
df_content['engagement_rate'] = ((df_content['likes'] + df_content['comments'] + df_content['shares']) / df_content['views']) * 100
df_content['engagement_rate'] = df_content['engagement_rate'].round(2)
df_content[['platform','views','likes','comments','shares','engagement_rate']].head(5)
platform views likes comments shares engagement_rate
0 Instagram 15895 2356 116 253 17.14
1 Instagram 960 94 16 26 14.17
2 Instagram 76920 3910 2879 1185 10.37
3 YouTube 54986 1827 488 1637 7.19
4 Instagram 6365 253 261 163 10.64

Core Social Media Metrics: Definitions and Pitfalls#

  • Views show exposure, but do not guarantee active interest.
  • Likes, comments, and shares measure different types of audience engagement.
  • The engagement rate helps normalize post performance across audiences.
  • Watch time and click-through rate (CTR) reveal audience attention and interaction quality.
  • A common mistake is comparing raw metrics across platforms without normalizing.
# BEGINNER EXAMPLE 4: Identify the top 3 posts with highest engagement rate
top_posts = df_content.sort_values('engagement_rate', ascending=False).head(3)
print('Top 3 most engaging posts:')
print(top_posts[['platform','views','likes','comments','shares','engagement_rate']])
Top 3 most engaging posts:
      platform  views  likes  comments  shares  engagement_rate
351  Instagram  74643  10653      3569    2036            21.78
10   Instagram  16123   2202       791     426            21.21
383     TikTok  36731   4965      1769     952            20.93
# BEGINNER EXAMPLE 5: Compare average engagement by platform
platform_avg_engagement = df_content.groupby('platform')['engagement_rate'].mean().round(2).sort_values(ascending=False)
print('Average engagement rate by platform:')
print(platform_avg_engagement)
Average engagement rate by platform:
platform
Instagram    13.04
TikTok       12.69
YouTube      12.16
Name: engagement_rate, dtype: float64
# BEGINNER EXAMPLE 6: Visualize the distribution of engagement rates
import matplotlib.pyplot as plt
plt.figure(figsize=(8,5))
df_content['engagement_rate'].hist(bins=25, color='skyblue')
plt.title('Distribution of Engagement Rates (%)')
plt.xlabel('Engagement Rate (%)')
plt.ylabel('Number of Posts')
plt.grid(False)
plt.tight_layout()
plt.show()
No description has been provided for this image
# INTERMEDIATE EXAMPLE 1: Load and examine YouTube analytics data with watch time and CTR
np.random.seed(42)
n_videos = 300
views = np.random.randint(100, 500000, n_videos)
df_analytics = pd.DataFrame({
    'video_id': range(1, n_videos+1),
    'publish_date': pd.date_range('2022-01-01', periods=n_videos, freq='D'),
    'views': views,
    'watch_time': np.random.randint(1000, 500000, n_videos),
    'likes': (views * np.random.uniform(0.01, 0.08, n_videos)).astype(int),
    'comments': (views * np.random.uniform(0.001, 0.02, n_videos)).astype(int),
    'ctr': np.round(np.random.uniform(2, 10, n_videos), 2)
})
print(df_analytics.shape)
print(df_analytics.head(3))
(300, 7)
   video_id publish_date   views  watch_time  likes  comments   ctr
0         1   2022-01-01  122058      158381   1890      1975  4.14
1         2   2022-01-02  146967      481671   1730      1900  6.99
2         3   2022-01-03  132032      195806  10217       337  5.28
# INTERMEDIATE EXAMPLE 2: Find videos with highest click-through rate (CTR)
top_ctr = df_analytics.sort_values('ctr', ascending=False).head(5)
print('Top 5 videos by click-through rate:')
print(top_ctr[['video_id','publish_date','views','ctr']])
Top 5 videos by click-through rate:
     video_id publish_date   views   ctr
298       299   2022-10-26  125757  9.98
143       144   2022-05-24  164331  9.98
193       194   2022-07-13  245410  9.96
177       178   2022-06-27  158438  9.93
163       164   2022-06-13   77473  9.90
# INTERMEDIATE EXAMPLE 3: Compare watch time by upload month
df_analytics['month'] = df_analytics['publish_date'].dt.to_period('M')
monthly_watch = df_analytics.groupby('month')['watch_time'].sum()
monthly_watch.plot(kind='bar', color='seagreen', figsize=(10,4), title='Total Watch Time by Month')
plt.xlabel('Month')
plt.ylabel('Total Watch Time (minutes)')
plt.tight_layout()
plt.show()
No description has been provided for this image
# INTERMEDIATE EXAMPLE 4: Detect months with above-average CTR
mean_ctr = df_analytics['ctr'].mean()
good_months = df_analytics.groupby('month')['ctr'].mean() > mean_ctr
print('Months with average CTR above dataset mean:')
print(good_months[good_months].index.tolist())
Months with average CTR above dataset mean:
[Period('2022-04', 'M'), Period('2022-05', 'M'), Period('2022-08', 'M'), Period('2022-09', 'M')]
# INTERMEDIATE EXAMPLE 5: Identify videos with more than twice average engagement
avg_likes = df_analytics['likes'].mean()
above2x = df_analytics[df_analytics['likes'] > 2 * avg_likes]
print(f'Videos with >2x average likes: {len(above2x)} out of {len(df_analytics)}')
Videos with >2x average likes: 43 out of 300
# INTERMEDIATE EXAMPLE 6: Explore correlation between CTR and watch time
corr = df_analytics[['ctr','watch_time']].corr().iloc[0,1]
print(f'Correlation between CTR and watch time: {corr:.2f}')
Correlation between CTR and watch time: -0.04
# ADVANCED EXAMPLE 1: Extract and rank content categories from trending videos
top_cats = df['category_id'].value_counts().head(5)
print('Most common content categories among trending videos:')
print(top_cats)
Most common content categories among trending videos:
category_id
20    147
10     30
24     11
1       4
22      4
Name: count, dtype: int64
# ADVANCED EXAMPLE 2: Time series trend of engagement rate for one year
np.random.seed(42)
n_days = 365
dates = pd.date_range('2023-01-01', periods=n_days, freq='D')
views = np.random.randint(1000, 50000, n_days)
likes = (views * np.random.uniform(0.03, 0.12, n_days)).astype(int)
df_ts = pd.DataFrame({
    'date': dates,
    'views': views,
    'likes': likes,
    'engagement_rate': np.round(likes / views * 100, 2)
})
plt.figure(figsize=(12,4))
plt.plot(df_ts['date'], df_ts['engagement_rate'], color='purple')
plt.title('Engagement Rate Trend Over One Year')
plt.ylabel('Engagement Rate (%)')
plt.xlabel('Date')
plt.tight_layout()
plt.show()
No description has been provided for this image
# ADVANCED EXAMPLE 3: Find best day-of-week to post for highest engagement rate
df_ts['dayofweek'] = df_ts['date'].dt.day_name()
dow_avg = df_ts.groupby('dayofweek')['engagement_rate'].mean().sort_values(ascending=False)
print('Average engagement rate by day of week:')
print(dow_avg)
Average engagement rate by day of week:
dayofweek
Friday       8.030192
Sunday       8.010943
Saturday     7.794423
Tuesday      7.784038
Thursday     7.315769
Monday       7.231731
Wednesday    6.723077
Name: engagement_rate, dtype: float64
# ERROR HANDLING EXAMPLE 1: Handling missing engagement values
df_content_missing = df_content.copy()
np.random.seed(42)
mask = np.random.rand(len(df_content_missing)) < 0.05
df_content_missing.loc[mask, 'likes'] = np.nan
missing_count = df_content_missing['likes'].isna().sum()
print(f'There are {missing_count} posts with missing like counts.')
There are 31 posts with missing like counts.
# ERROR HANDLING EXAMPLE 2: Incorrect aggregation - median vs. mean engagement
mean_eng = df_content['engagement_rate'].mean().round(2)
median_eng = df_content['engagement_rate'].median().round(2)
print(f'Mean engagement rate: {mean_eng}%')
print(f'Median engagement rate: {median_eng}%')
Mean engagement rate: 12.61%
Median engagement rate: 12.7%
# ERROR HANDLING EXAMPLE 3: Misinterpreted CTR as percent instead of ratio
df_analytics['ctr_ratio'] = df_analytics['ctr'] / 100
err_example = df_analytics[['ctr','ctr_ratio']].head(3)
print('Demonstrating correct conversion of CTR percent to decimal ratio:')
print(err_example)
Demonstrating correct conversion of CTR percent to decimal ratio:
    ctr  ctr_ratio
0  4.14     0.0414
1  6.99     0.0699
2  5.28     0.0528
# ERROR HANDLING EXAMPLE 4: Wrong grouping logic for viral content detection
top_channels = df.groupby('channel_title')['views'].sum().sort_values(ascending=False).head(3)
print('Total views by channel (sum across videos, not per video):')
print(top_channels)
Total views by channel (sum across videos, not per video):
channel_title
TREASURE (트레저)    3634391
Wemmbu            3619594
MoreSidemen       2628249
Name: views, dtype: int64

Best Practices and Common Analytics Patterns#

  • Benchmark performance using consistent metrics so your comparisons are fair.
  • Segment your audience to find top fans and periodic trends.
  • Use timelines to detect growth (or decline) in content engagement.
  • Compare across platforms only after normalizing scales and audience size.
  • Always check for missing values and unusual outliers before drawing conclusions.
# TINY END-TO-END EXAMPLE: Recommend a social media content strategy
high_eng_posts = df_content[df_content['engagement_rate'] > df_content['engagement_rate'].mean()]
most_pop_platform = high_eng_posts['platform'].mode()[0]
best_time = high_eng_posts['date'].dt.hour.mode()[0]
print('Content Strategy Recommendation:')
print(f'Focus new posts on {most_pop_platform} and schedule around {best_time}:00 for higher engagement.')
Content Strategy Recommendation:
Focus new posts on YouTube and schedule around 12:00 for higher engagement.
print('You have completed the API-driven Social Media Analytics lesson!')
You have completed the API-driven Social Media Analytics lesson!
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.