Lesson 47 · Social Media Content Analytics
Analyzing Growth in Views and Subscribers
In this lesson, we will explore how to analyze the growth of views and subscribers using real social media datasets. Tracking view and subscriber growth is…
- CourseSocial Media Content Analytics
- Lesson47 of 41
- Video29 min
- FormatJupyter notebook · 22 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbAnalyzing Growth in Views and Subscribers#
- In this lesson, we will explore how to analyze the growth of views and subscribers using real social media datasets.
- Tracking view and subscriber growth is crucial for creators and businesses to understand their audience and improve content strategy.
- You will learn to calculate growth trends, detect patterns, and make data-driven recommendations for content optimization.
- By the end, you will be able to use engagement data to support smarter decisions about what and when to post.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')
Understanding Social Media Analytics Concepts#
- Social media datasets represent posts, videos, accounts, and audience interactions.
- Key metrics include views (how many times content was seen), likes, comments, watch time, and click-through rate (CTR).
- Subscriber counts reflect audience growth and loyalty.
- Engagement rate measures how actively viewers interact with your content.
- Beginners often mistake high views for success without considering engagement or growth over time.
# We will use the Social Media Time Series Dataset for simple trend analysis
np.random.seed(42)
n_days = 365
dates = pd.date_range('2023-01-01', periods=n_days, freq='D')
views = np.random.randint(1000, 50000, n_days)
likes = (views * np.random.uniform(0.03, 0.12, n_days)).astype(int)
subscribers = 1000 + np.cumsum(np.random.randint(0, 50, n_days))
ts_df = pd.DataFrame({
'date': dates,
'views': views,
'likes': likes,
'subscribers': subscribers
})
print(ts_df.head(3))
# Basic: Plot views and subscribers over time to see growth visually
fig, ax1 = plt.subplots(figsize=(10, 5))
color = 'tab:blue'
ax1.set_xlabel('Date')
ax1.set_ylabel('Views', color=color)
ax1.plot(ts_df['date'], ts_df['views'], color=color, label='Views')
ax1.tick_params(axis='y', labelcolor=color)
ax2 = ax1.twinx()
color = 'tab:red'
ax2.set_ylabel('Subscribers', color=color)
ax2.plot(ts_df['date'], ts_df['subscribers'], color=color, label='Subscribers')
ax2.tick_params(axis='y', labelcolor=color)
plt.title('Daily Views and Cumulative Subscribers Over Time')
fig.tight_layout()
plt.show()
# Basic: Calculate daily percent growth in views and subscribers
ts_df['views_pct_change'] = ts_df['views'].pct_change().fillna(0) * 100
ts_df['subs_pct_change'] = ts_df['subscribers'].pct_change().fillna(0) * 100
print(ts_df[['date', 'views_pct_change', 'subs_pct_change']].head(7))
# Basic: Rolling average to smooth trends
ts_df['views_rolling'] = ts_df['views'].rolling(window=7, min_periods=1).mean()
ts_df['subs_rolling'] = ts_df['subscribers'].rolling(window=7, min_periods=1).mean()
plt.figure(figsize=(10,4))
plt.plot(ts_df['date'], ts_df['views_rolling'], label='7-day Avg Views')
plt.plot(ts_df['date'], ts_df['subs_rolling'], label='7-day Avg Subscribers')
plt.legend()
plt.title('Smoothed (7-day) Averages for Views and Subscribers')
plt.xlabel('Date')
plt.ylabel('Counts')
plt.show()
# Basic: Correlation between daily views and daily subscriber gain
ts_df['daily_subs_gain'] = ts_df['subscribers'].diff().fillna(0)
correlation = ts_df['views'].corr(ts_df['daily_subs_gain'])
print(f'Correlation between daily views and new subscribers: {correlation:.2f}')
# Intermediate: Identify top 10 days with highest view growth
top_growth = ts_df.sort_values('views_pct_change', ascending=False).head(10)
print(top_growth[['date', 'views', 'views_pct_change', 'subs_pct_change']])
# Intermediate: Calculate overall view and subscriber growth rates
total_views_growth = 100 * (ts_df['views'].iloc[-1] - ts_df['views'].iloc[0]) / ts_df['views'].iloc[0]
total_subs_growth = 100 * (ts_df['subscribers'].iloc[-1] - ts_df['subscribers'].iloc[0]) / ts_df['subscribers'].iloc[0]
print(f'Total views growth over the year: {total_views_growth:.1f}%')
print(f'Total subscriber growth over the year: {total_subs_growth:.1f}%')
# Intermediate: Weekly gain patterns
ts_df['week'] = ts_df['date'].dt.isocalendar().week
weekly_growth = ts_df.groupby('week').agg({'views': 'sum', 'daily_subs_gain': 'sum'})
weekly_growth['subs_per_1000_views'] = 1000 * weekly_growth['daily_subs_gain'] / weekly_growth['views']
plt.figure(figsize=(10,4))
plt.plot(weekly_growth.index, weekly_growth['subs_per_1000_views'])
plt.title('Weekly Subscribers Gained per 1,000 Views')
plt.xlabel('Week Number')
plt.ylabel('Subscribers per 1,000 Views')
plt.show()
# Intermediate: Highlight periods with negative growth
negative_growth = ts_df[ts_df['views_pct_change'] < 0]
print('Periods with negative daily view growth:')
print(negative_growth[['date', 'views', 'views_pct_change', 'subs_pct_change']].head())
# Advanced: Use YouTube Trending Videos Dataset to compare channel growth
try:
from pathlib import Path
import os, pickle
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request
SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
def get_yt_service():
api_key = os.environ.get('YOUTUBE_API_KEY')
if api_key:
return build('youtube', 'v3', developerKey=api_key)
if Path('client_secret.json').exists():
creds = None
if Path('token_ro.pickle').exists():
with open('token_ro.pickle', 'rb') as f:
creds = pickle.load(f)
if not creds or not creds.valid:
if creds and creds.expired and creds.refresh_token:
creds.refresh(Request())
else:
flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
creds = flow.run_local_server(port=0)
with open('token_ro.pickle', 'wb') as f:
pickle.dump(creds, f)
return build('youtube', 'v3', credentials=creds)
raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
def fetch_yt_trending(max_results=200, region='US'):
youtube = get_yt_service()
records, token = [], None
while len(records) < max_results:
resp = youtube.videos().list(
part='snippet,statistics',
chart='mostPopular',
regionCode=region,
maxResults=min(50, max_results - len(records)),
pageToken=token
).execute()
for item in resp.get('items', []):
s = item['snippet']; st = item.get('statistics', {})
records.append({
'video_id': item['id'],
'trending_date': pd.Timestamp.today().date(),
'title': s.get('title', ''),
'channel_title': s.get('channelTitle', ''),
'category_id': s.get('categoryId', ''),
'views': int(st.get('viewCount', 0)),
'likes': int(st.get('likeCount', 0)),
'comment_count': int(st.get('commentCount', 0)),
})
token = resp.get('nextPageToken')
if not token: break
return pd.DataFrame(records)
yt_df = fetch_yt_trending()
print('Live trending data:', yt_df.shape)
except Exception as e:
print(f'Falling back to synthetic: {e}')
np.random.seed(42)
n = 1000
yt_df = pd.DataFrame({
'video_id': [f'vid{i}' for i in range(n)],
'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
'title': [f'Video Title {i}' for i in range(n)],
'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
'category_id': np.random.choice([1,2,10,22,24,28], n),
'views': np.random.randint(10000, 5000000, n),
'likes': np.random.randint(100, 200000, n),
'comment_count': np.random.randint(10, 50000, n),
})
print('Synthetic fallback:', yt_df.shape)
print(yt_df.head(3))
# Advanced: Which channel is growing fastest in trending content?
growth = (yt_df.groupby('channel_title')['views'].sum().sort_values(ascending=False))
print('Channels by total trending video views:')
print(growth.head(5))
# Advanced: Identify fastest-growing trending video
yt_df['trending_rank'] = yt_df['views'].rank(ascending=False, method='min')
fastest_trend = yt_df.loc[yt_df['trending_rank'] == 1]
print(fastest_trend[['title', 'channel_title', 'views', 'likes']])
# Advanced: Calculate engagement rate for each trending channel
yt_df['engagement_rate'] = (yt_df['likes'] + yt_df['comment_count']) / yt_df['views'] * 100
ch_avg = yt_df.groupby('channel_title')['engagement_rate'].mean().sort_values(ascending=False)
print('Highest engagement rates by channel:')
print(ch_avg.head(5))
# Error Handling: What happens if data has missing values?
yt_df_missing = yt_df.copy()
yt_df_missing.loc[yt_df_missing.sample(frac=0.02, random_state=42).index, 'views'] = np.nan
missing_count = yt_df_missing['views'].isnull().sum()
print(f'Artificially introduced missing view values: {missing_count}')
yt_df_missing['views_filled'] = yt_df_missing['views'].fillna(yt_df_missing['views'].median())
print('Any remaining NA:', yt_df_missing['views_filled'].isnull().any())
# Error Handling: Wrong aggregationmixing up category.
cat_sum = yt_df.groupby('category_id')['views'].sum()
overall_sum = yt_df['views'].sum()
if np.isclose(cat_sum.sum(), overall_sum):
print('Aggregation sums correctly.')
else:
print('Aggregation mismatch detected!')
# Error Handling: Misinterpreting engagement ratios
yt_df['ctr'] = yt_df['likes'] / yt_df['views'] * 100
bad_ctrs = yt_df[yt_df['ctr'] > 100]
print('Any CTR > 100%?', not bad_ctrs.empty)
print(bad_ctrs[['title', 'likes', 'views', 'ctr']].head())
# Error Handling: Grouping by wrong content key
wrong_group = yt_df.groupby('video_id')['views'].sum().sum()
true_group = yt_df['views'].sum()
assert wrong_group == true_group, 'Grouping by video_id should not affect total sum.'
print('Grouping by video_id and summing is safe for total views in this dataset.')
Best Practices and Patterns in Content Analytics#
- Always benchmark performance against your historical growth and category averages.
- Segment your audience or content to reveal high-performing groups.
- Use rolling averages and normalizations to compare across time and channels.
- Clearly define each metric and use the same formulas for every analysis.
- Optimize your schedule and content style based on actual engagement data.
# Best Practice: Identify best and worst weekly periods
best_week = weekly_growth['subs_per_1000_views'].idxmax()
worst_week = weekly_growth['subs_per_1000_views'].idxmin()
print(f'Your best subscriber-per-view week: {best_week}')
print(f'Your worst subscriber-per-view week: {worst_week}')
# Pattern: Detect viral periods in the time series
viral = ts_df[ts_df['views_pct_change'] > ts_df['views_pct_change'].mean() + 2*ts_df['views_pct_change'].std()]
print(f'Number of viral view spikes: {viral.shape[0]}')
print(viral[['date', 'views', 'views_pct_change']].head())
# Pattern: Compare engagement rates week by week
ts_df['engagement_rate'] = ts_df['likes'] / ts_df['views'] * 100
weekly_engagement = ts_df.groupby('week')['engagement_rate'].mean()
plt.figure(figsize=(10,4))
plt.plot(weekly_engagement.index, weekly_engagement, '-o', color='purple')
plt.title('Weekly Average Engagement Rate (%)')
plt.xlabel('Week Number')
plt.ylabel('Engagement Rate (%)')
plt.show()
# End-to-End Example: Recommend a content strategy improvement
avg_engagement = ts_df['engagement_rate'].mean()
viral_weeks = weekly_engagement[weekly_engagement > avg_engagement].index.tolist()
if viral_weeks:
suggestion = f'Consider replicating your content or campaigns from weeks: {viral_weeks}.'
else:
suggestion = 'No above-average engagement weeks found. Try improving titles, thumbnails, or posting times.'
print('Strategy Recommendation:')
print(suggestion)
Congratulations, You Have Explored Growth Analytics!#
- You now know how to measure, plot, and interpret view and subscriber growth.
- Keep practicing by analyzing different segments or new data.
- For more lessons, subscribe to our YouTube channel and stay updated!
- Share your most interesting insight or visualization with us in the comments!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



