Mathew K Analytics

Lesson 47 · Social Media Content Analytics

Analyzing Growth in Views and Subscribers

In this lesson, we will explore how to analyze the growth of views and subscribers using real social media datasets. Tracking view and subscriber growth is…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Analyzing Growth in Views and Subscribers#

  • In this lesson, we will explore how to analyze the growth of views and subscribers using real social media datasets.
  • Tracking view and subscriber growth is crucial for creators and businesses to understand their audience and improve content strategy.
  • You will learn to calculate growth trends, detect patterns, and make data-driven recommendations for content optimization.
  • By the end, you will be able to use engagement data to support smarter decisions about what and when to post.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')

Understanding Social Media Analytics Concepts#

  • Social media datasets represent posts, videos, accounts, and audience interactions.
  • Key metrics include views (how many times content was seen), likes, comments, watch time, and click-through rate (CTR).
  • Subscriber counts reflect audience growth and loyalty.
  • Engagement rate measures how actively viewers interact with your content.
  • Beginners often mistake high views for success without considering engagement or growth over time.
# We will use the Social Media Time Series Dataset for simple trend analysis
np.random.seed(42)
n_days = 365
dates = pd.date_range('2023-01-01', periods=n_days, freq='D')
views = np.random.randint(1000, 50000, n_days)
likes = (views * np.random.uniform(0.03, 0.12, n_days)).astype(int)
subscribers = 1000 + np.cumsum(np.random.randint(0, 50, n_days))
ts_df = pd.DataFrame({
    'date': dates,
    'views': views,
    'likes': likes,
    'subscribers': subscribers
})
print(ts_df.head(3))
        date  views  likes  subscribers
0 2023-01-01  16795   1790         1025
1 2023-01-02   1860    108         1060
2 2023-01-03  39158   1772         1060
# Basic: Plot views and subscribers over time to see growth visually
fig, ax1 = plt.subplots(figsize=(10, 5))
color = 'tab:blue'
ax1.set_xlabel('Date')
ax1.set_ylabel('Views', color=color)
ax1.plot(ts_df['date'], ts_df['views'], color=color, label='Views')
ax1.tick_params(axis='y', labelcolor=color)

ax2 = ax1.twinx()
color = 'tab:red'
ax2.set_ylabel('Subscribers', color=color)
ax2.plot(ts_df['date'], ts_df['subscribers'], color=color, label='Subscribers')
ax2.tick_params(axis='y', labelcolor=color)
plt.title('Daily Views and Cumulative Subscribers Over Time')
fig.tight_layout()
plt.show()
No description has been provided for this image
# Basic: Calculate daily percent growth in views and subscribers
ts_df['views_pct_change'] = ts_df['views'].pct_change().fillna(0) * 100
ts_df['subs_pct_change'] = ts_df['subscribers'].pct_change().fillna(0) * 100
print(ts_df[['date', 'views_pct_change', 'subs_pct_change']].head(7))
        date  views_pct_change  subs_pct_change
0 2023-01-01          0.000000         0.000000
1 2023-01-02        -88.925275         3.414634
2 2023-01-03       2005.268817         0.000000
3 2023-01-04         16.788396         0.660377
4 2023-01-05        -73.139159         4.498594
5 2023-01-06        -40.858027         3.049327
6 2023-01-07        145.698555         1.218451
# Basic: Rolling average to smooth trends
ts_df['views_rolling'] = ts_df['views'].rolling(window=7, min_periods=1).mean()
ts_df['subs_rolling'] = ts_df['subscribers'].rolling(window=7, min_periods=1).mean()
plt.figure(figsize=(10,4))
plt.plot(ts_df['date'], ts_df['views_rolling'], label='7-day Avg Views')
plt.plot(ts_df['date'], ts_df['subs_rolling'], label='7-day Avg Subscribers')
plt.legend()
plt.title('Smoothed (7-day) Averages for Views and Subscribers')
plt.xlabel('Date')
plt.ylabel('Counts')
plt.show()
No description has been provided for this image
# Basic: Correlation between daily views and daily subscriber gain
ts_df['daily_subs_gain'] = ts_df['subscribers'].diff().fillna(0)
correlation = ts_df['views'].corr(ts_df['daily_subs_gain'])
print(f'Correlation between daily views and new subscribers: {correlation:.2f}')
Correlation between daily views and new subscribers: -0.02
# Intermediate: Identify top 10 days with highest view growth
top_growth = ts_df.sort_values('views_pct_change', ascending=False).head(10)
print(top_growth[['date', 'views', 'views_pct_change', 'subs_pct_change']])
          date  views  views_pct_change  subs_pct_change
360 2023-12-27  46236       3453.881630         0.490798
73  2023-03-15  38065       3178.639104         0.710179
189 2023-07-09  35766       2050.691521         0.311812
2   2023-01-03  39158       2005.268817         0.000000
227 2023-08-16  24509       1947.535505         0.559701
138 2023-05-19  22518       1767.164179         1.000931
30  2023-01-31  20118       1592.010093         0.787845
299 2023-10-27  47940       1398.125000         0.474741
34  2023-02-04  32551       1335.862373         0.479233
209 2023-07-29  25548       1277.993528         0.635133
# Intermediate: Calculate overall view and subscriber growth rates
total_views_growth = 100 * (ts_df['views'].iloc[-1] - ts_df['views'].iloc[0]) / ts_df['views'].iloc[0]
total_subs_growth = 100 * (ts_df['subscribers'].iloc[-1] - ts_df['subscribers'].iloc[0]) / ts_df['subscribers'].iloc[0]
print(f'Total views growth over the year: {total_views_growth:.1f}%')
print(f'Total subscriber growth over the year: {total_subs_growth:.1f}%')
Total views growth over the year: 118.8%
Total subscriber growth over the year: 863.6%
# Intermediate: Weekly gain patterns
ts_df['week'] = ts_df['date'].dt.isocalendar().week
weekly_growth = ts_df.groupby('week').agg({'views': 'sum', 'daily_subs_gain': 'sum'})
weekly_growth['subs_per_1000_views'] = 1000 * weekly_growth['daily_subs_gain'] / weekly_growth['views']
plt.figure(figsize=(10,4))
plt.plot(weekly_growth.index, weekly_growth['subs_per_1000_views'])
plt.title('Weekly Subscribers Gained per 1,000 Views')
plt.xlabel('Week Number')
plt.ylabel('Subscribers per 1,000 Views')
plt.show()
No description has been provided for this image
# Intermediate: Highlight periods with negative growth
negative_growth = ts_df[ts_df['views_pct_change'] < 0]
print('Periods with negative daily view growth:')
print(negative_growth[['date', 'views', 'views_pct_change', 'subs_pct_change']].head())
Periods with negative daily view growth:
         date  views  views_pct_change  subs_pct_change
1  2023-01-02   1860        -88.925275         3.414634
4  2023-01-05  12284        -73.139159         4.498594
5  2023-01-06   7265        -40.858027         3.049327
8  2023-01-09  22962        -39.880610         1.736973
10 2023-01-11  45131         -6.349733         2.011263
# Advanced: Use YouTube Trending Videos Dataset to compare channel growth
try:
    from pathlib import Path
    import os, pickle
    from googleapiclient.discovery import build
    from google_auth_oauthlib.flow import InstalledAppFlow
    from google.auth.transport.requests import Request
    SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
    def get_yt_service():
        api_key = os.environ.get('YOUTUBE_API_KEY')
        if api_key:
            return build('youtube', 'v3', developerKey=api_key)
        if Path('client_secret.json').exists():
            creds = None
            if Path('token_ro.pickle').exists():
                with open('token_ro.pickle', 'rb') as f:
                    creds = pickle.load(f)
            if not creds or not creds.valid:
                if creds and creds.expired and creds.refresh_token:
                    creds.refresh(Request())
                else:
                    flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
                    creds = flow.run_local_server(port=0)
                with open('token_ro.pickle', 'wb') as f:
                    pickle.dump(creds, f)
            return build('youtube', 'v3', credentials=creds)
        raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
    def fetch_yt_trending(max_results=200, region='US'):
        youtube = get_yt_service()
        records, token = [], None
        while len(records) < max_results:
            resp = youtube.videos().list(
                part='snippet,statistics',
                chart='mostPopular',
                regionCode=region,
                maxResults=min(50, max_results - len(records)),
                pageToken=token
            ).execute()
            for item in resp.get('items', []):
                s = item['snippet']; st = item.get('statistics', {})
                records.append({
                    'video_id':      item['id'],
                    'trending_date': pd.Timestamp.today().date(),
                    'title':         s.get('title', ''),
                    'channel_title': s.get('channelTitle', ''),
                    'category_id':   s.get('categoryId', ''),
                    'views':         int(st.get('viewCount', 0)),
                    'likes':         int(st.get('likeCount', 0)),
                    'comment_count': int(st.get('commentCount', 0)),
                })
            token = resp.get('nextPageToken')
            if not token: break
        return pd.DataFrame(records)
    yt_df = fetch_yt_trending()
    print('Live trending data:', yt_df.shape)
except Exception as e:
    print(f'Falling back to synthetic: {e}')
    np.random.seed(42)
    n = 1000
    yt_df = pd.DataFrame({
        'video_id':      [f'vid{i}' for i in range(n)],
        'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
        'title':         [f'Video Title {i}' for i in range(n)],
        'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
        'category_id':   np.random.choice([1,2,10,22,24,28], n),
        'views':         np.random.randint(10000, 5000000, n),
        'likes':         np.random.randint(100, 200000, n),
        'comment_count': np.random.randint(10, 50000, n),
    })
    print('Synthetic fallback:', yt_df.shape)
print(yt_df.head(3))
Live trending data: (199, 8)
      video_id trending_date  \
0  82-jTNka3uc    2026-06-02   
1  3oB9AxspVow    2026-06-02   
2  l-cyT28MFyk    2026-06-02   

                                               title     channel_title  \
0  Ariana Grande - hate that i made you love me (...  ArianaGrandeVevo   
1           The End of Oak Street | Official Trailer      Warner Bros.   
2        PLAYING HARDCORE MINECRAFT UNTIL WE BEAT IT            Jynxzi   

  category_id    views   likes  comment_count  
0          10  2729045  434354          27307  
1           1  2757746   34107           2836  
2          24   734463   16149            360  
# Advanced: Which channel is growing fastest in trending content?
growth = (yt_df.groupby('channel_title')['views'].sum().sort_values(ascending=False))
print('Channels by total trending video views:')
print(growth.head(5))
Channels by total trending video views:
channel_title
TREASURE (트레저)    5074720
Wemmbu            3677608
IShowSpeed        2772610
Warner Bros.      2757746
MoreSidemen       2736827
Name: views, dtype: int64
# Advanced: Identify fastest-growing trending video
yt_df['trending_rank'] = yt_df['views'].rank(ascending=False, method='min')
fastest_trend = yt_df.loc[yt_df['trending_rank'] == 1]
print(fastest_trend[['title', 'channel_title', 'views', 'likes']])
                    title   channel_title    views   likes
38  TREASURE - ‘IF I’ M/V  TREASURE (트레저)  5074720  242064
# Advanced: Calculate engagement rate for each trending channel
yt_df['engagement_rate'] = (yt_df['likes'] + yt_df['comment_count']) / yt_df['views'] * 100
ch_avg = yt_df.groupby('channel_title')['engagement_rate'].mean().sort_values(ascending=False)
print('Highest engagement rates by channel:')
print(ch_avg.head(5))
Highest engagement rates by channel:
channel_title
div_y              32.593081
bleood             26.097920
KOT4Q              25.664404
Bill McClintock    23.538197
Ivycomb Music      23.409439
Name: engagement_rate, dtype: float64
# Error Handling: What happens if data has missing values?
yt_df_missing = yt_df.copy()
yt_df_missing.loc[yt_df_missing.sample(frac=0.02, random_state=42).index, 'views'] = np.nan
missing_count = yt_df_missing['views'].isnull().sum()
print(f'Artificially introduced missing view values: {missing_count}')
yt_df_missing['views_filled'] = yt_df_missing['views'].fillna(yt_df_missing['views'].median())
print('Any remaining NA:', yt_df_missing['views_filled'].isnull().any())
Artificially introduced missing view values: 4
Any remaining NA: False
# Error Handling: Wrong aggregationmixing up category.
cat_sum = yt_df.groupby('category_id')['views'].sum()
overall_sum = yt_df['views'].sum()
if np.isclose(cat_sum.sum(), overall_sum):
    print('Aggregation sums correctly.')
else:
    print('Aggregation mismatch detected!')
Aggregation sums correctly.
# Error Handling: Misinterpreting engagement ratios
yt_df['ctr'] = yt_df['likes'] / yt_df['views'] * 100
bad_ctrs = yt_df[yt_df['ctr'] > 100]
print('Any CTR > 100%?', not bad_ctrs.empty)
print(bad_ctrs[['title', 'likes', 'views', 'ctr']].head())
Any CTR > 100%? False
Empty DataFrame
Columns: [title, likes, views, ctr]
Index: []
# Error Handling: Grouping by wrong content key
wrong_group = yt_df.groupby('video_id')['views'].sum().sum()
true_group = yt_df['views'].sum()
assert wrong_group == true_group, 'Grouping by video_id should not affect total sum.'
print('Grouping by video_id and summing is safe for total views in this dataset.')
Grouping by video_id and summing is safe for total views in this dataset.

Best Practices and Patterns in Content Analytics#

  • Always benchmark performance against your historical growth and category averages.
  • Segment your audience or content to reveal high-performing groups.
  • Use rolling averages and normalizations to compare across time and channels.
  • Clearly define each metric and use the same formulas for every analysis.
  • Optimize your schedule and content style based on actual engagement data.
# Best Practice: Identify best and worst weekly periods
best_week = weekly_growth['subs_per_1000_views'].idxmax()
worst_week = weekly_growth['subs_per_1000_views'].idxmin()
print(f'Your best subscriber-per-view week: {best_week}')
print(f'Your worst subscriber-per-view week: {worst_week}')
Your best subscriber-per-view week: 41
Your worst subscriber-per-view week: 14
# Pattern: Detect viral periods in the time series
viral = ts_df[ts_df['views_pct_change'] > ts_df['views_pct_change'].mean() + 2*ts_df['views_pct_change'].std()]
print(f'Number of viral view spikes: {viral.shape[0]}')
print(viral[['date', 'views', 'views_pct_change']].head())
Number of viral view spikes: 20
         date  views  views_pct_change
2  2023-01-03  39158       2005.268817
30 2023-01-31  20118       1592.010093
34 2023-02-04  32551       1335.862373
73 2023-03-15  38065       3178.639104
82 2023-03-24  25253       1152.628968
# Pattern: Compare engagement rates week by week
ts_df['engagement_rate'] = ts_df['likes'] / ts_df['views'] * 100
weekly_engagement = ts_df.groupby('week')['engagement_rate'].mean()
plt.figure(figsize=(10,4))
plt.plot(weekly_engagement.index, weekly_engagement, '-o', color='purple')
plt.title('Weekly Average Engagement Rate (%)')
plt.xlabel('Week Number')
plt.ylabel('Engagement Rate (%)')
plt.show()
No description has been provided for this image
# End-to-End Example: Recommend a content strategy improvement
avg_engagement = ts_df['engagement_rate'].mean()
viral_weeks = weekly_engagement[weekly_engagement > avg_engagement].index.tolist()
if viral_weeks:
    suggestion = f'Consider replicating your content or campaigns from weeks: {viral_weeks}.'
else:
    suggestion = 'No above-average engagement weeks found. Try improving titles, thumbnails, or posting times.'
print('Strategy Recommendation:')
print(suggestion)
Strategy Recommendation:
Consider replicating your content or campaigns from weeks: [2, 3, 4, 8, 11, 16, 19, 20, 21, 22, 23, 24, 28, 31, 33, 34, 36, 37, 38, 39, 40, 41, 42, 43, 46, 47, 52].

Congratulations, You Have Explored Growth Analytics!#

  • You now know how to measure, plot, and interpret view and subscriber growth.
  • Keep practicing by analyzing different segments or new data.
  • For more lessons, subscribe to our YouTube channel and stay updated!
  • Share your most interesting insight or visualization with us in the comments!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.