Mathew K Analytics

Lesson 22 · Social Media Content Analytics

Analyzing Views and Watch Time Distribution

This lesson explores how to analyze the distribution of video views and total watch time using real social media analytics data. Understanding how your…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Analyzing Views and Watch Time Distribution#

  • This lesson explores how to analyze the distribution of video views and total watch time using real social media analytics data.

  • Understanding how your audience engages with content helps creators and businesses tailor strategies and improve performance.

  • Learners will use Python to uncover which videos drive both high view counts and watch times, identify top performers, and avoid common analysis mistakes.

  • You will work hands-on with trending video and analytics datasets to gain actionable insights.

import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Core Concepts: Views, Watch Time, and Engagement#

  • Social media datasets often track each video or post as a row, with engagement metrics (like views, likes, comments) as columns.
  • Views show how many times a piece of content was loaded, but not how long people watched.
  • Watch time measures total minutes audiences spend watching: it can reveal what is really capturing attention.
  • Beginners often focus only on high views, forgetting that short watch time may mean low impact.
  • Ratios like engagement rate or average watch time give a better sense of content effectiveness.
# Load the YouTube Analytics Dataset (synthetic fallback if needed)
np.random.seed(42)
n_videos = 300
views = np.random.randint(100, 500000, n_videos)
df = pd.DataFrame({
    'video_id': range(1, n_videos+1),
    'publish_date': pd.date_range('2022-01-01', periods=n_videos, freq='D'),
    'views': views,
    'watch_time': np.random.randint(1000, 500000, n_videos),
    'likes': (views * np.random.uniform(0.01, 0.08, n_videos)).astype(int),
    'comments': (views * np.random.uniform(0.001, 0.02, n_videos)).astype(int),
    'ctr': np.round(np.random.uniform(2, 10, n_videos), 2)
})
print(df.shape)
print(df.head(3))
(300, 7)
   video_id publish_date   views  watch_time  likes  comments   ctr
0         1   2022-01-01  122058      158381   1890      1975  4.14
1         2   2022-01-02  146967      481671   1730      1900  6.99
2         3   2022-01-03  132032      195806  10217       337  5.28
# Beginner Example: List videos with the highest view counts
top_views = df.nlargest(5, 'views')[['video_id', 'views', 'watch_time']]
print(top_views)
     video_id   views  watch_time
249       250  499146      125862
286       287  496357      204687
257       258  491414      285821
173       174  491334      296972
166       167  489670      306628
# Beginner Example: List videos with the highest total watch time
top_watch_time = df.nlargest(5, 'watch_time')[['video_id', 'views', 'watch_time']]
print(top_watch_time)
     video_id   views  watch_time
217       218   43685      499863
186       187  393522      495495
198       199  185440      494415
125       126  401887      493738
268       269  447700      486988
# Beginner Example: Compute the average views and watch time across all videos
avg_views = df['views'].mean()
avg_watch_time = df['watch_time'].mean()
print(f'Average views per video: {int(avg_views):,}')
print(f'Average total watch time per video: {int(avg_watch_time):,} minutes')
Average views per video: 253,436
Average total watch time per video: 247,090 minutes
# Beginner Example: Calculate average watch time PER VIEW (engagement depth)
df['avg_watch_time_per_view'] = df['watch_time'] / df['views']
print(df[['video_id', 'views', 'watch_time', 'avg_watch_time_per_view']].head(3))
   video_id   views  watch_time  avg_watch_time_per_view
0         1  122058      158381                 1.297588
1         2  146967      481671                 3.277409
2         3  132032      195806                 1.483019
# Beginner Example: Visualize the distribution of views
import matplotlib.pyplot as plt
plt.figure(figsize=(7,4))
plt.hist(df['views'], bins=30, color='skyblue', edgecolor='black')
plt.title('Distribution of Video Views')
plt.xlabel('Views')
plt.ylabel('Number of Videos')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Intermediate Example: Visualize total watch time distribution
plt.figure(figsize=(7,4))
plt.hist(df['watch_time'], bins=30, color='coral', edgecolor='black')
plt.title('Distribution of Video Total Watch Time (minutes)')
plt.xlabel('Total Watch Time (minutes)')
plt.ylabel('Number of Videos')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Intermediate Example: Find videos with above-average views and watch time
above_avg = df[(df['views'] > avg_views) & (df['watch_time'] > avg_watch_time)]
print(f'Number of videos above average in both views and watch time: {len(above_avg)}')
print(above_avg[['video_id', 'views', 'watch_time']].head(5))
Number of videos above average in both views and watch time: 72
    video_id   views  watch_time
11        12  430510      462079
13        14  374971      446299
19        20  329465      343767
21        22  263013      419400
22        23  321979      356528
# Intermediate Example: Analyze correlation between views and watch time
correlation = df['views'].corr(df['watch_time'])
print(f'Correlation between views and watch time: {correlation:.2f}')
Correlation between views and watch time: 0.05
# Intermediate Example: Flag videos with low watch time per view (shallow engagement)
low_engage = df[df['avg_watch_time_per_view'] < df['avg_watch_time_per_view'].quantile(0.25)]
print('Example of lower-engagement videos:')
print(low_engage[['video_id', 'views', 'watch_time', 'avg_watch_time_per_view']].head(5))
Example of lower-engagement videos:
    video_id   views  watch_time  avg_watch_time_per_view
3          4  365938       71467                 0.195298
10        11  475702      169229                 0.355746
14        15  388568      107081                 0.275579
16        17  191435       90045                 0.470369
17        18  278267       35698                 0.128287
# Intermediate Example: Plot views vs. watch time as a scatter plot
plt.figure(figsize=(7,5))
plt.scatter(df['views'], df['watch_time'], alpha=0.6, edgecolors='w')
plt.xlabel('Views')
plt.ylabel('Total Watch Time (minutes)')
plt.title('Scatter Plot: Views vs. Total Watch Time')
plt.grid(alpha=0.3)
plt.tight_layout()
plt.show()
No description has been provided for this image
# Advanced Example: Calculate the top 5% of videos by average watch time per view
threshold = df['avg_watch_time_per_view'].quantile(0.95)
top_5pct = df[df['avg_watch_time_per_view'] >= threshold]
print(f'Top 5% high-engagement videos: {len(top_5pct)}')
print(top_5pct[['video_id', 'views', 'watch_time', 'avg_watch_time_per_view']].sort_values('avg_watch_time_per_view', ascending=False).head(5))
Top 5% high-engagement videos: 15
     video_id  views  watch_time  avg_watch_time_per_view
43         44   2847      478095               167.929399
223       224   2793      236362                84.626566
119       120   9368      481047                51.350021
147       148  12766      442620                34.671784
106       107  11634      284501                24.454272
# Advanced Example: Compare average watch time per view across time periods
df['month'] = df['publish_date'].dt.to_period('M')
monthly_engagement = df.groupby('month')['avg_watch_time_per_view'].mean()
monthly_engagement.plot(kind='line', marker='o', figsize=(10,4), title='Avg Watch Time per View by Month')
plt.ylabel('Avg Watch Time per View (min)')
plt.xlabel('Month')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Advanced Example: Benchmark videos by click-through rate vs. watch time
plt.figure(figsize=(8,5))
plt.scatter(df['ctr'], df['avg_watch_time_per_view'], c=df['watch_time'], cmap='viridis', s=40, alpha=0.7)
plt.colorbar(label='Total Watch Time (minutes)')
plt.xlabel('CTR (%)')
plt.ylabel('Avg Watch Time per View (minutes)')
plt.title('CTR vs. Avg Watch Time per View (colored by Watch Time)')
plt.grid(alpha=0.3)
plt.tight_layout()
plt.show()
No description has been provided for this image
# Error Handling: Handle missing values in views or watch_time
df.loc[np.random.choice(df.index, 5, replace=False), 'views'] = np.nan
df.loc[np.random.choice(df.index, 5, replace=False), 'watch_time'] = np.nan
missing = df[df['views'].isnull() | df['watch_time'].isnull()]
print(f'Rows with missing views or watch_time: {missing.shape[0]}')
Rows with missing views or watch_time: 10
# Error Handling: Fill missing values and recalculate average watch time per view
df['views'] = df['views'].fillna(0)
df['watch_time'] = df['watch_time'].fillna(0)
df['avg_watch_time_per_view'] = df.apply(
    lambda row: row['watch_time']/row['views'] if row['views'] > 0 else 0, axis=1)
print('Recalculated avg_watch_time_per_view with filled missing data:')
print(df[['video_id', 'views', 'watch_time', 'avg_watch_time_per_view']].head(3))
Recalculated avg_watch_time_per_view with filled missing data:
   video_id     views  watch_time  avg_watch_time_per_view
0         1  122058.0    158381.0                 1.297588
1         2  146967.0    481671.0                 3.277409
2         3  132032.0    195806.0                 1.483019
# Error Handling: Avoiding incorrect aggregation (should not sum ratios)
incorrect_total = df['avg_watch_time_per_view'].sum()
print(f'INCORRECT: Total of avg_watch_time_per_view (should not sum ratios): {incorrect_total:.2f}')
correct_mean = df['avg_watch_time_per_view'].mean()
print(f'CORRECT: Mean avg_watch_time_per_view: {correct_mean:.2f}')
INCORRECT: Total of avg_watch_time_per_view (should not sum ratios): 829.37
CORRECT: Mean avg_watch_time_per_view: 2.76
# Error Handling: Misinterpreting CTR/engagement rates (division by zero protections already shown)
zero_row = pd.DataFrame({'views':[0], 'watch_time':[5000]})
zero_row['avg_watch_time_per_view'] = zero_row.apply(lambda row: row['watch_time']/row['views'] if row['views'] > 0 else 0, axis=1)
print(zero_row)
   views  watch_time  avg_watch_time_per_view
0      0        5000                        0

Best Practices: Interpreting and Comparing Content Performance#

  • Benchmark new videos against the averages for views, watch time, and engagement ratios.
  • Segment content by type, genre, or publish time to spot strengths or gaps.
  • Track growth and trends by charting metrics over time.
  • Use ratios (like avg watch time per view) for apples-to-apples comparison.
  • Document how metrics are calculated so teams are consistent.
# Common Analytics Pattern: Identify top-performing videos by category (synthetic)
df['category'] = np.random.choice(['Education', 'Entertainment', 'HowTo', 'Gaming'], len(df))
category_grouped = df.groupby('category')['avg_watch_time_per_view'].mean().sort_values(ascending=False)
print('Average Watch Time per View by Category:')
print(category_grouped)
Average Watch Time per View by Category:
category
Entertainment    5.464230
Education        2.092068
HowTo            1.657434
Gaming           1.559866
Name: avg_watch_time_per_view, dtype: float64
# Common Analytics Pattern: Track monthly growth in total watch time
monthly_growth = df.groupby('month')['watch_time'].sum()
monthly_growth.plot(kind='bar', figsize=(10,4), color='teal', title='Monthly Total Watch Time Growth')
plt.ylabel('Total Watch Time (minutes)')
plt.xlabel('Month')
plt.tight_layout()
plt.show()
No description has been provided for this image

End-to-End Problem: From Raw Data to Strategy#

  • Analyze your video dataset to:
    • List the three videos with the highest avg watch time per view.
    • Check if they share a category, publish month, or pattern.
    • Suggest why these videos perform so well.
  • Make a simple recommendation for future content strategy.
# End-to-End Solution: Find top three videos and summarize patterns
top_videos = df.sort_values('avg_watch_time_per_view', ascending=False).head(3)
print('Top 3 videos (by avg watch time per view):')
print(top_videos[['video_id', 'category', 'month', 'avg_watch_time_per_view']])

categories = top_videos['category'].unique()
months = top_videos['month'].unique()
print(f'Pattern: Categories = {categories}, Months = {months}')
print('Recommendation: Focus more on these content categories in upcoming months for maximum viewer engagement.')
Top 3 videos (by avg watch time per view):
     video_id       category    month  avg_watch_time_per_view
43         44  Entertainment  2022-02               167.929399
223       224  Entertainment  2022-08                84.626566
119       120  Entertainment  2022-04                51.350021
Pattern: Categories = ['Entertainment'], Months = <PeriodArray>
['2022-02', '2022-08', '2022-04']
Length: 3, dtype: period[M]
Recommendation: Focus more on these content categories in upcoming months for maximum viewer engagement.

Lesson Wrap-Up#

  • You have explored how to analyze view and watch time distribution using real analytics data.

  • Key techniques included distribution plots, ratio calculations, error handling, and actionable strategy insights.

  • Remember: high views are great, but depth of engagement tells the real story.

  • Try these methods next time you review your content analytics!

  • Subscribe to our YouTube channel for video walkthroughs and more lessons.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.