Lesson 16 · Social Media Content Analytics
Identifying Missing and Inconsistent Social Media Data
In this lesson, we will solve real-world problems where social media data may be missing, incomplete, or inconsistent. Social media content creators and…
- CourseSocial Media Content Analytics
- Lesson16 of 41
- Video26 min
- FormatJupyter notebook · 8 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIdentifying Missing and Inconsistent Social Media Data#
- In this lesson, we will solve real-world problems where social media data may be missing, incomplete, or inconsistent.
- Social media content creators and businesses rely on high-quality data to make decisions. Missing or inconsistent data can cause incorrect analysis and poor strategies.
- You will learn how to spot, handle, and fix these data issues, ensuring your reports give accurate insights.
- By the end, you will be equipped to find and address data gaps in YouTube and other content datasets for better analytics.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
Core Social Media Analytics Concepts#
- Social media datasets typically track videos or posts, including metrics like views, likes, comments, and shares.
- Higher values in these metrics often indicate stronger audience engagement or content performance.
- Other important metrics may include CTR (click-through rate), watch time, and engagement rate (likes + comments relative to views).
- Beginners sometimes ignore missing values or incorrect numbers, which can distort analysis and decision making.
- Consistent metric definitions are required for reliable content insights.
# Let us load the YouTube Trending Videos dataset
try:
from pathlib import Path
import os, pickle
from googleapiclient.discovery import build
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request
SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
def get_yt_service():
api_key = os.environ.get('YOUTUBE_API_KEY')
if api_key:
return build('youtube', 'v3', developerKey=api_key)
if Path('client_secret.json').exists():
creds = None
if Path('token_ro.pickle').exists():
with open('token_ro.pickle', 'rb') as f:
creds = pickle.load(f)
if not creds or not creds.valid:
if creds and creds.expired and creds.refresh_token:
creds.refresh(Request())
else:
flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
creds = flow.run_local_server(port=0)
with open('token_ro.pickle', 'wb') as f:
pickle.dump(creds, f)
return build('youtube', 'v3', credentials=creds)
raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
def fetch_yt_trending(max_results=200, region='US'):
youtube = get_yt_service()
records, token = [], None
while len(records) < max_results:
resp = youtube.videos().list(
part='snippet,statistics',
chart='mostPopular',
regionCode=region,
maxResults=min(50, max_results - len(records)),
pageToken=token
).execute()
for item in resp.get('items', []):
s = item['snippet']; st = item.get('statistics', {})
records.append({
'video_id': item['id'],
'trending_date': pd.Timestamp.today().date(),
'title': s.get('title', ''),
'channel_title': s.get('channelTitle', ''),
'category_id': s.get('categoryId', ''),
'views': int(st.get('viewCount', 0)),
'likes': int(st.get('likeCount', 0)),
'comment_count': int(st.get('commentCount', 0)),
})
token = resp.get('nextPageToken')
if not token: break
return pd.DataFrame(records)
df = fetch_yt_trending()
print('Live trending data:', df.shape)
except Exception as e:
print(f'Falling back to synthetic: {e}')
np.random.seed(42)
n = 1000
df = pd.DataFrame({
'video_id': [f'vid{i}' for i in range(n)],
'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
'title': [f'Video Title {i}' for i in range(n)],
'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
'category_id': np.random.choice([1,2,10,22,24,28], n),
'views': np.random.randint(10000, 5000000, n),
'likes': np.random.randint(100, 200000, n),
'comment_count': np.random.randint(10, 50000, n),
})
print('Synthetic fallback:', df.shape)
print(df.head(3))
# Let us introduce missing values for demonstration
df_missing = df.copy()
np.random.seed(42)
missing_idx = np.random.choice(df_missing.index, size=30, replace=False)
df_missing.loc[missing_idx, 'likes'] = np.nan
df_missing.loc[missing_idx[:15], 'views'] = np.nan
print(df_missing.head(10))
Beginner Example 1: Detecting Missing Data#
- Missing values may cause total counts and averages to be wrong if they are not corrected.
- We need to check which columns have missing values, to avoid incorrect engagement insights.
# Basic: Count missing values per column
print(df_missing.isnull().sum())
# Basic: List rows with missing likes or views
print(df_missing[df_missing['likes'].isnull() | df_missing['views'].isnull()].head())
Beginner Example 2: Simple Imputation of Missing Values#
- One common approach is to fill missing values with the column average or median.
- This keeps your calculations valid, though it is best to record that data was changed.
# Fill missing likes with the median, views with the mean
df_filled = df_missing.copy()
likes_median = df_filled['likes'].median()
views_mean = df_filled['views'].mean()
df_filled['likes'].fillna(likes_median, inplace=True)
df_filled['views'].fillna(views_mean, inplace=True)
print('Missing after fill:', df_filled.isnull().sum())
Beginner Example 3: Visualizing Missing Data Patterns#
- Visual summaries help you see where missing values cluster in your dataset.
- This helps spot systematic gaps by channel, category, or posting date.
- Let us use a simple missing data bar plot.
import matplotlib.pyplot as plt
missing = df_missing.isnull().sum()
missing = missing[missing > 0]
plt.figure(figsize=(6,3))
missing.plot(kind='bar', color='tomato')
plt.title('Missing Value Count by Column')
plt.ylabel('Count')
plt.tight_layout()
plt.savefig('missing_pattern.png')
plt.close()
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



