Mathew K Analytics

Lesson 16 · Social Media Content Analytics

Identifying Missing and Inconsistent Social Media Data

In this lesson, we will solve real-world problems where social media data may be missing, incomplete, or inconsistent. Social media content creators and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Identifying Missing and Inconsistent Social Media Data#

  • In this lesson, we will solve real-world problems where social media data may be missing, incomplete, or inconsistent.
  • Social media content creators and businesses rely on high-quality data to make decisions. Missing or inconsistent data can cause incorrect analysis and poor strategies.
  • You will learn how to spot, handle, and fix these data issues, ensuring your reports give accurate insights.
  • By the end, you will be equipped to find and address data gaps in YouTube and other content datasets for better analytics.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Core Social Media Analytics Concepts#

  • Social media datasets typically track videos or posts, including metrics like views, likes, comments, and shares.
  • Higher values in these metrics often indicate stronger audience engagement or content performance.
  • Other important metrics may include CTR (click-through rate), watch time, and engagement rate (likes + comments relative to views).
  • Beginners sometimes ignore missing values or incorrect numbers, which can distort analysis and decision making.
  • Consistent metric definitions are required for reliable content insights.
# Let us load the YouTube Trending Videos dataset
try:
    from pathlib import Path
    import os, pickle
    from googleapiclient.discovery import build
    from google_auth_oauthlib.flow import InstalledAppFlow
    from google.auth.transport.requests import Request
    SCOPES = ['https://www.googleapis.com/auth/youtube.readonly']
    def get_yt_service():
        api_key = os.environ.get('YOUTUBE_API_KEY')
        if api_key:
            return build('youtube', 'v3', developerKey=api_key)
        if Path('client_secret.json').exists():
            creds = None
            if Path('token_ro.pickle').exists():
                with open('token_ro.pickle', 'rb') as f:
                    creds = pickle.load(f)
            if not creds or not creds.valid:
                if creds and creds.expired and creds.refresh_token:
                    creds.refresh(Request())
                else:
                    flow = InstalledAppFlow.from_client_secrets_file('client_secret.json', SCOPES)
                    creds = flow.run_local_server(port=0)
                with open('token_ro.pickle', 'wb') as f:
                    pickle.dump(creds, f)
            return build('youtube', 'v3', credentials=creds)
        raise EnvironmentError('Set YOUTUBE_API_KEY or provide client_secret.json')
    def fetch_yt_trending(max_results=200, region='US'):
        youtube = get_yt_service()
        records, token = [], None
        while len(records) < max_results:
            resp = youtube.videos().list(
                part='snippet,statistics',
                chart='mostPopular',
                regionCode=region,
                maxResults=min(50, max_results - len(records)),
                pageToken=token
            ).execute()
            for item in resp.get('items', []):
                s = item['snippet']; st = item.get('statistics', {})
                records.append({
                    'video_id':      item['id'],
                    'trending_date': pd.Timestamp.today().date(),
                    'title':         s.get('title', ''),
                    'channel_title': s.get('channelTitle', ''),
                    'category_id':   s.get('categoryId', ''),
                    'views':         int(st.get('viewCount', 0)),
                    'likes':         int(st.get('likeCount', 0)),
                    'comment_count': int(st.get('commentCount', 0)),
                })
            token = resp.get('nextPageToken')
            if not token: break
        return pd.DataFrame(records)
    df = fetch_yt_trending()
    print('Live trending data:', df.shape)
except Exception as e:
    print(f'Falling back to synthetic: {e}')
    np.random.seed(42)
    n = 1000
    df = pd.DataFrame({
        'video_id':      [f'vid{i}' for i in range(n)],
        'trending_date': pd.date_range('2023-01-01', periods=n, freq='D'),
        'title':         [f'Video Title {i}' for i in range(n)],
        'channel_title': np.random.choice(['ChannelA','ChannelB','ChannelC'], n),
        'category_id':   np.random.choice([1,2,10,22,24,28], n),
        'views':         np.random.randint(10000, 5000000, n),
        'likes':         np.random.randint(100, 200000, n),
        'comment_count': np.random.randint(10, 50000, n),
    })
    print('Synthetic fallback:', df.shape)
print(df.head(3))
Live trending data: (199, 8)
      video_id trending_date  \
0  v1t4MTqdfyI    2026-05-31   
1  0JlMjgqduVw    2026-05-31   
2  1lxFXWZ0IEU    2026-05-31   

                                               title  channel_title  \
0  Ariana Grande - hate that i made you love me (...  Ariana Grande   
1  House of the Dragon Season 3 | Official Final ...        HBO Max   
2                                            DAY 515         Jynxzi   

  category_id    views   likes  comment_count  
0          10  2877939  313793          21217  
1          24  3141203   48058           2744  
2          24   847110   13341            176  
# Let us introduce missing values for demonstration
df_missing = df.copy()
np.random.seed(42)
missing_idx = np.random.choice(df_missing.index, size=30, replace=False)
df_missing.loc[missing_idx, 'likes'] = np.nan
df_missing.loc[missing_idx[:15], 'views'] = np.nan
print(df_missing.head(10))
      video_id trending_date  \
0  v1t4MTqdfyI    2026-05-31   
1  0JlMjgqduVw    2026-05-31   
2  1lxFXWZ0IEU    2026-05-31   
3  83C3TZ4Zm_o    2026-05-31   
4  jLbst85USN8    2026-05-31   
5  _hyFGrcVv5U    2026-05-31   
6  t-1Z5hoWMp0    2026-05-31   
7  Xuhbhd5m-UA    2026-05-31   
8  dIA5f0yCL98    2026-05-31   
9  oY0F84RRMLI    2026-05-31   

                                               title      channel_title  \
0  Ariana Grande - hate that i made you love me (...      Ariana Grande   
1  House of the Dragon Season 3 | Official Final ...            HBO Max   
2                                            DAY 515             Jynxzi   
3                            aespa 에스파 'LEMONADE' MV             SMTOWN   
4    Call of Duty: Modern Warfare 4 | Reveal Trailer       Call of Duty   
5          I Went to WAR on a Hardcore Minecraft SMP             Wemmbu   
6              Yungeen Ace - Who Me (Official Video)        Yungeen Ace   
7  SPIDER-MAN BRAND NEW DAY TEASER: Hulk, Sadie S...  Emergency Awesome   
8  🔴LIVE | 24 HOUR BAGATHON | 7 Days To Die MARAT...     TheBurntPeanut   
9     mgk & Wiz Khalifa - mph (Official Music Video)                mgk   

  category_id       views     likes  comment_count  
0          10   2877939.0  313793.0          21217  
1          24   3141203.0   48058.0           2744  
2          24    847110.0   13341.0            176  
3          10  10024724.0  466424.0          28206  
4          20  27479238.0  184033.0          19655  
5          20    478863.0   81031.0          21503  
6          10    124480.0   11884.0            739  
7          24    123167.0    3693.0            306  
8          20    735404.0   11028.0             29  
9          10    211774.0       NaN           1470  

Beginner Example 1: Detecting Missing Data#

  • Missing values may cause total counts and averages to be wrong if they are not corrected.
  • We need to check which columns have missing values, to avoid incorrect engagement insights.
# Basic: Count missing values per column
print(df_missing.isnull().sum())
video_id          0
trending_date     0
title             0
channel_title     0
category_id       0
views            15
likes            30
comment_count     0
dtype: int64
# Basic: List rows with missing likes or views
print(df_missing[df_missing['likes'].isnull() | df_missing['views'].isnull()].head())
       video_id trending_date  \
9   oY0F84RRMLI    2026-05-31   
15  s9mt9HFU9bY    2026-05-31   
16  mfUtseK27pc    2026-05-31   
18  zpvGp5kOg18    2026-05-31   
30  icDuEHSxE-w    2026-05-31   

                                                title         channel_title  \
9      mgk & Wiz Khalifa - mph (Official Music Video)                   mgk   
15  Young Miko, Rauw Alejandro - Aquel diciembre (...         YoungMikoVEVO   
16  Marvel Animation’s X-Men ‘97 Season 2 | Offici...  Marvel Entertainment   
18               VV: ULTIMATUM - OFFICIAL TRAILER 2/2         VV: ULTIMATUM   
30                     Disclosure Day | Final Trailer    Universal Pictures   

   category_id      views  likes  comment_count  
9           10   211774.0    NaN           1470  
15          10        NaN    NaN           1109  
16          24  2953493.0    NaN           8408  
18          22    29313.0    NaN            562  
30          24        NaN    NaN           6420  

Beginner Example 2: Simple Imputation of Missing Values#

  • One common approach is to fill missing values with the column average or median.
  • This keeps your calculations valid, though it is best to record that data was changed.
# Fill missing likes with the median, views with the mean
df_filled = df_missing.copy()
likes_median = df_filled['likes'].median()
views_mean = df_filled['views'].mean()
df_filled['likes'].fillna(likes_median, inplace=True)
df_filled['views'].fillna(views_mean, inplace=True)
print('Missing after fill:', df_filled.isnull().sum())
Missing after fill: video_id         0
trending_date    0
title            0
channel_title    0
category_id      0
views            0
likes            0
comment_count    0
dtype: int64

Beginner Example 3: Visualizing Missing Data Patterns#

  • Visual summaries help you see where missing values cluster in your dataset.
  • This helps spot systematic gaps by channel, category, or posting date.
  • Let us use a simple missing data bar plot.
import matplotlib.pyplot as plt
missing = df_missing.isnull().sum()
missing = missing[missing > 0]
plt.figure(figsize=(6,3))
missing.plot(kind='bar', color='tomato')
plt.title('Missing Value Count by Column')
plt.ylabel('Count')
plt.tight_layout()
plt.savefig('missing_pattern.png')
plt.close()
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.