Mathew K Analytics

Lesson 30 · Social Media Content Analytics

Sentiment Analysis on Social Media Comments

In this lesson, we will solve the problem of detecting sentiment in social media comments. Sentiment analysis helps creators and brands understand how…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Sentiment Analysis on Social Media Comments#

  • In this lesson, we will solve the problem of detecting sentiment in social media comments.
  • Sentiment analysis helps creators and brands understand how audiences feel about their content.
  • By the end, you will analyze real comments, measure positive and negative feedback, and generate actionable insights for improving content strategy.
  • The focus is on practical tools and hands-on analytics, not complex AI theory.
import pandas as pd
import numpy as np
from textblob import TextBlob
import warnings
warnings.filterwarnings('ignore')

What is Sentiment Analysis in Social Media Analytics?#

  • Social media comments are feedback messages left by users on posts, videos, or images.
  • These comments can be positive, negative, or neutral.
  • Sentiment analysis assigns each comment a score that measures how positive or negative the language is.
  • Metrics like likes, views, and comments show engagement, but sentiment gives context on the quality of that engagement.
  • Beginners often forget that high comment numbers do not always mean positive reactions.
  • Always look at both engagement numbers and sentiment for true insight.
# Load a sample dataset of social media comments
np.random.seed(42)
comments_data = [
    {'post_id': 1, 'platform': 'YouTube',   'comment': 'Great video, learned a lot!'},
    {'post_id': 1, 'platform': 'YouTube',   'comment': 'Very clear explanation, thank you.'},
    {'post_id': 2, 'platform': 'Instagram', 'comment': 'Amazing content as always.'},
    {'post_id': 2, 'platform': 'Instagram', 'comment': 'Not what I expected, a bit confusing.'},
    {'post_id': 3, 'platform': 'TikTok',    'comment': 'This is so helpful!'},
    {'post_id': 3, 'platform': 'TikTok',    'comment': 'Could you make a follow-up video?'},
    {'post_id': 4, 'platform': 'YouTube',   'comment': 'Terrible content, waste of time.'},
    {'post_id': 4, 'platform': 'YouTube',   'comment': 'I disagree with this approach.'},
    {'post_id': 5, 'platform': 'Instagram', 'comment': 'Loved it, sharing with my friends.'},
    {'post_id': 5, 'platform': 'TikTok',    'comment': 'The best tutorial I have seen.'},
]
df = pd.DataFrame(comments_data)
df['timestamp'] = pd.date_range('2023-01-01', periods=len(df), freq='3h')
print(df.shape)
print(df.head(3))
(10, 4)
   post_id   platform                             comment           timestamp
0        1    YouTube         Great video, learned a lot! 2023-01-01 00:00:00
1        1    YouTube  Very clear explanation, thank you. 2023-01-01 03:00:00
2        2  Instagram          Amazing content as always. 2023-01-01 06:00:00
for i, row in df.iterrows():
    print(f"Comment: {row['comment']}")
Comment: Great video, learned a lot!
Comment: Very clear explanation, thank you.
Comment: Amazing content as always.
Comment: Not what I expected, a bit confusing.
Comment: This is so helpful!
Comment: Could you make a follow-up video?
Comment: Terrible content, waste of time.
Comment: I disagree with this approach.
Comment: Loved it, sharing with my friends.
Comment: The best tutorial I have seen.
def get_sentiment(text):
    analysis = TextBlob(text)
    # Polarity ranges from -1 (negative) to 1 (positive)
    return analysis.polarity

df['sentiment'] = df['comment'].apply(get_sentiment)
print(df[['comment', 'sentiment']])
                                 comment  sentiment
0            Great video, learned a lot!       1.00
1     Very clear explanation, thank you.       0.13
2             Amazing content as always.       0.60
3  Not what I expected, a bit confusing.      -0.20
4                    This is so helpful!       0.00
5      Could you make a follow-up video?       0.00
6       Terrible content, waste of time.      -0.60
7         I disagree with this approach.       0.00
8     Loved it, sharing with my friends.       0.70
9         The best tutorial I have seen.       1.00
# Classify comments as Positive, Neutral, or Negative based on sentiment value
def classify_sentiment(s):
    if s > 0.2:
        return 'Positive'
    elif s < -0.2:
        return 'Negative'
    else:
        return 'Neutral'

df['sentiment_label'] = df['sentiment'].apply(classify_sentiment)
print(df[['comment', 'sentiment', 'sentiment_label']])
                                 comment  sentiment sentiment_label
0            Great video, learned a lot!       1.00        Positive
1     Very clear explanation, thank you.       0.13         Neutral
2             Amazing content as always.       0.60        Positive
3  Not what I expected, a bit confusing.      -0.20         Neutral
4                    This is so helpful!       0.00         Neutral
5      Could you make a follow-up video?       0.00         Neutral
6       Terrible content, waste of time.      -0.60        Negative
7         I disagree with this approach.       0.00         Neutral
8     Loved it, sharing with my friends.       0.70        Positive
9         The best tutorial I have seen.       1.00        Positive
# Count the number of comments in each sentiment category
counts = df['sentiment_label'].value_counts()
print(counts)
sentiment_label
Neutral     5
Positive    4
Negative    1
Name: count, dtype: int64
# Plotting sentiment distribution
import matplotlib.pyplot as plt
plt.style.use('seaborn-v0_8-colorblind')
counts.plot(kind='bar', color=['green','grey','red'])
plt.title('Distribution of Comment Sentiment')
plt.xlabel('Sentiment')
plt.ylabel('Number of Comments')
plt.tight_layout()
plt.savefig('sentiment_distribution.png')
plt.show()
No description has been provided for this image
# Beginner Example: Find all negative comments
negatives = df[df['sentiment_label'] == 'Negative']
print(negatives[['platform', 'comment', 'sentiment']])
  platform                           comment  sentiment
6  YouTube  Terrible content, waste of time.       -0.6
# Beginner Example: Show only positive comments
positives = df[df['sentiment_label'] == 'Positive']
print(positives[['platform', 'comment', 'sentiment']])
    platform                             comment  sentiment
0    YouTube         Great video, learned a lot!        1.0
2  Instagram          Amazing content as always.        0.6
8  Instagram  Loved it, sharing with my friends.        0.7
9     TikTok      The best tutorial I have seen.        1.0
# Beginner Example: Calculate average sentiment per platform
avg_sent_by_platform = df.groupby('platform')['sentiment'].mean()
print(avg_sent_by_platform)
platform
Instagram    0.366667
TikTok       0.333333
YouTube      0.132500
Name: sentiment, dtype: float64
# Intermediate Example: Track sentiment over time
df['date'] = df['timestamp'].dt.date
daily_sentiment = df.groupby('date')['sentiment'].mean()
print(daily_sentiment)
date
2023-01-01    0.11625
2023-01-02    0.85000
Name: sentiment, dtype: float64
# Intermediate Example: Identify influencers of negative sentiment
neg_comment_counts = df[df['sentiment_label']=='Negative'].groupby('platform')['comment'].count()
print(neg_comment_counts)
platform
YouTube    1
Name: comment, dtype: int64
# Intermediate Example: Analyze sentiment by post_id
mean_sent_post = df.groupby('post_id')['sentiment'].mean()
print(mean_sent_post)
post_id
1    0.565
2    0.200
3    0.000
4   -0.300
5    0.850
Name: sentiment, dtype: float64
# Advanced Example: Word clouds for sentiment analysis
from wordcloud import WordCloud
neg_words = ' '.join(df[df['sentiment_label']=='Negative']['comment'])
pos_words = ' '.join(df[df['sentiment_label']=='Positive']['comment'])
wordcloud_neg = WordCloud(width=400, height=200, background_color='white').generate(neg_words)
wordcloud_pos = WordCloud(width=400, height=200, background_color='white', colormap='Greens').generate(pos_words)
plt.figure(figsize=(10,4))
plt.subplot(1,2,1); plt.imshow(wordcloud_neg, interpolation='bilinear'); plt.axis('off'); plt.title('Negative Comments')
plt.subplot(1,2,2); plt.imshow(wordcloud_pos, interpolation='bilinear'); plt.axis('off'); plt.title('Positive Comments')
plt.tight_layout()
plt.savefig('sentiment_wordclouds.png')
plt.show()
No description has been provided for this image
# Advanced Example: Platform vs sentiment crosstab analysis
cross = pd.crosstab(df['platform'], df['sentiment_label'])
print(cross)
sentiment_label  Negative  Neutral  Positive
platform                                    
Instagram               0        1         2
TikTok                  0        2         1
YouTube                 1        2         1
# Advanced Example: Sentiment over time per platform
sent_time_platform = df.groupby(['platform', 'date'])['sentiment'].mean().unstack('platform')
sent_time_platform.plot(marker='o')
plt.title('Average Sentiment Over Time by Platform')
plt.ylabel('Avg Sentiment')
plt.tight_layout()
plt.savefig('sent_trend_platform.png')
plt.show()
No description has been provided for this image
# Error Example: Handle missing comments
df_missing = df.copy()
df_missing.loc[3, 'comment'] = None
df_missing['sentiment'] = df_missing['comment'].apply(
    lambda x: get_sentiment(x) if pd.notnull(x) else np.nan)
print(df_missing[['comment','sentiment']])
                              comment  sentiment
0         Great video, learned a lot!       1.00
1  Very clear explanation, thank you.       0.13
2          Amazing content as always.       0.60
3                                None        NaN
4                 This is so helpful!       0.00
5   Could you make a follow-up video?       0.00
6    Terrible content, waste of time.      -0.60
7      I disagree with this approach.       0.00
8  Loved it, sharing with my friends.       0.70
9      The best tutorial I have seen.       1.00
# Error Example: Incorrect aggregation mistake
try:
    # This is incorrect: taking the mean across text fields
    test = df[['comment', 'sentiment']].mean()
except Exception as e:
    print('Error:', e)
Error: Could not convert ['Great video, learned a lot!Very clear explanation, thank you.Amazing content as always.Not what I expected, a bit confusing.This is so helpful!Could you make a follow-up video?Terrible content, waste of time.I disagree with this approach.Loved it, sharing with my friends.The best tutorial I have seen.'] to numeric
# Error Example: Confusing sentiment with comment count
grouped = df.groupby('platform').agg({'sentiment':'mean', 'comment':'count'})
print(grouped)
           sentiment  comment
platform                     
Instagram   0.366667        3
TikTok      0.333333        3
YouTube     0.132500        4

Best Practices for Sentiment Analytics#

  • Always combine sentiment scores with engagement metrics.
  • Check trends over time and across different platforms.
  • Use consistent thresholding and labeling.
  • Handle missing or ambiguous comments carefully.
  • Segment the audience when possible for focused feedback.
  • Prioritize actionable insights, not just the numbers.
  • Benchmark against earlier content to measure improvement.
  • Use visuals to help non-technical teams understand sentiment data.
# End-to-end Example: Recommend content action
max_post = mean_sent_post.idxmax()
min_post = mean_sent_post.idxmin()
print(f"Post with highest avg sentiment: {max_post} ({mean_sent_post[max_post]:.2f})")
print(f"Post with lowest avg sentiment: {min_post} ({mean_sent_post[min_post]:.2f})")
if mean_sent_post[min_post] < 0:
    print(f"Review post {min_post}. Consider responding or adjusting future content.")
else:
    print(f"All posts have positive or neutral feedback. Keep up the good work!")
Post with highest avg sentiment: 5 (0.85)
Post with lowest avg sentiment: 4 (-0.30)
Review post 4. Consider responding or adjusting future content.

Congratulations! You have completed sentiment analysis for social media comments.#

  • Try these next steps:

  • Import a larger comment dataset and repeat your analysis.

  • Tune the thresholds or experiment with different sentiment libraries like VADER or spaCy.

  • Share your best insights with your team or publish a content feedback report.

  • Want more in-depth lessons? Subscribe to the YouTube channel for weekly tutorials and real data analytics projects.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.