Lesson 30 · Social Media Content Analytics
Sentiment Analysis on Social Media Comments
In this lesson, we will solve the problem of detecting sentiment in social media comments. Sentiment analysis helps creators and brands understand how…
- CourseSocial Media Content Analytics
- Lesson30 of 41
- Video24 min
- FormatJupyter notebook · 20 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbSentiment Analysis on Social Media Comments#
- In this lesson, we will solve the problem of detecting sentiment in social media comments.
- Sentiment analysis helps creators and brands understand how audiences feel about their content.
- By the end, you will analyze real comments, measure positive and negative feedback, and generate actionable insights for improving content strategy.
- The focus is on practical tools and hands-on analytics, not complex AI theory.
import pandas as pd
import numpy as np
from textblob import TextBlob
import warnings
warnings.filterwarnings('ignore')
What is Sentiment Analysis in Social Media Analytics?#
- Social media comments are feedback messages left by users on posts, videos, or images.
- These comments can be positive, negative, or neutral.
- Sentiment analysis assigns each comment a score that measures how positive or negative the language is.
- Metrics like likes, views, and comments show engagement, but sentiment gives context on the quality of that engagement.
- Beginners often forget that high comment numbers do not always mean positive reactions.
- Always look at both engagement numbers and sentiment for true insight.
# Load a sample dataset of social media comments
np.random.seed(42)
comments_data = [
{'post_id': 1, 'platform': 'YouTube', 'comment': 'Great video, learned a lot!'},
{'post_id': 1, 'platform': 'YouTube', 'comment': 'Very clear explanation, thank you.'},
{'post_id': 2, 'platform': 'Instagram', 'comment': 'Amazing content as always.'},
{'post_id': 2, 'platform': 'Instagram', 'comment': 'Not what I expected, a bit confusing.'},
{'post_id': 3, 'platform': 'TikTok', 'comment': 'This is so helpful!'},
{'post_id': 3, 'platform': 'TikTok', 'comment': 'Could you make a follow-up video?'},
{'post_id': 4, 'platform': 'YouTube', 'comment': 'Terrible content, waste of time.'},
{'post_id': 4, 'platform': 'YouTube', 'comment': 'I disagree with this approach.'},
{'post_id': 5, 'platform': 'Instagram', 'comment': 'Loved it, sharing with my friends.'},
{'post_id': 5, 'platform': 'TikTok', 'comment': 'The best tutorial I have seen.'},
]
df = pd.DataFrame(comments_data)
df['timestamp'] = pd.date_range('2023-01-01', periods=len(df), freq='3h')
print(df.shape)
print(df.head(3))
for i, row in df.iterrows():
print(f"Comment: {row['comment']}")
def get_sentiment(text):
analysis = TextBlob(text)
# Polarity ranges from -1 (negative) to 1 (positive)
return analysis.polarity
df['sentiment'] = df['comment'].apply(get_sentiment)
print(df[['comment', 'sentiment']])
# Classify comments as Positive, Neutral, or Negative based on sentiment value
def classify_sentiment(s):
if s > 0.2:
return 'Positive'
elif s < -0.2:
return 'Negative'
else:
return 'Neutral'
df['sentiment_label'] = df['sentiment'].apply(classify_sentiment)
print(df[['comment', 'sentiment', 'sentiment_label']])
# Count the number of comments in each sentiment category
counts = df['sentiment_label'].value_counts()
print(counts)
# Plotting sentiment distribution
import matplotlib.pyplot as plt
plt.style.use('seaborn-v0_8-colorblind')
counts.plot(kind='bar', color=['green','grey','red'])
plt.title('Distribution of Comment Sentiment')
plt.xlabel('Sentiment')
plt.ylabel('Number of Comments')
plt.tight_layout()
plt.savefig('sentiment_distribution.png')
plt.show()
# Beginner Example: Find all negative comments
negatives = df[df['sentiment_label'] == 'Negative']
print(negatives[['platform', 'comment', 'sentiment']])
# Beginner Example: Show only positive comments
positives = df[df['sentiment_label'] == 'Positive']
print(positives[['platform', 'comment', 'sentiment']])
# Beginner Example: Calculate average sentiment per platform
avg_sent_by_platform = df.groupby('platform')['sentiment'].mean()
print(avg_sent_by_platform)
# Intermediate Example: Track sentiment over time
df['date'] = df['timestamp'].dt.date
daily_sentiment = df.groupby('date')['sentiment'].mean()
print(daily_sentiment)
# Intermediate Example: Identify influencers of negative sentiment
neg_comment_counts = df[df['sentiment_label']=='Negative'].groupby('platform')['comment'].count()
print(neg_comment_counts)
# Intermediate Example: Analyze sentiment by post_id
mean_sent_post = df.groupby('post_id')['sentiment'].mean()
print(mean_sent_post)
# Advanced Example: Word clouds for sentiment analysis
from wordcloud import WordCloud
neg_words = ' '.join(df[df['sentiment_label']=='Negative']['comment'])
pos_words = ' '.join(df[df['sentiment_label']=='Positive']['comment'])
wordcloud_neg = WordCloud(width=400, height=200, background_color='white').generate(neg_words)
wordcloud_pos = WordCloud(width=400, height=200, background_color='white', colormap='Greens').generate(pos_words)
plt.figure(figsize=(10,4))
plt.subplot(1,2,1); plt.imshow(wordcloud_neg, interpolation='bilinear'); plt.axis('off'); plt.title('Negative Comments')
plt.subplot(1,2,2); plt.imshow(wordcloud_pos, interpolation='bilinear'); plt.axis('off'); plt.title('Positive Comments')
plt.tight_layout()
plt.savefig('sentiment_wordclouds.png')
plt.show()
# Advanced Example: Platform vs sentiment crosstab analysis
cross = pd.crosstab(df['platform'], df['sentiment_label'])
print(cross)
# Advanced Example: Sentiment over time per platform
sent_time_platform = df.groupby(['platform', 'date'])['sentiment'].mean().unstack('platform')
sent_time_platform.plot(marker='o')
plt.title('Average Sentiment Over Time by Platform')
plt.ylabel('Avg Sentiment')
plt.tight_layout()
plt.savefig('sent_trend_platform.png')
plt.show()
# Error Example: Handle missing comments
df_missing = df.copy()
df_missing.loc[3, 'comment'] = None
df_missing['sentiment'] = df_missing['comment'].apply(
lambda x: get_sentiment(x) if pd.notnull(x) else np.nan)
print(df_missing[['comment','sentiment']])
# Error Example: Incorrect aggregation mistake
try:
# This is incorrect: taking the mean across text fields
test = df[['comment', 'sentiment']].mean()
except Exception as e:
print('Error:', e)
# Error Example: Confusing sentiment with comment count
grouped = df.groupby('platform').agg({'sentiment':'mean', 'comment':'count'})
print(grouped)
Best Practices for Sentiment Analytics#
- Always combine sentiment scores with engagement metrics.
- Check trends over time and across different platforms.
- Use consistent thresholding and labeling.
- Handle missing or ambiguous comments carefully.
- Segment the audience when possible for focused feedback.
- Prioritize actionable insights, not just the numbers.
- Benchmark against earlier content to measure improvement.
- Use visuals to help non-technical teams understand sentiment data.
# End-to-end Example: Recommend content action
max_post = mean_sent_post.idxmax()
min_post = mean_sent_post.idxmin()
print(f"Post with highest avg sentiment: {max_post} ({mean_sent_post[max_post]:.2f})")
print(f"Post with lowest avg sentiment: {min_post} ({mean_sent_post[min_post]:.2f})")
if mean_sent_post[min_post] < 0:
print(f"Review post {min_post}. Consider responding or adjusting future content.")
else:
print(f"All posts have positive or neutral feedback. Keep up the good work!")
Congratulations! You have completed sentiment analysis for social media comments.#
Try these next steps:
Import a larger comment dataset and repeat your analysis.
Tune the thresholds or experiment with different sentiment libraries like VADER or spaCy.
Share your best insights with your team or publish a content feedback report.
Want more in-depth lessons? Subscribe to the YouTube channel for weekly tutorials and real data analytics projects.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



