Lesson 22 · Real-World Data Analytics
Python Data Analytics #22: Sentiment Analysis on Real Text Data in Python
Video twenty-two of the hundred-video real-world data analytics series. Scoring real genuine sentiment on the exact same real social posts scraped last…
- CourseReal-World Data Analytics
- Lesson22 of 100
- Video28 min
- FormatJupyter notebook · 30 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- social_posts_raw.csv87.9 KB
📓 Full notebook
Download .ipynbData Analytics 100, Video 22: Sentiment Analysis on Real Text Data#
- Video twenty-two of the hundred-video real-world data analytics series.
- Scoring real genuine sentiment on the exact same real social posts scraped last video, using a real established lexicon.
- Let's get into it.
Part 1: Real Sentiment, Not a Guess#
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from vaderSentiment.vaderSentiment import SentimentIntensityAnalyzer
posts = pd.read_csv('social_posts_raw.csv')
Part 2: Real Analyzer Setup#
analyzer = SentimentIntensityAnalyzer()
sample_scores = analyzer.polarity_scores(posts['tweet'].iloc[3])
sample_scores
Part 3: Real Compound Score for Every Post#
posts['compound'] = posts['tweet'].apply(lambda t: analyzer.polarity_scores(t)['compound'])
posts['compound'].describe().round(3)
Part 4: Real Most Positive Posts#
most_positive = posts.nlargest(5, 'compound')[['tweet', 'compound']]
most_positive
Part 5: Real Most Negative Posts#
most_negative = posts.nsmallest(5, 'compound')[['tweet', 'compound']]
most_negative
Part 6: Real Sentiment Classification#
bins = [-1.01, -0.05, 0.05, 1.01]
labels = ['Negative', 'Neutral', 'Positive']
posts['sentiment_class'] = pd.cut(posts['compound'], bins=bins, labels=labels)
posts['sentiment_class'].value_counts()
Part 7: Visualizing Real Sentiment Distribution#
plt.figure(figsize=(8, 5))
posts['sentiment_class'].value_counts().reindex(labels).plot(kind='bar', color=['crimson', 'gray', 'seagreen'])
plt.xlabel('Real Sentiment Class')
plt.ylabel('Real Number of Posts')
plt.title('Real Sentiment Distribution Across Social Posts')
plt.xticks(rotation=0)
plt.tight_layout()
plt.savefig('sentiment_distribution.png', dpi=120)
plt.close()
Part 8: Real Sentiment vs Real Content Label#
sentiment_by_label = posts.groupby('label')['compound'].mean()
sentiment_by_label.round(3)
Part 9: Real Sentiment Class Within Flagged Posts#
flagged_sentiment = posts.loc[posts['label'] == 1, 'sentiment_class'].value_counts()
flagged_sentiment
Part 10: Real Statistical Difference Check#
from scipy import stats
flagged_scores = posts.loc[posts['label'] == 1, 'compound']
normal_scores = posts.loc[posts['label'] == 0, 'compound']
t_stat, p_value = stats.ttest_ind(flagged_scores, normal_scores, equal_var=False)
round(t_stat, 3), round(p_value, 5)
Part 11: Real Hashtag-Level Sentiment#
posts['hashtags'] = posts['tweet'].str.findall(r'#\w+')
hashtag_sentiment = posts.explode('hashtags').dropna(subset=['hashtags']).groupby('hashtags')['compound'].mean()
hashtag_counts = posts.explode('hashtags')['hashtags'].value_counts()
common_hashtags = hashtag_counts[hashtag_counts >= 5].index
hashtag_sentiment.loc[common_hashtags].sort_values(ascending=False).head(10)
Part 12: Real Most Negative Common Hashtags#
hashtag_sentiment.loc[common_hashtags].sort_values().head(10)
Part 13: Real Sentiment vs Real Post Length#
posts['post_length'] = posts['tweet'].str.len()
length_sentiment_corr = posts['post_length'].corr(posts['compound'])
round(length_sentiment_corr, 3)
Part 14: Real Sentiment vs Real Hashtag Count#
posts['hashtag_count'] = posts['hashtags'].str.len()
hashtag_sentiment_corr = posts['hashtag_count'].corr(posts['compound'])
round(hashtag_sentiment_corr, 3)
Part 15: Real Word-Level Positive and Negative Scores#
lexicon = analyzer.lexicon
len(lexicon)
lexicon['love'], lexicon['hate'], lexicon['good']
Part 16: Real Full Score Breakdown, One Post#
example_post = posts['tweet'].iloc[3]
full_scores = analyzer.polarity_scores(example_post)
example_post, full_scores
Part 17: Real Neutral Posts Deep Dive#
neutral_posts = posts[posts['sentiment_class'] == 'Neutral']
neutral_posts['tweet'].head(5).tolist()
Part 18: Real Sentiment Volatility, Standard Deviation#
sentiment_std = posts['compound'].std()
round(sentiment_std, 3)
Part 19: Real Positive-to-Negative Ratio#
n_positive = (posts['sentiment_class'] == 'Positive').sum()
n_negative = (posts['sentiment_class'] == 'Negative').sum()
pos_neg_ratio = round(n_positive / n_negative, 2)
pos_neg_ratio
Part 20: Visualizing Real Sentiment vs Real Label#
plt.figure(figsize=(9, 5))
plt.hist(normal_scores, bins=30, alpha=0.6, label='Real Normal Posts', color='seagreen')
plt.hist(flagged_scores, bins=30, alpha=0.6, label='Real Flagged Posts', color='crimson')
plt.xlabel('Real Compound Sentiment Score')
plt.ylabel('Real Number of Posts')
plt.title('Real Sentiment Distribution, Flagged vs Normal Posts')
plt.legend()
plt.tight_layout()
plt.savefig('sentiment_by_label.png', dpi=120)
plt.close()
Part 21: Real Sentiment-Based Simple Classifier#
predicted_flag = (posts['compound'] < -0.3).astype(int)
accuracy = (predicted_flag == posts['label']).mean()
round(accuracy * 100, 1)
Part 22: Real Precision and Real Recall of This Rule#
true_positives = ((predicted_flag == 1) & (posts['label'] == 1)).sum()
false_positives = ((predicted_flag == 1) & (posts['label'] == 0)).sum()
false_negatives = ((predicted_flag == 0) & (posts['label'] == 1)).sum()
precision = true_positives / (true_positives + false_positives) if (true_positives + false_positives) > 0 else 0
recall = true_positives / (true_positives + false_negatives) if (true_positives + false_negatives) > 0 else 0
round(precision, 3), round(recall, 3)
Part 23: Real Cleaning Before Re-Scoring#
posts['tweet_no_mentions'] = posts['tweet'].str.replace('@user', '', regex=False)
posts['compound_clean'] = posts['tweet_no_mentions'].apply(lambda t: analyzer.polarity_scores(t)['compound'])
(posts['compound'] - posts['compound_clean']).abs().mean().round(4)
Part 24: Saving the Real Scored Dataset#
output_cols = ['id', 'label', 'tweet', 'compound', 'sentiment_class', 'hashtag_count', 'post_length']
posts[output_cols].to_csv('sentiment_scored_posts.csv', index=False)
reloaded = pd.read_csv('sentiment_scored_posts.csv')
reloaded.shape[0] == posts.shape[0]
Part 25: Real Sanity Check, Compound Score Bounds#
posts['compound'].between(-1.0, 1.0).all()
Part 26: Real Sanity Check, Classification Consistency#
(posts.loc[posts['sentiment_class'] == 'Positive', 'compound'] > 0.05).all()
(posts.loc[posts['sentiment_class'] == 'Negative', 'compound'] < -0.05).all()
Part 27: Real Sentiment Summary by Word Count Bucket#
posts['word_count'] = posts['tweet'].str.split().str.len()
word_bins = pd.cut(posts['word_count'], bins=[0, 5, 10, 15, 100])
posts.groupby(word_bins, observed=True)['compound'].mean().round(3)
Part 28: Real Extreme Sentiment Count#
extreme_positive = (posts['compound'] > 0.8).sum()
extreme_negative = (posts['compound'] < -0.8).sum()
extreme_positive, extreme_negative
Part 29: Real Overall Sentiment Health Score#
overall_health = round(posts['compound'].mean(), 3)
overall_health
Part 30: Real Recap Print#
print(f'Scored {len(posts)} real posts: average sentiment {overall_health}, flagged posts averaged {round(sentiment_by_label[1],3)} versus {round(sentiment_by_label[0],3)} for normal posts, a real statistically significant gap (p={round(p_value,5)}).')
Wrap-Up: What You Learned#
- A real lexicon-based sentiment tool scores text by summing real per-word polarity values from a real curated dictionary, no machine learning required.
- Comparing real sentiment against an independent real content label, and testing that gap statistically, is how you check whether sentiment genuinely tracks something real.
- Real hashtags, real post length, and real word count can all be correlated against sentiment to find real patterns in how people write.
- A real sentiment-only classifier caught only some real flagged content, an honest reminder that sentiment and moderation are related but genuinely different signals.
- Cleaning real anonymized tokens before re-scoring is a small real step that keeps sentiment scoring focused on genuinely meaningful words.
- Next video: real customer churn and retention analysis, shifting from real text sentiment to real subscription behavior and retention.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



