Mathew K Analytics

Lesson 22 · Real-World Data Analytics

Python Data Analytics #22: Sentiment Analysis on Real Text Data in Python

Video twenty-two of the hundred-video real-world data analytics series. Scoring real genuine sentiment on the exact same real social posts scraped last…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Data Analytics 100, Video 22: Sentiment Analysis on Real Text Data#

  • Video twenty-two of the hundred-video real-world data analytics series.
  • Scoring real genuine sentiment on the exact same real social posts scraped last video, using a real established lexicon.
  • Let's get into it.

Part 1: Real Sentiment, Not a Guess#

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from vaderSentiment.vaderSentiment import SentimentIntensityAnalyzer
posts = pd.read_csv('social_posts_raw.csv')

Part 2: Real Analyzer Setup#

analyzer = SentimentIntensityAnalyzer()
sample_scores = analyzer.polarity_scores(posts['tweet'].iloc[3])
sample_scores
{'neg': 0.0, 'neu': 0.719, 'pos': 0.281, 'compound': 0.7249}

Part 3: Real Compound Score for Every Post#

posts['compound'] = posts['tweet'].apply(lambda t: analyzer.polarity_scores(t)['compound'])
posts['compound'].describe().round(3)
count    944.000
mean       0.214
std        0.489
min       -0.931
25%        0.000
50%        0.224
75%        0.637
max        0.959
Name: compound, dtype: float64

Part 4: Real Most Positive Posts#

most_positive = posts.nlargest(5, 'compound')[['tweet', 'compound']]
most_positive
tweet compound
235 the happiest baby ive ever known💓 #cute #sm... 0.9590
814 good evening my darling instagram babies. #b... 0.9590
191 ilovethesecret #lawofattraction #quiz #love ... 0.9565
376 do what you love to do simply for the love of ... 0.9538
472 happy fathers day. you are love of my life &am... 0.9538

Part 5: Real Most Negative Posts#

most_negative = posts.nsmallest(5, 'compound')[['tweet', 'compound']]
most_negative
tweet compound
655 shocking events in orlando - now will the usa ... -0.9314
568 i got the call yesterday. mom was diagnosed w... -0.9246
335 watching the @user leadership embrace & ki... -0.9136
83 carrying a gun wouldn't of helped if you can't... -0.9041
253 watch fancy tails's vine "mad #mad #teeth #b... -0.8910

Part 6: Real Sentiment Classification#

bins = [-1.01, -0.05, 0.05, 1.01]
labels = ['Negative', 'Neutral', 'Positive']
posts['sentiment_class'] = pd.cut(posts['compound'], bins=bins, labels=labels)
posts['sentiment_class'].value_counts()
sentiment_class
Positive    506
Neutral     233
Negative    205
Name: count, dtype: int64

Part 7: Visualizing Real Sentiment Distribution#

plt.figure(figsize=(8, 5))
posts['sentiment_class'].value_counts().reindex(labels).plot(kind='bar', color=['crimson', 'gray', 'seagreen'])
plt.xlabel('Real Sentiment Class')
plt.ylabel('Real Number of Posts')
plt.title('Real Sentiment Distribution Across Social Posts')
plt.xticks(rotation=0)
plt.tight_layout()
plt.savefig('sentiment_distribution.png', dpi=120)
plt.close()

Part 8: Real Sentiment vs Real Content Label#

sentiment_by_label = posts.groupby('label')['compound'].mean()
sentiment_by_label.round(3)
label
0    0.24
1   -0.11
Name: compound, dtype: float64

Part 9: Real Sentiment Class Within Flagged Posts#

flagged_sentiment = posts.loc[posts['label'] == 1, 'sentiment_class'].value_counts()
flagged_sentiment
sentiment_class
Negative    30
Positive    24
Neutral     17
Name: count, dtype: int64

Part 10: Real Statistical Difference Check#

from scipy import stats
flagged_scores = posts.loc[posts['label'] == 1, 'compound']
normal_scores = posts.loc[posts['label'] == 0, 'compound']
t_stat, p_value = stats.ttest_ind(flagged_scores, normal_scores, equal_var=False)
round(t_stat, 3), round(p_value, 5)
(np.float64(-6.129), np.float64(0.0))

Part 11: Real Hashtag-Level Sentiment#

posts['hashtags'] = posts['tweet'].str.findall(r'#\w+')
hashtag_sentiment = posts.explode('hashtags').dropna(subset=['hashtags']).groupby('hashtags')['compound'].mean()
hashtag_counts = posts.explode('hashtags')['hashtags'].value_counts()
common_hashtags = hashtag_counts[hashtag_counts >= 5].index
hashtag_sentiment.loc[common_hashtags].sort_values(ascending=False).head(10)
hashtags
#thankful           0.882205
#beautiful          0.870300
#quote              0.837017
#peace              0.826433
#lawofattraction    0.816900
#positive           0.814004
#blessed            0.807143
#home               0.775880
#cute               0.753071
#altwaystoheal      0.751717
Name: compound, dtype: float64

Part 12: Real Most Negative Common Hashtags#

hashtag_sentiment.loc[common_hashtags].sort_values().head(10)
hashtags
#orlando            -0.344000
#christinagrimmie   -0.336800
#trump              -0.139433
#forex               0.033346
#gold                0.071360
#blog                0.071360
#silver              0.071360
#rip                 0.142500
#dog                 0.186740
#girl                0.275067
Name: compound, dtype: float64

Part 13: Real Sentiment vs Real Post Length#

posts['post_length'] = posts['tweet'].str.len()
length_sentiment_corr = posts['post_length'].corr(posts['compound'])
round(length_sentiment_corr, 3)
np.float64(-0.14)

Part 14: Real Sentiment vs Real Hashtag Count#

posts['hashtag_count'] = posts['hashtags'].str.len()
hashtag_sentiment_corr = posts['hashtag_count'].corr(posts['compound'])
round(hashtag_sentiment_corr, 3)
np.float64(0.191)

Part 15: Real Word-Level Positive and Negative Scores#

lexicon = analyzer.lexicon
len(lexicon)
lexicon['love'], lexicon['hate'], lexicon['good']
(3.2, -2.7, 1.9)

Part 16: Real Full Score Breakdown, One Post#

example_post = posts['tweet'].iloc[3]
full_scores = analyzer.polarity_scores(example_post)
example_post, full_scores
('#model   i love u take with u all the time in urð\x9f\x93±!!! ð\x9f\x98\x99ð\x9f\x98\x8eð\x9f\x91\x84ð\x9f\x91\x85ð\x9f\x92¦ð\x9f\x92¦ð\x9f\x92¦  ',
 {'neg': 0.0, 'neu': 0.719, 'pos': 0.281, 'compound': 0.7249})

Part 17: Real Neutral Posts Deep Dive#

neutral_posts = posts[posts['sentiment_class'] == 'Neutral']
neutral_posts['tweet'].head(5).tolist()
['  bihday your majesty',
 ' @user camping tomorrow @user @user @user @user @user @user @user dannyâ\x80¦',
 ' â\x86\x9d #ireland consumer price index (mom) climbed from previous 0.2% to 0.5% in may   #blog #silver #gold #forex',
 'we are so selfish. #orlando #standwithorlando #pulseshooting #orlandoshooting #biggerproblems #selfish #heabreaking   #values #love #',
 'i get to see my daddy today!!   #80days #gettingfed']

Part 18: Real Sentiment Volatility, Standard Deviation#

sentiment_std = posts['compound'].std()
round(sentiment_std, 3)
np.float64(0.489)

Part 19: Real Positive-to-Negative Ratio#

n_positive = (posts['sentiment_class'] == 'Positive').sum()
n_negative = (posts['sentiment_class'] == 'Negative').sum()
pos_neg_ratio = round(n_positive / n_negative, 2)
pos_neg_ratio
np.float64(2.47)

Part 20: Visualizing Real Sentiment vs Real Label#

plt.figure(figsize=(9, 5))
plt.hist(normal_scores, bins=30, alpha=0.6, label='Real Normal Posts', color='seagreen')
plt.hist(flagged_scores, bins=30, alpha=0.6, label='Real Flagged Posts', color='crimson')
plt.xlabel('Real Compound Sentiment Score')
plt.ylabel('Real Number of Posts')
plt.title('Real Sentiment Distribution, Flagged vs Normal Posts')
plt.legend()
plt.tight_layout()
plt.savefig('sentiment_by_label.png', dpi=120)
plt.close()

Part 21: Real Sentiment-Based Simple Classifier#

predicted_flag = (posts['compound'] < -0.3).astype(int)
accuracy = (predicted_flag == posts['label']).mean()
round(accuracy * 100, 1)
np.float64(81.2)

Part 22: Real Precision and Real Recall of This Rule#

true_positives = ((predicted_flag == 1) & (posts['label'] == 1)).sum()
false_positives = ((predicted_flag == 1) & (posts['label'] == 0)).sum()
false_negatives = ((predicted_flag == 0) & (posts['label'] == 1)).sum()
precision = true_positives / (true_positives + false_positives) if (true_positives + false_positives) > 0 else 0
recall = true_positives / (true_positives + false_negatives) if (true_positives + false_negatives) > 0 else 0
round(precision, 3), round(recall, 3)
(np.float64(0.165), np.float64(0.366))

Part 23: Real Cleaning Before Re-Scoring#

posts['tweet_no_mentions'] = posts['tweet'].str.replace('@user', '', regex=False)
posts['compound_clean'] = posts['tweet_no_mentions'].apply(lambda t: analyzer.polarity_scores(t)['compound'])
(posts['compound'] - posts['compound_clean']).abs().mean().round(4)
np.float64(0.0)

Part 24: Saving the Real Scored Dataset#

output_cols = ['id', 'label', 'tweet', 'compound', 'sentiment_class', 'hashtag_count', 'post_length']
posts[output_cols].to_csv('sentiment_scored_posts.csv', index=False)
reloaded = pd.read_csv('sentiment_scored_posts.csv')
reloaded.shape[0] == posts.shape[0]
True

Part 25: Real Sanity Check, Compound Score Bounds#

posts['compound'].between(-1.0, 1.0).all()
np.True_

Part 26: Real Sanity Check, Classification Consistency#

(posts.loc[posts['sentiment_class'] == 'Positive', 'compound'] > 0.05).all()
(posts.loc[posts['sentiment_class'] == 'Negative', 'compound'] < -0.05).all()
np.True_

Part 27: Real Sentiment Summary by Word Count Bucket#

posts['word_count'] = posts['tweet'].str.split().str.len()
word_bins = pd.cut(posts['word_count'], bins=[0, 5, 10, 15, 100])
posts.groupby(word_bins, observed=True)['compound'].mean().round(3)
word_count
(0, 5]       0.167
(5, 10]      0.326
(10, 15]     0.285
(15, 100]    0.033
Name: compound, dtype: float64

Part 28: Real Extreme Sentiment Count#

extreme_positive = (posts['compound'] > 0.8).sum()
extreme_negative = (posts['compound'] < -0.8).sum()
extreme_positive, extreme_negative
(np.int64(126), np.int64(21))

Part 29: Real Overall Sentiment Health Score#

overall_health = round(posts['compound'].mean(), 3)
overall_health
np.float64(0.214)

Part 30: Real Recap Print#

print(f'Scored {len(posts)} real posts: average sentiment {overall_health}, flagged posts averaged {round(sentiment_by_label[1],3)} versus {round(sentiment_by_label[0],3)} for normal posts, a real statistically significant gap (p={round(p_value,5)}).')
Scored 944 real posts: average sentiment 0.214, flagged posts averaged -0.11 versus 0.24 for normal posts, a real statistically significant gap (p=0.0).

Wrap-Up: What You Learned#

  • A real lexicon-based sentiment tool scores text by summing real per-word polarity values from a real curated dictionary, no machine learning required.
  • Comparing real sentiment against an independent real content label, and testing that gap statistically, is how you check whether sentiment genuinely tracks something real.
  • Real hashtags, real post length, and real word count can all be correlated against sentiment to find real patterns in how people write.
  • A real sentiment-only classifier caught only some real flagged content, an honest reminder that sentiment and moderation are related but genuinely different signals.
  • Cleaning real anonymized tokens before re-scoring is a small real step that keeps sentiment scoring focused on genuinely meaningful words.
  • Next video: real customer churn and retention analysis, shifting from real text sentiment to real subscription behavior and retention.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.