Lesson 29 · Real-World Data Analytics
Python Data Analytics #29: Real Airline Customer Sentiment Analysis in Python
Video twenty-nine of the hundred-video real-world data analytics series. Real genuine customer tweets directed at a real airline, hunting for exactly what…
- CourseReal-World Data Analytics
- Lesson29 of 100
- FormatJupyter notebook · 30 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- airline_tweets_raw.csv preview78.7 KB
📓 Full notebook
Download .ipynbData Analytics 100, Video 29: Real Airline Customer Sentiment Analysis#
- Video twenty-nine of the hundred-video real-world data analytics series.
- Real genuine customer tweets directed at a real airline, hunting for exactly what drives negative sentiment.
- Let's get into it.
Part 1: Real Tweets, Real Complaints#
import pandas as pd
import matplotlib.pyplot as plt
tweets = pd.read_csv('airline_tweets_raw.csv')
tweets.shape
Part 2: Real Data Preview#
tweets[['airline_sentiment', 'text']].head(5)
tweets.columns.tolist()
Part 3: Real Overall Sentiment Breakdown#
sentiment_counts = tweets['airline_sentiment'].value_counts()
sentiment_counts
Part 4: Real Sentiment as Percentages#
sentiment_pct = round(sentiment_counts / len(tweets) * 100, 1)
sentiment_pct
Part 5: Visualizing Real Sentiment Breakdown#
plt.figure(figsize=(7, 5))
colors = {'negative': 'crimson', 'neutral': 'gray', 'positive': 'seagreen'}
sentiment_counts.plot(kind='bar', color=[colors[s] for s in sentiment_counts.index])
plt.ylabel('Real Tweet Count')
plt.title('Real Customer Sentiment Breakdown')
plt.xticks(rotation=0)
plt.tight_layout()
plt.savefig('sentiment_breakdown.png', dpi=120)
plt.close()
Part 6: Real Negative Reason Breakdown#
reason_counts = tweets['negativereason'].value_counts()
reason_counts
Part 7: Visualizing Real Complaint Reasons#
plt.figure(figsize=(9, 6))
reason_counts.plot(kind='barh', color='darkorange')
plt.xlabel('Real Tweet Count')
plt.title('Real Negative Tweet Reasons')
plt.gca().invert_yaxis()
plt.tight_layout()
plt.savefig('negative_reasons.png', dpi=120)
plt.close()
Part 8: Real Top Complaint Reason#
top_reason = reason_counts.idxmax()
top_reason_pct = round(reason_counts.max() / reason_counts.sum() * 100, 1)
top_reason, top_reason_pct
Part 9: Real Sentiment Confidence Distribution#
tweets['airline_sentiment_confidence'].describe().round(3)
Part 10: Real Low-Confidence Tweets#
low_confidence = tweets[tweets['airline_sentiment_confidence'] < 0.6]
len(low_confidence)
low_confidence[['airline_sentiment', 'airline_sentiment_confidence', 'text']].head(3)
Part 11: Real Negative Reason Confidence#
tweets['negativereason_confidence'].describe().round(3)
Part 12: Real Retweet Volume by Sentiment#
retweets_by_sentiment = tweets.groupby('airline_sentiment')['retweet_count'].mean().round(3)
retweets_by_sentiment
Part 13: Real Most-Retweeted Tweet#
most_retweeted = tweets.loc[tweets['retweet_count'].idxmax()]
most_retweeted[['airline_sentiment', 'retweet_count', 'text']]
Part 14: Real Word Count by Sentiment#
tweets['word_count'] = tweets['text'].apply(lambda t: len(str(t).split()))
wordcount_by_sentiment = tweets.groupby('airline_sentiment')['word_count'].mean().round(2)
wordcount_by_sentiment
Part 15: Real Timezone Distribution#
timezone_counts = tweets['user_timezone'].value_counts()
timezone_counts.head(10)
Part 16: Real Sentiment by Top Timezone#
top_timezones = timezone_counts.head(3).index
timezone_sentiment = tweets[tweets['user_timezone'].isin(top_timezones)].groupby(['user_timezone', 'airline_sentiment']).size().unstack(fill_value=0)
timezone_sentiment
Part 17: Real Tweet Location Field#
has_location = tweets['tweet_location'].notna().sum()
has_location_pct = round(has_location / len(tweets) * 100, 1)
has_location_pct
Part 18: Real Missing Data Audit#
missing_summary = tweets.isna().sum().sort_values(ascending=False)
missing_summary
Part 19: Real Negative Reason Given Confidence Threshold#
high_confidence_reasons = tweets[tweets['negativereason_confidence'] >= 0.8]
high_confidence_reasons['negativereason'].value_counts()
Part 20: Real Customer Service Complaints Deep Dive#
cs_complaints = tweets[tweets['negativereason'] == 'Customer Service Issue']
cs_complaints[['text']].head(5)
Part 21: Real Average Confidence by Complaint Type#
confidence_by_reason = tweets.groupby('negativereason')['negativereason_confidence'].mean().round(3).sort_values(ascending=False)
confidence_by_reason
Part 22: Real Positive Tweet Sample#
positive_tweets = tweets[tweets['airline_sentiment'] == 'positive']
positive_tweets[['text']].sample(5, random_state=42)
Part 23: Real Mention Pattern Check#
tweets['mentions_airline'] = tweets['text'].str.contains('@VirginAmerica', case=False, na=False)
mention_rate = round(tweets['mentions_airline'].mean() * 100, 1)
mention_rate
Part 24: Real Sentiment Among Direct Mentions#
mention_sentiment = tweets[tweets['mentions_airline']]['airline_sentiment'].value_counts(normalize=True).round(3) * 100
mention_sentiment
Part 25: Real Word Count vs Confidence Correlation#
wordcount_confidence_corr = tweets['word_count'].corr(tweets['airline_sentiment_confidence'])
round(wordcount_confidence_corr, 3)
Part 26: Saving the Real Complaint Summary Table#
complaint_summary = pd.DataFrame({'count': reason_counts, 'avg_confidence': confidence_by_reason}).round(3)
complaint_summary.to_csv('complaint_reason_summary.csv')
reloaded_summary = pd.read_csv('complaint_reason_summary.csv', index_col=0)
reloaded_summary.shape == complaint_summary.shape
Part 27: Real Sanity Check, Sentiment Counts Sum Correctly#
sentiment_counts.sum() == len(tweets)
Part 28: Real Sanity Check, Confidence Scores in Range#
all((0 <= tweets['airline_sentiment_confidence']) & (tweets['airline_sentiment_confidence'] <= 1))
Part 29: Real Priority Recommendation#
negative_share = sentiment_pct['negative']
print(f'With {negative_share}% of tweets negative and {top_reason} the leading real complaint reason at {top_reason_pct}% of negatives, that is the real single area most worth investigating first.')
Part 30: Real Recap Print#
print(f'Across {len(tweets)} real customer tweets, {negative_share}% were negative, with {top_reason} driving {top_reason_pct}% of those real complaints.')
Wrap-Up: What You Learned#
- Hand-labeled sentiment and complaint-reason fields, when available, are far richer than running sentiment analysis from scratch on raw real text.
- Confidence scores attached to real labels are worth auditing separately, low-confidence rows deserve real skepticism before trusting them.
- Retweet volume revealed which real sentiment category actually spreads furthest, a genuinely different question from which is most common.
- Cross-tabulating sentiment against fields like timezone or direct-mention status can surface real patterns a single value count would miss.
- A single real brand, deeply analyzed, can be just as instructive as a multi-brand comparison when the real data honestly supports that scope.
- Next video: capstone, a real social media campaign performance report pulling together the real domain from this entire section.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



