Mathew K Analytics

Lesson 29 · Real-World Data Analytics

Python Data Analytics #29: Real Airline Customer Sentiment Analysis in Python

Video twenty-nine of the hundred-video real-world data analytics series. Real genuine customer tweets directed at a real airline, hunting for exactly what…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Data Analytics 100, Video 29: Real Airline Customer Sentiment Analysis#

  • Video twenty-nine of the hundred-video real-world data analytics series.
  • Real genuine customer tweets directed at a real airline, hunting for exactly what drives negative sentiment.
  • Let's get into it.

Part 1: Real Tweets, Real Complaints#

import pandas as pd
import matplotlib.pyplot as plt
tweets = pd.read_csv('airline_tweets_raw.csv')
tweets.shape
(341, 15)

Part 2: Real Data Preview#

tweets[['airline_sentiment', 'text']].head(5)
tweets.columns.tolist()
['tweet_id',
 'airline_sentiment',
 'airline_sentiment_confidence',
 'negativereason',
 'negativereason_confidence',
 'airline',
 'airline_sentiment_gold',
 'name',
 'negativereason_gold',
 'retweet_count',
 'text',
 'tweet_coord',
 'tweet_created',
 'tweet_location',
 'user_timezone']

Part 3: Real Overall Sentiment Breakdown#

sentiment_counts = tweets['airline_sentiment'].value_counts()
sentiment_counts
airline_sentiment
negative    132
neutral     113
positive     96
Name: count, dtype: int64

Part 4: Real Sentiment as Percentages#

sentiment_pct = round(sentiment_counts / len(tweets) * 100, 1)
sentiment_pct
airline_sentiment
negative    38.7
neutral     33.1
positive    28.2
Name: count, dtype: float64

Part 5: Visualizing Real Sentiment Breakdown#

plt.figure(figsize=(7, 5))
colors = {'negative': 'crimson', 'neutral': 'gray', 'positive': 'seagreen'}
sentiment_counts.plot(kind='bar', color=[colors[s] for s in sentiment_counts.index])
plt.ylabel('Real Tweet Count')
plt.title('Real Customer Sentiment Breakdown')
plt.xticks(rotation=0)
plt.tight_layout()
plt.savefig('sentiment_breakdown.png', dpi=120)
plt.close()

Part 6: Real Negative Reason Breakdown#

reason_counts = tweets['negativereason'].value_counts()
reason_counts
negativereason
Customer Service Issue         42
Flight Booking Problems        22
Cancelled Flight               16
Can't Tell                     16
Late Flight                    13
Bad Flight                     10
Lost Luggage                    5
Flight Attendant Complaints     4
Damaged Luggage                 3
longlines                       1
Name: count, dtype: int64

Part 7: Visualizing Real Complaint Reasons#

plt.figure(figsize=(9, 6))
reason_counts.plot(kind='barh', color='darkorange')
plt.xlabel('Real Tweet Count')
plt.title('Real Negative Tweet Reasons')
plt.gca().invert_yaxis()
plt.tight_layout()
plt.savefig('negative_reasons.png', dpi=120)
plt.close()

Part 8: Real Top Complaint Reason#

top_reason = reason_counts.idxmax()
top_reason_pct = round(reason_counts.max() / reason_counts.sum() * 100, 1)
top_reason, top_reason_pct
('Customer Service Issue', np.float64(31.8))

Part 9: Real Sentiment Confidence Distribution#

tweets['airline_sentiment_confidence'].describe().round(3)
count    341.000
mean       0.888
std        0.166
min        0.348
25%        0.681
50%        1.000
75%        1.000
max        1.000
Name: airline_sentiment_confidence, dtype: float64

Part 10: Real Low-Confidence Tweets#

low_confidence = tweets[tweets['airline_sentiment_confidence'] < 0.6]
len(low_confidence)
low_confidence[['airline_sentiment', 'airline_sentiment_confidence', 'text']].head(3)
airline_sentiment airline_sentiment_confidence text
1 positive 0.3486 @VirginAmerica plus you've added commercials t...
114 positive 0.3482 @VirginAmerica come back to #PHL already. We n...
142 neutral 0.3550 @VirginAmerica Can you find us a flt out of LA...

Part 11: Real Negative Reason Confidence#

tweets['negativereason_confidence'].describe().round(3)
count    167.000
mean       0.577
std        0.359
min        0.000
25%        0.350
50%        0.664
75%        1.000
max        1.000
Name: negativereason_confidence, dtype: float64

Part 12: Real Retweet Volume by Sentiment#

retweets_by_sentiment = tweets.groupby('airline_sentiment')['retweet_count'].mean().round(3)
retweets_by_sentiment
airline_sentiment
negative    0.008
neutral     0.044
positive    0.042
Name: retweet_count, dtype: float64

Part 13: Real Most-Retweeted Tweet#

most_retweeted = tweets.loc[tweets['retweet_count'].idxmax()]
most_retweeted[['airline_sentiment', 'retweet_count', 'text']]
airline_sentiment                                             positive
retweet_count                                                        2
text                 Always have it together!!! You're welcome! RT ...
Name: 147, dtype: object

Part 14: Real Word Count by Sentiment#

tweets['word_count'] = tweets['text'].apply(lambda t: len(str(t).split()))
wordcount_by_sentiment = tweets.groupby('airline_sentiment')['word_count'].mean().round(2)
wordcount_by_sentiment
airline_sentiment
negative    19.10
neutral     13.25
positive    13.27
Name: word_count, dtype: float64

Part 15: Real Timezone Distribution#

timezone_counts = tweets['user_timezone'].value_counts()
timezone_counts.head(10)
user_timezone
Pacific Time (US & Canada)     96
Eastern Time (US & Canada)     79
Central Time (US & Canada)     25
Sydney                         11
Quito                           9
Mountain Time (US & Canada)     7
Atlantic Time (Canada)          7
Arizona                         7
Alaska                          3
Caracas                         3
Name: count, dtype: int64

Part 16: Real Sentiment by Top Timezone#

top_timezones = timezone_counts.head(3).index
timezone_sentiment = tweets[tweets['user_timezone'].isin(top_timezones)].groupby(['user_timezone', 'airline_sentiment']).size().unstack(fill_value=0)
timezone_sentiment
airline_sentiment negative neutral positive
user_timezone
Central Time (US & Canada) 7 11 7
Eastern Time (US & Canada) 35 22 22
Pacific Time (US & Canada) 42 22 32

Part 17: Real Tweet Location Field#

has_location = tweets['tweet_location'].notna().sum()
has_location_pct = round(has_location / len(tweets) * 100, 1)
has_location_pct
np.float64(75.7)

Part 18: Real Missing Data Audit#

missing_summary = tweets.isna().sum().sort_values(ascending=False)
missing_summary
airline_sentiment_gold          341
negativereason_gold             341
tweet_coord                     308
negativereason                  209
negativereason_confidence       174
tweet_location                   83
user_timezone                    82
tweet_id                          0
name                              0
airline                           0
airline_sentiment                 0
airline_sentiment_confidence      0
text                              0
retweet_count                     0
tweet_created                     0
word_count                        0
dtype: int64

Part 19: Real Negative Reason Given Confidence Threshold#

high_confidence_reasons = tweets[tweets['negativereason_confidence'] >= 0.8]
high_confidence_reasons['negativereason'].value_counts()
negativereason
Customer Service Issue     16
Can't Tell                  6
Cancelled Flight            6
Flight Booking Problems     4
Late Flight                 4
Lost Luggage                4
Bad Flight                  3
Damaged Luggage             2
Name: count, dtype: int64

Part 20: Real Customer Service Complaints Deep Dive#

cs_complaints = tweets[tweets['negativereason'] == 'Customer Service Issue']
cs_complaints[['text']].head(5)
text
24 @VirginAmerica you guys messed up my seating.....
25 @VirginAmerica status match program. I applie...
32 @VirginAmerica help, left expensive headphones...
33 @VirginAmerica awaiting my return phone call, ...
39 @VirginAmerica Your chat support is not workin...

Part 21: Real Average Confidence by Complaint Type#

confidence_by_reason = tweets.groupby('negativereason')['negativereason_confidence'].mean().round(3).sort_values(ascending=False)
confidence_by_reason
negativereason
Damaged Luggage                0.883
Lost Luggage                   0.867
Can't Tell                     0.766
Customer Service Issue         0.759
Cancelled Flight               0.756
Bad Flight                     0.717
longlines                      0.682
Flight Booking Problems        0.673
Late Flight                    0.652
Flight Attendant Complaints    0.507
Name: negativereason_confidence, dtype: float64

Part 22: Real Positive Tweet Sample#

positive_tweets = tweets[tweets['airline_sentiment'] == 'positive']
positive_tweets[['text']].sample(5, random_state=42)
text
273 @VirginAmerica cutest salt and pepper shaker e...
264 @VirginAmerica thanks for gate checking my bag...
248 @VirginAmerica love the 90s music blasting at ...
332 @VirginAmerica thanks for taking care of @Suup...
117 @VirginAmerica and again! Another rep kicked b...

Part 23: Real Mention Pattern Check#

tweets['mentions_airline'] = tweets['text'].str.contains('@VirginAmerica', case=False, na=False)
mention_rate = round(tweets['mentions_airline'].mean() * 100, 1)
mention_rate
np.float64(100.0)

Part 24: Real Sentiment Among Direct Mentions#

mention_sentiment = tweets[tweets['mentions_airline']]['airline_sentiment'].value_counts(normalize=True).round(3) * 100
mention_sentiment
airline_sentiment
negative    38.7
neutral     33.1
positive    28.2
Name: proportion, dtype: float64

Part 25: Real Word Count vs Confidence Correlation#

wordcount_confidence_corr = tweets['word_count'].corr(tweets['airline_sentiment_confidence'])
round(wordcount_confidence_corr, 3)
np.float64(0.053)

Part 26: Saving the Real Complaint Summary Table#

complaint_summary = pd.DataFrame({'count': reason_counts, 'avg_confidence': confidence_by_reason}).round(3)
complaint_summary.to_csv('complaint_reason_summary.csv')
reloaded_summary = pd.read_csv('complaint_reason_summary.csv', index_col=0)
reloaded_summary.shape == complaint_summary.shape
True

Part 27: Real Sanity Check, Sentiment Counts Sum Correctly#

sentiment_counts.sum() == len(tweets)
np.True_

Part 28: Real Sanity Check, Confidence Scores in Range#

all((0 <= tweets['airline_sentiment_confidence']) & (tweets['airline_sentiment_confidence'] <= 1))
True

Part 29: Real Priority Recommendation#

negative_share = sentiment_pct['negative']
print(f'With {negative_share}% of tweets negative and {top_reason} the leading real complaint reason at {top_reason_pct}% of negatives, that is the real single area most worth investigating first.')
With 38.7% of tweets negative and Customer Service Issue the leading real complaint reason at 31.8% of negatives, that is the real single area most worth investigating first.

Part 30: Real Recap Print#

print(f'Across {len(tweets)} real customer tweets, {negative_share}% were negative, with {top_reason} driving {top_reason_pct}% of those real complaints.')
Across 341 real customer tweets, 38.7% were negative, with Customer Service Issue driving 31.8% of those real complaints.

Wrap-Up: What You Learned#

  • Hand-labeled sentiment and complaint-reason fields, when available, are far richer than running sentiment analysis from scratch on raw real text.
  • Confidence scores attached to real labels are worth auditing separately, low-confidence rows deserve real skepticism before trusting them.
  • Retweet volume revealed which real sentiment category actually spreads furthest, a genuinely different question from which is most common.
  • Cross-tabulating sentiment against fields like timezone or direct-mention status can surface real patterns a single value count would miss.
  • A single real brand, deeply analyzed, can be just as instructive as a multi-brand comparison when the real data honestly supports that scope.
  • Next video: capstone, a real social media campaign performance report pulling together the real domain from this entire section.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.