Mathew K Analytics

Lesson 44 · Data Mining

Sentiment Analysis with Text Mining: A Practical Data Science Capstone Project

We will load a dataset and preview it. Welcome! In this lesson, we will use Python to perform sentiment analysis on text data. We will learn how to…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Data setup (Telecom Customer Churn Dataset)#

We will load a dataset and preview it.

import numpy as np
np.random.seed(42)
import pandas as pd
url = 'https://raw.githubusercontent.com/IBM/telco-customer-churn-on-icp4d/master/data/Telco-Customer-Churn.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(7043, 21)
   customerID  gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  7590-VHVEG  Female              0     Yes         No       1           No   
1  5575-GNVDE    Male              0      No         No      34          Yes   
2  3668-QPYBK    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity  ... DeviceProtection  \
0  No phone service             DSL             No  ...               No   
1                No             DSL            Yes  ...              Yes   
2                No             DSL            Yes  ...               No   

  TechSupport StreamingTV StreamingMovies        Contract PaperlessBilling  \
0          No          No              No  Month-to-month              Yes   
1          No          No              No        One year               No   
2          No          No              No  Month-to-month              Yes   

      PaymentMethod MonthlyCharges  TotalCharges Churn  
0  Electronic check          29.85         29.85    No  
1      Mailed check          56.95        1889.5    No  
2      Mailed check          53.85        108.15   Yes  

[3 rows x 21 columns]

Week 12 Capstone Project: Sentiment Analysis using Text Mining#

Welcome! In this lesson, we will use Python to perform sentiment analysis on text data.

We will learn how to preprocess, explore, and model text for mining opinions and emotions.

No experience is needed just follow along, try examples, and have fun learning!

Let us get started.

# Suppress warnings for clean output
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings('ignore')

What is Sentiment Analysis?#

Sentiment analysis is a way to figure out if a piece of text has a positive, negative, or neutral feeling.

It is used in many real-world settings, like finding out if customer reviews are happy or not.

Text mining turns words into numbers so we can use computers to analyze them.

# Data setup (We will use a small sample of movie reviews for fast learning)
import pandas as pd
data = {
    'review': [
        "This movie was amazing! I loved it.",
        "Terrible movie. It was boring and slow.",
        "It was okay, not bad but not great.",
        "Fantastic performance and beautiful music.",
        "Waste of time. Do not watch it.",
        "Best movie ever! Highly recommend.",
        "I did not enjoy the film.",
        "The plot was predictable but the acting was good.",
        "Mediocre experience. Seen better movies.",
        "Loved the cinematography and the story."
    ],
    'sentiment': [
        "positive", "negative", "neutral", "positive", "negative",
        "positive", "negative", "neutral", "neutral", "positive"
    ]
}
df = pd.DataFrame(data)
print(df.shape)
print(df.head())
(10, 2)
                                       review sentiment
0         This movie was amazing! I loved it.  positive
1     Terrible movie. It was boring and slow.  negative
2         It was okay, not bad but not great.   neutral
3  Fantastic performance and beautiful music.  positive
4             Waste of time. Do not watch it.  negative
# Check for missing values
df.isnull().sum()
review       0
sentiment    0
dtype: int64
# Remove punctuation and lowercase all text
import string
df['clean_review'] = df['review'].apply(lambda x: x.lower().translate(str.maketrans('', '', string.punctuation)))
print(df[['review', 'clean_review']].head())
                                       review  \
0         This movie was amazing! I loved it.   
1     Terrible movie. It was boring and slow.   
2         It was okay, not bad but not great.   
3  Fantastic performance and beautiful music.   
4             Waste of time. Do not watch it.   

                                clean_review  
0          this movie was amazing i loved it  
1      terrible movie it was boring and slow  
2          it was okay not bad but not great  
3  fantastic performance and beautiful music  
4              waste of time do not watch it  
# Tokenize the words (split each review into a list of words)
df['tokens'] = df['clean_review'].apply(lambda x: x.split())
print(df[['clean_review', 'tokens']].head())
                                clean_review  \
0          this movie was amazing i loved it   
1      terrible movie it was boring and slow   
2          it was okay not bad but not great   
3  fantastic performance and beautiful music   
4              waste of time do not watch it   

                                            tokens  
0        [this, movie, was, amazing, i, loved, it]  
1    [terrible, movie, it, was, boring, and, slow]  
2       [it, was, okay, not, bad, but, not, great]  
3  [fantastic, performance, and, beautiful, music]  
4            [waste, of, time, do, not, watch, it]  
# Remove common stopwords
from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS
df['tokens_nostop'] = df['tokens'].apply(lambda words: [w for w in words if w not in ENGLISH_STOP_WORDS])
print(df[['tokens', 'tokens_nostop']].head())
                                            tokens  \
0        [this, movie, was, amazing, i, loved, it]   
1    [terrible, movie, it, was, boring, and, slow]   
2       [it, was, okay, not, bad, but, not, great]   
3  [fantastic, performance, and, beautiful, music]   
4            [waste, of, time, do, not, watch, it]   

                                tokens_nostop  
0                     [movie, amazing, loved]  
1             [terrible, movie, boring, slow]  
2                          [okay, bad, great]  
3  [fantastic, performance, beautiful, music]  
4                        [waste, time, watch]  
# Count the most common words across all reviews
from collections import Counter
all_words = [word for tokens in df['tokens_nostop'] for word in tokens]
word_counts = Counter(all_words)
print(word_counts.most_common(5))
[('movie', 3), ('loved', 2), ('amazing', 1), ('terrible', 1), ('boring', 1)]
# Turn the text into numbers using Bag of Words
from sklearn.feature_extraction.text import CountVectorizer
vectorizer = CountVectorizer(stop_words='english')
X = vectorizer.fit_transform(df['clean_review'])
print(X.toarray()[:3])
print(vectorizer.get_feature_names_out())
[[0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0]
 [0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 1 0 1 0 0 0]
 [0 0 1 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0]]
['acting' 'amazing' 'bad' 'beautiful' 'best' 'better' 'boring'
 'cinematography' 'did' 'enjoy' 'experience' 'fantastic' 'film' 'good'
 'great' 'highly' 'loved' 'mediocre' 'movie' 'movies' 'music' 'okay'
 'performance' 'plot' 'predictable' 'recommend' 'seen' 'slow' 'story'
 'terrible' 'time' 'waste' 'watch']
# Split our data into training and test sets
from sklearn.model_selection import train_test_split
y = df['sentiment']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
print('Training set size:', X_train.shape[0])
print('Test set size:', X_test.shape[0])
Training set size: 7
Test set size: 3
# Build a simple Logistic Regression classifier
from sklearn.linear_model import LogisticRegression
clf = LogisticRegression(max_iter=200)
clf.fit(X_train, y_train)
train_score = clf.score(X_train, y_train)
test_score = clf.score(X_test, y_test)
print(f'Training accuracy: {train_score:.2f}')
print(f'Test accuracy: {test_score:.2f}')
Training accuracy: 1.00
Test accuracy: 0.33
# Make predictions on new reviews
example_reviews = [
    "Absolutely wonderful! Will watch again.",
    "Not worth my money or my time.",
    "It was fine. Some good, some bad.",
    "Spectacular visuals and gripping story!"
]
X_new = vectorizer.transform([r.lower().translate(str.maketrans('', '', string.punctuation)) for r in example_reviews])
predictions = clf.predict(X_new)
for review, pred in zip(example_reviews, predictions):
    print(f"Review: {review}\n Predicted sentiment: {pred}\n")
    
Review: Absolutely wonderful! Will watch again.
 Predicted sentiment: negative

Review: Not worth my money or my time.
 Predicted sentiment: negative

Review: It was fine. Some good, some bad.
 Predicted sentiment: neutral

Review: Spectacular visuals and gripping story!
 Predicted sentiment: positive

Visualizing Word Frequencies#

Let us make a simple word cloud to see which words show up most.

This makes it easy to spot common themes in the reviews.

# Draw a word cloud of review words
from wordcloud import WordCloud
import matplotlib.pyplot as plt
wordcloud = WordCloud(width=600, height=300, background_color='white').generate(' '.join(all_words))
plt.figure(figsize=(10,5))
plt.imshow(wordcloud, interpolation='bilinear')
plt.axis('off')
plt.show()
---------------------------------------------------------------------------
ModuleNotFoundError                       Traceback (most recent call last)
Cell In[13], line 2
      1 # Draw a word cloud of review words
----> 2 from wordcloud import WordCloud
      3 import matplotlib.pyplot as plt
      4 wordcloud = WordCloud(width=600, height=300, background_color='white').generate(' '.join(all_words))

ModuleNotFoundError: No module named 'wordcloud'
# Try inputting your own review to test the model
user_review = input("Enter a review to predict sentiment: ")
clean_user = user_review.lower().translate(str.maketrans('', '', string.punctuation))
X_user = vectorizer.transform([clean_user])
user_pred = clf.predict(X_user)[0]
print(f'The predicted sentiment is: {user_pred}')
The predicted sentiment is: negative
# Practice: Change a word in this review and see how the sentiment changes
prac_review = "Horrible acting but the music was good."
# Clean and predict
prac_clean = prac_review.lower().translate(str.maketrans('', '', string.punctuation))
X_prac = vectorizer.transform([prac_clean])
prac_result = clf.predict(X_prac)[0]
print(f'Original review: {prac_review}')
print(f'Predicted sentiment: {prac_result}')
Original review: Horrible acting but the music was good.
Predicted sentiment: neutral
# Tips for better sentiment models
# 1. Use a bigger dataset for more accurate results
# 2. Try more powerful models: RandomForest, SVM, or deep learning
# 3. Clean the text more (remove emojis, correct misspellings, etc.)
# 4. Try out TF-IDF features, which weigh rare words more
# 5. Explore libraries like spaCy or NLTK for advanced processing
# Challenge: Classify these reviews (try to guess before running!)
reviews = [
    "I will skip this movie. Nothing memorable.",
    "Super fun and creative. Loved every minute!",
    "Not my cup of tea. Fell asleep halfway.",
    "Surprisingly touching and heartfelt."
    "It was just average, nothing special."
]
X_chall = vectorizer.transform([
    r.lower().translate(str.maketrans('', '', string.punctuation)) for r in reviews
])
results = clf.predict(X_chall)
for r, p in zip(reviews, results):
    print(f'Review: {r} -- Sentiment: {p}')
    
Review: I will skip this movie. Nothing memorable. -- Sentiment: positive
Review: Super fun and creative. Loved every minute! -- Sentiment: positive
Review: Not my cup of tea. Fell asleep halfway. -- Sentiment: positive
Review: Surprisingly touching and heartfelt.It was just average, nothing special. -- Sentiment: positive

Recap and Next Steps#

Today, you learned how to

  • Clean and process text for analysis
  • Turn words into numbers
  • Build a simple sentiment classifier
  • Visualize word patterns
  • Test and experiment with your own reviews

Sentiment analysis is powerful for understanding language online.

Keep experimenting and let your curiosity guide you!

Thank you for learning with us!

If you liked this lesson, please subscribe and share your results.

Try more exercises and keep exploring the world of data mining in Python.

See you in the next video!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.