Lesson 44 · Data Mining
Sentiment Analysis with Text Mining: A Practical Data Science Capstone Project
We will load a dataset and preview it. Welcome! In this lesson, we will use Python to perform sentiment analysis on text data. We will learn how to…
- CourseData Mining
- Lesson44 of 31
- Video21 min
- FormatJupyter notebook · 17 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbData setup (Telecom Customer Churn Dataset)#
We will load a dataset and preview it.
import numpy as np
np.random.seed(42)
import pandas as pd
url = 'https://raw.githubusercontent.com/IBM/telco-customer-churn-on-icp4d/master/data/Telco-Customer-Churn.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
Week 12 Capstone Project: Sentiment Analysis using Text Mining#
Welcome! In this lesson, we will use Python to perform sentiment analysis on text data.
We will learn how to preprocess, explore, and model text for mining opinions and emotions.
No experience is needed just follow along, try examples, and have fun learning!
Let us get started.
# Suppress warnings for clean output
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings('ignore')
What is Sentiment Analysis?#
Sentiment analysis is a way to figure out if a piece of text has a positive, negative, or neutral feeling.
It is used in many real-world settings, like finding out if customer reviews are happy or not.
Text mining turns words into numbers so we can use computers to analyze them.
# Data setup (We will use a small sample of movie reviews for fast learning)
import pandas as pd
data = {
'review': [
"This movie was amazing! I loved it.",
"Terrible movie. It was boring and slow.",
"It was okay, not bad but not great.",
"Fantastic performance and beautiful music.",
"Waste of time. Do not watch it.",
"Best movie ever! Highly recommend.",
"I did not enjoy the film.",
"The plot was predictable but the acting was good.",
"Mediocre experience. Seen better movies.",
"Loved the cinematography and the story."
],
'sentiment': [
"positive", "negative", "neutral", "positive", "negative",
"positive", "negative", "neutral", "neutral", "positive"
]
}
df = pd.DataFrame(data)
print(df.shape)
print(df.head())
# Check for missing values
df.isnull().sum()
# Remove punctuation and lowercase all text
import string
df['clean_review'] = df['review'].apply(lambda x: x.lower().translate(str.maketrans('', '', string.punctuation)))
print(df[['review', 'clean_review']].head())
# Tokenize the words (split each review into a list of words)
df['tokens'] = df['clean_review'].apply(lambda x: x.split())
print(df[['clean_review', 'tokens']].head())
# Remove common stopwords
from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS
df['tokens_nostop'] = df['tokens'].apply(lambda words: [w for w in words if w not in ENGLISH_STOP_WORDS])
print(df[['tokens', 'tokens_nostop']].head())
# Count the most common words across all reviews
from collections import Counter
all_words = [word for tokens in df['tokens_nostop'] for word in tokens]
word_counts = Counter(all_words)
print(word_counts.most_common(5))
# Turn the text into numbers using Bag of Words
from sklearn.feature_extraction.text import CountVectorizer
vectorizer = CountVectorizer(stop_words='english')
X = vectorizer.fit_transform(df['clean_review'])
print(X.toarray()[:3])
print(vectorizer.get_feature_names_out())
# Split our data into training and test sets
from sklearn.model_selection import train_test_split
y = df['sentiment']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
print('Training set size:', X_train.shape[0])
print('Test set size:', X_test.shape[0])
# Build a simple Logistic Regression classifier
from sklearn.linear_model import LogisticRegression
clf = LogisticRegression(max_iter=200)
clf.fit(X_train, y_train)
train_score = clf.score(X_train, y_train)
test_score = clf.score(X_test, y_test)
print(f'Training accuracy: {train_score:.2f}')
print(f'Test accuracy: {test_score:.2f}')
# Make predictions on new reviews
example_reviews = [
"Absolutely wonderful! Will watch again.",
"Not worth my money or my time.",
"It was fine. Some good, some bad.",
"Spectacular visuals and gripping story!"
]
X_new = vectorizer.transform([r.lower().translate(str.maketrans('', '', string.punctuation)) for r in example_reviews])
predictions = clf.predict(X_new)
for review, pred in zip(example_reviews, predictions):
print(f"Review: {review}\n Predicted sentiment: {pred}\n")
Visualizing Word Frequencies#
Let us make a simple word cloud to see which words show up most.
This makes it easy to spot common themes in the reviews.
# Draw a word cloud of review words
from wordcloud import WordCloud
import matplotlib.pyplot as plt
wordcloud = WordCloud(width=600, height=300, background_color='white').generate(' '.join(all_words))
plt.figure(figsize=(10,5))
plt.imshow(wordcloud, interpolation='bilinear')
plt.axis('off')
plt.show()
# Try inputting your own review to test the model
user_review = input("Enter a review to predict sentiment: ")
clean_user = user_review.lower().translate(str.maketrans('', '', string.punctuation))
X_user = vectorizer.transform([clean_user])
user_pred = clf.predict(X_user)[0]
print(f'The predicted sentiment is: {user_pred}')
# Practice: Change a word in this review and see how the sentiment changes
prac_review = "Horrible acting but the music was good."
# Clean and predict
prac_clean = prac_review.lower().translate(str.maketrans('', '', string.punctuation))
X_prac = vectorizer.transform([prac_clean])
prac_result = clf.predict(X_prac)[0]
print(f'Original review: {prac_review}')
print(f'Predicted sentiment: {prac_result}')
# Tips for better sentiment models
# 1. Use a bigger dataset for more accurate results
# 2. Try more powerful models: RandomForest, SVM, or deep learning
# 3. Clean the text more (remove emojis, correct misspellings, etc.)
# 4. Try out TF-IDF features, which weigh rare words more
# 5. Explore libraries like spaCy or NLTK for advanced processing
# Challenge: Classify these reviews (try to guess before running!)
reviews = [
"I will skip this movie. Nothing memorable.",
"Super fun and creative. Loved every minute!",
"Not my cup of tea. Fell asleep halfway.",
"Surprisingly touching and heartfelt."
"It was just average, nothing special."
]
X_chall = vectorizer.transform([
r.lower().translate(str.maketrans('', '', string.punctuation)) for r in reviews
])
results = clf.predict(X_chall)
for r, p in zip(reviews, results):
print(f'Review: {r} -- Sentiment: {p}')
Recap and Next Steps#
Today, you learned how to
- Clean and process text for analysis
- Turn words into numbers
- Build a simple sentiment classifier
- Visualize word patterns
- Test and experiment with your own reviews
Sentiment analysis is powerful for understanding language online.
Keep experimenting and let your curiosity guide you!
Thank you for learning with us!
If you liked this lesson, please subscribe and share your results.
Try more exercises and keep exploring the world of data mining in Python.
See you in the next video!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



