Mathew K Analytics

Lesson 14 · Scikit-learn deep dive

Scikit-learn Tutorial #14: Text Data & Vectorisation

Video fourteen of the eighteen-part series: turning raw text into numbers a model can use. CountVectorizer, TfidfVectorizer, and a full text classification…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Scikit-learn Deep-Dive, Video 14: Text Data - Vectorization#

  • Video fourteen of the eighteen-part series: turning raw text into numbers a model can use.
  • CountVectorizer, TfidfVectorizer, and a full text classification pipeline.
  • Let's get into it.

Part 1: The Problem - Models Need Numbers, Not Raw Text#

documents = [
    'the product quality is excellent and shipping was fast',
    'terrible experience, the item arrived broken and late',
    'great value for the price, very happy with this purchase',
    'awful quality, would not recommend this product at all',
    'fast shipping and excellent customer service overall',
    'broken item, poor quality, very disappointing purchase'
]
labels = [1, 0, 1, 0, 1, 0]
print(len(documents), len(labels))
6 6

Part 2: CountVectorizer Basics - Bag of Words#

from sklearn.feature_extraction.text import CountVectorizer
count_vec = CountVectorizer()
X_counts = count_vec.fit_transform(documents)
print(X_counts.shape)
print(X_counts.toarray()[0][:10])
(6, 35)
[0 1 0 0 0 0 0 0 1 0]

Part 3: vocabulary_ and get_feature_names_out#

print(count_vec.vocabulary_['quality'])
feature_names = count_vec.get_feature_names_out()
print(len(feature_names))
print(sorted(feature_names)[:10])
23
35
['all', 'and', 'arrived', 'at', 'awful', 'broken', 'customer', 'disappointing', 'excellent', 'experience']

Part 4: stop_words and min_df / max_df#

filtered_vec = CountVectorizer(stop_words='english', min_df=1, max_df=0.9)
X_filtered = filtered_vec.fit_transform(documents)
print(X_filtered.shape[1], X_counts.shape[1])
print('the' in filtered_vec.vocabulary_)
23 35
False

Part 5: ngram_range - Capturing Word Order Context#

bigram_vec = CountVectorizer(stop_words='english', ngram_range=(1, 2))
X_bigram = bigram_vec.fit_transform(documents)
bigram_features = [f for f in bigram_vec.get_feature_names_out() if ' ' in f]
print(bigram_features[:5])
print(X_bigram.shape[1], X_filtered.shape[1])
['arrived broken', 'awful quality', 'broken item', 'broken late', 'customer service']
49 23

Part 6: TfidfVectorizer - Term Frequency-Inverse Document Frequency#

from sklearn.feature_extraction.text import TfidfVectorizer
tfidf_vec = TfidfVectorizer(stop_words='english')
X_tfidf = tfidf_vec.fit_transform(documents)
print(X_tfidf.shape)
print(X_tfidf.toarray()[0].round(3)[:8])
(6, 23)
[0.    0.    0.    0.    0.    0.461 0.    0.461]

Part 7: Comparing Count vs Tfidf on the Same Corpus#

quality_idx = filtered_vec.vocabulary_['quality']
tfidf_quality_idx = tfidf_vec.vocabulary_['quality']
print('count weight:', X_filtered.toarray()[:, quality_idx])
print('tfidf weight:', X_tfidf.toarray()[:, tfidf_quality_idx].round(3))
count weight: [1 0 0 1 0 1]
tfidf weight: [0.389 0.    0.    0.39  0.    0.326]

Part 8: transform on New, Unseen Text#

new_reviews = ['excellent quality and fast shipping', 'this is a completely novel gibberish sentence']
X_new = tfidf_vec.transform(new_reviews)
print(X_new.shape)
print(round(X_new[1].sum(), 3))
(2, 23)
0.0

Part 9: A Text Classification Pipeline#

from sklearn.pipeline import Pipeline
from sklearn.linear_model import LogisticRegression
text_pipe = Pipeline([
    ('tfidf', TfidfVectorizer(stop_words='english')),
    ('model', LogisticRegression(max_iter=1000))
])
text_pipe.fit(documents, labels)
print(text_pipe.predict(['fast shipping and excellent quality']))
[1]

Part 10: A Real Pattern - a Reusable build_text_classifier Function#

def build_text_classifier(ngram_range=(1, 1), max_features=None):
    return Pipeline([
        ('tfidf', TfidfVectorizer(stop_words='english', ngram_range=ngram_range, max_features=max_features)),
        ('model', LogisticRegression(max_iter=1000))
    ])
clf = build_text_classifier(ngram_range=(1, 2))
clf.fit(documents, labels)
print(clf.predict(['terrible broken item, very poor']))
[0]

Wrap-Up: What You Learned#

  • Every scikit-learn estimator expects numeric input, so raw text has to be vectorized first.
  • CountVectorizer builds a vocabulary and represents each document as a vector of raw word counts.
  • vocabulary_ and get_feature_names_out expose the learned vocabulary and its column ordering.
  • stop_words drops low-information common words; min_df/max_df filter out words that are too rare or too common.
  • ngram_range extends the vocabulary to multi-word phrases, capturing some word-order context.
  • TfidfVectorizer weights words by frequency in a document, scaled down by how common they are across the corpus.
  • Words appearing in nearly every document get shrunk toward zero weight under tfidf, unlike raw counts.
  • transform (never fit_transform) on new text reuses the learned vocabulary; unseen words are silently ignored.
  • A vectorizer drops directly into a Pipeline, taking raw text straight to a prediction in one object.
  • That wraps up text vectorization. Next up: Ensembling Meta-Estimators - Voting, Stacking, and Bagging.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.