Mathew K Analytics

Lesson 69 · Python Fundamentals

Customer Segmentation Using PCA and KMeans Clustering in Python: A Step-by-Step Guide

Welcome! Today you will learn how to group customers using real data. You will use important tools called PCA and KMeans. These help discover patterns and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Customer Segmentation with PCA and KMeans in Python#

Welcome! Today you will learn how to group customers using real data.

You will use important tools called PCA and KMeans.

These help discover patterns and clusters in messy information.

No experience is needed! Let us explore hands-on.

# Let us start by making sure warnings do not distract us
import warnings
warnings.filterwarnings('ignore')

# Now, import pandas for easy table data
import pandas as pd

What is customer segmentation?#

Companies often want to group customers by habits or traits.

Grouping helps them target special offers and better understand shoppers.

Today, you will use real shopper data to see clusters and patterns.

# Data setup
url = "https://raw.githubusercontent.com/mwaskom/seaborn-data/master/mall_customers.csv"
df = pd.read_csv(url)
print("Data shape:", df.shape)
df.head()
---------------------------------------------------------------------------
HTTPError                                 Traceback (most recent call last)
Cell In[2], line 3
      1 # Data setup
      2 url = "https://raw.githubusercontent.com/mwaskom/seaborn-data/master/mall_customers.csv"
----> 3 df = pd.read_csv(url)
      4 print("Data shape:", df.shape)
      5 df.head()

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\io\parsers\readers.py:1026, in read_csv(filepath_or_buffer, sep, delimiter, header, names, index_col, usecols, dtype, engine, converters, true_values, false_values, skipinitialspace, skiprows, skipfooter, nrows, na_values, keep_default_na, na_filter, verbose, skip_blank_lines, parse_dates, infer_datetime_format, keep_date_col, date_parser, date_format, dayfirst, cache_dates, iterator, chunksize, compression, thousands, decimal, lineterminator, quotechar, quoting, doublequote, escapechar, comment, encoding, encoding_errors, dialect, on_bad_lines, delim_whitespace, low_memory, memory_map, float_precision, storage_options, dtype_backend)
   1013 kwds_defaults = _refine_defaults_read(
   1014     dialect,
   1015     delimiter,
   (...)
   1022     dtype_backend=dtype_backend,
   1023 )
   1024 kwds.update(kwds_defaults)
-> 1026 return _read(filepath_or_buffer, kwds)

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\io\parsers\readers.py:620, in _read(filepath_or_buffer, kwds)
    617 _validate_names(kwds.get("names", None))
    619 # Create the parser.
--> 620 parser = TextFileReader(filepath_or_buffer, **kwds)
    622 if chunksize or iterator:
    623     return parser

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\io\parsers\readers.py:1620, in TextFileReader.__init__(self, f, engine, **kwds)
   1617     self.options["has_index_names"] = kwds["has_index_names"]
   1619 self.handles: IOHandles | None = None
-> 1620 self._engine = self._make_engine(f, self.engine)

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\io\parsers\readers.py:1880, in TextFileReader._make_engine(self, f, engine)
   1878     if "b" not in mode:
   1879         mode += "b"
-> 1880 self.handles = get_handle(
   1881     f,
   1882     mode,
   1883     encoding=self.options.get("encoding", None),
   1884     compression=self.options.get("compression", None),
   1885     memory_map=self.options.get("memory_map", False),
   1886     is_text=is_text,
   1887     errors=self.options.get("encoding_errors", "strict"),
   1888     storage_options=self.options.get("storage_options", None),
   1889 )
   1890 assert self.handles is not None
   1891 f = self.handles.handle

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\io\common.py:728, in get_handle(path_or_buf, mode, encoding, compression, memory_map, is_text, errors, storage_options)
    725     codecs.lookup_error(errors)
    727 # open URLs
--> 728 ioargs = _get_filepath_or_buffer(
    729     path_or_buf,
    730     encoding=encoding,
    731     compression=compression,
    732     mode=mode,
    733     storage_options=storage_options,
    734 )
    736 handle = ioargs.filepath_or_buffer
    737 handles: list[BaseBuffer]

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\io\common.py:384, in _get_filepath_or_buffer(filepath_or_buffer, encoding, compression, mode, storage_options)
    382 # assuming storage_options is to be interpreted as headers
    383 req_info = urllib.request.Request(filepath_or_buffer, headers=storage_options)
--> 384 with urlopen(req_info) as req:
    385     content_encoding = req.headers.get("Content-Encoding", None)
    386     if content_encoding == "gzip":
    387         # Override compression based on Content-Encoding header

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\io\common.py:289, in urlopen(*args, **kwargs)
    283 """
    284 Lazy-import wrapper for stdlib urlopen, as that imports a big chunk of
    285 the stdlib.
    286 """
    287 import urllib.request
--> 289 return urllib.request.urlopen(*args, **kwargs)

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\urllib\request.py:215, in urlopen(url, data, timeout, cafile, capath, cadefault, context)
    213 else:
    214     opener = _opener
--> 215 return opener.open(url, data, timeout)

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\urllib\request.py:521, in OpenerDirector.open(self, fullurl, data, timeout)
    519 for processor in self.process_response.get(protocol, []):
    520     meth = getattr(processor, meth_name)
--> 521     response = meth(req, response)
    523 return response

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\urllib\request.py:630, in HTTPErrorProcessor.http_response(self, request, response)
    627 # According to RFC 2616, "2xx" code indicates that the client's
    628 # request was successfully received, understood, and accepted.
    629 if not (200 <= code < 300):
--> 630     response = self.parent.error(
    631         'http', request, response, code, msg, hdrs)
    633 return response

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\urllib\request.py:559, in OpenerDirector.error(self, proto, *args)
    557 if http_err:
    558     args = (dict, 'default', 'http_error_default') + orig_args
--> 559     return self._call_chain(*args)

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\urllib\request.py:492, in OpenerDirector._call_chain(self, chain, kind, meth_name, *args)
    490 for handler in handlers:
    491     func = getattr(handler, meth_name)
--> 492     result = func(*args)
    493     if result is not None:
    494         return result

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\urllib\request.py:639, in HTTPDefaultErrorHandler.http_error_default(self, req, fp, code, msg, hdrs)
    638 def http_error_default(self, req, fp, code, msg, hdrs):
--> 639     raise HTTPError(req.full_url, code, msg, hdrs, fp)

HTTPError: HTTP Error 404: Not Found

Exploring the dataset#

Understanding the info inside your data is a huge first step.

Check out the features:

  • CustomerID: Unique number for each shopper
  • Genre: Male or Female
  • Age: Years old
  • Annual Income (k$): Yearly earnings, in thousands
  • Spending Score (1-100): How much they spend

Let us look a bit deeper.

# Describe gives us basic stats for numbers
df.describe()
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[3], line 2
      1 # Describe gives us basic stats for numbers
----> 2 df.describe()

NameError: name 'df' is not defined
# Let us peek at unique genres
print(df['Genre'].value_counts())
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[4], line 2
      1 # Let us peek at unique genres
----> 2 print(df['Genre'].value_counts())

NameError: name 'df' is not defined

Getting data ready for analysis#

Some analysis tools only understand numbers, not words.

So, we must turn 'Genre' into a number first.

This is called encoding.

# Convert Genre: Female -> 0, Male -> 1
df['Gender_Code'] = df['Genre'].map({'Female': 0, 'Male': 1})
df[['Genre', 'Gender_Code']].head()
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[5], line 2
      1 # Convert Genre: Female -> 0, Male -> 1
----> 2 df['Gender_Code'] = df['Genre'].map({'Female': 0, 'Male': 1})
      3 df[['Genre', 'Gender_Code']].head()

NameError: name 'df' is not defined
# Pick only numeric features for our clustering
features = [
    'Age',
    'Annual Income (k$)',
    'Spending Score (1-100)',
    'Gender_Code'
]
X = df[features]
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[6], line 8
      1 # Pick only numeric features for our clustering
      2 features = [
      3     'Age',
      4     'Annual Income (k$)',
      5     'Spending Score (1-100)',
      6     'Gender_Code'
      7 ]
----> 8 X = df[features]

NameError: name 'df' is not defined

What is PCA?#

PCA stands for Principal Component Analysis.

It helps us turn messy features into simple patterns.

PCA finds the big trends and lets us see data in fewer dimensions.

This helps with exploring and clustering.

# Let us run PCA to shrink our features to 2 principal components
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA

# Scale our data so every column is fair
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

pca = PCA(n_components=2, random_state=0)
X_pca = pca.fit_transform(X_scaled)
print("PCA-shape:", X_pca.shape)
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[7], line 7
      5 # Scale our data so every column is fair
      6 scaler = StandardScaler()
----> 7 X_scaled = scaler.fit_transform(X)
      9 pca = PCA(n_components=2, random_state=0)
     10 X_pca = pca.fit_transform(X_scaled)

NameError: name 'X' is not defined
# Plot the two main PCA axes to see any clusters
import matplotlib.pyplot as plt

plt.figure(figsize=(7,6))
plt.scatter(X_pca[:, 0], X_pca[:, 1], c='skyblue', alpha=0.5)
plt.xlabel('PC1')
plt.ylabel('PC2')
plt.title('Customers after PCA')
plt.show()
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[8], line 5
      2 import matplotlib.pyplot as plt
      4 plt.figure(figsize=(7,6))
----> 5 plt.scatter(X_pca[:, 0], X_pca[:, 1], c='skyblue', alpha=0.5)
      6 plt.xlabel('PC1')
      7 plt.ylabel('PC2')

NameError: name 'X_pca' is not defined
<Figure size 700x600 with 0 Axes>

Time for clustering! What is KMeans?#

KMeans finds groups (clusters) by guessing where group centers might be.

KMeans keeps moving the centers to fit real patterns.

We will start with 5 clusters, just for practice.

# Group customers into 5 clusters using KMeans
from sklearn.cluster import KMeans

kmeans = KMeans(n_clusters=5, random_state=0)
clusters = kmeans.fit_predict(X_pca)

df['Cluster'] = clusters
df[['CustomerID', 'Cluster']].head()
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[9], line 5
      2 from sklearn.cluster import KMeans
      4 kmeans = KMeans(n_clusters=5, random_state=0)
----> 5 clusters = kmeans.fit_predict(X_pca)
      7 df['Cluster'] = clusters
      8 df[['CustomerID', 'Cluster']].head()

NameError: name 'X_pca' is not defined
# Visualize clusters in PCA space
plt.figure(figsize=(8,6))
scatter = plt.scatter(X_pca[:, 0], X_pca[:, 1], c=clusters, cmap='Set2', alpha=0.7)
plt.xlabel('PC1')
plt.ylabel('PC2')
plt.title('Customer Clusters (KMeans, k=5)')
plt.colorbar(scatter, label='Cluster')
plt.show()
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[10], line 3
      1 # Visualize clusters in PCA space
      2 plt.figure(figsize=(8,6))
----> 3 scatter = plt.scatter(X_pca[:, 0], X_pca[:, 1], c=clusters, cmap='Set2', alpha=0.7)
      4 plt.xlabel('PC1')
      5 plt.ylabel('PC2')

NameError: name 'X_pca' is not defined
<Figure size 800x600 with 0 Axes>

Real-world use: What are clusters good for?#

Customer groups can help design special offers.

Managers might focus on high spenders or certain ages.

Not all clusters mean profit always explore with a goal in mind.

# See what an average customer looks like in each cluster
df.groupby('Cluster')[features].mean()
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[11], line 2
      1 # See what an average customer looks like in each cluster
----> 2 df.groupby('Cluster')[features].mean()

NameError: name 'df' is not defined
# Let us pick a cluster to analyze up-close
cluster_num = int(input("Enter a cluster number (0-4) to inspect: "))
sample = df[df['Cluster'] == cluster_num].head()
sample
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[12], line 3
      1 # Let us pick a cluster to analyze up-close
      2 cluster_num = int(input("Enter a cluster number (0-4) to inspect: "))
----> 3 sample = df[df['Cluster'] == cluster_num].head()
      4 sample

NameError: name 'df' is not defined
 
# Mini-project: Build your own segmentation rule
print("Let us make a group for young, high-spending customers.")
mask = (df['Age'] < 30) & (df['Spending Score (1-100)'] > 60)
young_high_spenders = df[mask]
print("There are", len(young_high_spenders), "customers in this group.")
young_high_spenders.head()
Let us make a group for young, high-spending customers.
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[13], line 3
      1 # Mini-project: Build your own segmentation rule
      2 print("Let us make a group for young, high-spending customers.")
----> 3 mask = (df['Age'] < 30) & (df['Spending Score (1-100)'] > 60)
      4 young_high_spenders = df[mask]
      5 print("There are", len(young_high_spenders), "customers in this group.")

NameError: name 'df' is not defined
# Troubleshooting: What if you get an error?
# Common issues: missing quotes, wrong spelling, wrong column names.
# If an error appears, read the message. It tells you what is broken.
# Try to match the error's words to your code to find the fix.

# Let us try an error. Uncomment the next line to see one.
# print(df['NotAColumn'].head())
# Best practices: Always check for missing data
missing = df.isnull().sum()
print("Missing values per column:")
print(missing)
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[15], line 2
      1 # Best practices: Always check for missing data
----> 2 missing = df.isnull().sum()
      3 print("Missing values per column:")
      4 print(missing)

NameError: name 'df' is not defined
# Extra tip: Try more or fewer clusters
for k in [2, 3, 6]:
    kmeans = KMeans(n_clusters=k, random_state=0)
    clusters_try = kmeans.fit_predict(X_pca)
    print(f"Number of clusters: {k} - Unique labels: {len(set(clusters_try))}")
    
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[16], line 4
      2 for k in [2, 3, 6]:
      3     kmeans = KMeans(n_clusters=k, random_state=0)
----> 4     clusters_try = kmeans.fit_predict(X_pca)
      5     print(f"Number of clusters: {k} - Unique labels: {len(set(clusters_try))}")

NameError: name 'X_pca' is not defined

Recap and next steps#

You have learned how to load, prepare, and group real customer data.

You tried PCA for simplifying features.

KMeans found pattern groups to help businesses take action.

Keep exploring and practicing new datasets!

Try more on your own!#

Can you:

  • Plot age versus spending score for one cluster?
  • Try other real datasets like Titanic or Iris?
  • Change PCA to 3 components and plot?

If you enjoyed this, subscribe for more Python, and share your clusters below!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.