Mathew K Analytics

Lesson 56 · Python for Data Science

1 - Supervised vs Unsupervised Learning in Python

In this lesson, we will discover what supervised and unsupervised learning are. We will see how Python can help you use these powerful machine learning…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb
 

Welcome to Beginner Python: Supervised vs Unsupervised Learning#

In this lesson, we will discover what supervised and unsupervised learning are. We will see how Python can help you use these powerful machine learning techniques.

By the end, you will know the basic differences and try out some practical code yourself!

# Let's begin with machine learning.
# Machine learning is how computers learn from data.
 
print("With supervised learning, we know the answer for each example.")
print("With unsupervised learning, we only have the data itself.")
With supervised learning, we know the answer for each example.
With unsupervised learning, we only have the data itself.

What is supervised learning?#

In supervised learning, you give the computer data and the correct answers.

For example: a list of fruits with their names, or pictures with the things in them.

The computer learns to predict answers for new data.

# Let's make a simple dataset for supervised learning.
fruits = ['apple', 'orange', 'banana', 'apple', 'orange']
labels = ['red', 'orange', 'yellow', 'red', 'orange']

print('Fruit:', fruits[0], '| Color:', labels[0])
print('Fruit:', fruits[2], '| Color:', labels[2])
Fruit: apple | Color: red
Fruit: banana | Color: yellow

What is unsupervised learning?#

Unsupervised learning is when you only have the data and no answers.

The computer tries to find patterns, groups, or structures in the data on its own.

For example: grouping customers by their shopping habits.

# Here is some shopping data for unsupervised learning.
shopping = [5, 8, 6, 50, 54, 48]

print('Shopping amounts:', shopping)
Shopping amounts: [5, 8, 6, 50, 54, 48]
# Let's compare supervised and unsupervised data.
print('Supervised has answers:')
print('fruits:', fruits)
print('labels:', labels)
print('Unsupervised just has data:')
print('shopping:', shopping)
Supervised has answers:
fruits: ['apple', 'orange', 'banana', 'apple', 'orange']
labels: ['red', 'orange', 'yellow', 'red', 'orange']
Unsupervised just has data:
shopping: [5, 8, 6, 50, 54, 48]

How do we use Python for supervised learning?#

Supervised learning uses labeled data to train a model.

Common models: classification and regression.

We will try the scikit-learn library for easy examples.

# Let's import the library we need.
from sklearn.tree import DecisionTreeClassifier
# Here is a simple supervised learning example.
X = [[1], [2], [3], [4]]  # features, like fruit size
y = ['apple', 'apple', 'orange', 'orange']  # labels
 
clf = DecisionTreeClassifier()
clf.fit(X, y)  # train the model
 
print('Predict size 1:', clf.predict([[1]]))
print('Predict size 4:', clf.predict([[4]]))
Predict size 1: ['apple']
Predict size 4: ['orange']

How do we use Python for unsupervised learning?#

Unsupervised learning lets the computer find new patterns in data without labels.

We will use the KMeans method from scikit-learn to group numbers.

# Import KMeans for clustering.
from sklearn.cluster import KMeans
# Now we find groups in our shopping data.
import numpy as np
shopping_array = np.array(shopping).reshape(-1, 1)

kmeans = KMeans(n_clusters=2, random_state=0)
kmeans.fit(shopping_array)

print('Predicted group for numbers:', kmeans.labels_)
Predicted group for numbers: [1 1 1 0 0 0]
# What happens if you give new data?
new_shopping = np.array([[7], [55]])
group = kmeans.predict(new_shopping)

print('Groups for new data points:', group)
Groups for new data points: [1 0]

Summary for supervised and unsupervised learning#

  • Supervised learning uses data and answers.

  • Unsupervised learning uses only data.

Python makes both easy with the right tools!

# Challenge: classify animals by size (supervised learning)!
animals = [1, 10, 15]
labels = ['mouse', 'dog', 'sheep']

X_new = [[2], [12]]

animal_model = DecisionTreeClassifier()
animal_model.fit([[a] for a in animals], labels)

print('Classify size 2:', animal_model.predict([[2]]))
print('Classify size 12:', animal_model.predict([[12]]))
Classify size 2: ['mouse']
Classify size 12: ['dog']
# Challenge: group people by height (unsupervised learning)!
heights = [145, 150, 160, 185, 190, 200]
heights_array = np.array(heights).reshape(-1, 1)
height_groups = KMeans(n_clusters=2, random_state=10)
height_groups.fit(heights_array)
print('Groups by height:', height_groups.labels_)
print('Which cluster is 170 in?', height_groups.predict([[170]]))
Groups by height: [1 1 1 0 0 0]
Which cluster is 170 in? [1]

Mini-Project Part 1: Fruit Color Predictor#

Imagine you work in a fruit market. You want a program that predicts the fruit's color from its size.

Let's prepare a dataset together.

# Sample data: fruit size and colors
sizes = [[5], [7], [8], [11]]
colors = ['green', 'orange', 'yellow', 'red']

color_model = DecisionTreeClassifier()
color_model.fit(sizes, colors)

print('Predict for size 7:', color_model.predict([[7]]))
print('Predict for size 10:', color_model.predict([[10]]))
Predict for size 7: ['orange']
Predict for size 10: ['red']

Mini-Project Part 2: Find Unusual Shopping Behavior#

Your store wants to spot shoppers who buy way more than others.

Let's use unsupervised learning to discover these shoppers.

# Store shopping amounts for customers.
store_shopping = [10, 12, 13, 100, 110, 105]
shopping_data = np.array(store_shopping).reshape(-1, 1)

shopper_kmeans = KMeans(n_clusters=2, random_state=15)
shopper_kmeans.fit(shopping_data)
 
for i, amount in enumerate(store_shopping):
    print('Amount $' + str(amount), '-> Group', shopper_kmeans.labels_[i])
    
Amount $10 -> Group 1
Amount $12 -> Group 1
Amount $13 -> Group 1
Amount $100 -> Group 0
Amount $110 -> Group 0
Amount $105 -> Group 0
# Check a new customer's spending group.
amount_input = input('Enter new amount: ')
amount = float(amount_input)
predicted = shopper_kmeans.predict([[amount]])
print('This customer is in group:', predicted[0])
This customer is in group: 0
 

Troubleshooting and Common Mistakes#

  • Check that your data and labels line up in supervised learning.

  • In unsupervised learning, remember there are no labels.

  • Watch for misspelled variable names.

If you see an error, read it slowly to find the cause.

# Some useful tips and tricks!
print('You can try different models, like RandomForest or LogisticRegression!')
print('Try making your own datasets from hobbies or schoolwork.')
You can try different models, like RandomForest or LogisticRegression!
Try making your own datasets from hobbies or schoolwork.
# Final challenge: Can you sort your favorite movies by rating?
movies = ['Movie A', 'Movie B', 'Movie C']
ratings = [7, 9, 6]
sorted_movies = [m for _, m in sorted(zip(ratings, movies), reverse=True)]
print('Your movies from highest to lowest rating:')
for movie in sorted_movies:
    print(movie)
    
Your movies from highest to lowest rating:
Movie B
Movie A
Movie C

Recap: What did we learn?#

  • Supervised learning = learning with answers and labels.
  • Unsupervised learning = finding patterns with no answers.

With Python and scikit-learn, you can try both easily.

Great job sticking through the lesson!

Thanks for watching!#

If this was helpful, like the video, leave a comment, and subscribe for more Python tips.

Share with a friend who wants to learn machine learning, too!

See you in the next lesson!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.