Mathew K Analytics

Lesson 37 · Python for Data Science

5 - Feature Engineering Basics in Python

Welcome! In this lesson, we will explore how feature engineering helps machines learn from data. We will use simple Python code you can run, change, and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Feature Engineering Basics in Python#

Welcome! In this lesson, we will explore how feature engineering helps machines learn from data.

We will use simple Python code you can run, change, and play with.

Feature engineering means creating or changing data, so it is more useful for analysis or machine learning.

Let us get started and see why this is so important!

What is Feature Engineering?#

Think of feature engineering as giving the computer better clues. We take raw data and turn it into new shapes and forms.

This helps computer programs learn faster and make better decisions!

# Let us start with a simple data list.
ages = [17, 21, 34, 42, 15, 29]

# Print the data.
print('Original ages data:', ages)
Original ages data: [17, 21, 34, 42, 15, 29]
# Feature engineering often means creating new features.
# For example, let us find out who is an adult.

is_adult = []

for age in ages:
    if age >= 18:
        is_adult.append(True)
    else:
        is_adult.append(False)

print('Is each person an adult?', is_adult)
Is each person an adult? [False, True, True, True, False, True]

Why Create New Features?#

Often, original data does not give us all the clues we need.

By adding new features, we help the computer find patterns and make smarter choices.

This is like giving someone better hints to solve a puzzle.

# Let us add two lists together to make a dataset.
names = ['Anna', 'Ben', 'Cleo', 'Dan', 'Elle', 'Felix']

for name, age, adult in zip(names, ages, is_adult):
    print(name, age, adult)
    
Anna 17 False
Ben 21 True
Cleo 34 True
Dan 42 True
Elle 15 False
Felix 29 True
# Let us try engineering a feature from text.

favorite_fruits = ['apple', 'banana', 'apple', 'orange', 'banana', 'apple']
# Let us count how many times 'apple' appears.
apple_count = favorite_fruits.count('apple')

print('Number of times apple appears:', apple_count)
Number of times apple appears: 3
# Handling missing or empty data is important.
# Let us see how to fill in missing ages.

ages_with_missing = [17, 21, None, 42, None, 29]

filled_ages = []
for age in ages_with_missing:
    if age is None:
        filled_ages.append(18)  # We use eighteen as a default age.
    else:
        filled_ages.append(age)

print('Filled ages:', filled_ages)
Filled ages: [17, 21, 18, 42, 18, 29]
# Categorical features are often easier to use as numbers.
# Let us turn fruits into numbers.

fruit_to_number = {'apple': 0, 'banana': 1, 'orange': 2}
fruit_numbers = []
for fruit in favorite_fruits:
    fruit_numbers.append(fruit_to_number[fruit])

print('Fruit as numbers:', fruit_numbers)
Fruit as numbers: [0, 1, 0, 2, 1, 0]
# Creating features from math or logic is common.
# Let us find the difference from the oldest age.

oldest = max(filled_ages)
age_gaps = [oldest - age for age in filled_ages]

print('Difference from oldest:', age_gaps)
Difference from oldest: [25, 21, 24, 0, 24, 13]

Making Features More Useful#

Sometimes, features work better when scaled down or changed.

For example, big numbers can be harder for computer models to use.

Let us normalize our ages so they fit between zero and one.

# Scaling numbers helps machine learning.
min_age = min(filled_ages)
max_age = max(filled_ages)

normalized_ages = []
for age in filled_ages:
    normalized = (age - min_age) / (max_age - min_age)
    normalized_ages.append(normalized)

print('Normalized ages:', normalized_ages)
Normalized ages: [0.0, 0.16, 0.04, 1.0, 0.04, 0.48]
# Sometimes we want features to show a relationship.
siblings = [1, 2, 0, 3, 1, 2]

# Let us make a new feature: age per sibling.
age_per_sibling = []
for age, sib in zip(filled_ages, siblings):
    if sib == 0:
        age_per_sibling.append(age)
    else:
        age_per_sibling.append(age // sib)

print('Age per sibling:', age_per_sibling)
Age per sibling: [17, 10, 18, 14, 18, 14]
# Feature engineering often means removing outliers.
all_ages = [17, 21, 201, 42, 18, 29]  # 201 is probably a data error.
clean_ages = []

for age in all_ages:
    if age < 120:
        clean_ages.append(age)

print('Ages with outliers removed:', clean_ages)
Ages with outliers removed: [17, 21, 42, 18, 29]
# Sometimes features are combined to make even better clues.
# Let us combine age and siblings into one string.
combined_feat = []
for age, sib in zip(filled_ages, siblings):
    result = str(age) + '_sib_' + str(sib)
    combined_feat.append(result)

print('Combined features:', combined_feat)
Combined features: ['17_sib_1', '21_sib_2', '18_sib_0', '42_sib_3', '18_sib_1', '29_sib_2']
# Feature selection means keeping only helpful features.
features = {'age': filled_ages, 'adult': is_adult, 'siblings': siblings}

selected_features = {}
for key in features:
    if key != 'siblings':
        selected_features[key] = features[key]

print('Selected features:', selected_features)
Selected features: {'age': [17, 21, 18, 42, 18, 29], 'adult': [False, True, True, True, False, True]}
# Let us ask the user for a new feature value.
city = input('Which city do you live in? ')

print('City entered:', city)
City entered: Boston
 
# Practice: Can you fill in your favorite animal and make a new feature?
animal = input('What is your favorite animal? ')
fav_animals = [animal] * len(names)

print('Everyone loves:', fav_animals)
Everyone loves: ['penguin', 'penguin', 'penguin', 'penguin', 'penguin', 'penguin']
 
# Mini Project Part 1: Build a tiny dataset.
student_data = [
    {'name': 'Anna', 'age': 17, 'score': 88, 'city': 'Boston'},
    {'name': 'Ben', 'age': 21, 'score': 92, 'city': 'Dallas'},
    {'name': 'Cleo', 'age': 34, 'score': 76, 'city': 'Boston'},
    {'name': 'Dan', 'age': 42, 'score': 81, 'city': 'Austin'},
]

for student in student_data:
    print(student)
    
{'name': 'Anna', 'age': 17, 'score': 88, 'city': 'Boston'}
{'name': 'Ben', 'age': 21, 'score': 92, 'city': 'Dallas'}
{'name': 'Cleo', 'age': 34, 'score': 76, 'city': 'Boston'}
{'name': 'Dan', 'age': 42, 'score': 81, 'city': 'Austin'}
# Mini Project Part 2: Engineer new features for our students.

for student in student_data:
    student['is_adult'] = student['age'] >= 18
    student['passed'] = student['score'] >= 80

# Print with new features
for student in student_data:
    print(student)
    
{'name': 'Anna', 'age': 17, 'score': 88, 'city': 'Boston', 'is_adult': False, 'passed': True}
{'name': 'Ben', 'age': 21, 'score': 92, 'city': 'Dallas', 'is_adult': True, 'passed': True}
{'name': 'Cleo', 'age': 34, 'score': 76, 'city': 'Boston', 'is_adult': True, 'passed': False}
{'name': 'Dan', 'age': 42, 'score': 81, 'city': 'Austin', 'is_adult': True, 'passed': True}
# Let us filter our student data for adults who passed.
adults_passed = []
for student in student_data:
    if student['is_adult'] and student['passed']:
        adults_passed.append(student['name'])

print('Adults who passed:', adults_passed)
Adults who passed: ['Ben', 'Dan']
# Sorting is simple but powerful.
sorted_students = sorted(student_data, key=lambda s: s['score'], reverse=True)

for student in sorted_students:
    print(student['name'], student['score'])
    
Ben 92
Anna 88
Dan 81
Cleo 76
# Final tip: See what features are available.
print('All features for first student:', list(student_data[0].keys()))
All features for first student: ['name', 'age', 'score', 'city', 'is_adult', 'passed']

Recap: What did we learn?#

We saw how new features help computers understand data.

We made and changed features using math, logic, and even input data.

Practice these skills to help any dataset shine!

You made it! Want to keep learning?#

Like this lesson? Please hit Like, leave a Comment, and Subscribe for more easy Python videos!

Share this with friends who want to level up their Python skills!

Thank you, and see you in the next video!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.