Mathew K Analytics

Lesson 45 · Python Fundamentals

2 - Removing Duplicates in Python: Data Cleaning Techniques

Learn why duplicate values can cause problems and how Python helps us solve them. Let us make your code cleaner, data easier to use, and your programs…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb
 

Removing Duplicates in Python#

Learn why duplicate values can cause problems and how Python helps us solve them.

Let us make your code cleaner, data easier to use, and your programs smarter!

Why Removing Duplicates Matters#

  • Duplicates can make your data confusing.
  • Searches take longer.
  • Results might not be accurate.

It is important to keep data neat!

# Here is a list with duplicate names
names = ['Alice', 'Bob', 'Alice', 'Eve', 'Bob', 'Dave']

print('Original list:')
print(names)
Original list:
['Alice', 'Bob', 'Alice', 'Eve', 'Bob', 'Dave']
# Remove duplicates using a set
unique_names = set(names)

print('Names with duplicates removed:')
print(unique_names)
Names with duplicates removed:
{'Bob', 'Alice', 'Eve', 'Dave'}
# Convert the set back to a list to keep working with it as a list
unique_names_list = list(unique_names)

print('Unique names as a list:')
print(unique_names_list)
Unique names as a list:
['Bob', 'Alice', 'Eve', 'Dave']
# Remove duplicates and keep the original order
numbers = [4, 2, 4, 3, 2, 1, 5, 3, 1]
unique_numbers = []

for number in numbers:
    if number not in unique_numbers:
        unique_numbers.append(number)

print('Numbers with duplicates removed, order kept:')
print(unique_numbers)
Numbers with duplicates removed, order kept:
[4, 2, 3, 1, 5]
# Use dictionary keys to remove duplicates and keep order (Python 3.7+)
colors = ['red', 'blue', 'red', 'green', 'blue', 'yellow']
unique_colors = list(dict.fromkeys(colors))

print('Unique colors, order kept:')
print(unique_colors)
Unique colors, order kept:
['red', 'blue', 'green', 'yellow']
# Get input from the user and remove duplicates
user_input = input('Enter a few items separated by commas: ')
items = [item.strip() for item in user_input.split(',')]
unique_items = list(dict.fromkeys(items))

print('Here are your unique items:')
print(unique_items)
Here are your unique items:
['apple', 'banana', 'orange']
 
# Challenge: Remove duplicates from a list with different types of data
mixed_list = [1, '1', 2, '2', 2, 1, 'one', 'One', 'ONE']
unique_mixed = list(dict.fromkeys(mixed_list))
print('Unique values from a mixed list:')
print(unique_mixed)
Unique values from a mixed list:
[1, '1', 2, '2', 'one', 'One', 'ONE']
# Remove duplicates using a list comprehension
animals = ['cat', 'dog', 'cat', 'bird', 'dog', 'cat', 'fish']
seen = set()
unique_animals = [animal for animal in animals if not (animal in seen or seen.add(animal))]

print('Animals with no duplicates:')
print(unique_animals)
Animals with no duplicates:
['cat', 'dog', 'bird', 'fish']
# Remove duplicates from user input numbers, sorted
number_input = input('Enter numbers separated by spaces: ')
numbers_list = [int(num) for num in number_input.split()]
unique_sorted = sorted(set(numbers_list))

print('Sorted set of unique numbers:')
print(unique_sorted)
Sorted set of unique numbers:
[1, 2, 4, 5, 9]
 

Handling Common Duplication Mistakes#

  • Sets do not keep the order.
  • String '1' and number 1 are different.
  • Upper case and lower case can matter.

Think about what counts as a duplicate for your problem.

# Remove duplicates ignoring upper/lowercase letters
cars = ['Toyota', 'toyota', 'Honda', 'HONDA', 'Ford', 'ford']
unique_cars = []
seen = set()
for car in cars:
    lower = car.lower()
    if lower not in seen:
        seen.add(lower)
        unique_cars.append(car)

print('Unique cars, case-insensitive:')
print(unique_cars)
Unique cars, case-insensitive:
['Toyota', 'Honda', 'Ford']
# Mini-project part 1: Clean a contact list
contacts = [
    {'name': 'Amy', 'phone': '123'},
    {'name': 'Bob', 'phone': '234'},
    {'name': 'Amy', 'phone': '123'},
    {'name': 'Eve', 'phone': '345'},
    {'name': 'bob', 'phone': '234'},
]

unique_contacts = []
seen_contacts = set()

for contact in contacts:
    key = (contact['name'].lower(), contact['phone'])
    if key not in seen_contacts:
        unique_contacts.append(contact)
        seen_contacts.add(key)

print('Unique contacts:')
print(unique_contacts)
Unique contacts:
[{'name': 'Amy', 'phone': '123'}, {'name': 'Bob', 'phone': '234'}, {'name': 'Eve', 'phone': '345'}]
# Mini-project part 2: Count how many items were removed
before = len(contacts)
after = len(unique_contacts)
removed = before - after
print(f'Removed {removed} duplicates from the contact list.')
Removed 2 duplicates from the contact list.
# Troubleshooting: Why did this not remove all duplicates?
bad_list = ['apple', 'Apple', 'APPLE', 'apple']
unique_bad = list(set(bad_list))
print('Result:', unique_bad)
Result: ['Apple', 'apple', 'APPLE']
# Extra tip: Combine two lists and remove duplicates
list1 = ['A', 'B', 'C', 'C']
list2 = ['C', 'D', 'E', 'A']
combined = list(set(list1 + list2))

print('Combined unique elements:')
print(combined)
Combined unique elements:
['A', 'E', 'C', 'D', 'B']
# Challenge: Remove duplicates from a text file (lines)
filename = 'sample.txt'

# First, make a small sample file to work with
with open(filename, 'w') as f:
    f.write('cat\n')
    f.write('dog\n')
    f.write('cat\n')
    f.write('fish\n')
    f.write('dog\n')

# Now load lines and remove duplicates
with open(filename, 'r') as f:
    lines = f.readlines()
unique_lines = list(dict.fromkeys(line.strip() for line in lines))

print('Unique lines from file:')
print(unique_lines)
Unique lines from file:
['cat', 'dog', 'fish']
# Recap: Different ways to remove duplicates
print('1. Use set() to remove all duplicates (order does not matter)')
print('2. Use dict.fromkeys() to keep order (Python 3.7+)')
print('3. Use a for-loop to filter duplicates the way you want')
1. Use set() to remove all duplicates (order does not matter)
2. Use dict.fromkeys() to keep order (Python 3.7+)
3. Use a for-loop to filter duplicates the way you want

Challenge Yourself!#

  • Try removing duplicates from a list of cities, ignoring case.
  • Remove repeated numbers, but keep only the highest value of each.
  • Create a function to do this, and share your version!

Thanks for working through removing duplicates. You have taken a big step towards cleaner, better code.

Thanks for watching!#

Like, comment, and subscribe if you learned something new.

Keep practicing, and share with a friend who needs clean data.

See you in the next Python lesson!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.