Lesson 40 · Python for Data Science
8 - Movie Dataset Cleanup in Python: Data Preparation Techniques
Ever wondered how streaming apps organize thousands of movies? They use data cleanup! Today, we will learn to tidy up messy movie info using Python. No…
- CoursePython for Data Science
- Lesson40 of 38
- Video11 min
- FormatJupyter notebook · 24 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Welcome to Movie Dataset Cleanup in Python!#
Ever wondered how streaming apps organize thousands of movies? They use data cleanup! Today, we will learn to tidy up messy movie info using Python. No experience needed: just curiosity and your keyboard.
We will start with the basics and soon you will clean your own tiny movie dataset.
Grab your snack and let's begin!
# First, let us create a simple list with some movie titles.
movies = [
' the matrix ',
'inception',
'Titanic ',
'The Godfather',
' avatar',
'spider-man',
' star wars '
]
print('Here are our raw movie titles:')
print(movies)
Why clean up the data?#
Datasets are often messy: extra spaces, different spelling, or strange symbols. Cleaning makes it easier to search, compare, and show movies.
Let us see how we can tidy up our movie titles.
# Let us fix extra spaces using the strip() function.
cleaned_movies = []
for title in movies:
cleaned_title = title.strip()
cleaned_movies.append(cleaned_title)
print(cleaned_movies)
# Let us make all the titles have the same style with title().
uniform_titles = []
for name in cleaned_movies:
uniform_titles.append(name.title())
print(uniform_titles)
# Imagine you got a typo in your list. Let us fix it.
uniform_titles[2] = 'Titanic'
print(uniform_titles)
# What if there is a duplicate?
movies_with_duplicates = uniform_titles + ['Inception']
print(movies_with_duplicates)
# Let us remove duplicates by making a set.
unique_movies = list(set(movies_with_duplicates))
print(unique_movies)
# Now let us ask the user to add a movie title.
new_movie = input('Type a movie to add: ')
new_movie_clean = new_movie.strip().title()
if new_movie_clean not in unique_movies:
unique_movies.append(new_movie_clean)
print('Added:', new_movie_clean)
else:
print('That movie is already in the list!')
print(unique_movies)
# Maybe we decide we do not want 'Avatar' on our list.
if 'Avatar' in unique_movies:
unique_movies.remove('Avatar')
print('Removed Avatar!')
else:
print('Avatar was not in the list.')
print(unique_movies)
# What if a movie is missing? Let us try to find it.
search_title = input('Type a movie to search for: ').strip().title()
if search_title in unique_movies:
print(search_title, 'is in our cleaned movie list!')
else:
print(search_title, 'is not found.')
# What about sorting our movies?
sorted_movies = sorted(unique_movies)
print('Sorted movie titles:')
print(sorted_movies)
# Sometimes we want just the movies that start with 'T'.
t_movies = [movie for movie in sorted_movies if movie.startswith('T')]
print('Movies starting with T:', t_movies)
# You can count how many movies you have.
print('Total number of cleaned movies:', len(unique_movies))
# Let us add extra info: a release year for each movie!
movie_years = {
'The Matrix': 1999,
'Inception': 2010,
'Titanic': 1997,
'The Godfather': 1972,
'Spider-Man': 2002,
'Star Wars': 1977,
'Jurassic Park': 1993
}
print(movie_years)
# Let us show all movies released after 2000.
for m, y in movie_years.items():
if y > 2000:
print(m, '-', y)
# Let us merge two small movie lists.
old_movies = ['E.T.', 'Back to the Future']
all_movies = sorted_movies + old_movies
print('Combined list:', all_movies)
# Let us try a tiny real-world style dataset with errors.
messy_data = [
{'title': 'Inception ', 'year': '2010'},
{'title': ' the godfather', 'year': '1972'},
{'title': 'Spider-man', 'year': None},
{'title': 'titanic', 'year': ''},
{'title': 'Avatar', 'year': '2009'},
{'title': '', 'year': '1982'}
]
for entry in messy_data:
print(entry)
# Let us clean up the messy movie dataset.
fixed_movies = []
for entry in messy_data:
title = entry['title'].strip().title()
year_str = entry['year']
try:
year = int(year_str)
except:
year = None
if title != '' and year:
fixed_movies.append({'title': title, 'year': year})
print('Cleaned dataset:')
for m in fixed_movies:
print(m)
# Mini-project: Find all movies between two years.
start = int(input('Start year: '))
end = int(input('End year: '))
matching = [m for m in fixed_movies if start <= m['year'] <= end]
print('Movies in that range:')
for m in matching:
print(m['title'], '-', m['year'])
# Extra: Find oldest and newest movie in the clean dataset.
if fixed_movies:
oldest = min(fixed_movies, key=lambda m: m['year'])
newest = max(fixed_movies, key=lambda m: m['year'])
print('Oldest:', oldest['title'], '-', oldest['year'])
print('Newest:', newest['title'], '-', newest['year'])
# Sometimes, users enter numbers or funny symbols instead of letters.
test_titles = ['Shrek!', 'Up#', ' the 40-Year-Old Virgin ', 'Toy-Story']
for t in test_titles:
clean_t = ''.join([c for c in t if c.isalnum() or c.isspace()]).strip().title()
print('Clean title:', clean_t)
# Tip: If you want to undo a change, use list slices.
original = unique_movies.copy()
removed = unique_movies.pop()
print('Latest list:', unique_movies)
unique_movies = original # Revert
print('Reverted:', unique_movies)
Recap: Youre a Movie Data Cleaner!#
We started from a rough list of titles, cleaned spacing and style, learned to update and remove entries, filtered and sorted, used dictionaries, merged lists, handled real-world errors, and finished with a mini project.
With a few lines of Python, you can prepare movie data for any app, website, or project.
Curious to go deeper? Experiment with more messy data!
Thanks for learning Like, Subscribe, and Share!#
If you enjoyed cleaning movie data in Python, please like this video, leave a comment with your favorite movie, and subscribe for more Python lessons.
Share with friends who want to learn coding in a fun way.
See you in the next lesson!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



