Mathew K Analytics

Lesson 20 · Data Science Projects

Analyzing Uber Trips Data: A Step-by-Step Guide to Data Science Techniques

Learn the basics of data mining with Python and Jupyter. Explore Uber trip data using real world examples. Perform simple statistics and visualizations with…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Uber NYC Trip Data Analysis for Beginners#

  • Learn the basics of data mining with Python and Jupyter.

  • Explore Uber trip data using real world examples.

  • Perform simple statistics and visualizations with pandas and seaborn.

  • No prior experience is needed for this lesson.

  • By the end, you will understand how to explore and plot real data sets.

What is Data Mining?#

  • Data mining means finding useful information in piles of data.

  • We spot patterns, detect trends, and answer questions with facts.

  • It is a big part of data science, business, and more.

  • Python makes this process easier, faster, and more fun.

  • We use data mining any time we explore, clean, or visualize data.

Our Mini Project: Uber Trips in New York#

  • We will use real Uber trip data collected from New York City.
  • This is a small sample, so it runs quickly even on slow laptops.
  • You will see how to find busy hours, popular spots, and more.
  • These skills are useful in almost any job or school project.
# Always suppress warnings for a cleaner notebook
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
# Data setup: Load Uber trip data from New York City (April 2014)
import pandas as pd
url = 'https://raw.githubusercontent.com/fivethirtyeight/uber-tlc-foil-response/master/uber-trip-data/uber-raw-data-apr14.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(564516, 4)
          Date/Time      Lat      Lon    Base
0  4/1/2014 0:11:00  40.7690 -73.9549  B02512
1  4/1/2014 0:17:00  40.7267 -74.0345  B02512
2  4/1/2014 0:21:00  40.7316 -73.9873  B02512
# What do our Uber data columns mean?
print(df.columns.tolist())
['Date/Time', 'Lat', 'Lon', 'Base']
# Clean up column names if needed
df.columns = [col.strip() for col in df.columns]
print(df.columns.tolist())
['Date/Time', 'Lat', 'Lon', 'Base']

Exploring the Data: What Do These Trips Show?#

  • The data shows Uber pickups in New York for each trip.
  • Each row gives us the date, the precise location (latitude and longitude), and a base code.
  • We can learn about city patterns like when people ride and where.
  • Before we visualize, let us make sure everything is clean and easy to use.
# Check for missing values in the table
print(df.isnull().sum())
Date/Time    0
Lat          0
Lon          0
Base         0
dtype: int64
# Change date column to a datetime type
df['Date/Time'] = pd.to_datetime(df['Date/Time'])
print(df['Date/Time'].dtype)
print(df['Date/Time'].head(2))
datetime64[ns]
0   2014-04-01 00:11:00
1   2014-04-01 00:17:00
Name: Date/Time, dtype: datetime64[ns]
# Add new columns for hour, day of week, and day of month
df['hour'] = df['Date/Time'].dt.hour
df['weekday'] = df['Date/Time'].dt.dayofweek
df['day'] = df['Date/Time'].dt.day
print(df[['hour', 'weekday', 'day']].head(3))
   hour  weekday  day
0     0        1    1
1     0        1    1
2     0        1    1

Time to Visualize: Let Us See the Trends#

  • Charts make patterns clear and easy to spot.
  • We will use seaborn and matplotlibtwo popular Python libraries.
  • Let us find out which hours are busiest for Uber pickups.
# Import plotting libraries
import matplotlib.pyplot as plt
import seaborn as sns
sns.set(style="whitegrid")
# Plot the number of pickups by hour
plt.figure(figsize=(10, 5))
sns.countplot(x='hour', data=df, color='skyblue')
plt.title('Number of Uber Pickups by Hour')
plt.xlabel('Hour of Day (0 to 23)')
plt.ylabel('Number of Trips')
plt.show()
No description has been provided for this image
# Compare pickups on different weekdays
plt.figure(figsize=(10, 5))
sns.countplot(x='weekday', data=df, color='lightgreen')
plt.title('Uber Pickups by Weekday (0=Monday, 6=Sunday)')
plt.xlabel('Day of Week')
plt.ylabel('Number of Trips')
plt.show()
No description has been provided for this image
# Make a simple time series plot: Total Uber trips per day
trips_by_day = df.groupby('day').size()
plt.figure(figsize=(12, 5))
trips_by_day.plot(kind='line', marker='o', color='coral')
plt.title('Number of Uber Trips per Day (April 2014)')
plt.xlabel('Day of Month')
plt.ylabel('Number of Trips')
plt.xticks(range(1, 31, 2))
plt.grid(True)
plt.show()
No description has been provided for this image
# Where do pickups happen the most? Plot pickup locations on a map
plt.figure(figsize=(8,8))
plt.scatter(df['Lon'], df['Lat'], s=1, alpha=0.03)
plt.title('Uber Pickup Spots in New York City (April 2014)')
plt.xlabel('Longitude')
plt.ylabel('Latitude')
plt.xlim(-74.1, -73.7)
plt.ylim(40.6, 40.9)
plt.show()
No description has been provided for this image
# Quick stats: mean, minimum, and maximum hour of pickups
print("Earliest pickup hour:", df['hour'].min())
print("Latest pickup hour:", df['hour'].max())
print("Mean pickup hour:", df['hour'].mean())
Earliest pickup hour: 0
Latest pickup hour: 23
Mean pickup hour: 14.46504262058117

Mini Project: What Other Questions Can We Ask?#

  • Try plotting the number of Uber trips per "Base" code.

  • Check if the location spread changes by hour or weekday.

  • For a challenge, plot pickups for just one very busy day.

  • Pause the video and experiment! Share your findings in the comments.

# Let us try a bar plot: Trips per Uber base location
plt.figure(figsize=(8,5))
sns.countplot(x='Base', data=df, palette='Set2')
plt.title('Number of Pickups by Uber Base Code')
plt.xlabel('Base Code')
plt.ylabel('Number of Trips')
plt.show()
No description has been provided for this image
# Try filtering: Show only trips from a specific base (input)
base = input("Enter an Uber base code (e.g., B02512): ")
filtered = df[df['Base'] == base]
print(filtered.shape)
print(filtered.head())
(35536, 7)
            Date/Time      Lat      Lon    Base  hour  weekday  day
0 2014-04-01 00:11:00  40.7690 -73.9549  B02512     0        1    1
1 2014-04-01 00:17:00  40.7267 -74.0345  B02512     0        1    1
2 2014-04-01 00:21:00  40.7316 -73.9873  B02512     0        1    1
3 2014-04-01 00:28:00  40.7588 -73.9776  B02512     0        1    1
4 2014-04-01 00:33:00  40.7594 -73.9722  B02512     0        1    1
# Save your filtered results to a new file
filtered.to_csv('uber_base_filtered.csv', index=False)
print("Filtered data saved to 'uber_base_filtered.csv'")
Filtered data saved to 'uber_base_filtered.csv'

Key Takeaways and Next Steps#

  • You loaded, cleaned, and explored a real Uber trip dataset.

  • You learned to summarize and visualize time and place with charts.

  • Try running the same steps on your local travel or commute data!

  • Did you find any surprises in the patterns? Share your results below.

  • Like and subscribe if you want more beginner data science lessons.

  • Practice is the fastest way to master these skills.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.