Lesson 34 · Mastering Pandas
Mastering Dates and Times in Pandas for Effective Data Analysis
In modern data analysis, handling dates and times is essential. Pandas offers powerful DateTime tools for filtering, resampling, plotting, and more. Let us…
- CourseMastering Pandas
- Lesson34 of 44
- Video12 min
- FormatJupyter notebook · 15 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbPandas DatetimeIndex and Timestamps: An Intermediate Beginner's Guide#
In modern data analysis, handling dates and times is essential. Pandas offers powerful DateTime tools for filtering, resampling, plotting, and more. Let us explore how to master time data in practical, friendly steps.
Real-World Time Series Example: U.S. Airline Passengers#
We will use a built-in flights dataset. First, let us see the raw data.
import warnings; warnings.filterwarnings("ignore")
# Data setup (Airline Passenger Flights Dataset)
import seaborn as sns
import numpy as np
np.random.seed(42)
df = sns.load_dataset('flights')
print(df.shape)
print(df.head(3))
Looking at the Data Structure#
Current columns are year, month, and passengers.
True time series need a real date column. Let us create one.
# Combine year and month columns into a single date string
df['date_str'] = df['year'].astype(str) + '-' + df['month'] + '-01'
print(df[['year','month','date_str']].head())
# Convert date string to pandas datetime (Timestamp) object
df['date'] = pd.to_datetime(df['date_str'], format='%Y-%b-%d')
print(df[['date_str', 'date']].head())
# Set the datetime as the DataFrame index (DatetimeIndex)
df_ts = df.set_index('date')
print(df_ts.head(3))
Why Use a DatetimeIndex?#
- Fast filtering and slicing by time
- Makes plotting and resampling easier
- Friendly with rolling averages and time math
# Plot the time series: passengers over time
import matplotlib.pyplot as plt
df_ts['passengers'].plot(figsize=(10,4), title='Monthly Airline Passengers')
plt.ylabel('Passengers')
plt.xlabel('Date')
plt.tight_layout()
plt.show()
# Filter for flights in the 1950s only (using the DatetimeIndex)
df_1950s = df_ts.loc['1950-01-01':'1959-12-31']
print(df_1950s.head(3))
print(f"Total records for 1950s: {len(df_1950s)}")
# Quick monthly summary: mean number of passengers by year
yearly_avg = df_ts.resample('Y')['passengers'].mean()
print(yearly_avg.head(3))
# Rolling average: passengers per month, window=6
df_ts['passengers_rolling6'] = df_ts['passengers'].rolling(window=6).mean()
df_ts[['passengers', 'passengers_rolling6']].head(10)
# Plot comparison: original vs. rolling average
df_ts[['passengers', 'passengers_rolling6']].plot(figsize=(10,4))
plt.title('Monthly Passengers and 6-Month Rolling Average')
plt.ylabel('Number of Passengers')
plt.xlabel('Date')
plt.tight_layout()
plt.show()
# Fast slice: Show passengers trend for summer months only
df_summer = df_ts[df_ts.index.month.isin([6,7,8])]
print(df_summer.head(6))
# Example: Find the month with the most passengers ever
row_max = df_ts['passengers'].idxmax()
print(f'Maximum passengers was in: {row_max}')
# Convert index back to a column (reset index)
flights_with_date = df_ts.reset_index()
print(flights_with_date[['date','passengers']].head())
Practice Prompt: Real Data, Real Dates#
Try changing the rolling window to 12 months. Which year had the slowest average growth? You can explore with just a few lines of code.
More DateTime Tricks#
You can use .dt accessor on date columns to get parts like year, month, or weekday.
Let us see examples:
# Use .dt accessor for date parts
years = flights_with_date['date'].dt.year.unique()
print('Years in the dataset:', years)
first_weekdays = flights_with_date['date'].dt.day_name().unique()
print('Days of week present:', first_weekdays)
Common Pitfalls in Datetime Handling#
- Watch out for string vs. datetime types
- Always check timezone handling with real-world data
- Missing or duplicate dates may cause subtle errors
If you get errors, try using pd.to_datetime again, or inspect the column type with type(df['column'].iloc[0]).
# Quick troubleshooting: check data types
print(flights_with_date.dtypes)
Extensions: Working with Date Ranges and Timestamps#
pd.date_rangelets you quickly make a time index- Pandas Timestamps behave a lot like Python datetimes, but add more power
Let us see a quick example:
# Generate a series of business days in January 2024
rng = pd.date_range(start='2024-01-01', end='2024-01-31', freq='B')
print(rng[:5])
# Example: work with single pandas Timestamp
stamp = pd.Timestamp('2024-01-05 10:30')
print(f'Year: {stamp.year}, Month: {stamp.month}, Hour: {stamp.hour}')
Recap: What Have We Learned?#
- How to create DatetimeIndex and Timestamps
- How to filter, roll, and plot with time series
- How to troubleshoot, extract parts, and extend with ranges
Pandas makes handling dates approachable. With a little practice, time series unlocks many new insights.
Next Steps: Keep Exploring!#
Try running experiments with your own CSVs. Practice rolling means, month selections, and custom date parsing.
Like this walkthrough? Subscribe for more hands-on pandas lessons. Let us keep learning together!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



