Mathew K Analytics

Lesson 5 · Python For Time Series

Comprehensive Guide to Pandas Series and DataFrames for Data Analysis in Python

Welcome! In this lesson, we will explore Pandas, a popular tool for working with data. We will learn about Series and DataFrames, two key building blocks in…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Introduction to Pandas Series and DataFrames#

Welcome! In this lesson, we will explore Pandas, a popular tool for working with data.

We will learn about Series and DataFrames, two key building blocks in Pandas.

By the end, you will know how to read data, explore it, and solve mini data problems with Pandas.

import warnings; warnings.filterwarnings("ignore")

# Let us import Pandas.
import pandas as pd
# Also import numpy for numerical calculations.
import numpy as np

What is a Pandas Series?#

A Series is like a smart list with labels.

Each value in a Series has a label called the index.

Let us make our first Series next.

# Create a simple Series of fruit counts
fruits = pd.Series([3, 5, 2], index=["apples", "bananas", "cherries"])

print(fruits)
apples      3
bananas     5
cherries    2
dtype: int64
# Accessing Series values by label
print("Bananas count:", fruits["bananas"])
Bananas count: 5

Introduction to DataFrames#

A DataFrame is like a spreadsheet or table.

It has rows and columns, with labels for each.

We can hold more complex data in a DataFrame.

# Make a small DataFrame for a class
data = {
    "name": ["Ana", "Ben", "Cleo"],
    "score": [90, 85, 88],
    "passed": [True, True, True]
}
students = pd.DataFrame(data)
print(students)
   name  score  passed
0   Ana     90    True
1   Ben     85    True
2  Cleo     88    True
# Show the shape and column names
print("Shape:", students.shape)
print("Columns:", students.columns.tolist())
Shape: (3, 3)
Columns: ['name', 'score', 'passed']

Data setup#

Next, let us work with a real-world dataset: Shampoo Sales.

We will load it from a web URL using Pandas.

# Load data from a CSV file
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/shampoo.csv"
sales = pd.read_csv(url)
print("Shape:", sales.shape)
print(sales.head())
Shape: (36, 2)
  Month  Sales
0  1-01  266.0
1  1-02  145.9
2  1-03  183.1
3  1-04  119.3
4  1-05  180.3
# Plot the shampoo sales data
import matplotlib.pyplot as plt

plt.plot(sales["Month"], sales["Sales"])
plt.title("Monthly Shampoo Sales")
plt.xlabel("Month")
plt.ylabel("Sales")
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
No description has been provided for this image
# Accessing a single column
monthly_sales = sales["Sales"]
print(monthly_sales.head())
0    266.0
1    145.9
2    183.1
3    119.3
4    180.3
Name: Sales, dtype: float64
# Describing the numbers
print(monthly_sales.describe())
count     36.000000
mean     312.600000
std      148.937164
min      119.300000
25%      192.450000
50%      280.150000
75%      411.100000
max      682.000000
Name: Sales, dtype: float64

Handling missing values#

Sometimes, data is missing or empty. We call this a missing value.

Let us see how to find and fill missing data.

# Check for missing values
print(sales.isnull().sum())
Month    0
Sales    0
dtype: int64
# Fill missing values, if any, with 0
sales_filled = sales.fillna(0)
print(sales_filled.head())
  Month  Sales
0  1-01  266.0
1  1-02  145.9
2  1-03  183.1
3  1-04  119.3
4  1-05  180.3
# Making a new column from other columns
sales_filled["LogSales"] = sales_filled["Sales"].apply(lambda x: 0 if x == 0 else np.log(x))
print(sales_filled.head())
  Month  Sales  LogSales
0  1-01  266.0  5.583496
1  1-02  145.9  4.982921
2  1-03  183.1  5.210032
3  1-04  119.3  4.781641
4  1-05  180.3  5.194622
# Filtering rows: sales over 500
high_sales = sales_filled[sales_filled["Sales"] > 500]
print(high_sales)
   Month  Sales  LogSales
30  3-07  575.5  6.355239
32  3-09  682.0  6.525030
34  3-11  581.3  6.365267
35  3-12  646.9  6.472192
# Sorting the DataFrame by sales
sorted_sales = sales_filled.sort_values(by="Sales", ascending=False)
print(sorted_sales.head())
   Month  Sales  LogSales
32  3-09  682.0  6.525030
35  3-12  646.9  6.472192
34  3-11  581.3  6.365267
30  3-07  575.5  6.355239
33  3-10  475.3  6.163946
# Grouping sales by the first year
sales_filled["Year"] = sales_filled["Month"].apply(lambda x: str(x).split("-")[0])
yearly_sales = sales_filled.groupby("Year")["Sales"].sum()
print(yearly_sales)
Year
1    2357.5
2    3153.5
3    5742.6
Name: Sales, dtype: float64
# Iterating over a DataFrame, row by row
for i, row in sales_filled.iterrows():
    print("Month:", row["Month"], "Sales:", row["Sales"])
    if i > 2:
        break
    
Month: 1-01 Sales: 266.0
Month: 1-02 Sales: 145.9
Month: 1-03 Sales: 183.1
Month: 1-04 Sales: 119.3
# Mini-project: Find the best month
best_month = sales_filled.loc[sales_filled["Sales"].idxmax(), "Month"]
print("Best month for shampoo sales was:", best_month)
Best month for shampoo sales was: 3-09
# Mini-project: Simple growth calculation
first = sales_filled["Sales"].iloc[0]
last = sales_filled["Sales"].iloc[-1]
growth = (last - first) / first * 100
print("Total sales growth over the period:", round(growth, 2), "%")
Total sales growth over the period: 143.2 %
# Catching errors when accessing columns
try:
    print(sales["profit"])
except KeyError:
    print("There is no profit column in our data.")
    
There is no profit column in our data.
# User input with a DataFrame
column = input("Type a column name to see (Month or Sales): ")
if column in sales.columns:
    print(sales[column].head())
else:
    print("That column is not in our data.")
    
0    266.0
1    145.9
2    183.1
3    119.3
4    180.3
Name: Sales, dtype: float64
 

Recap: What have we learned?#

  • How to use Pandas Series and DataFrames
  • How to load and explore real-world data
  • How to clean, analyze, and visualize the data
  • Simple ways to group, filter, and sort information

Great job learning these tools!

Try it yourself & Subscribe!#

Try changing numbers, labels, or filters in any code cell.

Share your mini-project results in the comments.

Like and Subscribe for more Python data science lessons!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.