Mathew K Analytics

Lesson 13 · Mastering Pandas

How to Rename Columns and Indexes in pandas DataFrames for Clear Data Analysis

Learning to rename columns and indexes helps make your data easier to understand and work with. We will use the Titanic Dataset for this lesson. Let us…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Renaming Columns and Indexes in Pandas#

Learning to rename columns and indexes helps make your data easier to understand and work with.

We will use the Titanic Dataset for this lesson.

Let us explore how and why to rename columns and indexes.

import warnings; warnings.filterwarnings('ignore')
# Data setup (Titanic Dataset)
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  

Why Rename Columns or Indexes?#

  • Column names may be long, unclear, or have typos
  • You may want to match column names across datasets
  • Clean names make analysis easier
print(df.columns)
Index(['PassengerId', 'Survived', 'Pclass', 'Name', 'Sex', 'Age', 'SibSp',
       'Parch', 'Ticket', 'Fare', 'Cabin', 'Embarked'],
      dtype='object')

Basic Renaming with rename()#

To rename columns or index values, use the .rename() method.

You provide a mapping: the current name, and the new name you want.

# Rename 'Sex' column to 'Gender' and 'Pclass' to 'PassengerClass'
df_renamed = df.rename(columns={ 'Sex': 'Gender', 'Pclass': 'PassengerClass' })
print(df_renamed.columns[:7])
Index(['PassengerId', 'Survived', 'PassengerClass', 'Name', 'Gender', 'Age',
       'SibSp'],
      dtype='object')
# The original DataFrame is unchanged unless inplace=True
print('Sex' in df.columns)
print('Gender' in df.columns)
True
False
# Rename in place to update the DataFrame itself
df.rename(columns={ 'Sex': 'Gender', 'Pclass': 'PassengerClass' }, inplace=True)
print(df.columns[:7])
Index(['PassengerId', 'Survived', 'PassengerClass', 'Name', 'Gender', 'Age',
       'SibSp'],
      dtype='object')
# Rename multiple columns: abbreviate Age and update Embarked
df.rename(columns={'Age': 'A', 'Embarked': 'PortOfEmbarkation'}, inplace=True)
print(df.columns[:7])
Index(['PassengerId', 'Survived', 'PassengerClass', 'Name', 'Gender', 'A',
       'SibSp'],
      dtype='object')

Renaming Columns with str Methods#

When columns have patterns such as extra spaces, uppercase letters, or special characters, use string methods.

For example, you can make all columns lowercase or strip whitespace.

Let us see this in action.

# Make all columns lowercase and remove spaces
df.columns = df.columns.str.lower().str.replace(' ', '', regex=False)
print(df.columns[:7])
Index(['passengerid', 'survived', 'passengerclass', 'name', 'gender', 'a',
       'sibsp'],
      dtype='object')
# Replace underscores with spaces and capitalize words
df.columns = df.columns.str.replace('_', ' ').str.title()
print(df.columns[:7])
Index(['Passengerid', 'Survived', 'Passengerclass', 'Name', 'Gender', 'A',
       'Sibsp'],
      dtype='object')

Renaming the DataFrame Index#

To rename DataFrame row indexes, use the rename() method with the index argument.

This is helpful when your index is not just numbers, for example after grouping.

# Example: set 'PassengerId' as index, then rename some index values
df_indexed = df.set_index('Passengerid')
df_indexed = df_indexed.rename(index={1: 'FirstPassenger', 2: 'SecondPassenger'})
print(df_indexed.head(3))
                 Survived  Passengerclass  \
Passengerid                                 
FirstPassenger          0               3   
SecondPassenger         1               1   
3                       1               3   

                                                              Name  Gender  \
Passengerid                                                                  
FirstPassenger                             Braund, Mr. Owen Harris    male   
SecondPassenger  Cumings, Mrs. John Bradley (Florence Briggs Th...  female   
3                                           Heikkinen, Miss. Laina  female   

                    A  Sibsp  Parch            Ticket     Fare Cabin  \
Passengerid                                                            
FirstPassenger   22.0      1      0         A/5 21171   7.2500   NaN   
SecondPassenger  38.0      1      0          PC 17599  71.2833   C85   
3                26.0      0      0  STON/O2. 3101282   7.9250   NaN   

                Portofembarkation  
Passengerid                        
FirstPassenger                  S  
SecondPassenger                 C  
3                               S  
# Reset index to go back to default numbering
df_reset = df_indexed.reset_index()
print(df_reset.head(2))
       Passengerid  Survived  Passengerclass  \
0   FirstPassenger         0               3   
1  SecondPassenger         1               1   

                                                Name  Gender     A  Sibsp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   

   Parch     Ticket     Fare Cabin Portofembarkation  
0      0  A/5 21171   7.2500   NaN                 S  
1      0   PC 17599  71.2833   C85                 C  

Common Pitfall: inplace vs Not inplace#

  • Using inplace=True directly edits your DataFrame and you cannot undo
  • Without inplace=True, your original stays unchanged unless you save the result

Use caution: decide if you want to keep the changes or not.

# Quick practice: Rename 'Fare' to 'TicketPrice' using input()
current = input('Type the current column name you want to rename: ')
new = input('Type the new name: ')
df2 = df.rename(columns={current: new})
print(df2.columns)
Index(['Passengerid', 'Survived', 'Passengerclass', 'Name', 'Gender', 'A',
       'Sibsp', 'Parch', 'Ticket', 'TicketPrice', 'Cabin',
       'Portofembarkation'],
      dtype='object')
# SCENARIO: Standardize all column names to lower case and no spaces
cols_standard = df.columns.str.lower().str.replace(' ', '', regex=False)
print(cols_standard)
Index(['passengerid', 'survived', 'passengerclass', 'name', 'gender', 'a',
       'sibsp', 'parch', 'ticket', 'fare', 'cabin', 'portofembarkation'],
      dtype='object')
# Undo column name changes by restoring from original data
df_restore = pd.read_csv(url)
print(df_restore.columns)
Index(['PassengerId', 'Survived', 'Pclass', 'Name', 'Sex', 'Age', 'SibSp',
       'Parch', 'Ticket', 'Fare', 'Cabin', 'Embarked'],
      dtype='object')

Use Case: Renaming Before Merging#

When combining datasets, names must match exactly.

Renaming ahead of merging helps ensure the join works right.

# Mini-challenge: Rename all columns to be snake_case
df_snake = df.copy()
df_snake.columns = df_snake.columns.str.lower().str.replace(' ', '_').str.replace('__', '_')
print(df_snake.columns[:7])
Index(['passengerid', 'survived', 'passengerclass', 'name', 'gender', 'a',
       'sibsp'],
      dtype='object')
# Quick tip: Rename all columns using a list
columns_new = [f'feature_{i}' for i in range(len(df.columns))]
df.columns = columns_new
print(df.head(2))
   feature_0  feature_1  feature_2  \
0          1          0          3   
1          2          1          1   

                                           feature_3 feature_4  feature_5  \
0                            Braund, Mr. Owen Harris      male       22.0   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...    female       38.0   

   feature_6  feature_7  feature_8  feature_9 feature_10 feature_11  
0          1          0  A/5 21171     7.2500        NaN          S  
1          1          0   PC 17599    71.2833        C85          C  

Practice and Apply#

  1. Make all column names lower case.
  2. Remove all spaces, replace with underscores.
  3. Try renaming 'Survived' to 'Outcome'.

These steps will help you solidify your new skills!

Recap: Renaming Columns and Indexes#

  • Use .rename() for specific changes, use string methods for patterns
  • Always check your results after renaming
  • Clean names save time and reduce mistakes
  • Practice and you will master tidy DataFrames!

Thank you for learning with us!#

For more Python and pandas lessons, subscribe and click the bell.

See you in the next video!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.