Mathew K Analytics

Lesson 14 · Mastering Pandas

How to Change Data Types and Convert Units in Pandas for Accurate Data Analysis

In this lesson, you will learn practical ways to change data types and convert measurement units in pandas DataFrames. This skill is essential for cleaning…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Changing Data Types and Converting Units with Pandas#

In this lesson, you will learn practical ways to change data types and convert measurement units in pandas DataFrames. This skill is essential for cleaning real-world data, fixing mismatches, and preparing for analysis.

Let us explore why and how to do these conversions, using the Titanic dataset as our guide.

import warnings; warnings.filterwarnings('ignore')
# Data setup (Titanic Dataset)
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  

Why do data types matter?#

Each column in a DataFrame can store numbers, text, dates, or categories. If a column has the wrong type, your analysis and calculations might be wrong or throw errors. Let us check the data types of our Titanic dataset.

# View data types of each column
print(df.dtypes)
PassengerId      int64
Survived         int64
Pclass           int64
Name            object
Sex             object
Age            float64
SibSp            int64
Parch            int64
Ticket          object
Fare           float64
Cabin           object
Embarked        object
dtype: object
# Check if 'Age' column has any missing values
print(df['Age'].isnull().sum())
177
# Convert Age to integer (if possible)
df['Age_int'] = df['Age'].fillna(-1).astype(int)
print(df[['Age', 'Age_int']].head(8))
    Age  Age_int
0  22.0       22
1  38.0       38
2  26.0       26
3  35.0       35
4  35.0       35
5   NaN       -1
6  54.0       54
7   2.0        2
# Convert 'Sex' column to a category for space and speed
df['Sex_cat'] = df['Sex'].astype('category')
print(df['Sex_cat'].dtype)
print(df['Sex_cat'].head())
category
0      male
1    female
2    female
3    female
4      male
Name: Sex_cat, dtype: category
Categories (2, object): ['female', 'male']
# Convert 'Pclass' to category, since there are only 3 ticket classes
df['Pclass'] = df['Pclass'].astype('category')
print(df['Pclass'].dtype)
print(df['Pclass'].unique())
category
[3, 1, 2]
Categories (3, int64): [1, 2, 3]
# Convert 'Cabin' column to string type (new in pandas 1.0+)
df['Cabin_str'] = df['Cabin'].astype('string')
print(df['Cabin_str'].dtype)
print(df[['Cabin', 'Cabin_str']].head(7))
string
  Cabin Cabin_str
0   NaN      <NA>
1   C85       C85
2   NaN      <NA>
3  C123      C123
4   NaN      <NA>
5   NaN      <NA>
6   E46       E46
# Convert 'Fare' from British pounds to US dollars (approximate, 1 GBP = 1.25 USD)
df['Fare_usd'] = df['Fare'] * 1.25
print(df[['Fare', 'Fare_usd']].head(7))
      Fare   Fare_usd
0   7.2500   9.062500
1  71.2833  89.104125
2   7.9250   9.906250
3  53.1000  66.375000
4   8.0500  10.062500
5   8.4583  10.572875
6  51.8625  64.828125
# Convert 'Age' from years to months (super easy!)
df['Age_months'] = df['Age'] * 12
print(df[['Age', 'Age_months']].head(10))
    Age  Age_months
0  22.0       264.0
1  38.0       456.0
2  26.0       312.0
3  35.0       420.0
4  35.0       420.0
5   NaN         NaN
6  54.0       648.0
7   2.0        24.0
8  27.0       324.0
9  14.0       168.0
# Convert 'Name' to uppercase (string operation)
df['Name_up'] = df['Name'].str.upper()
print(df[['Name', 'Name_up']].head(5))
                                                Name  \
0                            Braund, Mr. Owen Harris   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...   
2                             Heikkinen, Miss. Laina   
3       Futrelle, Mrs. Jacques Heath (Lily May Peel)   
4                           Allen, Mr. William Henry   

                                             Name_up  
0                            BRAUND, MR. OWEN HARRIS  
1  CUMINGS, MRS. JOHN BRADLEY (FLORENCE BRIGGS TH...  
2                             HEIKKINEN, MISS. LAINA  
3       FUTRELLE, MRS. JACQUES HEATH (LILY MAY PEEL)  
4                           ALLEN, MR. WILLIAM HENRY  
# Convert 'Embarked' codes from text to category codes (for ML or stats)
df['Embarked_code'] = df['Embarked'].astype('category').cat.codes
print(df[['Embarked', 'Embarked_code']].head(8))
  Embarked  Embarked_code
0        S              2
1        C              0
2        S              2
3        S              2
4        S              2
5        Q              1
6        S              2
7        S              2
# Try a bulk conversion of data types with .astype()
dtype_map = {'Fare': 'float32', 'Age': 'float32', 'Survived': 'int8'}
df = df.astype(dtype_map)
print(df[['Fare', 'Age', 'Survived']].dtypes)
Fare        float32
Age         float32
Survived       int8
dtype: object
# Convert 'Ticket' numbers to strings, even if they look numeric
df['Ticket_str'] = df['Ticket'].astype(str)
print(df[['Ticket', 'Ticket_str']].head(6))
             Ticket        Ticket_str
0         A/5 21171         A/5 21171
1          PC 17599          PC 17599
2  STON/O2. 3101282  STON/O2. 3101282
3            113803            113803
4            373450            373450
5            330877            330877
# Detect non-numeric values in 'Fare' (should be none, but a good habit!)
is_numeric = pd.to_numeric(df['Fare'], errors='coerce').notnull()
print('Non-numeric fares:', (~is_numeric).sum())
Non-numeric fares: 0
# Use pd.to_datetime() to convert 'Sex' (for demo, this fails!)
try:
    df['Sex_dt'] = pd.to_datetime(df['Sex'])
except Exception as e:
    print('Conversion failed:', e)
    
Conversion failed: Unknown datetime string format, unable to parse: male, at position 0
# Use input() to enter height in inches and convert to centimeters
inches = float(input('Enter height in inches: '))
centimeters = inches * 2.54
print(f'Height: {centimeters} cm')
Height: 172.72 cm
# Detect and convert floats stored as strings in a new column
df['FakeFare'] = df['Fare'].astype(str)
df['FakeFare_num'] = pd.to_numeric(df['FakeFare'], errors='coerce')
print(df[['FakeFare', 'FakeFare_num']].head(7))
  FakeFare  FakeFare_num
0     7.25        7.2500
1  71.2833       71.2833
2    7.925        7.9250
3     53.1       53.1000
4     8.05        8.0500
5   8.4583        8.4583
6  51.8625       51.8625
# Mini-project: How many adults and children? (Assume adult is 18 or older)
df['Is_adult'] = df['Age'] >= 18
print(df['Is_adult'].value_counts(dropna=False))
Is_adult
True     601
False    290
Name: count, dtype: int64
# What percent of each sex survived? (using type conversions)
result = df.groupby('Sex_cat')['Survived'].mean() * 100
print(result)
Sex_cat
female    74.203822
male      18.890815
Name: Survived, dtype: float64

Recap: Data Type Changes and Unit Conversion Matter#

You have now seen how to:

  • Check and change column data types
  • Safely convert units for analysis
  • Handle user input and errors
  • Prepare data for machine learning models

These skills help make your analysis accurate and reliable.

If you enjoyed this hands-on walkthrough, please like the video and subscribe for more friendly Python and pandas lessons!

Give these techniques a try on your own data, and let us know in the comments what new problems you have solved.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.