Mathew K Analytics

Lesson 27 · Python for Data Science

Mastering Pandas DataFrames: Essential Techniques for Data Science in Python

Pandas DataFrames are a must-have skill for anyone working with data in Python. They help you organize and analyze your information easily. Let us see why…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Welcome to Your First Lesson on Pandas DataFrames!#

Pandas DataFrames are a must-have skill for anyone working with data in Python.

They help you organize and analyze your information easily.

Let us see why learning them will open many doors for you.

# First, let us import pandas.
import pandas as pd

What is a DataFrame?#

A DataFrame is like a table with rows and columns.

Think of it as a super-powered spreadsheet in Python.

# Let us build our first DataFrame.
data = {
    'Name': ['Sam', 'Alex', 'Jordan'],
    'Age': [22, 35, 58],
    'City': ['New York', 'Paris', 'London']
}
df = pd.DataFrame(data)

# Look at the DataFrame.
df
Name Age City
0 Sam 22 New York
1 Alex 35 Paris
2 Jordan 58 London
# Checking the data type.
type(df)
pandas.core.frame.DataFrame

How to See Rows and Columns#

DataFrames have a shape rows go across, columns go up and down.

Let us look at how to see how big our table is.

# Find out how many rows and columns we have.
df.shape
(3, 3)
# Look at just the column names.
df.columns
Index(['Name', 'Age', 'City'], dtype='object')
# Look at the first two rows only.
df.head(2)
Name Age City
0 Sam 22 New York
1 Alex 35 Paris

Selecting Columns and Rows#

You can pick out columns or specific rows from your DataFrame.

Let us practice selecting data just like picking a column in a spreadsheet.

# Select a single column.
df['Name']
0       Sam
1      Alex
2    Jordan
Name: Name, dtype: object
# Get row number 1 (remember: counting starts at 0).
df.iloc[1]
Name     Alex
Age        35
City    Paris
Name: 1, dtype: object
# Select multiple columns.
df[['Name', 'Age']]
Name Age
0 Sam 22
1 Alex 35
2 Jordan 58
# Filter rows based on a condition.
df[df['Age'] > 30]
Name Age City
1 Alex 35 Paris
2 Jordan 58 London

Adding and Changing Data#

DataFrames are easy to updatejust like editing a spreadsheet.

Let us add a new column.

# Add a new column with values.
df['Country'] = ['USA', 'France', 'UK']
df
Name Age City Country
0 Sam 22 New York USA
1 Alex 35 Paris France
2 Jordan 58 London UK
# Change a value in the DataFrame.
df.at[2, 'City'] = 'Berlin'
df
Name Age City Country
0 Sam 22 New York USA
1 Alex 35 Paris France
2 Jordan 58 Berlin UK
# Remove a column using drop.
df = df.drop('Country', axis=1)
df
Name Age City
0 Sam 22 New York
1 Alex 35 Paris
2 Jordan 58 Berlin
# Let us delete a row too.
df = df.drop(1, axis=0)
df
Name Age City
0 Sam 22 New York
2 Jordan 58 Berlin

Describe and Summarize Your Data#

You often want a quick summary of what is in your table.

Pandas can do this quickly.

# Describe gives you useful stats for numbers.
df.describe()
Age
count 2.000000
mean 40.000000
std 25.455844
min 22.000000
25% 31.000000
50% 40.000000
75% 49.000000
max 58.000000
# Use info to see column types and missing values.
df.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 2 entries, 0 to 2
Data columns (total 3 columns):
 #   Column  Non-Null Count  Dtype 
---  ------  --------------  ----- 
 0   Name    2 non-null      object
 1   Age     2 non-null      int64 
 2   City    2 non-null      object
dtypes: int64(1), object(2)
memory usage: 180.0+ bytes

Reading and Writing Data#

DataFrames make it simple to load data from files, like CSV spreadsheets.

Let us see how to save and reload your work.

# Save your DataFrame to a CSV file.
df.to_csv('my_table.csv', index=False)
# Read your data back into Python.
df_loaded = pd.read_csv('my_table.csv')
df_loaded
Name Age City
0 Sam 22 New York
1 Jordan 58 Berlin
# Work with user input to create a new DataFrame.
names = input('Enter three names, separated by commas: ').split(',')
ages = input('Enter their ages, separated by commas: ').split(',')
cities = input('Enter their cities, separated by commas: ').split(',')
user_df = pd.DataFrame({
    'Name': [n.strip() for n in names],
    'Age': [int(a.strip()) for a in ages],
    'City': [c.strip() for c in cities]
})
user_df
})
  Cell In[19], line 11
    })
    ^
SyntaxError: unmatched '}'
bm
## Mini-Project: Favorite Movies Table Part 1

Let us build a table of movies and their ratings.

We will rate a few movies and then sort by score.
  Cell In[20], line 4
    Let us build a table of movies and their ratings.
        ^
SyntaxError: invalid syntax
# Make a DataFrame of movies and ratings.
movie_data = {
    'Movie': ['Inception', 'Moana', 'Avengers'],
    'Rating': [8.7, 7.6, 8.4]
}
movies_df = pd.DataFrame(movie_data)
movies_df
Movie Rating
0 Inception 8.7
1 Moana 7.6
2 Avengers 8.4
# Sort the table by the best rating at the top.
movies_df.sort_values('Rating', ascending=False)
Movie Rating
0 Inception 8.7
2 Avengers 8.4
1 Moana 7.6
# Calculate the average rating.
movies_df['Rating'].mean()
np.float64(8.233333333333333)
# Add a column for how many times you have watched each movie.
movies_df['Watches'] = [2, 10, 3]
movies_df
Movie Rating Watches
0 Inception 8.7 2
1 Moana 7.6 10
2 Avengers 8.4 3
# Find movies you watched more than three times.
movies_df[movies_df['Watches'] > 3]
Movie Rating Watches
1 Moana 7.6 10

Troubleshooting: What if You Get an Error?#

If you type a column name wrong or your lists are not the same length, pandas will warn you.

Read error messages carefullythey often tell you what went wrong.

Do not let mistakes stop you. Everyone makes them at first!

# Bonus: Create a DataFrame from a list of lists.
grades = [
    ['Science', 88],
    ['Math', 95],
    ['English', 79]
]
grades_df = pd.DataFrame(grades, columns=['Subject', 'Score'])
grades_df
Subject Score
0 Science 88
1 Math 95
2 English 79
# Challenge: Create a table for your weekly chores.
chores = input('Enter chores, separated by commas: ').split(',')
minutes = input('Enter minutes for each, separated by commas: ').split(',')
chores_df = pd.DataFrame({
    'Chore': [c.strip() for c in chores],
    'Minutes': [int(m.strip()) for m in minutes]
})
chores_df
})
  Cell In[27], line 9
    })
    ^
SyntaxError: unmatched '}'
bm
## Recap and Next Steps

Congratulationsyou have learned the basics of pandas DataFrames!

You now know how to create tables, select data, edit values, and even load from files.

With practice, these tools will help you understand any dataset.
  Cell In[28], line 4
    Congratulationsyou have learned the basics of pandas DataFrames!
                       ^
SyntaxError: invalid syntax

Thanks for Learning with Us!#

If this helped you, please like, subscribe, and share the video.

Let us learn Python togetheryour journey has just begun.

See you in the next lesson!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.