Mathew K Analytics

Lesson 4 · Mastering Pandas

Mastering DataFrames in Python Pandas: Creating & Inspecting for Data Science

Welcome to this hands-on session! We will start by introducing DataFrames, the heart of pandas. You will learn to create, explore, and understand tabular…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Creating and Inspecting DataFrames in Pandas#

Welcome to this hands-on session! We will start by introducing DataFrames, the heart of pandas. You will learn to create, explore, and understand tabular data using simple, real-world examples. Let us dive in!

import warnings
warnings.filterwarnings('ignore')

# Data setup (Iris Dataset)
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(150, 5)
   sepal_length  sepal_width  petal_length  petal_width species
0           5.1          3.5           1.4          0.2  setosa
1           4.9          3.0           1.4          0.2  setosa
2           4.7          3.2           1.3          0.2  setosa

What is a DataFrame?#

A DataFrame is like a table or spreadsheet in your notebook. Each column has a name and can hold different types of data. You can analyze, filter, sort, and visualize data easily with DataFrames.

# Quick tour of DataFrame info
df.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 150 entries, 0 to 149
Data columns (total 5 columns):
 #   Column        Non-Null Count  Dtype  
---  ------        --------------  -----  
 0   sepal_length  150 non-null    float64
 1   sepal_width   150 non-null    float64
 2   petal_length  150 non-null    float64
 3   petal_width   150 non-null    float64
 4   species       150 non-null    object 
dtypes: float64(4), object(1)
memory usage: 6.0+ KB
# Examining column names and data types
print('Column names:', list(df.columns))
print('Data types:')
print(df.dtypes)
Column names: ['sepal_length', 'sepal_width', 'petal_length', 'petal_width', 'species']
Data types:
sepal_length    float64
sepal_width     float64
petal_length    float64
petal_width     float64
species          object
dtype: object
# Creating a DataFrame from scratch
data = {
    'name': ['Elena', 'Ravi', 'Lee'],
    'score': [90, 88, 70],
    'passed': [True, True, False]
}
df_simple = pd.DataFrame(data)
print(df_simple)
    name  score  passed
0  Elena     90    True
1   Ravi     88    True
2    Lee     70   False

Summarizing DataFrames#

Pandas can help summarize columns using built-in functions. This helps you quickly see the range of the numeric columns or counts of unique values for text.

# Get summary statistics for numeric columns
print(df.describe())
       sepal_length  sepal_width  petal_length  petal_width
count    150.000000   150.000000    150.000000   150.000000
mean       5.843333     3.054000      3.758667     1.198667
std        0.828066     0.433594      1.764420     0.763161
min        4.300000     2.000000      1.000000     0.100000
25%        5.100000     2.800000      1.600000     0.300000
50%        5.800000     3.000000      4.350000     1.300000
75%        6.400000     3.300000      5.100000     1.800000
max        7.900000     4.400000      6.900000     2.500000
# Value counts for a text column
print(df['species'].value_counts())
species
setosa        50
versicolor    50
virginica     50
Name: count, dtype: int64

Indexing and Selecting Data#

Pandas lets you grab specific rows or columns using labels or numbers. This way, you can focus on just the part of the data you need.

# Select a single column (becomes a Series)
sepal_lengths = df['sepal_length']
print(sepal_lengths.head())
0    5.1
1    4.9
2    4.7
3    4.6
4    5.0
Name: sepal_length, dtype: float64
# Select multiple columns by name
subset = df[['sepal_length', 'species']]
print(subset.head())
   sepal_length species
0           5.1  setosa
1           4.9  setosa
2           4.7  setosa
3           4.6  setosa
4           5.0  setosa
# Selecting rows by position using iloc
row5 = df.iloc[4]
print(row5)
sepal_length       5.0
sepal_width        3.6
petal_length       1.4
petal_width        0.2
species         setosa
Name: 4, dtype: object
# Filtering rows with a condition
long_flowers = df[df['sepal_length'] > 7.0]
print(long_flowers.head())
     sepal_length  sepal_width  petal_length  petal_width    species
102           7.1          3.0           5.9          2.1  virginica
105           7.6          3.0           6.6          2.1  virginica
107           7.3          2.9           6.3          1.8  virginica
109           7.2          3.6           6.1          2.5  virginica
117           7.7          3.8           6.7          2.2  virginica

Modifying DataFrames#

You can add new columns, change values, or drop data as you clean and transform your data.

# Add a new column based on a calculation
df['sepal_ratio'] = df['sepal_length'] / df['sepal_width']
print(df[['sepal_length', 'sepal_width', 'sepal_ratio']].head())
   sepal_length  sepal_width  sepal_ratio
0           5.1          3.5     1.457143
1           4.9          3.0     1.633333
2           4.7          3.2     1.468750
3           4.6          3.1     1.483871
4           5.0          3.6     1.388889
# Deleting a column
df = df.drop('sepal_ratio', axis=1)
print(df.head(2))
   sepal_length  sepal_width  petal_length  petal_width species
0           5.1          3.5           1.4          0.2  setosa
1           4.9          3.0           1.4          0.2  setosa

Quick Practice: Exploring DataFrames#

Try using shape, head, and info on the df_simple DataFrame you made earlier. What do you notice compared to the main iris DataFrame?

# Input practice: Name a column you see in df
column_guess = input('Type any column name from the main iris DataFrame: ')
if column_guess in df.columns:
    print('Nice! That column exists.')
else:
    print('Check spelling and try again.')
    
Nice! That column exists.

Recap: What We Covered#

  • Creating DataFrames from real files and from scratch
  • Inspecting and summarizing columns
  • Selecting and filtering with labels and numbers
  • Adding and deleting columns to transform your data

This is the foundation for all pandas analysis!

What next? Try these:#

  • Experiment with loading different datasets.
  • Practice selecting and filtering columns and rows.
  • Visualize your DataFrame using .plot().

If you enjoyed this lesson, subscribe for more pandas tips and tutorials!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.