Lesson 1 · Full-Length Courses
Pandas for Absolute Beginners (with Real Data)
Pandas is the most popular Python library for working with tables of data, similar to a spreadsheet or a database table. This lesson runs entirely inside a…
- CourseFull-Length Courses
- Lesson1 of 7
- Video1 h 30 min
- FormatJupyter notebook · 97 code cells
- Data2 datasets
What you'll learn
- Beginner: The Series A Single Labelled Column
- Beginner: The DataFrame A Full Table
- Beginner: Building a DataFrame From a Nested Dictionary
- Intermediate: Combining Tables With merge()
- Intermediate: Stacking Tables With concat()
- Beginner: Working With a Real Dataset mtcars.csv
- Intermediate: Grouping Data With groupby()
- Intermediate: Arithmetic Between DataFrames and Series (Broadcasting)
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- mtcars.csv1.7 KB
- mtcars_efficiency_summary.csv116 B
📓 Full notebook
Download .ipynbPandas for Absolute Beginners (with Real Data)#
- Pandas is the most popular Python library for working with tables of data, similar to a spreadsheet or a database table.
- This lesson runs entirely inside a Jupyter Notebook in VS Code, one cell at a time.
- We start from the very basics, so no prior pandas experience is required.
- We then build up to intermediate skills like grouping, merging, and cleaning messy data, using a real dataset of cars called mtcars.csv.
- By the end, you will be comfortable loading, exploring, filtering, grouping, and summarizing real-world data with pandas.
Before You Start#
- Make sure Python, VS Code, and the Python and Jupyter extensions are already installed.
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- Place mtcars.csv in the same folder as your notebook so pandas can find it with just its file name.
- If pandas is not installed yet, open a terminal in VS Code and run: pip install pandas
import pandas as pd
import numpy as np
print(pd.__version__)
Beginner: The Series A Single Labelled Column#
- A Series is a one-dimensional, labelled array think of it as a single column in a spreadsheet.
- Every value in a Series has a label, called its index, even if you never set one yourself.
- Series are the building block that DataFrames (full tables) are made of.
obj = pd.Series([4, 7, -5, 3])
obj
obj.values
obj.index
obj2 = pd.Series([4, 7, -5, 3], index=['d', 'b', 'a', 'c'])
obj2
obj2['a']
obj2['d'] = 6
obj2
obj2[obj2 > 0]
obj2 * 2
print('b' in obj2)
print('e' in obj2)
Building a Series From a Dictionary#
- Python dictionaries map naturally onto a Series: keys become the index, and values become the data.
sdata = {'Ohio': 35000, 'Texas': 71000, 'Oregon': 16000, 'Utah': 5000}
obj3 = pd.Series(sdata)
obj3
obj3['Kansas'] = 18000
obj3
del obj3['Kansas']
obj3
Automatic Index Alignment#
- When you combine two Series with math operators, pandas lines up the values by their labels first, not by position.
- Labels that only exist in one of the two Series become missing values, shown as NaN (Not a Number).
data = {'California': None, 'Ohio': 35000, 'Oregon': 16000, 'Texas': 71000}
obj4 = pd.Series(data)
obj4
obj3 + obj4
result = obj4.add(obj3, fill_value=0)
result
Beginner: The DataFrame A Full Table#
- A DataFrame is a two-dimensional table made up of rows and columns, exactly like a spreadsheet.
- Every column in a DataFrame is really a Series, and all the columns share the same row index.
- Most real-world pandas work happens on DataFrames rather than single Series.
data = {'state': ['Ohio', 'Ohio', 'Ohio', 'Nevada', 'Nevada'],
'year': [2000, 2001, 2002, 2001, 2002],
'pop': [1.5, 1.7, 3.6, 2.4, 2.9]}
frame = pd.DataFrame(data)
frame
frame2 = pd.DataFrame(data, columns=['year', 'state', 'pop', 'debt'],
index=['one', 'two', 'three', 'four', 'five'])
frame2
Selecting Columns#
frame2['state']
frame2.year
Selecting Rows With loc and iloc#
frame2.loc['three']
frame2.iloc[2]
Adding, Changing, and Removing Columns#
frame2['debt'] = np.arange(5.)
frame2
val = pd.Series([-1.2, -1.5, -1.7], index=['two', 'four', 'five'])
frame2['debt'] = val
frame2
frame2['eastern'] = frame2.state == 'Ohio'
frame2
del frame2['state']
frame2.columns
Dropping Rows and Columns With .drop()#
data2 = pd.DataFrame(np.arange(16).reshape((4, 4)),
index=['Ohio', 'Colorado', 'Utah', 'New York'],
columns=['one', 'two', 'three', 'four'])
data2
data2.drop(['Colorado', 'Ohio'])
data2.drop(['two', 'four'], axis=1)
Beginner: Building a DataFrame From a Nested Dictionary#
pop = {'Nevada': {2001: 2.4, 2002: 2.9},
'Ohio': {2000: 1.5, 2001: 1.7, 2002: 3.6}}
frame3 = pd.DataFrame(pop)
frame3
frame3.index.name = 'year'
frame3.columns.name = 'state'
frame3
Intermediate: Combining Tables With merge()#
- pd.merge() joins two DataFrames together based on a shared column, exactly like a SQL JOIN.
- The 'how' argument controls which rows survive when keys do not perfectly match on both sides.
left_frame = pd.DataFrame({'key': range(5), 'left_value': ['a', 'b', 'c', 'd', 'e']})
right_frame = pd.DataFrame({'key': range(2, 7), 'right_value': ['f', 'g', 'h', 'i', 'j']})
print(left_frame)
print(right_frame)
pd.merge(left_frame, right_frame, on='key', how='inner')
pd.merge(left_frame, right_frame, on='key', how='left')
pd.merge(left_frame, right_frame, on='key', how='right')
pd.merge(left_frame, right_frame, on='key', how='outer')
Intermediate: Stacking Tables With concat()#
- pd.concat() stacks DataFrames together, either on top of each other or side by side.
- Unlike merge, concat does not match on a key column it simply lines things up by position or by index.
pd.concat([left_frame, right_frame])
pd.concat([left_frame, right_frame], axis=1)
Beginner: Working With a Real Dataset mtcars.csv#
- mtcars is a classic dataset listing specifications for 32 different car models.
- mpg: Miles per US gallon (fuel efficiency)
- cyl: Number of cylinders in the engine
- disp: Engine displacement in cubic inches
- hp: Gross horsepower
- drat: Rear axle ratio
- wt: Weight, in thousands of pounds
- qsec: Quarter-mile time in seconds
- vs: Engine shape, 0 means V-shaped, 1 means straight
- am: Transmission, 0 means automatic, 1 means manual
- gear: Number of forward gears
- carb: Number of carburetors
mtcars = pd.read_csv('mtcars.csv')
mtcars.head()
mtcars.shape
mtcars.describe()
mtcars.mean(numeric_only=True)
mtcars[mtcars['am'] == 0]
mtcars['mpg'].hist()
Intermediate: Grouping Data With groupby()#
- .groupby() splits a DataFrame into groups based on the values in one or more columns, then lets you summarize each group separately.
- This is one of the single most useful tools in all of pandas, often described as 'split, apply, combine'.
grouped_by_carb = mtcars.groupby('carb')
grouped_by_carb.mean(numeric_only=True)
grouped_by_carb_am = mtcars.groupby(['carb', 'am'])
grouped_by_carb_am.mean(numeric_only=True)
counts = grouped_by_carb_am['carb'].count()
counts
grouped_stats = mtcars.groupby('carb')[['mpg', 'hp']].agg(['mean', 'std'])
grouped_stats
import matplotlib.pyplot as plt
df = counts.unstack()
ax = df.plot(kind='bar', stacked=True, figsize=(10, 5), colormap='viridis')
ax.set_ylabel('Count')
plt.show()
Try It Yourself: Search by Car Model#
- We can combine string methods with boolean filtering to search for cars by name.
search_term = input('Enter part of a car model name to search for: ')
matches = mtcars[mtcars['model'].str.contains(search_term, case=False)]
matches[['model', 'mpg', 'hp']]
Intermediate: Arithmetic Between DataFrames and Series (Broadcasting)#
- When you subtract a Series from a DataFrame, pandas 'broadcasts' the Series across every row automatically.
- This avoids writing manual loops to repeat an operation across many rows or columns.
frame = pd.DataFrame(np.arange(12.).reshape((4, 3)),
columns=list('bde'),
index=['Utah', 'Ohio', 'Texas', 'Oregon'])
frame
series = frame.iloc[0]
frame - series
series3 = frame['d']
frame.sub(series3, axis=0)
Intermediate: A Closer Look at Indexing#
- We have already used square brackets, .loc, and .iloc a little now let's understand exactly how each one behaves.
obj = pd.Series([4.5, 7.2, -5.3, 3.6], index=['d', 'b', 'a', 'c'])
obj
obj.reindex(list('abcde'))
obj.reindex(list('abcde'), fill_value=0)
obj3 = pd.Series(['blue', 'purple', 'yellow'], index=[0, 2, 4])
obj3.reindex(range(6), method='ffill')
Series Indexing and Slicing#
obj = pd.Series(np.arange(4.), index=['a', 'b', 'c', 'd'])
print(obj.iloc[1])
print(obj['b'])
obj['b':'c']
DataFrame Indexing and Boolean Filtering#
data2['two']
data2[['two', 'four']]
data2['Colorado':'Utah']
data2[:2]
data2[data2['three'] > 5]
data2[data2 < 5] = 0
data2
loc and iloc, Side by Side#
- .loc[label] selects by label; .loc[[label1, label2]] selects a list of labels; .loc[start:end] selects an inclusive label slice.
- .iloc[index] selects by integer position; .iloc[[i, j]] selects a list of positions; .iloc[start:end] selects an exclusive position slice, just like normal Python.
- Both .loc and .iloc also accept a boolean condition to filter rows.
data2.loc['Colorado', ['two', 'three']]
data2.iloc[2]
data2.loc[data2.three > 5, data2.columns[:3]]
Advanced (Optional): Hierarchical Indexing#
- Hierarchical indexing lets a Series or DataFrame have more than one level of row or column labels at once.
- This section is optional for absolute beginners feel free to skip ahead to Missing Data if you prefer, and come back later.
data = pd.Series(np.random.randn(10),
index=[['a', 'a', 'a', 'b', 'b', 'b', 'c', 'c', 'd', 'd'],
[1, 2, 3, 1, 2, 3, 1, 2, 2, 3]])
data
data['b']
data[:, 2]
data.unstack()
frame = pd.DataFrame(np.arange(12).reshape((4, 3)), index=[['a', 'a', 'b', 'b'], [1, 2, 1, 2]],
columns=[['Ohio', 'Ohio', 'Colorado'], ['Green', 'Red', 'Green']])
frame.index.names = ['key1', 'key2']
frame.columns.names = ['state', 'colour']
frame
frame['Ohio']
Intermediate: Applying Your Own Functions#
- Built-in pandas methods cover a lot, but sooner or later you will need a calculation pandas does not already provide.
- .apply() runs your own function once per column (or per row), while .map() runs it on every individual value.
frame = pd.DataFrame(np.random.randn(4, 3), columns=list('bde'), index=['Utah', 'Ohio', 'Texas', 'Oregon'])
frame
f = lambda x: x.max() - x.min()
frame.apply(f)
format_fn = lambda x: '%.2f' % x
frame.map(format_fn)
Beginner: Handling Missing Data#
- Real datasets almost always have gaps. Pandas represents a missing value as NaN.
- Knowing how to detect, fill, and drop missing values is an essential, everyday pandas skill.
data = pd.Series([1, np.nan, 3.5, np.nan, 7])
data
data.isnull()
data.dropna()
data.fillna(0)
data.fillna(data.mean())
data.ffill()
data.bfill()
Missing Data in a Full DataFrame#
missing_df = pd.DataFrame([[1., 6.5, 3.], [1., np.nan, np.nan], [np.nan, np.nan, np.nan], [np.nan, 6.5, 3.]])
missing_df
missing_df.dropna()
missing_df.dropna(how='all')
df = pd.DataFrame(np.random.randn(7, 3))
df.iloc[:4, 1] = np.nan
df.iloc[:2, 2] = np.nan
df
df.dropna(thresh=2)
A Common Beginner Mistake: Index Objects Cannot Be Edited#
- Unlike a list, a pandas Index is immutable once created, individual labels inside it cannot be changed directly.
- The next cell deliberately triggers this error on purpose, so you recognize it immediately if you ever see it in your own code.
obj = pd.Series(range(3), index=list('abc'))
obj.index[1] = 'd'
Intermediate: Catching Errors Gracefully#
- Sometimes you want to detect a problem and handle it in your code, rather than letting the whole notebook stop.
- Python's try/except lets you catch a specific error type and respond to it instead of crashing.
try:
invalid = mtcars['not_a_real_column']
except KeyError as e:
print('KeyError:', e)
try:
bad_value = float('not a number')
except ValueError as e:
print('ValueError:', e)
Best Practice: Wrapping Repeated Work in a Function#
- If you find yourself writing the same few lines of pandas code repeatedly, turn them into a function.
- This avoids copy-paste mistakes and makes your analysis easy to reuse on a different column or dataset.
def summarize_by_group(df, group_col, value_col):
grouped = df.groupby(group_col)[value_col]
summary = grouped.agg(['mean', 'std', 'count'])
return summary.sort_values('mean', ascending=False)
summarize_by_group(mtcars, 'cyl', 'mpg')
End-to-End Mini Project: Which Cars Are Most Fuel-Efficient?#
- Let's combine everything from this lesson loading, filtering, grouping, and sorting into one small real analysis.
- Question: is a manual or automatic transmission more fuel-efficient in this dataset, and how does cylinder count affect the answer?
mtcars['transmission'] = mtcars['am'].map({0: 'automatic', 1: 'manual'})
mtcars[['model', 'am', 'transmission']].head()
efficiency_summary = mtcars.groupby('transmission')['mpg'].agg(['mean', 'min', 'max', 'count'])
efficiency_summary
cyl_transmission_summary = mtcars.groupby(['cyl', 'transmission'])['mpg'].mean().round(1)
cyl_transmission_summary
most_efficient = mtcars.sort_values('mpg', ascending=False)[['model', 'mpg', 'cyl', 'transmission']].head(5)
most_efficient
output_path = 'mtcars_efficiency_summary.csv'
efficiency_summary.to_csv(output_path)
print(f'Summary saved to {output_path}')
Wrap-Up: What You Learned#
- Series and DataFrames, pandas' core building blocks for labelled, one- and two-dimensional data.
- Selecting, filtering, adding, and dropping rows and columns using square brackets, .loc, and .iloc.
- Combining tables with merge() and concat(), and grouping and summarizing data with groupby() and .agg().
- Reshaping data with hierarchical indexes, unstack(), and stack().
- Detecting, filling, and dropping missing data with isnull(), fillna(), ffill(), bfill(), and dropna().
- Applying your own custom functions with .apply() and .map().
- Loading and saving real data with read_csv() and to_csv(), applied to the real mtcars dataset.
- Practice prompt: pick a dataset of your own even a simple CSV export from a spreadsheet you already have and try repeating the mini project's question-and-answer approach on it.
Enjoyed This Lesson?#
- If this pandas walkthrough helped you, consider subscribing it is free, and it is the best way to support more real-data tutorials like this one.
- Drop a comment below with what you would like to see covered next: NumPy, data visualization, or something else entirely.
- If you made it all the way to the end, give the video a like you just leveled up your pandas skills!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



