Lesson 7 · Data analytics zero to hero
Pandas Series & DataFrames from Scratch | Data Analytics #7
Video seven of the 30-part series, and the start of a multi-part pandas block: Series, DataFrames, and reading a real CSV file. We'll load the real mtcars…
- CourseData analytics zero to hero
- Lesson7 of 30
- Video15 min
- FormatJupyter notebook · 15 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- mtcars.csv1.7 KB
📓 Full notebook
Download .ipynbData Analytics Zero to Hero, Video 7: Pandas Series and DataFrames#
- Video seven of the 30-part series, and the start of a multi-part pandas block: Series, DataFrames, and reading a real CSV file.
- We'll load the real mtcars dataset, 1974 Motor Trend car road tests, straight into pandas.
- Let's jump straight in.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- Install pandas if you haven't already:
pip install pandas. - Place mtcars.csv in the same folder as this notebook.
Part 1: The Series#
import pandas as pd
mpg = pd.Series([21.0, 22.8, 21.4, 18.7, 18.1])
print(mpg)
print(mpg.values)
print(mpg.index)
mpg_named = pd.Series([21.0, 22.8, 21.4], index=['Mazda RX4', 'Datsun 710', 'Hornet 4 Drive'])
print(mpg_named)
print(mpg_named['Datsun 710'])
Part 2: Building a DataFrame#
data = {
'model': ['Mazda RX4', 'Datsun 710', 'Hornet 4 Drive'],
'mpg': [21.0, 22.8, 21.4],
'cyl': [6, 4, 6]
}
df = pd.DataFrame(data)
print(df)
Part 3: Reading a Real CSV File#
df = pd.read_csv('mtcars.csv')
print(df.head())
print(df.tail(3))
print(df.shape)
print(df.columns)
print(df.dtypes)
df.info()
print(df.describe())
Part 4: Selecting Columns#
mpg_col = df['mpg']
print(type(mpg_col))
print(mpg_col.head())
subset = df[['model', 'mpg', 'hp']]
print(type(subset))
print(subset.head())
Part 5: Selecting Rows with loc and iloc#
print(df.iloc[0])
print(df.iloc[0:3])
df_named = df.set_index('model')
print(df_named.loc['Fiat 128'])
print(df_named.loc['Fiat 128', 'mpg'])
Part 6: Boolean Filtering#
efficient = df[df['mpg'] > 25]
print(efficient[['model', 'mpg']])
v8_efficient = df[(df['cyl'] == 8) & (df['mpg'] > 15)]
print(v8_efficient[['model', 'cyl', 'mpg']])
Part 7: Adding Columns and Sorting#
df['kml'] = df['mpg'] * 0.425
print(df[['model', 'mpg', 'kml']].head())
top_mpg = df.sort_values('mpg', ascending=False)
print(top_mpg[['model', 'mpg']].head())
Wrap-Up: What You Learned#
- The Series: a single labeled column, and the building block of every DataFrame.
- Building a DataFrame from a dictionary, and reading one directly from a real CSV file with read_csv.
- Exploring a new DataFrame: head, tail, shape, columns, dtypes, info, and describe.
- Selecting columns with square brackets, and rows with loc and iloc.
- Boolean filtering, including combining multiple conditions.
- Adding a new computed column, and sorting with sort_values.
- This is the pandas foundation everything else in this series builds on. Video eight moves into real-world data cleaning: missing values, duplicates, and data types. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



