Lesson 26 · Mastering Pandas
How to Concatenate DataFrames Vertically and Horizontally Using Pandas in Python
This lesson will help you master combining DataFrames using pandas. You will learn by hands-on typing, merging data from realistic sources. Let us explore…
- CourseMastering Pandas
- Lesson26 of 44
- Video15 min
- FormatJupyter notebook · 13 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbConcatenating DataFrames Vertically and Horizontally in Pandas#
This lesson will help you master combining DataFrames using pandas.
You will learn by hands-on typing, merging data from realistic sources.
Let us explore real-world use cases and unlock new possibilities with your data!
# Suppress pandas warnings for a cleaner notebook
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
Why Concatenation?#
Imagine you receive monthly Titanic passenger data as separate sheets. Or maybe you want to add more columns with new passenger info.
Concatenation helps you stack tables together or join side by side.
It is vital for data cleaning, tracking updates, and analysis.
# Split the Titanic data into two DataFrames (simulate monthly arrivals)
df_jan = df.iloc[:400].copy()
df_feb = df.iloc[400:].copy()
print('January shape:', df_jan.shape)
print('February shape:', df_feb.shape)
# Concatenate DataFrames vertically (one below the other)
df_vertical = pd.concat([df_jan, df_feb], axis=0)
print('Combined shape:', df_vertical.shape)
# Reset index after concatenation for clean index numbers
df_vertical = df_vertical.reset_index(drop=True)
df_vertical.head(3)
Vertical Concatenation Details#
Vertical (row-wise) concatenation stacks data tables on top of each other. All columns must match or missing columns will get NaN values.
This is helpful when combining multiple files or appending new records.
# Concatenate with some columns missing: add an 'AgeGroup' column to one DataFrame
df_jan['AgeGroup'] = pd.cut(df_jan['Age'], bins=[0, 12, 18, 50, 80], labels=['Child', 'Teen', 'Adult', 'Senior'])
df_mixed = pd.concat([df_jan, df_feb], axis=0)
print(df_mixed[['Name', 'Age', 'AgeGroup']].head(5))
print(df_mixed[['Name', 'Age', 'AgeGroup']].tail(5))
Horizontal Concatenation (Column-wise)#
Next, let us see how to add extra columns that belong together.
This is useful for 'joining' details from two sourceslike adding ticket or contact info for each passenger.
# Create a new DataFrame with alternative contact information
contact_info = pd.DataFrame({
'PassengerId': df_jan['PassengerId'],
'Contact': ['contact'+str(i)+'@mail.com' for i in df_jan['PassengerId']]
})
contact_info.head(3)
# Add contact information columns to the original January DataFrame (horizontal concat)
df_horiz = pd.concat([df_jan.reset_index(drop=True), contact_info.drop('PassengerId', axis=1).reset_index(drop=True)], axis=1)
df_horiz[['Name', 'Age', 'Contact']].head(3)
# Horizontal concat WITHOUT resetting index causes mismatches if DataFrames differ in rows
short_contacts = contact_info.iloc[:10]
horiz_bad = pd.concat([df_jan, short_contacts.drop('PassengerId',axis=1)], axis=1)
print(horiz_bad[['Name', 'Contact']].head(12))
Common Parameters and Pitfalls#
- axis=0: stack rows (default).
- axis=1: stack columns.
- ignore_index=True: reset row labels.
- join='outer': union of columns (default for concat).
- join='inner': only columns (or rows) shared by all.
Always check your index, shape, and column line-up after concatenation.
# Using ignore_index for fresh row numbers after vertical stacking
df_stack = pd.concat([df_jan, df_feb], axis=0, ignore_index=True)
print(df_stack.index[:8])
# Inner join: keep ONLY columns present in both DataFrames
small_jan = df_jan[['PassengerId', 'Name', 'Age']]
small_feb = df_feb[['PassengerId', 'Name', 'Fare']]
joined = pd.concat([small_jan, small_feb], axis=0, join='inner', ignore_index=True)
print(joined.head(3))
print(joined.tail(3))
# Mini-project: Merge two datasets on matching column (merge vs. concat demo)
df_ticket = df[['PassengerId', 'Ticket', 'Cabin']].sample(10, random_state=42).reset_index(drop=True)
df_sample = df[['PassengerId', 'Name', 'Age']].sample(10, random_state=24).reset_index(drop=True)
merged = pd.merge(df_sample, df_ticket, on='PassengerId', how='left')
print(merged)
Challenge: Practice Problem#
Suppose you have two DataFrames:
- One holds the first 5 passengers' names and ages.
- The other has their genders and fares.
Challenge yourself: Vertically stack both DataFrames, then horizontally join as well.
What differences do you notice?
# Try: Concatenate first 5 names/ages with first 5 sexes/fares vertically and horizontally
A = df[['Name', 'Age']].head(5)
B = df[['Sex', 'Fare']].head(5)
vert = pd.concat([A, B], axis=0, ignore_index=True)
horiz = pd.concat([A.reset_index(drop=True), B.reset_index(drop=True)], axis=1)
print('Vertical stack:\n', vert)
print('Horizontal stack:\n', horiz)
Recap: Concatenation Basics#
In this lesson, you practiced combining DataFrames using vertical and horizontal concatenation.
You learned to split, append, and align tablesessential for real-world data work.
Keep experimenting and double-check your index, shape, and columns as you merge!
Next: Try using concat on your own datasets.
Thanks for Learning!#
If you found this helpful, subscribe to our channel for more pandas coding tutorials.
Type your favorite tip from this lesson in the comments below!
Happy coding and see you in the next Jupyter lesson!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



