Mathew K Analytics

Lesson 44 · Mastering Pandas

Efficient Data Analysis with Pandas: Master Method Chaining & Pipelines in Python

In this lesson, you will master chaining pandas methods and building readable pipelines. We will use the Titanic dataset to see how chaining makes data…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Chaining Operations & Method Pipelines in Pandas#

In this lesson, you will master chaining pandas methods and building readable pipelines.

We will use the Titanic dataset to see how chaining makes data analysis clearer and more efficient.

You will also learn common patterns for creating intermediate steps, debugging, and best practices.

import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  

Why Chain Methods?#

With chaining, you perform a sequence of operations step-by-step in a single line.

This helps with fast data exploration and reduces bugs from intermediate variables.

Lets see how this works with a simple filter.

# Select women who survived
women_survivors = df[df['Sex'] == 'female'][df['Survived'] == 1]
print(women_survivors[['Name', 'Age', 'Survived']].head(3))
                                                Name   Age  Survived
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  38.0         1
2                             Heikkinen, Miss. Laina  26.0         1
3       Futrelle, Mrs. Jacques Heath (Lily May Peel)  35.0         1

Problems With Direct Chaining#

Chaining square brackets like above can sometimes cause errors or warnings in pandas.

For reliable pipelines, use .loc or angle for method-based chaining.

# Safer chained filter using .loc
women_survivors = df.loc[(df['Sex'] == 'female') & (df['Survived'] == 1)]
print(women_survivors[['Name', 'Age', 'Survived']].head(3))
                                                Name   Age  Survived
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  38.0         1
2                             Heikkinen, Miss. Laina  26.0         1
3       Futrelle, Mrs. Jacques Heath (Lily May Peel)  35.0         1
# Start your first real method chain
age_stats = (df.query("Embarked == 'S'").groupby('Pclass')[['Age']].mean().sort_values('Age', ascending=False))
print(age_stats)
              Age
Pclass           
1       38.152037
2       30.386731
3       25.696552
# Chaining with .assign() to add columns
df_chain = (df.assign(Title = df['Name'].str.extract(r'(\w+)\.') ).assign(Child = df['Age'] < 16).query("Fare > 50")[['Name', 'Title', 'Child', 'Fare']]
)
print(df_chain.head(3))
                                                Name Title  Child     Fare
1  Cumings, Mrs. John Bradley (Florence Briggs Th...   Mrs  False  71.2833
3       Futrelle, Mrs. Jacques Heath (Lily May Peel)   Mrs  False  53.1000
6                            McCarthy, Mr. Timothy J    Mr  False  51.8625
# .pipe for custom functions
def add_age_bins(df):
    bins = [0, 12, 18, 35, 60, 100]
    labels = ['child', 'teen', 'young_adult', 'adult', 'senior']
    return df.assign(AgeGroup = pd.cut(df['Age'], bins, labels=labels))

df_piped = (df
.pipe(add_age_bins)
.groupby(['Pclass', 'AgeGroup'])
.agg({'Survived':'mean', 'Fare':'median'})
)
print(df_piped.head())
                    Survived       Fare
Pclass AgeGroup                        
1      child        0.750000  135.77500
       teen         0.916667  108.90000
       young_adult  0.757576   72.79585
       adult        0.611111   58.68960
       senior       0.214286   34.07710

Dealing With Missing Values in a Chain#

You often want to clean data right in the method flow.

Let us fill missing ages with the median and drop any unused columns.

# Fill NA and drop
clean_chain = (df
.assign(Age = lambda d: d['Age'].fillna(d['Age'].median()))
.drop(columns=['Cabin', 'Ticket'])
)
print(clean_chain.head(2))
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   

   Parch     Fare Embarked  
0      0   7.2500        S  
1      0  71.2833        C  
# Sort, filter, and reset index in one chain
chain_result = (df
.sort_values('Fare', ascending=False)
.query('Age >= 18')
.reset_index(drop=True)
[['Name', 'Age', 'Fare']]
)
print(chain_result.head(5))
                                 Name   Age      Fare
0                    Ward, Miss. Anna  35.0  512.3292
1              Lesurer, Mr. Gustave J  35.0  512.3292
2  Cardeza, Mr. Thomas Drake Martinez  36.0  512.3292
3          Fortune, Miss. Mabel Helen  23.0  263.0000
4      Fortune, Mr. Charles Alexander  19.0  263.0000

Chaining GroupBy, Aggregation, and Custom Transforms#

You can combine grouping, aggregating, and custom changes in one flow for deep analysis.

Let us see the survival rates and mean fares, split by embarkation point and class.

# Analyze by Embarked and Pclass
summary = (df
.dropna(subset=['Embarked'])
.groupby(['Embarked', 'Pclass'])
.agg(survival_rate = ('Survived', 'mean'), mean_fare = ('Fare', 'mean'))
.reset_index()
)
print(summary.head())
  Embarked  Pclass  survival_rate   mean_fare
0        C       1       0.694118  104.718529
1        C       2       0.529412   25.358335
2        C       3       0.378788   11.214083
3        Q       1       0.500000   90.000000
4        Q       2       0.666667   12.350000
# Chain reshaping techniques: pivot
pivot_df = (df
.dropna(subset=['Embarked'])
.pivot_table(index='Pclass', columns='Embarked', values='Fare', aggfunc='median')
)
print(pivot_df)
Embarked        C      Q      S
Pclass                         
1         78.2667  90.00  52.00
2         24.0000  12.35  13.50
3          7.8958   7.75   8.05
# Chain merging: join summary back to passengers
full_merge = (df
.dropna(subset=['Embarked'])
.merge(summary, how='left', on=['Embarked', 'Pclass'])
[['Name', 'Embarked', 'Pclass', 'Fare', 'survival_rate', 'mean_fare']]
)
print(full_merge.head(3))
                                                Name Embarked  Pclass  \
0                            Braund, Mr. Owen Harris        S       3   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...        C       1   
2                             Heikkinen, Miss. Laina        S       3   

      Fare  survival_rate   mean_fare  
0   7.2500       0.189802   14.644083  
1  71.2833       0.694118  104.718529  
2   7.9250       0.189802   14.644083  
# Visualization in a chain: plot survival by class
(df
.groupby('Pclass')['Survived']
.mean()
.plot(kind='bar', title='Survival Rate by Passenger Class'))
<Axes: title={'center': 'Survival Rate by Passenger Class'}, xlabel='Pclass'>
No description has been provided for this image

Mini Project: Pipeline for Fare and Survival Analysis#

Now it is your turn! Let us build an end-to-end analysis that:

  • Cleans age and embarked columns
  • Bins Fare into quartiles
  • Groups by Fare bin and sex for survival rate
  • Shows results as a quick bar chart
# Clean, bin, and analyze with a long method chain
fare_bins = (df
.dropna(subset=['Fare', 'Embarked'])
.assign(Age = lambda d: d['Age'].fillna(d['Age'].median()))
.assign(FareBin = lambda d: pd.qcut(d['Fare'], 4, labels=['Low', 'Mid', 'High', 'VeryHigh']))
.groupby(['FareBin', 'Sex'])['Survived']
.mean()
.unstack()
)
print(fare_bins)
Sex         female      male
FareBin                     
Low       0.697674  0.077778
Mid       0.641791  0.159236
High      0.698925  0.279070
VeryHigh  0.853211  0.306306
# Visualize results in the same pipeline
fare_bins.plot(kind='bar', title='Survival Rate by Fare Bin and Sex')
<Axes: title={'center': 'Survival Rate by Fare Bin and Sex'}, xlabel='FareBin'>
No description has been provided for this image

Best Practices for Chaining#

  • Break long chains for readability using parentheses.
  • Use .assign or .pipe for custom logic steps.
  • Keep chains focuseddo not mix too many unrelated tasks.
  • Always check intermediary results if something looks wrong.
  • Prefer method chains over chaining only with square brackets.
# Debugging: use .head() mid-chain for sanity checks
(df.assign(name_up = lambda d: d['Name'].str.upper())
.head(2)
)
PassengerId Survived Pclass Name Sex Age SibSp Parch Ticket Fare Cabin Embarked name_up
0 1 0 3 Braund, Mr. Owen Harris male 22.0 1 0 A/5 21171 7.2500 NaN S BRAUND, MR. OWEN HARRIS
1 2 1 1 Cumings, Mrs. John Bradley (Florence Briggs Th... female 38.0 1 0 PC 17599 71.2833 C85 C CUMINGS, MRS. JOHN BRADLEY (FLORENCE BRIGGS TH...
# Avoid mistakes: forget parentheses by accident
# result = df.assign(NewAge = df['Age'] + 5).head
# print(result)
# Challenge: Type your own chain!
name_start = input("Type a letter: ").upper()
subset = (df[df['Name'].str.startswith(name_start)]
.sort_values('Fare', ascending=False)
.head(3)
)
print(subset[['Name', 'Age', 'Fare']])
                                                Name   Age      Fare
118                         Baxter, Mr. Quigg Edmond  24.0  247.5208
299  Baxter, Mrs. James (Helene DeLaudeniere Chaput)  50.0  247.5208
380                            Bidois, Miss. Rosalie  42.0  227.5250
# Spot the chain bug: will this fail?
bad_chain = (df.query('Pclass == 1')
# .dropna(subset='Cabin')  # Oops, missing brackets on subset!
.sort_values('Fare', ascending=True)
.head(2)
)
print(bad_chain[['Name', 'Cabin', 'Fare']])
                                Name Cabin  Fare
633    Parr, Mr. William Henry Marsh   NaN   0.0
822  Reuchlin, Jonkheer. John George   NaN   0.0

Lesson Recap#

  • Chaining lets you express complex workflows in one smooth series of steps.
  • Use method chaining for filters, cleaning, grouping, and even plotting.
  • .pipe and .assign bring extra power for custom logic.
  • Check output often, and break chains for readability.
  • Now you can craft clear, repeatable data pipelines in pandas!

Keep Practicing and Join the Community!#

Try chaining your own analysis, share your cool pipelines in the comments, and stay tuned for more Python data tips.

Subscribe if you found this helpful!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.