Mathew K Analytics

Lesson 10 · Data Science Projects

Comprehensive Airbnb Data Analysis Using Python for Biblical Stewardship Insights

Welcome to this beginner friendly lesson on analyzing Airbnb data. We will use a real life dataset of New York City Airbnb listings. You will learn…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Airbnb Data Analysis: Exploring NYC Listings#

  • Welcome to this beginner friendly lesson on analyzing Airbnb data.
  • We will use a real life dataset of New York City Airbnb listings.
  • You will learn important steps for loading, cleaning, analyzing, and visualizing data.
  • By the end, you will know how to get insights from a real world dataset.
  • Let us dive in!
# Always suppress warnings at the start for a cleaner experience
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)

# This makes sure you only see output you want while learning

What is Airbnb Data and Why Analyze It?#

  • Airbnb is an online marketplace for short term rentals and stays.
  • The public dataset contains listings, prices, neighbourhoods, reviews, and more for New York City.
  • Learning to analyze such data is a valuable job skill for both students and professionals.
  • Data analysis helps answer real questions, like:
  • What is the most expensive area?
  • How are prices distributed?
  • Which neighborhoods are most popular?
  • You will learn to answer these using Python!
# Data setup: Load the Airbnb NYC Listings dataset using pandas
import kagglehub, os, pandas as pd
path = kagglehub.dataset_download('dgomonov/new-york-city-airbnb-open-data')
files = os.listdir(path)
df = pd.read_csv(os.path.join(path, 'AB_NYC_2019.csv'))
print(df.shape)
print(df.head(3))
(48895, 16)
     id                                 name  host_id  host_name  \
0  2539   Clean & quiet apt home by the park     2787       John   
1  2595                Skylit Midtown Castle     2845   Jennifer   
2  3647  THE VILLAGE OF HARLEM....NEW YORK !     4632  Elisabeth   

  neighbourhood_group neighbourhood  latitude  longitude        room_type  \
0            Brooklyn    Kensington  40.64749  -73.97237     Private room   
1           Manhattan       Midtown  40.75362  -73.98377  Entire home/apt   
2           Manhattan        Harlem  40.80902  -73.94190     Private room   

   price  minimum_nights  number_of_reviews last_review  reviews_per_month  \
0    149               1                  9  2018-10-19               0.21   
1    225               1                 45  2019-05-21               0.38   
2    150               3                  0         NaN                NaN   

   calculated_host_listings_count  availability_365  
0                               6               365  
1                               2               355  
2                               1               365  
# Take a closer look at column names and basic info
print(df.columns)

print(df.info())
Index(['id', 'name', 'host_id', 'host_name', 'neighbourhood_group',
       'neighbourhood', 'latitude', 'longitude', 'room_type', 'price',
       'minimum_nights', 'number_of_reviews', 'last_review',
       'reviews_per_month', 'calculated_host_listings_count',
       'availability_365'],
      dtype='object')
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 48895 entries, 0 to 48894
Data columns (total 16 columns):
 #   Column                          Non-Null Count  Dtype  
---  ------                          --------------  -----  
 0   id                              48895 non-null  int64  
 1   name                            48879 non-null  object 
 2   host_id                         48895 non-null  int64  
 3   host_name                       48874 non-null  object 
 4   neighbourhood_group             48895 non-null  object 
 5   neighbourhood                   48895 non-null  object 
 6   latitude                        48895 non-null  float64
 7   longitude                       48895 non-null  float64
 8   room_type                       48895 non-null  object 
 9   price                           48895 non-null  int64  
 10  minimum_nights                  48895 non-null  int64  
 11  number_of_reviews               48895 non-null  int64  
 12  last_review                     38843 non-null  object 
 13  reviews_per_month               38843 non-null  float64
 14  calculated_host_listings_count  48895 non-null  int64  
 15  availability_365                48895 non-null  int64  
dtypes: float64(3), int64(7), object(6)
memory usage: 6.0+ MB
None

Understanding the Data Columns#

  • Here is what some important columns represent:
  • id: listing unique identifier
  • name: title of the listing
  • host_name: name of the seller
  • neighbourhood_group: major NYC area like Manhattan or Brooklyn
  • neighbourhood: more detailed local name
  • latitude and longitude: location coordinates
  • room_type: type such as entire home/apartment or private room
  • price: price per night in dollars
  • minimum_nights: minimum nights required to book
  • number_of_reviews: how many times reviewed
  • availability_365: how many days per year available
  • Knowing what each column represents helps make analysis more meaningful!
# Check if there are missing values in any column
print(df.isnull().sum())
id                                    0
name                                 16
host_id                               0
host_name                            21
neighbourhood_group                   0
neighbourhood                         0
latitude                              0
longitude                             0
room_type                             0
price                                 0
minimum_nights                        0
number_of_reviews                     0
last_review                       10052
reviews_per_month                 10052
calculated_host_listings_count        0
availability_365                      0
dtype: int64
# Basic data cleaning: fill missing host_name with 'Unknown', drop rows missing price
df['host_name'].fillna('Unknown', inplace=True)
df.dropna(subset=['price'], inplace=True)

# Now let us check again for missing values
print(df.isnull().sum())
id                                    0
name                                 16
host_id                               0
host_name                             0
neighbourhood_group                   0
neighbourhood                         0
latitude                              0
longitude                             0
room_type                             0
price                                 0
minimum_nights                        0
number_of_reviews                     0
last_review                       10052
reviews_per_month                 10052
calculated_host_listings_count        0
availability_365                      0
dtype: int64

Your First Insights: Listings Count#

  • Let us answer some simple questions:
  • How many listings are there in total?
  • How many hosts are offering places?
  • How many unique neighborhoods are listed?
# Count number of listings, unique hosts, and neighborhoods
print("Total listings:", len(df))
print("Unique hosts:", df['host_name'].nunique())
print("Neighborhoods:", df['neighbourhood'].nunique())
Total listings: 48895
Unique hosts: 11453
Neighborhoods: 221
# Explore price statistics: average, min, max, median
print("Average price:", df['price'].mean())
print("Minimum price:", df['price'].min())
print("Maximum price:", df['price'].max())
print("Median price:", df['price'].median())
Average price: 152.7206871868289
Minimum price: 0
Maximum price: 10000
Median price: 106.0

Visualizing Data: Price Distribution#

  • A visualization can quickly show how prices are distributed.
  • You will draw a histogram, which is a bar chart showing how many listings fall into each price range.
  • This helps you see if most prices are low, high, or somewhere in the middle.
# Draw a histogram of price: focus on listings under $500 for visibility
import matplotlib.pyplot as plt
plt.figure(figsize=(8,4))
df[df['price'] < 500]['price'].hist(bins=50, color='skyblue')
plt.title('NYC Airbnb Price Distribution (Below $500)')
plt.xlabel('Price ($ per night)')
plt.ylabel('Number of Listings')
plt.show()
No description has been provided for this image
# Let us see which neighbourhood group has the most listings
counts = df['neighbourhood_group'].value_counts()
print(counts)
neighbourhood_group
Manhattan        21661
Brooklyn         20104
Queens            5666
Bronx             1091
Staten Island      373
Name: count, dtype: int64

Plotting Listings by Area#

  • It is much easier to compare popularity visually.
  • Let us make a bar chart showing the count per neighbourhood group.
# Horizontal bar chart for listings count per neighbourhood group
import seaborn as sns
plt.figure(figsize=(7,4))
sns.barplot(x=counts.values, y=counts.index, palette='pastel')
plt.title('Listings by Neighbourhood Group')
plt.xlabel('Number of Listings')
plt.ylabel('Neighbourhood Group')
plt.show()
No description has been provided for this image
# Find the average price for each neighbourhood group
group_avg = df.groupby('neighbourhood_group')['price'].mean().sort_values(ascending=False)
print(group_avg)
neighbourhood_group
Manhattan        196.875814
Brooklyn         124.383207
Staten Island    114.812332
Queens            99.517649
Bronx             87.496792
Name: price, dtype: float64
# Bar chart: average price per neighbourhood group
plt.figure(figsize=(7,4))
sns.barplot(x=group_avg.values, y=group_avg.index, palette='summer')
plt.title('Average Airbnb Price by Neighbourhood Group')
plt.xlabel('Average Price')
plt.ylabel('Neighbourhood Group')
plt.show()
No description has been provided for this image

Mapping Listings: Price and Location#

  • Visualizing listings on a map can be powerful.
  • We can use latitude and longitude to see where expensive and cheap listings are.
  • Scatter plots can reveal clusters and price hot spots.
# Scatter plot: listings by location, colored by price (under $500 for clarity)
plt.figure(figsize=(8,6))
sample = df[df['price'] < 500].sample(2000, random_state=42)
plt.scatter(sample['longitude'], sample['latitude'], c=sample['price'], cmap='cool', alpha=0.7)
plt.colorbar(label='Price ($)')
plt.xlabel('Longitude')
plt.ylabel('Latitude')
plt.title('NYC Airbnb Listings: Location and Price (Sample under $500)')
plt.show()
No description has been provided for this image
# What are the most common room types?
print(df['room_type'].value_counts())
room_type
Entire home/apt    25409
Private room       22326
Shared room         1160
Name: count, dtype: int64
# Bar plot: room type counts
plt.figure(figsize=(6,3))
sns.countplot(y='room_type', data=df, palette='Set2', order=df['room_type'].value_counts().index)
plt.title('Airbnb Room Types in NYC')
plt.xlabel('Number of Listings')
plt.ylabel('Room Type')
plt.show()
No description has been provided for this image

Time for a Mini Project!#

  • Try these questions on your own code:
    1. Who are the top 5 hosts by number of listings?
    1. What is the median price for each room type?
    1. How does minimum night requirement differ across neighbourhoods?
  • Pause the video and give these a try.
  • When you are ready, continue for more visual fun and insights.
# Correlation: Which features go together?
corr = df[['price', 'minimum_nights', 'number_of_reviews', 'availability_365']].corr()
print(corr)
                      price  minimum_nights  number_of_reviews  \
price              1.000000        0.042799          -0.047954   
minimum_nights     0.042799        1.000000          -0.080116   
number_of_reviews -0.047954       -0.080116           1.000000   
availability_365   0.081829        0.144303           0.172028   

                   availability_365  
price                      0.081829  
minimum_nights             0.144303  
number_of_reviews          0.172028  
availability_365           1.000000  
# Visual heatmap of correlations
plt.figure(figsize=(5,4))
sns.heatmap(corr, annot=True, cmap='YlGnBu')
plt.title('Feature Correlations')
plt.show()
No description has been provided for this image

Share Your Results and Keep Exploring!#

  • You just analyzed real world data, made visualizations, and discovered trends.
  • Try changing filters or columns to answer your own questions.
  • If you enjoyed this lesson, subscribe to the channel and leave a comment with your favorite insight.
  • There is much more to learn in data science. Imagine what you will create next!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.