Mathew K Analytics

Lesson 32 · Data visualisation in python

Geospatial Analysis Case Study: Mapping and Visualizing Real-World Location Data

In this lesson, you will learn how to work with Python, explore real-world geospatial data, and visualize it using friendly tools. You do not need any…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Welcome to Python Geospatial Mapping!#

In this lesson, you will learn how to work with Python, explore real-world geospatial data, and visualize it using friendly tools.

You do not need any programming experience, just curiosity!

We will build up step by step, from basic variables to mapping airline passenger flows.

import warnings; warnings.filterwarnings('ignore')

# This line tells Python to suppress warning messages.
# Warnings often show up when using data science tools.
# Keeping them hidden makes things less distracting for new users.

What is Geospatial Data?#

Geospatial data means information that has locations or coordinates attached to it.

Think of airline routes, city locations, or weather measurements. It helps us answer 'where' questions using data!

# Let us start with a simple variable.
city = "London"
print(city)

# Python stores text inside quote marks.
London
# You can also store numbers.
passengers = 250
print(passengers)
250
# Lists help us group many items. Let us list some major cities.
cities = ["London", "Paris", "New York", "Tokyo"]
print(cities)
['London', 'Paris', 'New York', 'Tokyo']
# Access an item from a list using its position (starting from 0).
print(cities[1])
Paris

Real-World Data#

Now let us try working with real airline passenger data.

We will use a CSV file (a table of data) with months and the number of airline passengers. Then we will learn how to explore and plot it!

# Data setup
import pandas as pd
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/airline-passengers.csv"
df = pd.read_csv(url)
print("Shape:", df.shape)
print(df.head())
Shape: (144, 2)
     Month  Passengers
0  1949-01         112
1  1949-02         118
2  1949-03         132
3  1949-04         129
4  1949-05         121
# Plot airline passenger counts to see patterns over time.
import matplotlib.pyplot as plt
df.plot(x="Month", y="Passengers", legend=False, title="Airline Passengers Over Time")
plt.ylabel("Passengers")
plt.show()
No description has been provided for this image
# Let us print the month with the highest passenger number.
max_row = df[df["Passengers"] == df["Passengers"].max()]
print("Busiest month:", max_row["Month"].values[0])
Busiest month: 1960-07
# Create a list of all months.
months = df["Month"].tolist()
print(months[:5])  # Show the first five months
['1949-01', '1949-02', '1949-03', '1949-04', '1949-05']
# Calculate how many months had more than 400 passengers.
over_400 = df[df["Passengers"] > 400]
print("Months with over 400 passengers:", over_400.shape[0])
Months with over 400 passengers: 28
# Safe access: What if a month is missing in the data?
search_month = input("Enter a month (like 1950-08): ")
if search_month in months:
    print("Yes,", search_month, "is in the airline data.")
else:
    print("Sorry,", search_month, "is not in our list.")
    
Yes, 1952-08 is in the airline data.
# Add a new column to categorize each month as High or Low traffic.
df["TrafficLevel"] = ["High" if p > 400 else "Low" for p in df["Passengers"]]
print(df[["Month", "Passengers", "TrafficLevel"]].head())
     Month  Passengers TrafficLevel
0  1949-01         112          Low
1  1949-02         118          Low
2  1949-03         132          Low
3  1949-04         129          Low
4  1949-05         121          Low
# Remove a column if it is not needed anymore.
df = df.drop(columns=["TrafficLevel"])
print(df.head())
     Month  Passengers
0  1949-01         112
1  1949-02         118
2  1949-03         132
3  1949-04         129
4  1949-05         121
# Find and print basic statistics about the passenger numbers.
print("Average:", df["Passengers"].mean())
print("Minimum:", df["Passengers"].min())
print("Maximum:", df["Passengers"].max())
Average: 280.2986111111111
Minimum: 104
Maximum: 622
# Loop through each month and print when passengers were above 450.
for idx, row in df.iterrows():
    if row['Passengers'] > 450:
        print(row['Month'], 'had over 450 passengers.')
        
1957-07 had over 450 passengers.
1957-08 had over 450 passengers.
1958-07 had over 450 passengers.
1958-08 had over 450 passengers.
1959-06 had over 450 passengers.
1959-07 had over 450 passengers.
1959-08 had over 450 passengers.
1959-09 had over 450 passengers.
1960-04 had over 450 passengers.
1960-05 had over 450 passengers.
1960-06 had over 450 passengers.
1960-07 had over 450 passengers.
1960-08 had over 450 passengers.
1960-09 had over 450 passengers.
1960-10 had over 450 passengers.
# Create a list of months that had less than 350 passengers.
low_months = [row['Month'] for idx, row in df.iterrows() if row['Passengers'] < 350]
print(low_months)
['1949-01', '1949-02', '1949-03', '1949-04', '1949-05', '1949-06', '1949-07', '1949-08', '1949-09', '1949-10', '1949-11', '1949-12', '1950-01', '1950-02', '1950-03', '1950-04', '1950-05', '1950-06', '1950-07', '1950-08', '1950-09', '1950-10', '1950-11', '1950-12', '1951-01', '1951-02', '1951-03', '1951-04', '1951-05', '1951-06', '1951-07', '1951-08', '1951-09', '1951-10', '1951-11', '1951-12', '1952-01', '1952-02', '1952-03', '1952-04', '1952-05', '1952-06', '1952-07', '1952-08', '1952-09', '1952-10', '1952-11', '1952-12', '1953-01', '1953-02', '1953-03', '1953-04', '1953-05', '1953-06', '1953-07', '1953-08', '1953-09', '1953-10', '1953-11', '1953-12', '1954-01', '1954-02', '1954-03', '1954-04', '1954-05', '1954-06', '1954-07', '1954-08', '1954-09', '1954-10', '1954-11', '1954-12', '1955-01', '1955-02', '1955-03', '1955-04', '1955-05', '1955-06', '1955-08', '1955-09', '1955-10', '1955-11', '1955-12', '1956-01', '1956-02', '1956-03', '1956-04', '1956-05', '1956-10', '1956-11', '1956-12', '1957-01', '1957-02', '1957-04', '1957-10', '1957-11', '1957-12', '1958-01', '1958-02', '1958-04', '1958-11', '1958-12', '1959-02']
# Sort the data so busiest months appear first.
sorted_df = df.sort_values("Passengers", ascending=False)
print(sorted_df.head())
       Month  Passengers
138  1960-07         622
139  1960-08         606
127  1959-08         559
126  1959-07         548
137  1960-06         535
# Add together all passenger counts from 1955 onward.
from datetime import datetime
recent_mask = df["Month"].apply(lambda x: int(x.split('-')[0]) >= 1955)
recent_sum = df[recent_mask]["Passengers"].sum()
print("Total passengers from 1955 onward:", recent_sum)
Total passengers from 1955 onward: 27194
# Mini-project Part 1: Rough map of monthly traffic ups and downs.
import numpy as np
plt.figure(figsize=(12,5))
plt.plot(df["Month"], df["Passengers"], label="Monthly Traffic", color="steelblue")
plt.scatter(df["Month"][::12], df["Passengers"][::12], color="firebrick", label="Year starts")
plt.xlabel("Month")
plt.ylabel("Passengers")
plt.xticks(df["Month"][::12], rotation=45)
plt.legend()
plt.title("Airline Traffic Seasonality (Year Starts Highlighted)")
plt.tight_layout()
plt.show()
No description has been provided for this image
# Mini-project Part 2: Map busiest months using world cities and simple coordinates.
world_cities = {
    "London": (51.5074, -0.1278),
    "Paris": (48.8566, 2.3522),
    "New York": (40.7128, -74.0060),
    "Tokyo": (35.6895, 139.6917),
}
busiest = sorted_df.iloc[0:4]["Month"].tolist()
print("Busiest months:", busiest)

# In a real map, we would plot world city coordinates!
for city, coords in world_cities.items():
    print(f"{city}: Latitude {coords[0]}, Longitude {coords[1]}")
    
Busiest months: ['1960-07', '1960-08', '1959-08', '1959-07']
London: Latitude 51.5074, Longitude -0.1278
Paris: Latitude 48.8566, Longitude 2.3522
New York: Latitude 40.7128, Longitude -74.006
Tokyo: Latitude 35.6895, Longitude 139.6917
# Best practices: Always check for missing or strange numbers.
print("Any missing data?", df.isna().any().any())
print("Example data types:", df.dtypes.to_dict())
Any missing data? False
Example data types: {'Month': dtype('O'), 'Passengers': dtype('int64')}
# Troubleshooting: Code sometimes gives an error.
try:
    print(df["NonexistentColumn"].head())
except KeyError:
    print('Oops! That column does not exist in this data.')
    
Oops! That column does not exist in this data.
# Tips: Use the .info() and .describe() functions for a quick overview.
print(df.info())
print(df.describe())
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 144 entries, 0 to 143
Data columns (total 2 columns):
 #   Column      Non-Null Count  Dtype 
---  ------      --------------  ----- 
 0   Month       144 non-null    object
 1   Passengers  144 non-null    int64 
dtypes: int64(1), object(1)
memory usage: 2.4+ KB
None
       Passengers
count  144.000000
mean   280.298611
std    119.966317
min    104.000000
25%    180.000000
50%    265.500000
75%    360.500000
max    622.000000
# Challenge: Find the first month with fewer than 150 passengers.
for idx, row in df.iterrows():
    if row['Passengers'] < 150:
        print('First month with under 150 passengers:', row['Month'])
        break
    
First month with under 150 passengers: 1949-01

Recap: What Have You Learned?#

You loaded real geospatial data, explored lists and tables, made charts, filtered and mapped real numbers to real locations. You built mini projects and learned to troubleshoot along the way!

Keep building, and you will soon be a data explorer!

Thank You and Next Steps!#

If you enjoyed this lesson, like and subscribe for more beginner Python projects on this channel. Try out the bonus challenge: pick a new city or dataset and make your own plots!

See you in the next video!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.