Lesson 69 · Data Science Projects
Building a COVID-19 Data Tracker with API Integration and Time-Series Visualization in Python
Welcome! Let us explore real-world data mining using Python. We will track COVID-19 global cases, learn basic analysis, and visualize changes over time. You…
- CourseData Science Projects
- Lesson69 of 33
- Video20 min
- FormatJupyter notebook · 12 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 12: Data Mining Introduction with COVID-19 Data#
Welcome! Let us explore real-world data mining using Python.
We will track COVID-19 global cases, learn basic analysis, and visualize changes over time.
You will also get first-hand practice loading datasets like Titanic and Supermarket Sales.
By the end, you will be able to analyze public data and see trends with simple charts.
# Always suppress warnings for a beginner-friendly experience
import warnings
warnings.filterwarnings('ignore')
# Let us import our basic analysis tools
import pandas as pd
import matplotlib.pyplot as plt
import numpy as np
np.random.seed(42)
COVID-19 Global Cases: Data Setup#
We will use daily worldwide COVID-19 case data published by Johns Hopkins CSSE.
This data helps governments and researchers track the virus over time and by country.
# Data setup (COVID-19 Global Cases)
import pandas as pd
url = 'https://raw.githubusercontent.com/CSSEGISandData/COVID-19/master/csse_covid_19_data/csse_covid_19_time_series/time_series_covid19_confirmed_global.csv'
df_covid = pd.read_csv(url)
print(df_covid.shape)
print(df_covid.head(3))
Basic Exploration: Columns and Missing Data#
Let us check which columns are present and see if we have missing values.
Understanding the structure is the first step in data cleaning.
# Look at the first few column names and check for missing values
print('Columns:', df_covid.columns[:10].tolist())
print('Any missing:', df_covid.isnull().any().any())
Simple Data Cleaning: Working with Country Names#
Country names in public data are not always consistent.
Let us inspect unique country entries next.
# See all unique country names
print(sorted(df_covid['Country/Region'].unique())[:10])
print('Total countries:', df_covid['Country/Region'].nunique())
# Convert date columns from text to useable format (pandas DateTime)
date_cols = df_covid.columns[4:]
df_long = df_covid.melt(id_vars=['Province/State', 'Country/Region', 'Lat', 'Long'], var_name='Date', value_name='Confirmed')
df_long['Date'] = pd.to_datetime(df_long['Date'])
Visualizing COVID-19: Total Cases for a Country#
Let us plot how total COVID-19 cases have changed for one country.
This helps spot waves and important events.
# Plot line chart of confirmed cases in Italy
country = 'Italy'
country_data = df_long[df_long['Country/Region'] == country].groupby('Date')['Confirmed'].sum()
plt.figure(figsize=(10,5))
plt.plot(country_data.index, country_data.values)
plt.title(f'COVID-19 Confirmed Cases in {country}')
plt.xlabel('Date')
plt.ylabel('Confirmed Cases')
plt.grid(True)
plt.show()
# Practice: Let the user pick a country to plot COVID-19 trends interactively
chosen_country = input('Enter a country to visualize: ')
country_data2 = df_long[df_long['Country/Region'] == chosen_country].groupby('Date')['Confirmed'].sum()
plt.figure(figsize=(10,5))
plt.plot(country_data2.index, country_data2.values, color='orange')
plt.title(f'COVID-19 Confirmed Cases in {chosen_country}')
plt.xlabel('Date')
plt.ylabel('Confirmed Cases')
plt.grid(True)
plt.show()
Mini-Exploration: Monthly New Cases#
Viewing new cases by month helps spot waves easily.
We will compute monthly new cases for Italy.
# Visualize monthly new COVID-19 cases for Italy
monthly = country_data.diff().resample('M').sum()
plt.figure(figsize=(10,5))
plt.bar(monthly.index.strftime('%Y-%m'), monthly.values)
plt.title('Italy Monthly New COVID-19 Cases')
plt.xlabel('Month')
plt.ylabel('New Cases')
plt.xticks(rotation=45)
plt.grid(True, axis='y')
plt.tight_layout()
plt.show()
# Compare two countries side by side for 2020
country1 = 'Italy'
country2 = 'India'
data1 = df_long[(df_long['Country/Region']==country1) & (df_long['Date']<'2021-01-01')].groupby('Date')['Confirmed'].sum()
data2 = df_long[(df_long['Country/Region']==country2) & (df_long['Date']<'2021-01-01')].groupby('Date')['Confirmed'].sum()
plt.figure(figsize=(10,5))
plt.plot(data1.index, data1.values, label=country1, color='red')
plt.plot(data2.index, data2.values, label=country2, color='blue')
plt.title('COVID-19: Italy vs India (2020)')
plt.xlabel('Date')
plt.ylabel('Confirmed Cases')
plt.legend()
plt.show()
# Export cleaned tidy COVID data to CSV for future use
df_long.to_csv('tidy_covid_data.csv', index=False)
print('Saved as tidy_covid_data.csv')
Introducing Another Dataset: Titanic Passengers#
The Titanic dataset is famous for teaching data cleaning and survival predictions.
Let us load a sample and compare its structure to COVID-19 data.
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df_titanic = pd.read_csv(url)
print(df_titanic.shape)
print(df_titanic.head(3))
# Practice: Find number of missing entries in Age column
missing_age = df_titanic['Age'].isnull().sum()
print('Missing Age values:', missing_age)
Recap: What We Learned#
We practiced loading real-world datasets and transforming COVID-19 daily data for time series charts.
You learned to clean, reshape, visualize, and compare trends across different nations.
Missing data is everywhere, so we always check for it, as we did with the Titanic dataset.
Thanks and Next Steps!#
Want more project guides like this? Practice with our notebook and subscribe for updates.
In the next part, we will try grouping and comparing data from retail and Airbnb.
See you next time!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



