Mathew K Analytics

Lesson 69 · Data Science Projects

Building a COVID-19 Data Tracker with API Integration and Time-Series Visualization in Python

Welcome! Let us explore real-world data mining using Python. We will track COVID-19 global cases, learn basic analysis, and visualize changes over time. You…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 12: Data Mining Introduction with COVID-19 Data#

Welcome! Let us explore real-world data mining using Python.

We will track COVID-19 global cases, learn basic analysis, and visualize changes over time.

You will also get first-hand practice loading datasets like Titanic and Supermarket Sales.

By the end, you will be able to analyze public data and see trends with simple charts.

# Always suppress warnings for a beginner-friendly experience
import warnings
warnings.filterwarnings('ignore')

# Let us import our basic analysis tools
import pandas as pd
import matplotlib.pyplot as plt
import numpy as np
np.random.seed(42)

COVID-19 Global Cases: Data Setup#

We will use daily worldwide COVID-19 case data published by Johns Hopkins CSSE.

This data helps governments and researchers track the virus over time and by country.

# Data setup (COVID-19 Global Cases)
import pandas as pd
url = 'https://raw.githubusercontent.com/CSSEGISandData/COVID-19/master/csse_covid_19_data/csse_covid_19_time_series/time_series_covid19_confirmed_global.csv'
df_covid = pd.read_csv(url)
print(df_covid.shape)
print(df_covid.head(3))
(289, 1147)
  Province/State Country/Region       Lat       Long  1/22/20  1/23/20  \
0            NaN    Afghanistan  33.93911  67.709953        0        0   
1            NaN        Albania  41.15330  20.168300        0        0   
2            NaN        Algeria  28.03390   1.659600        0        0   

   1/24/20  1/25/20  1/26/20  1/27/20  ...  2/28/23  3/1/23  3/2/23  3/3/23  \
0        0        0        0        0  ...   209322  209340  209358  209362   
1        0        0        0        0  ...   334391  334408  334408  334427   
2        0        0        0        0  ...   271441  271448  271463  271469   

   3/4/23  3/5/23  3/6/23  3/7/23  3/8/23  3/9/23  
0  209369  209390  209406  209436  209451  209451  
1  334427  334427  334427  334427  334443  334457  
2  271469  271477  271477  271490  271494  271496  

[3 rows x 1147 columns]

Basic Exploration: Columns and Missing Data#

Let us check which columns are present and see if we have missing values.

Understanding the structure is the first step in data cleaning.

# Look at the first few column names and check for missing values
print('Columns:', df_covid.columns[:10].tolist())
print('Any missing:', df_covid.isnull().any().any())
 
 
Columns: ['Province/State', 'Country/Region', 'Lat', 'Long', '1/22/20', '1/23/20', '1/24/20', '1/25/20', '1/26/20', '1/27/20']
Any missing: True

Simple Data Cleaning: Working with Country Names#

Country names in public data are not always consistent.

Let us inspect unique country entries next.

# See all unique country names
print(sorted(df_covid['Country/Region'].unique())[:10])
print('Total countries:', df_covid['Country/Region'].nunique())
['Afghanistan', 'Albania', 'Algeria', 'Andorra', 'Angola', 'Antarctica', 'Antigua and Barbuda', 'Argentina', 'Armenia', 'Australia']
Total countries: 201
# Convert date columns from text to useable format (pandas DateTime)
date_cols = df_covid.columns[4:]
df_long = df_covid.melt(id_vars=['Province/State', 'Country/Region', 'Lat', 'Long'], var_name='Date', value_name='Confirmed')
df_long['Date'] = pd.to_datetime(df_long['Date'])

Visualizing COVID-19: Total Cases for a Country#

Let us plot how total COVID-19 cases have changed for one country.

This helps spot waves and important events.

# Plot line chart of confirmed cases in Italy
country = 'Italy'
country_data = df_long[df_long['Country/Region'] == country].groupby('Date')['Confirmed'].sum()
plt.figure(figsize=(10,5))
plt.plot(country_data.index, country_data.values)
plt.title(f'COVID-19 Confirmed Cases in {country}')
plt.xlabel('Date')
plt.ylabel('Confirmed Cases')
plt.grid(True)
plt.show()
No description has been provided for this image
# Practice: Let the user pick a country to plot COVID-19 trends interactively
chosen_country = input('Enter a country to visualize: ')
country_data2 = df_long[df_long['Country/Region'] == chosen_country].groupby('Date')['Confirmed'].sum()
plt.figure(figsize=(10,5))
plt.plot(country_data2.index, country_data2.values, color='orange')
plt.title(f'COVID-19 Confirmed Cases in {chosen_country}')
plt.xlabel('Date')
plt.ylabel('Confirmed Cases')
plt.grid(True)
plt.show()
No description has been provided for this image

Mini-Exploration: Monthly New Cases#

Viewing new cases by month helps spot waves easily.

We will compute monthly new cases for Italy.

# Visualize monthly new COVID-19 cases for Italy
monthly = country_data.diff().resample('M').sum()
plt.figure(figsize=(10,5))
plt.bar(monthly.index.strftime('%Y-%m'), monthly.values)
plt.title('Italy Monthly New COVID-19 Cases')
plt.xlabel('Month')
plt.ylabel('New Cases')
plt.xticks(rotation=45)
plt.grid(True, axis='y')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Compare two countries side by side for 2020
country1 = 'Italy'
country2 = 'India'
data1 = df_long[(df_long['Country/Region']==country1) & (df_long['Date']<'2021-01-01')].groupby('Date')['Confirmed'].sum()
data2 = df_long[(df_long['Country/Region']==country2) & (df_long['Date']<'2021-01-01')].groupby('Date')['Confirmed'].sum()
plt.figure(figsize=(10,5))
plt.plot(data1.index, data1.values, label=country1, color='red')
plt.plot(data2.index, data2.values, label=country2, color='blue')
plt.title('COVID-19: Italy vs India (2020)')
plt.xlabel('Date')
plt.ylabel('Confirmed Cases')
plt.legend()
plt.show()
No description has been provided for this image
# Export cleaned tidy COVID data to CSV for future use
df_long.to_csv('tidy_covid_data.csv', index=False)
print('Saved as tidy_covid_data.csv')
Saved as tidy_covid_data.csv

Introducing Another Dataset: Titanic Passengers#

The Titanic dataset is famous for teaching data cleaning and survival predictions.

Let us load a sample and compare its structure to COVID-19 data.

# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df_titanic = pd.read_csv(url)
print(df_titanic.shape)
print(df_titanic.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  
# Practice: Find number of missing entries in Age column
missing_age = df_titanic['Age'].isnull().sum()
print('Missing Age values:', missing_age)
Missing Age values: 177

Recap: What We Learned#

We practiced loading real-world datasets and transforming COVID-19 daily data for time series charts.

You learned to clean, reshape, visualize, and compare trends across different nations.

Missing data is everywhere, so we always check for it, as we did with the Titanic dataset.

Thanks and Next Steps!#

Want more project guides like this? Practice with our notebook and subscribe for updates.

In the next part, we will try grouping and comparing data from retail and Airbnb.

See you next time!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.