Lesson 6 · Python for Data Science
6 - APIs and Web Scraping in Python
Have you ever wanted to get data from websites or apps automatically? APIs and web scraping make this possible in Python. In this lesson, we will learn the…
- CoursePython for Data Science
- Lesson6 of 38
- Video11 min
- FormatJupyter notebook · 20 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Introduction: APIs and Web Scraping in Python#
Have you ever wanted to get data from websites or apps automatically?
APIs and web scraping make this possible in Python.
In this lesson, we will learn the basics and explore simple, practical examples.
By the end, you will know how to pull real data into your Python programs!
What is an API?#
API stands for Application Programming Interface.
It lets different programs talk to each other safely and share data.
Many sites, like weather or news, provide APIs for public use.
# Let us import the requests library, which helps us make API calls.
import requests
# Let us check what happens when we access a real API.
response = requests.get("https://api.github.com")
print(response.status_code)
# Let us see the data we got back.
print(response.text[:250]) # Show the first 250 characters
APIs often return JSON#
Most APIs send data in JSON format.
JSON looks like Python dictionaries, and stores data in pairs of labels and values.
Luckily, requests can help us handle JSON easily.
# Let us get structured JSON data from the API.
data = response.json()
print(type(data))
# Let us explore a value inside the API response.
print(data["current_user_url"])
What is Web Scraping?#
What if there is no API available?
We can collect data from websites by reading the HTML directly.
This technique is called web scraping.
Python makes this possible with the BeautifulSoup library.
# Let us import BeautifulSoup for our scraping tasks.
from bs4 import BeautifulSoup
# Let us download a simple web page to scrape.
url = "https://quotes.toscrape.com/"
page = requests.get(url)
soup = BeautifulSoup(page.text, "html.parser")
# Let us find all the quote text on the page.
quotes = soup.find_all("span", class_="text")
for q in quotes:
print(q.text)
# Let us grab authors of these quotes.
authors = soup.find_all("small", class_="author")
for i in range(len(quotes)):
print(f"{quotes[i].text} - {authors[i].text}")
# Sometimes, websites ask for your name!
name = input("What is your first name? ")
print(f"Welcome, {name}! Ready to scrape more data?")
# Be gentle! Let us pause between requests to avoid bothering servers.
import time
time.sleep(2)
print("Waited 2 seconds before moving on.")
Handling errors the safe way#
Not every web request will succeed.
Servers can be down, or you could have typed the address wrong.
We will now learn how to check if things are working, and handle any problems.
# Safe requests with error handling.
bad_url = "https://thisdoesnotexist.toscrape.com/"
try:
bad_response = requests.get(bad_url)
bad_response.raise_for_status()
except requests.exceptions.RequestException as e:
print("Something went wrong:", e)
# Mini-project: Collect the first page of quotes and authors and save them to a file.
all_quotes = []
for i in range(len(quotes)):
all_quotes.append(f"{quotes[i].text} - {authors[i].text}")
with open("quotes_list.txt", "w", encoding="utf-8") as file:
for line in all_quotes:
file.write(line + "\n")
print("Saved quotes to quotes_list.txt!")
# Mini-project, part two: Count how many quotes on the page contain the word "life".
life_quotes = [q.text for q in quotes if "life" in q.text.lower()]
print(f"Found {len(life_quotes)} quotes about life!")
# Sorting: Let us find the longest quote.
longest_quote = max(quotes, key=lambda q: len(q.text))
print("The longest quote is:")
print(longest_quote.text)
# Troubleshooting tip: What if nothing prints?
if not quotes:
print("No quotes found! Try checking the website address or your internet connection.")
else:
print(f"We found {len(quotes)} quotes on the page.")
# Extra tip: Adding headers to look like a real browser.
headers = {"User-Agent": "Mozilla/5.0"}
custom_page = requests.get(url, headers=headers)
print(f"Custom request status: {custom_page.status_code}")
# Challenge: Try scraping the tags for each quote.
tags = soup.find_all("div", class_="tags")
for i in range(len(tags)):
tag_texts = [t.text for t in tags[i].find_all("a", class_="tag")]
print(f"{quotes[i].text} | Tags: {', '.join(tag_texts)}")
Lesson Recap#
- You learned how to use Python to request and process data from APIs.
- You discovered web scraping for data without an API.
- You used BeautifulSoup to find and print quotes and authors.
- You practiced saving, sorting, and handling errors with web data.
Now you have real tools for diving into information from all over the internet!
Thank you! What is next?#
Keep exploring different APIs and websites.
Comment below what data you want to scrape next!
If you enjoyed this lesson, like, subscribe, and share with a friend.
Happy coding!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



