Lesson 4 · Data Science Projects
Building a News Data Scraping and Analysis Project Using Python
Learn how to extract headlines from news websites. See why web scraping is useful for data mining projects. Get hands on with real news data using Python.…
- CourseData Science Projects
- Lesson4 of 33
- Video19 min
- FormatJupyter notebook · 15 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbNews Scraping and Analysis#
- Learn how to extract headlines from news websites.
- See why web scraping is useful for data mining projects.
- Get hands on with real news data using Python.
- No coding experience needed. We will explain each step.
- By the end, you will try your own simple news mining challenge.
- If you enjoy this lesson, please like and subscribe to our YouTube channel!
What is Web Scraping?#
- Web scraping means using code to collect content or data from websites.
- It helps us gather news articles, product prices, or social media posts quickly.
- Companies, researchers, and journalists use scraping for new insights.
- We must scrape only public data that is legal and safe.
- Today, we will scrape live BBC News headlines and explore them together.
# First, suppress warnings to keep our notebook tidy
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
# Data setup
import requests
from bs4 import BeautifulSoup
url = 'https://www.bbc.com/news'
html = requests.get(url).text
soup = BeautifulSoup(html, 'html.parser')
headlines = [h.get_text(strip=True) for h in soup.select('h2') if h.get_text(strip=True)]
print(headlines[:10])
# How many headlines did we scrape?
print("Total headlines scraped:", len(headlines))
# Let us preview the first five headlines
for i, h in enumerate(headlines[:5], 1):
print(f"{i}. {h}")
Why scrape the news?#
- Scraping allows real time analysis of trends and breaking stories.
- You can monitor keywords or track news about a topic.
- Data mining on headlines helps spot big events faster.
- Scraping also supports projects like news aggregators and dashboards.
# Let us put our headlines into a pandas DataFrame
import pandas as pd
df = pd.DataFrame({'headline': headlines})
print(df.shape)
df.head()
# Remove duplicate headlines if there are any
df = df.drop_duplicates('headline').reset_index(drop=True)
print("Unique headlines:", len(df))
# Check for any empty or very short headlines
short_headlines = df[df['headline'].str.len() < 8]
print(short_headlines.shape)
short_headlines.head()
# Drop all headlines less than 8 characters
df = df[df['headline'].str.len() >= 8].reset_index(drop=True)
print("Cleaned headline count:", len(df))
What comes next?#
- Now we will analyze our cleaned news data.
- Let us discover what topics are trending today.
- We will look for popular words and create a simple visualization.
- The next steps will use pandas and seaborn, two beginner friendly Python libraries.
# Get word frequency from all headlines
from collections import Counter
all_words = ' '.join(df['headline']).lower().split()
common_words = Counter(all_words).most_common(10)
for word, count in common_words:
print(f"{word}: {count}")
# Plot the word frequency as a bar chart
import matplotlib.pyplot as plt
words, counts = zip(*common_words)
plt.figure(figsize=(8,4))
plt.bar(words, counts, color='skyblue')
plt.title('Top 10 Words in BBC News Headlines')
plt.ylabel('Count')
plt.xlabel('Word')
plt.show()
# Are there any headlines with certain keywords?
key = input("Type a word to search for in the headlines: ").strip().lower()
matches = df[df['headline'].str.lower().str.contains(key)]
print(f'Headlines with "{key}":', len(matches))
if not matches.empty:
print(matches.head(5))
# Let us count headlines by length for more insights
df['len'] = df['headline'].str.len()
print(df['len'].describe())
# Plot a histogram of headline lengths
import seaborn as sns
plt.figure(figsize=(7,3))
sns.histplot(df['len'], bins=12, color='purple')
plt.title('Distribution of Headline Lengths')
plt.xlabel('Number of characters')
plt.ylabel('Count')
plt.show()
# Find headlines with numbers (like years or stats)
import re
num_headlines = df[df['headline'].str.contains(r'\d')]
print("Total with numbers:", len(num_headlines))
num_headlines.head(5)
Mini challenge: Explore on your own!#
- Try changing the search word to find headlines about sports or politics.
- Tweak the minimum headline length and see how your counts change.
- What is the most surprising trend in today's word frequencies?
- Try running the code again in a few days for new results.
- Share your favorite discovery in the YouTube comment section!
# Optional: Save collected headlines for future analysis
df[['headline']].to_csv('bbc_headlines.csv', index=False)
print('Saved to bbc_headlines.csv!')
Congratulations!#
- You completed a real web scraping and news mining workflow.
- You learned about scraping, cleaning, analyzing, and visualizing data.
- We hope you feel confident to try more scraping and data mining tasks.
- If you liked this lesson, please support us by subscribing to our YouTube channel.
- See you in the next beginner friendly data mining tutorial!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



