Mathew K Analytics

Lesson 4 · Data Science Projects

Building a News Data Scraping and Analysis Project Using Python

Learn how to extract headlines from news websites. See why web scraping is useful for data mining projects. Get hands on with real news data using Python.…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

News Scraping and Analysis#

  • Learn how to extract headlines from news websites.
  • See why web scraping is useful for data mining projects.
  • Get hands on with real news data using Python.
  • No coding experience needed. We will explain each step.
  • By the end, you will try your own simple news mining challenge.
  • If you enjoy this lesson, please like and subscribe to our YouTube channel!

What is Web Scraping?#

  • Web scraping means using code to collect content or data from websites.
  • It helps us gather news articles, product prices, or social media posts quickly.
  • Companies, researchers, and journalists use scraping for new insights.
  • We must scrape only public data that is legal and safe.
  • Today, we will scrape live BBC News headlines and explore them together.
# First, suppress warnings to keep our notebook tidy
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
# Data setup
import requests
from bs4 import BeautifulSoup
url = 'https://www.bbc.com/news'
html = requests.get(url).text
soup = BeautifulSoup(html, 'html.parser')
headlines = [h.get_text(strip=True) for h in soup.select('h2') if h.get_text(strip=True)]
print(headlines[:10])
['Scores killed as floods sweep several Asian nations', 'At least 193 dead in Sri Lanka flood, many more missing', "Venezuela condemns Trump's threat to close country's airspace", "Nigerian villagers 'too scared to speak' after hundreds of schoolchildren kidnapped", "Four killed in shooting at child's birthday party in California", "Three days of mourning begin after Hong Kong's deadliest fire in decades", 'How the deadly Hong Kong fire spread in minutes', "Forgotten photos reveal women who powered India's freedom struggle", "Invincible no more: Why India's Test cricket reputation lies in ruins", 'More than 70,000 killed in Gaza since Israel offensive began, Hamas-run health ministry says']
# How many headlines did we scrape?
print("Total headlines scraped:", len(headlines))
Total headlines scraped: 54
# Let us preview the first five headlines
for i, h in enumerate(headlines[:5], 1):
    print(f"{i}. {h}")
1. Scores killed as floods sweep several Asian nations
2. At least 193 dead in Sri Lanka flood, many more missing
3. Venezuela condemns Trump's threat to close country's airspace
4. Nigerian villagers 'too scared to speak' after hundreds of schoolchildren kidnapped
5. Four killed in shooting at child's birthday party in California

Why scrape the news?#

  • Scraping allows real time analysis of trends and breaking stories.
  • You can monitor keywords or track news about a topic.
  • Data mining on headlines helps spot big events faster.
  • Scraping also supports projects like news aggregators and dashboards.
# Let us put our headlines into a pandas DataFrame
import pandas as pd
df = pd.DataFrame({'headline': headlines})
print(df.shape)
df.head()
(54, 1)
headline
0 Scores killed as floods sweep several Asian na...
1 At least 193 dead in Sri Lanka flood, many mor...
2 Venezuela condemns Trump's threat to close cou...
3 Nigerian villagers 'too scared to speak' after...
4 Four killed in shooting at child's birthday pa...
# Remove duplicate headlines if there are any
df = df.drop_duplicates('headline').reset_index(drop=True)
print("Unique headlines:", len(df))
Unique headlines: 41
# Check for any empty or very short headlines
short_headlines = df[df['headline'].str.len() < 8]
print(short_headlines.shape)
short_headlines.head()
(1, 1)
headline
34 Sport
# Drop all headlines less than 8 characters
df = df[df['headline'].str.len() >= 8].reset_index(drop=True)
print("Cleaned headline count:", len(df))
Cleaned headline count: 40

What comes next?#

  • Now we will analyze our cleaned news data.
  • Let us discover what topics are trending today.
  • We will look for popular words and create a simple visualization.
  • The next steps will use pandas and seaborn, two beginner friendly Python libraries.
# Get word frequency from all headlines
from collections import Counter
all_words = ' '.join(df['headline']).lower().split()
common_words = Counter(all_words).most_common(10)
for word, count in common_words:
    print(f"{word}: {count}")
in: 16
to: 12
of: 8
the: 6
and: 5
hong: 4
-: 4
from: 4
killed: 3
more: 3
# Plot the word frequency as a bar chart
import matplotlib.pyplot as plt
words, counts = zip(*common_words)
plt.figure(figsize=(8,4))
plt.bar(words, counts, color='skyblue')
plt.title('Top 10 Words in BBC News Headlines')
plt.ylabel('Count')
plt.xlabel('Word')
plt.show()
No description has been provided for this image
# Are there any headlines with certain keywords?
key = input("Type a word to search for in the headlines: ").strip().lower()
matches = df[df['headline'].str.lower().str.contains(key)]
print(f'Headlines with "{key}":', len(matches))
if not matches.empty:
    print(matches.head(5))
Headlines with "climate": 0
# Let us count headlines by length for more insights
df['len'] = df['headline'].str.len()
print(df['len'].describe())
count    40.000000
mean     56.300000
std      18.929288
min       9.000000
25%      50.750000
50%      60.000000
75%      67.000000
max      92.000000
Name: len, dtype: float64
# Plot a histogram of headline lengths
import seaborn as sns
plt.figure(figsize=(7,3))
sns.histplot(df['len'], bins=12, color='purple')
plt.title('Distribution of Headline Lengths')
plt.xlabel('Number of characters')
plt.ylabel('Count')
plt.show()
No description has been provided for this image
# Find headlines with numbers (like years or stats)
import re
num_headlines = df[df['headline'].str.contains(r'\d')]
print("Total with numbers:", len(num_headlines))
num_headlines.head(5)
Total with numbers: 3
headline len
1 At least 193 dead in Sri Lanka flood, many mor... 55
9 More than 70,000 killed in Gaza since Israel o... 92
19 The five things that set the 2025 Atlantic hur... 65

Mini challenge: Explore on your own!#

  • Try changing the search word to find headlines about sports or politics.
  • Tweak the minimum headline length and see how your counts change.
  • What is the most surprising trend in today's word frequencies?
  • Try running the code again in a few days for new results.
  • Share your favorite discovery in the YouTube comment section!
# Optional: Save collected headlines for future analysis
df[['headline']].to_csv('bbc_headlines.csv', index=False)
print('Saved to bbc_headlines.csv!')
Saved to bbc_headlines.csv!

Congratulations!#

  • You completed a real web scraping and news mining workflow.
  • You learned about scraping, cleaning, analyzing, and visualizing data.
  • We hope you feel confident to try more scraping and data mining tasks.
  • If you liked this lesson, please support us by subscribing to our YouTube channel.
  • See you in the next beginner friendly data mining tutorial!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.