Lesson 6 · Data Science Projects
How to Extract YouTube Channel Video Data Using Python and YouTube Data API
We will load a dataset and preview it. What is web scraping? Why would you want to scrape YouTube channel data? What skills will you learn in this lesson?…
- CourseData Science Projects
- Lesson6 of 33
- Video14 min
- FormatJupyter notebook · 16 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbData setup (Quote Scraping)#
We will load a dataset and preview it.
import numpy as np
np.random.seed(42)
import requests
from bs4 import BeautifulSoup
url = 'https://quotes.toscrape.com/'
html = requests.get(url).text
soup = BeautifulSoup(html, 'html.parser')
quotes = [q.get_text(strip=True) for q in soup.select('.quote span.text')][:5]
print(quotes)
YouTube Channel Videos Scraping#
- What is web scraping?
- Why would you want to scrape YouTube channel data?
- What skills will you learn in this lesson?
- Safety reminders for beginners.
Lesson Plan#
- Downloading public data about YouTube videos.
- Using RSS feeds to avoid breaking YouTube rules.
- Parsing and cleaning text data.
- Analyzing and visualizing video titles.
- Discovering trends using counts and word clouds.
# Suppress warnings for a smooth beginner experience
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
Data setup#
- We will scrape video titles from a YouTube channel using RSS.
- This approach is safe and follows YouTube's guidelines.
- No login or API key needed.
# Data setup: Scraping YouTube channel video titles using RSS
import feedparser
channel_id = 'UCYO_jab_esuFRV4b17AJtAw' # 3Blue1Brown, an educational channel
feed_url = f'https://www.youtube.com/feeds/videos.xml?channel_id={channel_id}'
feed = feedparser.parse(feed_url)
titles = [entry.title for entry in feed.entries[:5]]
print(titles)
# Previewing our data
for i, title in enumerate(titles, 1):
print(f"{i}. {title}")
What is RSS and why do we use it?#
- RSS stands for Really Simple Syndication.
- It is a standard way to collect updates from websites.
- Most YouTube channels provide video lists via RSS.
- RSS is ideal for beginners because it is reliable and safe.
# Why do we only take five titles?
# This is for speed and simplicity.
# With more experience, you can explore larger samples.
Let us explore our scraped text#
- Scraping gives us raw text from the web.
- We should always inspect our text for encoding or character issues.
- The next step is to convert titles into a pandas DataFrame.
# Turn titles list into a DataFrame for easy analysis
import pandas as pd
df = pd.DataFrame({'Title': titles})
print(df.shape)
print(df.head())
# Check for empty or broken titles
print(df[ df['Title'].isnull() | (df['Title'].str.strip() == "") ])
# Clean up extra spaces or weird characters
df['Title'] = df['Title'].str.strip()
Next, let us explore the video titles#
- Counting words is a common first analysis step.
- This can show you the channel's style or focus.
- We will split all the words from the titles.
- Then we will see which words appear most often.
# Make all words lowercase and split into words
all_words = ' '.join(df['Title']).lower().split()
print(all_words)
# Count the frequency of each word in the titles
from collections import Counter
word_counts = Counter(all_words)
print(word_counts.most_common(5))
Data visualization: Bar plot of top words#
- Simple charts can make word counts easy to see.
- We will draw a bar chart of the five most frequent words.
- Visualizations give us insights at a glance.
# Show bar plot of frequent words using matplotlib
import matplotlib.pyplot as plt
top_words = word_counts.most_common(5)
words, counts = zip(*top_words)
plt.figure(figsize=(8,4))
plt.bar(words, counts, color='skyblue')
plt.title('Most Common Words in Video Titles')
plt.xlabel('Word')
plt.ylabel('Count')
plt.show()
# Find the longest video title
longest_idx = df['Title'].str.len().idxmax()
print('Longest title:', df.loc[longest_idx, 'Title'])
# Interactive: Try searching for a word in video titles
search_word = input("Type a word to search in titles:").strip().lower()
matches = df[ df['Title'].str.lower().str.contains(search_word) ]
print(matches)
Mini project: Try a different YouTube channel#
- Change the channel_id to another favorite YouTube channel.
- Repeat the steps to load, clean, and analyze new video titles.
- What words are most common for that channel?
- Share your findings in the comments if you are watching on YouTube!
# Extra: How to find a YouTube channel's ID
print("Go to the YouTube channel. In the URL, look for 'channel/UC...'. Copy the long code after '/channel/'.")
# Let us celebrate! You have just built your first YouTube scraper.
print("Great job! You collected and analyzed public YouTube video data.")
print("Consider subscribing if you want more beginner friendly Python and data lessons!")
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



