Mathew K Analytics

Lesson 6 · Data Science Projects

How to Extract YouTube Channel Video Data Using Python and YouTube Data API

We will load a dataset and preview it. What is web scraping? Why would you want to scrape YouTube channel data? What skills will you learn in this lesson?…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Data setup (Quote Scraping)#

We will load a dataset and preview it.

import numpy as np
np.random.seed(42)
import requests
from bs4 import BeautifulSoup
url = 'https://quotes.toscrape.com/'
html = requests.get(url).text
soup = BeautifulSoup(html, 'html.parser')
quotes = [q.get_text(strip=True) for q in soup.select('.quote span.text')][:5]
print(quotes)
['“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', '“It is our choices, Harry, that show what we truly are, far more than our abilities.”', '“There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle.”', '“The person, be it gentleman or lady, who has not pleasure in a good novel, must be intolerably stupid.”', "“Imperfection is beauty, madness is genius and it's better to be absolutely ridiculous than absolutely boring.”"]

YouTube Channel Videos Scraping#

  • What is web scraping?
  • Why would you want to scrape YouTube channel data?
  • What skills will you learn in this lesson?
  • Safety reminders for beginners.

Lesson Plan#

  • Downloading public data about YouTube videos.
  • Using RSS feeds to avoid breaking YouTube rules.
  • Parsing and cleaning text data.
  • Analyzing and visualizing video titles.
  • Discovering trends using counts and word clouds.
# Suppress warnings for a smooth beginner experience
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
 

Data setup#

  • We will scrape video titles from a YouTube channel using RSS.
  • This approach is safe and follows YouTube's guidelines.
  • No login or API key needed.
# Data setup: Scraping YouTube channel video titles using RSS
import feedparser
channel_id = 'UCYO_jab_esuFRV4b17AJtAw'  # 3Blue1Brown, an educational channel
feed_url = f'https://www.youtube.com/feeds/videos.xml?channel_id={channel_id}'
feed = feedparser.parse(feed_url)
titles = [entry.title for entry in feed.entries[:5]]
print(titles)
 
["The most absurd product I've made", 'Why Laplace transforms are so useful', 'The dynamics of e^(πi)', 'But what is a Laplace Transform?', "The Physics of Euler's Formula | Laplace Transform Prelude"]
# Previewing our data
for i, title in enumerate(titles, 1):
    print(f"{i}. {title}")
 
1. The most absurd product I've made
2. Why Laplace transforms are so useful
3. The dynamics of e^(πi)
4. But what is a Laplace Transform?
5. The Physics of Euler's Formula | Laplace Transform Prelude

What is RSS and why do we use it?#

  • RSS stands for Really Simple Syndication.
  • It is a standard way to collect updates from websites.
  • Most YouTube channels provide video lists via RSS.
  • RSS is ideal for beginners because it is reliable and safe.
# Why do we only take five titles?
# This is for speed and simplicity.
# With more experience, you can explore larger samples.
 

Let us explore our scraped text#

  • Scraping gives us raw text from the web.
  • We should always inspect our text for encoding or character issues.
  • The next step is to convert titles into a pandas DataFrame.
# Turn titles list into a DataFrame for easy analysis
import pandas as pd
df = pd.DataFrame({'Title': titles})
print(df.shape)
print(df.head())
 
(5, 1)
                                               Title
0                  The most absurd product I've made
1               Why Laplace transforms are so useful
2                             The dynamics of e^(πi)
3                   But what is a Laplace Transform?
4  The Physics of Euler's Formula | Laplace Trans...
# Check for empty or broken titles
print(df[ df['Title'].isnull() | (df['Title'].str.strip() == "") ])
 
Empty DataFrame
Columns: [Title]
Index: []
# Clean up extra spaces or weird characters
df['Title'] = df['Title'].str.strip()
 

Next, let us explore the video titles#

  • Counting words is a common first analysis step.
  • This can show you the channel's style or focus.
  • We will split all the words from the titles.
  • Then we will see which words appear most often.
# Make all words lowercase and split into words
all_words = ' '.join(df['Title']).lower().split()
print(all_words)
 
['the', 'most', 'absurd', 'product', "i've", 'made', 'why', 'laplace', 'transforms', 'are', 'so', 'useful', 'the', 'dynamics', 'of', 'e^(πi)', 'but', 'what', 'is', 'a', 'laplace', 'transform?', 'the', 'physics', 'of', "euler's", 'formula', '|', 'laplace', 'transform', 'prelude']
# Count the frequency of each word in the titles
from collections import Counter
word_counts = Counter(all_words)
print(word_counts.most_common(5))
 
[('the', 3), ('laplace', 3), ('of', 2), ('most', 1), ('absurd', 1)]

Data visualization: Bar plot of top words#

  • Simple charts can make word counts easy to see.
  • We will draw a bar chart of the five most frequent words.
  • Visualizations give us insights at a glance.
# Show bar plot of frequent words using matplotlib
import matplotlib.pyplot as plt
top_words = word_counts.most_common(5)
words, counts = zip(*top_words)
plt.figure(figsize=(8,4))
plt.bar(words, counts, color='skyblue')
plt.title('Most Common Words in Video Titles')
plt.xlabel('Word')
plt.ylabel('Count')
plt.show()
 
No description has been provided for this image
# Find the longest video title
longest_idx = df['Title'].str.len().idxmax()
print('Longest title:', df.loc[longest_idx, 'Title'])
 
Longest title: The Physics of Euler's Formula | Laplace Transform Prelude
# Interactive: Try searching for a word in video titles
search_word = input("Type a word to search in titles:").strip().lower()
matches = df[ df['Title'].str.lower().str.contains(search_word) ]
print(matches)
 
Empty DataFrame
Columns: [Title]
Index: []

Mini project: Try a different YouTube channel#

  • Change the channel_id to another favorite YouTube channel.
  • Repeat the steps to load, clean, and analyze new video titles.
  • What words are most common for that channel?
  • Share your findings in the comments if you are watching on YouTube!
# Extra: How to find a YouTube channel's ID
print("Go to the YouTube channel. In the URL, look for 'channel/UC...'. Copy the long code after '/channel/'.")
 
Go to the YouTube channel. In the URL, look for 'channel/UC...'. Copy the long code after '/channel/'.
# Let us celebrate! You have just built your first YouTube scraper.
print("Great job! You collected and analyzed public YouTube video data.")
print("Consider subscribing if you want more beginner friendly Python and data lessons!")
 
Great job! You collected and analyzed public YouTube video data.
Consider subscribing if you want more beginner friendly Python and data lessons!
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.