Mathew K Analytics

Python library centre

Comprehensive BeautifulSoup Tutorial for Web Scraping with Python

BeautifulSoup is a Python library for working with HTML and XML files. It helps you extract information from websites by parsing HTML code. Many people use…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Introduction to BeautifulSoup#

  • BeautifulSoup is a Python library for working with HTML and XML files.
  • It helps you extract information from websites by parsing HTML code.
  • Many people use BeautifulSoup for web scraping and data extraction.
  • You can use it to find, navigate, and modify elements in web pages.
  • Real-world uses include downloading quotes, news headlines, and product details from online stores.
  • BeautifulSoup makes handling messy web data much easier.
import warnings; warnings.filterwarnings("ignore")
 
# Install BeautifulSoup if you do not have it
# pip install beautifulsoup4
from bs4 import BeautifulSoup

Core Concepts in BeautifulSoup#

  • BeautifulSoup mainly works by creating a "soup" object from HTML content.
  • The soup object lets you search and navigate the web page structure.
  • Tags represent HTML elements like
    and .
  • You can search for tags using methods like find() and find_all().
  • BeautifulSoup helps you access text and attributes easily.
# Example HTML to use
html_doc = '\n'.join([
    '<html>',
    '  <head><title>My First Web Page</title></head>',
    '  <body>',
    '    <p class="intro">Hello, World!</p>',
    '  </body>',
    '</html>'
])
print('Sample HTML ready!')
# Make a BeautifulSoup object
soup = BeautifulSoup(html_doc, 'html.parser')
print(type(soup))
# Print the prettified HTML
print(soup.prettify())

Beginner Example 1: Get the Title#

  • You can find the title of the HTML easily.
  • The title tag shows the page title in the browser's tab.
  • Use soup.title to access it.
# Get the <title> tag
title_tag = soup.title
print(title_tag)
print(title_tag.string)

Beginner Example 2: Find a Paragraph#

  • You can access the first paragraph with soup.p.
  • This returns the first

    tag in the HTML.

  • Use .string to get just the text.
# Get the first <p> tag
first_p = soup.p
print(first_p)
print(first_p.string)

Beginner Example 3: Get Attribute Value#

  • Tags can have attributes like class, id, or href.
  • Use .get() to extract an attribute value safely.
  • In our example the

    tag has a class.

# Get the 'class' attribute from <p>
p_class = first_p.get('class')
print(p_class)

Intermediate Example 1: Find All Tags#

  • find_all() gets all tags of a certain type.
  • You can use it to make a list of paragraphs.
  • It returns a list even if there is only one paragraph.
# Find all <p> tags
paragraphs = soup.find_all('p')
print(paragraphs)
for para in paragraphs:
    print(para.get_text())

Intermediate Example 2: Searching by Class#

  • Sometimes, you want to find tags by their class name.
  • Pass a dictionary to find_all() with key 'class_' and the class value.
  • This helps filter out exactly what you want.
# Find all paragraphs with class 'intro'
class_paragraphs = soup.find_all('p', class_='intro')
print(class_paragraphs)

Intermediate Example 3: Nested Tags#

  • HTML can have tags inside other tags.
  • You can use .find() on tags too, not just soup.
  • Here is how you get the from the <head>.</li> </ul> </div> </div> </div> <div id="cell-id=5a72c74c" class="cell border-box-sizing code_cell rendered"> <div class="input"> <div class="inner_cell"> <div class="input_area"> <div class=" highlight hl-python"><pre><span></span><span class="c1"># Get the <head> tag and then its <title></span> <span class="n">head_tag</span> <span class="o">=</span> <span class="n">soup</span><span class="o">.</span><span class="n">head</span> <span class="n">head_title</span> <span class="o">=</span> <span class="n">head_tag</span><span class="o">.</span><span class="n">find</span><span class="p">(</span><span class="s1">'title'</span><span class="p">)</span> <span class="nb">print</span><span class="p">(</span><span class="n">head_title</span><span class="o">.</span><span class="n">string</span><span class="p">)</span> </pre></div> </div> </div> </div> </div> <div id="cell-id=043fe5ab" class="cell border-box-sizing text_cell rendered"><div class="inner_cell"> <div class="text_cell_render border-box-sizing rendered_html"> <h1 id="Advanced-Example-1:-Navigating-the-DOM-Tree">Advanced Example 1: Navigating the DOM Tree<a class="anchor-link" href="#Advanced-Example-1:-Navigating-the-DOM-Tree">#</a></h1><ul> <li>BeautifulSoup lets you move up, down, and sideways in the tag tree.</li> <li>You can use .parent to go one level up.</li> <li>.children and .descendants let you loop through elements inside.</li> </ul> </div> </div> </div> <div id="cell-id=582d1a9d" class="cell border-box-sizing code_cell rendered"> <div class="input"> <div class="inner_cell"> <div class="input_area"> <div class=" highlight hl-python"><pre><span></span><span class="c1"># Navigate from <p> up to its parent</span> <span class="nb">print</span><span class="p">(</span><span class="n">first_p</span><span class="o">.</span><span class="n">parent</span><span class="o">.</span><span class="n">name</span><span class="p">)</span> </pre></div> </div> </div> </div> </div> <div id="cell-id=79b5aa8d" class="cell border-box-sizing code_cell rendered"> <div class="input"> <div class="inner_cell"> <div class="input_area"> <div class=" highlight hl-python"><pre><span></span><span class="c1"># Loop through all direct children of <body></span> <span class="n">body_tag</span> <span class="o">=</span> <span class="n">soup</span><span class="o">.</span><span class="n">body</span> <span class="k">for</span> <span class="n">child</span> <span class="ow">in</span> <span class="n">body_tag</span><span class="o">.</span><span class="n">children</span><span class="p">:</span> <span class="nb">print</span><span class="p">(</span><span class="n">child</span><span class="p">)</span> </pre></div> </div> </div> </div> </div> <div id="cell-id=935907b6" class="cell border-box-sizing text_cell rendered"><div class="inner_cell"> <div class="text_cell_render border-box-sizing rendered_html"> <h1 id="Advanced-Example-2:-Searching-with-CSS-Selectors">Advanced Example 2: Searching with CSS Selectors<a class="anchor-link" href="#Advanced-Example-2:-Searching-with-CSS-Selectors">#</a></h1><ul> <li>select() lets you use CSS selectors like in web design.</li> <li>For example, use 'p.intro' to find <p> elements with class intro.</li> <li>This makes it powerful for complex searches.</li> </ul> </div> </div> </div> <div id="cell-id=d86c8ba2" class="cell border-box-sizing code_cell rendered"> <div class="input"> <div class="inner_cell"> <div class="input_area"> <div class=" highlight hl-python"><pre><span></span><span class="c1"># Use select() with a CSS selector</span> <span class="n">intro_ps</span> <span class="o">=</span> <span class="n">soup</span><span class="o">.</span><span class="n">select</span><span class="p">(</span><span class="s1">'p.intro'</span><span class="p">)</span> <span class="k">for</span> <span class="n">tag</span> <span class="ow">in</span> <span class="n">intro_ps</span><span class="p">:</span> <span class="nb">print</span><span class="p">(</span><span class="n">tag</span><span class="o">.</span><span class="n">text</span><span class="p">)</span> </pre></div> </div> </div> </div> </div> <div id="cell-id=90a29e3a" class="cell border-box-sizing text_cell rendered"><div class="inner_cell"> <div class="text_cell_render border-box-sizing rendered_html"> <h1 id="Error-Handling-Example:-Tag-Not-Found">Error Handling Example: Tag Not Found<a class="anchor-link" href="#Error-Handling-Example:-Tag-Not-Found">#</a></h1><ul> <li>Sometimes you ask for a tag that does not exist.</li> <li>BeautifulSoup returns None, not an error.</li> <li>You should always check before using attributes of a found tag.</li> </ul> </div> </div> </div> <div id="cell-id=fe24c70b" class="cell border-box-sizing code_cell rendered"> <div class="input"> <div class="inner_cell"> <div class="input_area"> <div class=" highlight hl-python"><pre><span></span><span class="c1"># Try to find a tag that does not exist</span> <span class="n">nonexistent</span> <span class="o">=</span> <span class="n">soup</span><span class="o">.</span><span class="n">find</span><span class="p">(</span><span class="s1">'h2'</span><span class="p">)</span> <span class="k">if</span> <span class="n">nonexistent</span> <span class="ow">is</span> <span class="kc">None</span><span class="p">:</span> <span class="nb">print</span><span class="p">(</span><span class="s1">'Tag not found.'</span><span class="p">)</span> <span class="k">else</span><span class="p">:</span> <span class="nb">print</span><span class="p">(</span><span class="n">nonexistent</span><span class="p">)</span> </pre></div> </div> </div> </div> </div> <div id="cell-id=b96746eb" class="cell border-box-sizing text_cell rendered"><div class="inner_cell"> <div class="text_cell_render border-box-sizing rendered_html"> <h1 id="Debugging-Example:-Printing-the-Soup">Debugging Example: Printing the Soup<a class="anchor-link" href="#Debugging-Example:-Printing-the-Soup">#</a></h1><ul> <li>If your searches do not return what you expect, print soup.prettify().</li> <li>This can help you see what HTML you are working with.</li> <li>Printing out tags and their attributes is a good way to debug.</li> </ul> </div> </div> </div> <div id="cell-id=3bcd74ad" class="cell border-box-sizing text_cell rendered"><div class="inner_cell"> <div class="text_cell_render border-box-sizing rendered_html"> <h1 id="Best-Practices-for-BeautifulSoup">Best Practices for BeautifulSoup<a class="anchor-link" href="#Best-Practices-for-BeautifulSoup">#</a></h1><ul> <li>Always check that your tag searches were successful before using them.</li> <li>Use prettify() for debugging complex HTML.</li> <li>Prefer select() for complex or flexible searches.</li> <li>Remember to respect website rules when scraping live pages.</li> </ul> </div> </div> </div> <div id="cell-id=6b95a284" class="cell border-box-sizing text_cell rendered"><div class="inner_cell"> <div class="text_cell_render border-box-sizing rendered_html"> <h1 id="Common-Patterns">Common Patterns<a class="anchor-link" href="#Common-Patterns">#</a></h1><ul> <li>Loop over find_all results with for.</li> <li>Use .get() to avoid errors with missing attributes.</li> <li>Use .text or .get_text() to get all text inside a tag.</li> <li>Use conditional checks for missing tags.</li> </ul> </div> </div> </div> <div id="cell-id=6fdb81fb" class="cell border-box-sizing code_cell rendered"> <div class="input"> <div class="inner_cell"> <div class="input_area"> <div class=" highlight hl-python"><pre><span></span><span class="c1"># Mini Project: Extract Quotes from HTML</span> <span class="n">quotes_html</span> <span class="o">=</span> <span class="s2">"""</span> <span class="s2"><html></span> <span class="s2"><body></span> <span class="s2"> <div class="quote"></span> <span class="s2"> <span class="text">"The best way to get started is to quit talking and begin doing."</span></span> <span class="s2"> <span class="author">Walt Disney</span></span> <span class="s2"> </div></span> <span class="s2"> <div class="quote"></span> <span class="s2"> <span class="text">"Life is what happens when you are busy making other plans."</span></span> <span class="s2"> <span class="author">John Lennon</span></span> <span class="s2"> </div></span> <span class="s2"></body></span> <span class="s2"></html></span> <span class="s2">"""</span> <span class="n">soup2</span> <span class="o">=</span> <span class="n">BeautifulSoup</span><span class="p">(</span><span class="n">quotes_html</span><span class="p">,</span> <span class="s1">'html.parser'</span><span class="p">)</span> <span class="n">quote_divs</span> <span class="o">=</span> <span class="n">soup2</span><span class="o">.</span><span class="n">find_all</span><span class="p">(</span><span class="s1">'div'</span><span class="p">,</span> <span class="n">class_</span><span class="o">=</span><span class="s1">'quote'</span><span class="p">)</span> <span class="k">for</span> <span class="n">div</span> <span class="ow">in</span> <span class="n">quote_divs</span><span class="p">:</span> <span class="n">text</span> <span class="o">=</span> <span class="n">div</span><span class="o">.</span><span class="n">find</span><span class="p">(</span><span class="s1">'span'</span><span class="p">,</span> <span class="n">class_</span><span class="o">=</span><span class="s1">'text'</span><span class="p">)</span><span class="o">.</span><span class="n">get_text</span><span class="p">()</span> <span class="n">author</span> <span class="o">=</span> <span class="n">div</span><span class="o">.</span><span class="n">find</span><span class="p">(</span><span class="s1">'span'</span><span class="p">,</span> <span class="n">class_</span><span class="o">=</span><span class="s1">'author'</span><span class="p">)</span><span class="o">.</span><span class="n">get_text</span><span class="p">()</span> <span class="nb">print</span><span class="p">(</span><span class="sa">f</span><span class="s1">'</span><span class="si">{</span><span class="n">text</span><span class="si">}</span><span class="s1"> - </span><span class="si">{</span><span class="n">author</span><span class="si">}</span><span class="s1">'</span><span class="p">)</span> </pre></div> </div> </div> </div> </div> <div id="cell-id=456366b1" class="cell border-box-sizing text_cell rendered"><div class="inner_cell"> <div class="text_cell_render border-box-sizing rendered_html"> <h1 id="Want-More-Python-Tutorials?">Want More Python Tutorials?<a class="anchor-link" href="#Want-More-Python-Tutorials?">#</a></h1><ul> <li>Subscribe to our YouTube channel for more beginner-friendly Python lessons!</li> <li>Like and share this video if you found it helpful.</li> </ul> </div> </div> </div>

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.