Mathew K Analytics

Lesson 7 · Python standard library deep dive

Python re Explained: Regular Expressions Made Simple | Standard Library #7

Video seven of the twenty-five-part series: re, Python's pattern-matching engine for text. Matching, searching, groups, substitution, compiling, and flags.…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Python Standard Library Deep-Dive, Video 7: re (Regular Expressions)#

  • Video seven of the twenty-five-part series: re, Python's pattern-matching engine for text.
  • Matching, searching, groups, substitution, compiling, and flags.
  • Let's get into it.

Part 1: What re Offers#

import re
text = 'The year 2026 was eventful'
match = re.search(r'\d+', text)
print(match)
print(match.group())
<re.Match object; span=(9, 13), match='2026'>
2026

Part 2: match(), search(), fullmatch()#

text = 'hello world'
print(re.match(r'hello', text))
print(re.match(r'world', text))
print(re.search(r'world', text))
print(re.fullmatch(r'hello world', text))
print(re.fullmatch(r'hello', text))
<re.Match object; span=(0, 5), match='hello'>
None
<re.Match object; span=(6, 11), match='world'>
<re.Match object; span=(0, 11), match='hello world'>
None

Part 3: findall() and finditer()#

text = 'Order 12 has 5 items, order 45 has 2 items'
numbers = re.findall(r'\d+', text)
print(numbers)
for m in re.finditer(r'\d+', text):
    print(m.group(), m.start(), m.end())
['12', '5', '45', '2']
12 6 8
5 13 14
45 28 30
2 35 36

Part 4: Character Classes and Quantifiers#

text = 'Call 555-1234 or email me@test.com, id_007 works too'
print(re.findall(r'\d+', text))
print(re.findall(r'\w+', text))
print(re.findall(r'\s+', text))
print(re.findall(r'[aeiou]', 'hello world'))
print(re.findall(r'[^aeiou\s]', 'hello world'))
['555', '1234', '007']
['Call', '555', '1234', 'or', 'email', 'me', 'test', 'com', 'id_007', 'works', 'too']
[' ', ' ', ' ', ' ', ' ', ' ', ' ']
['e', 'o', 'o']
['h', 'l', 'l', 'w', 'r', 'l', 'd']
print(re.findall(r'ab?c', 'ac abc abbc'))
print(re.findall(r'ab*c', 'ac abc abbbc'))
print(re.findall(r'ab+c', 'ac abc abbbc'))
print(re.findall(r'\d{3}', '12 123 1234 12345'))
print(re.findall(r'\d{2,4}', '12 123 1234 12345'))
['ac', 'abc']
['ac', 'abc', 'abbbc']
['abc', 'abbbc']
['123', '123', '123']
['12', '123', '1234', '1234']

Part 5: Anchors and Boundaries#

print(re.findall(r'^\d+', '123 abc 456'))
print(re.findall(r'\d+$', '123 abc 456'))
print(re.findall(r'\bcat\b', 'cat category cats cat'))
print(re.findall(r'cat', 'cat category cats cat'))
['123']
['456']
['cat', 'cat']
['cat', 'cat', 'cat', 'cat']

Part 6: Groups and Capturing#

text = 'Date: 2026-03-15'
match = re.search(r'(\d{4})-(\d{2})-(\d{2})', text)
print(match.group())
print(match.group(0))
print(match.group(1))
print(match.groups())
2026-03-15
2026-03-15
2026
('2026', '03', '15')
match = re.search(r'(?P<year>\d{4})-(?P<month>\d{2})-(?P<day>\d{2})', text)
print(match.group('year'))
print(match.group('month'))
print(match.groupdict())
2026
03
{'year': '2026', 'month': '03', 'day': '15'}

Part 7: sub() and subn()#

text = 'Contact: 555-1234 or 555-5678'
redacted = re.sub(r'\d{3}-\d{4}', 'XXX-XXXX', text)
print(redacted)
result, count = re.subn(r'\d{3}-\d{4}', 'XXX-XXXX', text)
print(result, count)
Contact: XXX-XXXX or XXX-XXXX
Contact: XXX-XXXX or XXX-XXXX 2
text = 'John Smith, Jane Doe'
swapped = re.sub(r'(\w+) (\w+)', r'\2 \1', text)
print(swapped)
def uppercase_match(m):
    return m.group().upper()
shouted = re.sub(r'\w+', uppercase_match, text)
print(shouted)
Smith John, Doe Jane
JOHN SMITH, JANE DOE

Part 8: split() with Regex#

text = 'apple, banana;  cherry,date'
print(re.split(r'[,;]\s*', text))
print(text.split(','))
text2 = 'one1two22three333four'
print(re.split(r'\d+', text2))
['apple', 'banana', 'cherry', 'date']
['apple', ' banana;  cherry', 'date']
['one', 'two', 'three', 'four']

Part 9: Compiling Patterns for Reuse#

phone_pattern = re.compile(r'\d{3}-\d{3}-\d{4}')
text1 = 'Call 555-123-4567'
text2 = 'Or reach 555-987-6543 instead'
print(phone_pattern.search(text1).group())
print(phone_pattern.search(text2).group())
print(phone_pattern.findall('555-111-2222 and 555-333-4444'))
555-123-4567
555-987-6543
['555-111-2222', '555-333-4444']

Part 10: Flags: IGNORECASE, MULTILINE, DOTALL#

print(re.findall(r'python', 'Python PYTHON python', re.IGNORECASE))
text = 'line one\nline two\nline three'
print(re.findall(r'^line', text))
print(re.findall(r'^line', text, re.MULTILINE))
print(re.findall(r'one.two', 'one\ntwo'))
print(re.findall(r'one.two', 'one\ntwo', re.DOTALL))
['Python', 'PYTHON', 'python']
['line']
['line', 'line', 'line']
[]
['one\ntwo']

Part 11: Common Patterns#

email_pattern = re.compile(r'^[\w.+-]+@[\w-]+(?:\.[\w-]+)+$')
candidates = ['user@example.com', 'not-an-email', 'a.b+c@sub.domain.co']
for candidate in candidates:
    is_valid = bool(email_pattern.match(candidate))
    print(candidate, is_valid)
user@example.com True
not-an-email False
a.b+c@sub.domain.co True
report = 'Revenue: $45,231.50 up from $38,900.00 last quarter, a 16.28% increase'
numbers = re.findall(r'\d[\d,]*\.?\d*', report)
print(numbers)
cleaned = [float(n.replace(',', '')) for n in numbers]
print(cleaned)
['45,231.50', '38,900.00', '16.28']
[45231.5, 38900.0, 16.28]

Wrap-Up: What You Learned#

  • match, search, and fullmatch, each anchoring differently.
  • findall for a plain list of matches, finditer for match objects with position data.
  • Character classes, quantifiers, anchors, and word boundaries.
  • Capturing groups, both positional and named, and how to pull them apart.
  • sub and subn for substitution, including function-based and backreference replacements.
  • re.split for pattern-based splitting, and re.compile for reusable, faster patterns.
  • Flags: IGNORECASE, MULTILINE, DOTALL.
  • Two real patterns: email validation and extracting numbers from mixed text.
  • That wraps up re. Next up: json and csv, for structured data serialization.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.