Mathew K Analytics

Lesson 11 · Python data types deep dive

Python Bytes Explained: Binary Data, Encoding & Decoding | Data Types #11

Video eleven of the twelve-part series: bytes, Python's immutable sequence of raw binary data. Construction, indexing, the text-binary boundary, and all…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Python Data Types Deep-Dive, Video 11: Bytes (bytes)#

  • Video eleven of the twelve-part series: bytes, Python's immutable sequence of raw binary data.
  • Construction, indexing, the text-binary boundary, and all forty-two public bytes methods.
  • Let's get into it.

Part 1: What Makes bytes Different#

data = b'hello'
print(type(data))
print(len(data))
print(data[0])
print(list(data))
<class 'bytes'>
5
104
[104, 101, 108, 108, 111]

Part 2: Bytes Literals and Construction#

literal = b'hello world'
empty = b''
also_empty = bytes()
from_ints = bytes([72, 101, 108, 108, 111])
zero_filled = bytes(5)
from_string = bytes('hello', encoding='utf-8')
print(literal, empty, from_ints, zero_filled, from_string)
b'hello world' b'' b'Hello' b'\x00\x00\x00\x00\x00' b'hello'
try:
    bytes('hello')
except TypeError as e:
    print(f'Caught: {e}')
hex_string = '48656c6c6f'
from_hex = bytes.fromhex(hex_string)
print(from_hex)
print(from_hex.hex())
Caught: string argument without an encoding
b'Hello'
48656c6c6f

Part 3: Indexing and Slicing#

data = b'python'
print(data[0])
print(data[0:1])
print(type(data[0]))
print(type(data[0:1]))
print(data[1:4])
print(data[::-1])
112
b'p'
<class 'int'>
<class 'bytes'>
b'yth'
b'nohtyp'

Part 4: decode() and the Text-Binary Boundary#

raw = b'caf\xc3\xa9'
text = raw.decode('utf-8')
print(text)
print(type(text))
print(raw.decode('utf-8', errors='replace'))
try:
    raw.decode('ascii')
except UnicodeDecodeError as e:
    print(f'Caught: {e}')
café
<class 'str'>
café
Caught: 'ascii' codec can't decode byte 0xc3 in position 3: ordinal not in range(128)

Part 5: hex() and fromhex() in Practice#

mac_bytes = bytes([0x00, 0x1A, 0x2B, 0x3C, 0x4D, 0x5E])
print(mac_bytes.hex())
print(mac_bytes.hex(':'))
print(mac_bytes.hex('-', 1))
import hashlib
digest = hashlib.sha256(b'hello world').digest()
print(digest.hex())
001a2b3c4d5e
00:1a:2b:3c:4d:5e
00-1a-2b-3c-4d-5e
b94d27b9934d3e08a52e52d7da7dabfac484efe37a5380ee9088f7ace2efcde9

Part 6: Case-Changing Methods#

text = b'the Quick BROWN fox'
print(text.upper())
print(text.lower())
print(text.capitalize())
print(text.title())
print(text.swapcase())
b'THE QUICK BROWN FOX'
b'the quick brown fox'
b'The quick brown fox'
b'The Quick Brown Fox'
b'THE qUICK brown FOX'

Part 7: Case-Checking Methods#

print(b'HELLO'.isupper())
print(b'Hello'.isupper())
print(b'hello'.islower())
print(b'Hello World'.istitle())
print(b'123'.isupper())
True
False
True
True
False

Part 8: Padding: center, ljust, rjust, zfill#

word = b'hi'
print(word.center(10, b'*'))
print(word.ljust(10, b'.'))
print(word.rjust(10, b'.'))
print(b'42'.zfill(5))
print(b'-42'.zfill(5))
b'****hi****'
b'hi........'
b'........hi'
b'00042'
b'-0042'

Part 9: Stripping: strip, lstrip, rstrip#

messy = b'   padded   '
print(messy.strip())
print(messy.lstrip())
print(messy.rstrip())
print(b'xxxhelloxxx'.strip(b'x'))
b'padded'
b'padded   '
b'   padded'
b'hello'

Part 10: Searching: find, rfind, index, rindex#

data = b'the quick brown fox'
print(data.find(b'quick'))
print(data.find(b'slow'))
print(data.rfind(b'o'))
print(data.index(b'brown'))
try:
    data.index(b'slow')
except ValueError as e:
    print(f'Caught: {e}')
4
-1
17
10
Caught: subsection not found

Part 11: count, startswith, endswith#

data = b'mississippi'
print(data.count(b's'))
print(data.count(b'ss'))
filename = b'archive.tar.gz'
print(filename.startswith(b'archive'))
print(filename.endswith(b'.gz'))
print(filename.endswith((b'.gz', b'.zip', b'.tar')))
4
2
True
True
True

Part 12: Splitting: split, rsplit, splitlines, partition, rpartition#

csv_row = b'apple,banana,cherry,date'
print(csv_row.split(b','))
print(csv_row.split(b',', 2))
print(csv_row.rsplit(b',', 1))
print(b'  spaced   words  '.split())
[b'apple', b'banana', b'cherry', b'date']
[b'apple', b'banana', b'cherry,date']
[b'apple,banana,cherry', b'date']
[b'spaced', b'words']
multiline = b'line one\nline two\r\nline three'
print(multiline.splitlines())
print(multiline.splitlines(keepends=True))
email = b'user@example.com'
print(email.partition(b'@'))
print(email.rpartition(b'@'))
[b'line one', b'line two', b'line three']
[b'line one\n', b'line two\r\n', b'line three']
(b'user', b'@', b'example.com')
(b'user', b'@', b'example.com')

Part 13: join()#

parts = [b'the', b'quick', b'brown', b'fox']
print(b' '.join(parts))
print(b','.join(parts))
print(b''.join(parts))
b'the quick brown fox'
b'the,quick,brown,fox'
b'thequickbrownfox'

Part 14: replace()#

data = b'the cat sat on the mat'
print(data.replace(b'the', b'THE'))
print(data.replace(b'the', b'THE', 1))
b'THE cat sat on THE mat'
b'THE cat sat on the mat'

Part 15: Character Classification Methods#

print(b'Python3'.isalnum())
print(b'Python 3'.isalnum())
print(b'Python'.isalpha())
print(b'123'.isdigit())
print(b'   '.isspace())
print(b'hello'.isascii())
print(bytes([200]).isascii())
True
False
True
True
True
True
False

Part 16: translate() and maketrans()#

table = bytes.maketrans(b'aeiou', b'AEIOU')
print(b'the quick brown fox'.translate(table))
delete_table = bytes.maketrans(b'', b'')
print(b'the quick brown fox'.translate(delete_table, delete=b'aeiou'))
b'thE qUIck brOwn fOx'
b'th qck brwn fx'

Part 17: expandtabs()#

row = b'name\tage\tcity'
print(repr(row.expandtabs()))
print(repr(row.expandtabs(4)))
b'name    age     city'
b'name    age city'

Part 18: removeprefix() and removesuffix()#

filename = b'IMG_final.png'
print(filename.removeprefix(b'IMG_'))
print(filename.removeprefix(b'nope_'))
print(filename.removesuffix(b'.png'))
b'final.png'
b'IMG_final.png'
b'IMG_final'

Part 19: Immutability in Practice#

data = b'hello'
try:
    data[0] = 72
except TypeError as e:
    print(f'Caught: {e}')
new_data = bytes([72]) + data[1:]
print(new_data)
Caught: 'bytes' object does not support item assignment
b'Hello'

Part 20: Common Patterns#

with open('binary_demo.bin', 'wb') as f:
    f.write(b'\x00\x01\x02\x03hello')
with open('binary_demo.bin', 'rb') as f:
    contents = f.read()
print(contents)
print(type(contents))
b'\x00\x01\x02\x03hello'
<class 'bytes'>
import struct
packed = struct.pack('>I', 1024)
print(packed)
print(packed.hex())
unpacked = struct.unpack('>I', packed)
print(unpacked)
b'\x00\x00\x04\x00'
00000400
(1024,)

Wrap-Up: What You Learned#

  • bytes is an immutable sequence of integers 0-255, representing raw binary data.
  • Construction from literals, iterables of ints, zero-fill, and fromhex; indexing returns int, slicing returns bytes.
  • decode() bridges bytes back to str; hex() and fromhex() bridge bytes to readable hex text.
  • Case methods, padding, stripping, searching, splitting, joining, and replacing, all mirroring str's behavior on raw bytes.
  • Five ASCII-only classification methods, translate/maketrans, expandtabs, removeprefix, removesuffix.
  • Immutability, and real patterns: binary file I/O and the struct module.
  • That's all forty-two public bytes methods covered. Next up: bytearray, bytes' mutable sibling.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.