Mathew K Analytics

Lesson 3 · Python for Data Analysts

NumPy Zero to Hero: The Complete Beginner-to-Advanced Course

One notebook, start to finish: from your first array to broadcasting, vectorized math, and basic linear algebra. No prior NumPy experience needed. Let's get…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

NumPy Zero to Hero: The Complete Beginner-to-Advanced Course#

  • One notebook, start to finish: from your first array to broadcasting, vectorized math, and basic linear algebra.
  • No prior NumPy experience needed. Let's get straight into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • If NumPy isn't installed yet, open a terminal in VS Code and run: pip install numpy

Part 1: Arrays#

import numpy as np
print(np.__version__)
2.2.6

Creating Arrays#

a = np.array([1, 2, 3, 4, 5])
print(a)
print(type(a))
print(a.dtype)
[1 2 3 4 5]
<class 'numpy.ndarray'>
int64
zeros = np.zeros(5)
ones = np.ones((2, 3))
full = np.full((2, 2), 7)
identity = np.eye(3)
print(zeros)
print(ones)
print(full)
print(identity)
[0. 0. 0. 0. 0.]
[[1. 1. 1.]
 [1. 1. 1.]]
[[7 7]
 [7 7]]
[[1. 0. 0.]
 [0. 1. 0.]
 [0. 0. 1.]]
sequence = np.arange(0, 10, 2)
spaced = np.linspace(0, 1, 5)
print(sequence)
print(spaced)
[0 2 4 6 8]
[0.   0.25 0.5  0.75 1.  ]

Array Shape and Dimensions#

grid = np.array([[1, 2, 3], [4, 5, 6]])
print(grid.shape)
print(grid.ndim)
print(grid.size)
(2, 3)
2
6
flat = np.arange(12)
reshaped = flat.reshape(3, 4)
print(flat)
print(reshaped)
print(reshaped.reshape(-1))
[ 0  1  2  3  4  5  6  7  8  9 10 11]
[[ 0  1  2  3]
 [ 4  5  6  7]
 [ 8  9 10 11]]
[ 0  1  2  3  4  5  6  7  8  9 10 11]

Part 2: Indexing, Slicing, and Math#

Indexing and Slicing#

arr = np.array([10, 20, 30, 40, 50])
print(arr[0])
print(arr[-1])
print(arr[1:4])
print(arr[::2])
10
50
[20 30 40]
[10 30 50]
grid = np.array([[1, 2, 3], [4, 5, 6], [7, 8, 9]])
print(grid[1, 2])
print(grid[0])
print(grid[:, 1])
print(grid[0:2, 0:2])
6
[1 2 3]
[2 5 8]
[[1 2]
 [4 5]]

Boolean Masking and Fancy Indexing#

scores = np.array([55, 90, 72, 40, 88, 65])
mask = scores >= 70
print(mask)
print(scores[mask])
print(scores[scores >= 70])
[False  True  True False  True False]
[90 72 88]
[90 72 88]
values = np.array([10, 20, 30, 40, 50])
print(values[[0, 2, 4]])
values[values > 25] = 0
print(values)
[10 30 50]
[10 20  0  0  0]

Vectorized Math#

a = np.array([1, 2, 3])
b = np.array([10, 20, 30])
print(a + b)
print(a * b)
print(b / a)
print(a ** 2)
print(np.sqrt(b))
[11 22 33]
[10 40 90]
[10. 10. 10.]
[1 4 9]
[3.16227766 4.47213595 5.47722558]

Aggregations#

data = np.array([[1, 2, 3], [4, 5, 6]])
print(data.sum())
print(data.mean())
print(data.max())
print(data.min())
print(data.std())
21
3.5
6
1
1.707825127659933
print(data.sum(axis=0))
print(data.sum(axis=1))
[5 7 9]
[ 6 15]

Part 3: Broadcasting and Combining Arrays#

Broadcasting#

prices = np.array([10.0, 20.0, 30.0])
print(prices * 1.1)
[11. 22. 33.]
matrix = np.array([[1, 2, 3], [4, 5, 6], [7, 8, 9]])
row_addition = np.array([10, 20, 30])
print(matrix + row_addition)
[[11 22 33]
 [14 25 36]
 [17 28 39]]

Sorting and Searching#

unsorted = np.array([5, 2, 9, 1, 7])
print(np.sort(unsorted))
print(unsorted)
print(np.argsort(unsorted))
[1 2 5 7 9]
[5 2 9 1 7]
[3 1 0 4 2]
values = np.array([3, 8, 1, 9, 4])
positions = np.where(values > 4)
print(positions)
print(values[positions])
(array([1, 3]),)
[8 9]

Stacking and Splitting#

a = np.array([1, 2, 3])
b = np.array([4, 5, 6])
print(np.vstack([a, b]))
print(np.hstack([a, b]))
[[1 2 3]
 [4 5 6]]
[1 2 3 4 5 6]
combined = np.arange(12)
parts = np.split(combined, 3)
print(parts)
[array([0, 1, 2, 3]), array([4, 5, 6, 7]), array([ 8,  9, 10, 11])]

Part 4: Random Numbers, Linear Algebra, and Performance#

The Random Module#

rng = np.random.default_rng(seed=42)
print(rng.integers(1, 100, size=5))
print(rng.random(3))
print(rng.normal(loc=0, scale=1, size=4))
[ 9 77 65 44 43]
[0.69736803 0.09417735 0.97562235]
[ 0.1278404  -0.31624259 -0.01680116 -0.85304393]

Basic Linear Algebra#

v1 = np.array([1, 2, 3])
v2 = np.array([4, 5, 6])
print(np.dot(v1, v2))

m1 = np.array([[1, 2], [3, 4]])
m2 = np.array([[5, 6], [7, 8]])
print(m1 @ m2)
32
[[19 22]
 [43 50]]
square = np.array([[4, 2], [1, 3]])
print(np.linalg.det(square))
print(np.linalg.inv(square))
10.000000000000002
[[ 0.3 -0.2]
 [-0.1  0.4]]

Performance: NumPy vs. Plain Python#

import time

python_list = list(range(1_000_000))
numpy_array = np.arange(1_000_000)

start = time.time()
doubled_list = [x * 2 for x in python_list]
list_time = time.time() - start

start = time.time()
doubled_array = numpy_array * 2
array_time = time.time() - start

print(f'Python list comprehension: {list_time:.5f}s')
print(f'NumPy vectorized: {array_time:.5f}s')
Python list comprehension: 0.06207s
NumPy vectorized: 0.00000s

Saving and Loading Arrays#

data_to_save = np.arange(20).reshape(4, 5)
np.save('my_array.npy', data_to_save)
loaded = np.load('my_array.npy')
print(np.array_equal(data_to_save, loaded))
True

Capstone Project: Class Grades Analyzer#

rng = np.random.default_rng(seed=7)
students = ['Amir', 'Bianca', 'Carlos', 'Deepa', 'Elin', 'Farid']
assignments = ['Quiz 1', 'Quiz 2', 'Project', 'Final']
scores = rng.integers(55, 101, size=(len(students), len(assignments)))
print(scores)
[[98 83 86 96]
 [81 90 93 65]
 [57 68 68 95]
 [96 55 77 92]
 [61 91 60 76]
 [92 68 70 67]]
student_averages = scores.mean(axis=1)
for name, avg in zip(students, student_averages):
    print(f'{name}: {avg:.1f}')
Amir: 90.8
Bianca: 82.2
Carlos: 72.0
Deepa: 80.0
Elin: 72.0
Farid: 74.2
assignment_averages = scores.mean(axis=0)
for name, avg in zip(assignments, assignment_averages):
    print(f'{name}: {avg:.1f}')
Quiz 1: 80.8
Quiz 2: 75.8
Project: 75.7
Final: 81.8
curve = 5
curved_scores = np.clip(scores + curve, 0, 100)
print(curved_scores)
[[100  88  91 100]
 [ 86  95  98  70]
 [ 62  73  73 100]
 [100  60  82  97]
 [ 66  96  65  81]
 [ 97  73  75  72]]
final_averages = curved_scores.mean(axis=1)
top_student_index = np.argmax(final_averages)
print(f'Top student: {students[top_student_index]} ({final_averages[top_student_index]:.1f})')

at_risk_mask = final_averages < 70
print('Students below 70:', np.array(students)[at_risk_mask])
Top student: Amir (94.8)
Students below 70: []
def letter_grade(avg):
    if avg >= 90:
        return 'A'
    elif avg >= 80:
        return 'B'
    elif avg >= 70:
        return 'C'
    else:
        return 'D'

letter_grades = np.array([letter_grade(a) for a in final_averages])
ranking_order = np.argsort(final_averages)[::-1]

print('Final Ranking:')
for rank, idx in enumerate(ranking_order, start=1):
    print(f'{rank}. {students[idx]}: {final_averages[idx]:.1f} ({letter_grades[idx]})')
Final Ranking:
1. Amir: 94.8 (A)
2. Bianca: 87.2 (B)
3. Deepa: 84.8 (B)
4. Farid: 79.2 (C)
5. Elin: 77.0 (C)
6. Carlos: 77.0 (C)

Wrap-Up: What You Learned#

  • Arrays: creating them, checking their shape and dimensions, and reshaping them.
  • Indexing and slicing, in both one and two dimensions, plus boolean masking and fancy indexing.
  • Vectorized math and aggregations, including the axis parameter.
  • Broadcasting, sorting, searching with where, and stacking or splitting arrays.
  • The random module, basic linear algebra with dot/matmul, a real performance comparison against plain Python, and saving/loading arrays.
  • A capstone gradebook analyzer combining nearly all of it into one project.
  • You went from creating your first array to running real broadcasting, linear algebra, and a full sorted analysis. If you want the next build to land in your feed automatically, subscribing is the move see you in the next one.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.