Lesson 3 · Python for Data Analysts
NumPy Zero to Hero: The Complete Beginner-to-Advanced Course
One notebook, start to finish: from your first array to broadcasting, vectorized math, and basic linear algebra. No prior NumPy experience needed. Let's get…
- CoursePython for Data Analysts
- Lesson3 of 12
- Video33 min
- FormatJupyter notebook · 30 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbNumPy Zero to Hero: The Complete Beginner-to-Advanced Course#
- One notebook, start to finish: from your first array to broadcasting, vectorized math, and basic linear algebra.
- No prior NumPy experience needed. Let's get straight into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- If NumPy isn't installed yet, open a terminal in VS Code and run: pip install numpy
Part 1: Arrays#
import numpy as np
print(np.__version__)
Creating Arrays#
a = np.array([1, 2, 3, 4, 5])
print(a)
print(type(a))
print(a.dtype)
zeros = np.zeros(5)
ones = np.ones((2, 3))
full = np.full((2, 2), 7)
identity = np.eye(3)
print(zeros)
print(ones)
print(full)
print(identity)
sequence = np.arange(0, 10, 2)
spaced = np.linspace(0, 1, 5)
print(sequence)
print(spaced)
Array Shape and Dimensions#
grid = np.array([[1, 2, 3], [4, 5, 6]])
print(grid.shape)
print(grid.ndim)
print(grid.size)
flat = np.arange(12)
reshaped = flat.reshape(3, 4)
print(flat)
print(reshaped)
print(reshaped.reshape(-1))
Part 2: Indexing, Slicing, and Math#
Indexing and Slicing#
arr = np.array([10, 20, 30, 40, 50])
print(arr[0])
print(arr[-1])
print(arr[1:4])
print(arr[::2])
grid = np.array([[1, 2, 3], [4, 5, 6], [7, 8, 9]])
print(grid[1, 2])
print(grid[0])
print(grid[:, 1])
print(grid[0:2, 0:2])
Boolean Masking and Fancy Indexing#
scores = np.array([55, 90, 72, 40, 88, 65])
mask = scores >= 70
print(mask)
print(scores[mask])
print(scores[scores >= 70])
values = np.array([10, 20, 30, 40, 50])
print(values[[0, 2, 4]])
values[values > 25] = 0
print(values)
Vectorized Math#
a = np.array([1, 2, 3])
b = np.array([10, 20, 30])
print(a + b)
print(a * b)
print(b / a)
print(a ** 2)
print(np.sqrt(b))
Aggregations#
data = np.array([[1, 2, 3], [4, 5, 6]])
print(data.sum())
print(data.mean())
print(data.max())
print(data.min())
print(data.std())
print(data.sum(axis=0))
print(data.sum(axis=1))
Part 3: Broadcasting and Combining Arrays#
Broadcasting#
prices = np.array([10.0, 20.0, 30.0])
print(prices * 1.1)
matrix = np.array([[1, 2, 3], [4, 5, 6], [7, 8, 9]])
row_addition = np.array([10, 20, 30])
print(matrix + row_addition)
Sorting and Searching#
unsorted = np.array([5, 2, 9, 1, 7])
print(np.sort(unsorted))
print(unsorted)
print(np.argsort(unsorted))
values = np.array([3, 8, 1, 9, 4])
positions = np.where(values > 4)
print(positions)
print(values[positions])
Stacking and Splitting#
a = np.array([1, 2, 3])
b = np.array([4, 5, 6])
print(np.vstack([a, b]))
print(np.hstack([a, b]))
combined = np.arange(12)
parts = np.split(combined, 3)
print(parts)
Part 4: Random Numbers, Linear Algebra, and Performance#
The Random Module#
rng = np.random.default_rng(seed=42)
print(rng.integers(1, 100, size=5))
print(rng.random(3))
print(rng.normal(loc=0, scale=1, size=4))
Basic Linear Algebra#
v1 = np.array([1, 2, 3])
v2 = np.array([4, 5, 6])
print(np.dot(v1, v2))
m1 = np.array([[1, 2], [3, 4]])
m2 = np.array([[5, 6], [7, 8]])
print(m1 @ m2)
square = np.array([[4, 2], [1, 3]])
print(np.linalg.det(square))
print(np.linalg.inv(square))
Performance: NumPy vs. Plain Python#
import time
python_list = list(range(1_000_000))
numpy_array = np.arange(1_000_000)
start = time.time()
doubled_list = [x * 2 for x in python_list]
list_time = time.time() - start
start = time.time()
doubled_array = numpy_array * 2
array_time = time.time() - start
print(f'Python list comprehension: {list_time:.5f}s')
print(f'NumPy vectorized: {array_time:.5f}s')
Saving and Loading Arrays#
data_to_save = np.arange(20).reshape(4, 5)
np.save('my_array.npy', data_to_save)
loaded = np.load('my_array.npy')
print(np.array_equal(data_to_save, loaded))
Capstone Project: Class Grades Analyzer#
rng = np.random.default_rng(seed=7)
students = ['Amir', 'Bianca', 'Carlos', 'Deepa', 'Elin', 'Farid']
assignments = ['Quiz 1', 'Quiz 2', 'Project', 'Final']
scores = rng.integers(55, 101, size=(len(students), len(assignments)))
print(scores)
student_averages = scores.mean(axis=1)
for name, avg in zip(students, student_averages):
print(f'{name}: {avg:.1f}')
assignment_averages = scores.mean(axis=0)
for name, avg in zip(assignments, assignment_averages):
print(f'{name}: {avg:.1f}')
curve = 5
curved_scores = np.clip(scores + curve, 0, 100)
print(curved_scores)
final_averages = curved_scores.mean(axis=1)
top_student_index = np.argmax(final_averages)
print(f'Top student: {students[top_student_index]} ({final_averages[top_student_index]:.1f})')
at_risk_mask = final_averages < 70
print('Students below 70:', np.array(students)[at_risk_mask])
def letter_grade(avg):
if avg >= 90:
return 'A'
elif avg >= 80:
return 'B'
elif avg >= 70:
return 'C'
else:
return 'D'
letter_grades = np.array([letter_grade(a) for a in final_averages])
ranking_order = np.argsort(final_averages)[::-1]
print('Final Ranking:')
for rank, idx in enumerate(ranking_order, start=1):
print(f'{rank}. {students[idx]}: {final_averages[idx]:.1f} ({letter_grades[idx]})')
Wrap-Up: What You Learned#
- Arrays: creating them, checking their shape and dimensions, and reshaping them.
- Indexing and slicing, in both one and two dimensions, plus boolean masking and fancy indexing.
- Vectorized math and aggregations, including the axis parameter.
- Broadcasting, sorting, searching with where, and stacking or splitting arrays.
- The random module, basic linear algebra with dot/matmul, a real performance comparison against plain Python, and saving/loading arrays.
- A capstone gradebook analyzer combining nearly all of it into one project.
- You went from creating your first array to running real broadcasting, linear algebra, and a full sorted analysis. If you want the next build to land in your feed automatically, subscribing is the move see you in the next one.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



