Mathew K Analytics

Lesson 6 · Data analytics zero to hero

NumPy Arrays & Operations for Beginners | Data Analytics #6

Video six of the 30-part series, and the start of the tools block: NumPy, the array library nearly every other data tool in Python is built on top of. We'll…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Data Analytics Zero to Hero, Video 6: NumPy#

  • Video six of the 30-part series, and the start of the tools block: NumPy, the array library nearly every other data tool in Python is built on top of.
  • We'll load real numeric columns from mtcars, the 1974 Motor Trend car road test data, into NumPy arrays.
  • Let's jump straight in.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • Install NumPy if you haven't already: pip install numpy.
  • Place mtcars.csv in the same folder as this notebook.

Part 1: Why NumPy? Arrays vs. Lists#

import numpy as np

mpg_list = [21.0, 22.8, 21.4, 18.7, 18.1]
mpg_array = np.array(mpg_list)
print(mpg_array)
print(type(mpg_array))
[21.  22.8 21.4 18.7 18.1]
<class 'numpy.ndarray'>
print(mpg_list * 2)
print(mpg_array * 2)
[21.0, 22.8, 21.4, 18.7, 18.1, 21.0, 22.8, 21.4, 18.7, 18.1]
[42.  45.6 42.8 37.4 36.2]

Part 2: Creating Arrays#

zeros = np.zeros(5)
ones = np.ones(5)
sequence = np.arange(0, 10, 2)
evenly_spaced = np.linspace(0, 1, 5)
print(zeros)
print(ones)
print(sequence)
print(evenly_spaced)
[0. 0. 0. 0. 0.]
[1. 1. 1. 1. 1.]
[0 2 4 6 8]
[0.   0.25 0.5  0.75 1.  ]
grid = np.array([[1, 2, 3], [4, 5, 6]])
print(grid)
print(grid.shape)
print(grid.ndim)
[[1 2 3]
 [4 5 6]]
(2, 3)
2

Part 3: Indexing and Slicing#

mpg_array = np.array([21.0, 22.8, 21.4, 18.7, 18.1, 14.3])
print(mpg_array[0])
print(mpg_array[-1])
print(mpg_array[1:4])
print(mpg_array[:3])
21.0
14.3
[22.8 21.4 18.7]
[21.  22.8 21.4]
grid = np.array([[21.0, 6, 160], [22.8, 4, 108], [21.4, 6, 258]])
print(grid[0, 1])
print(grid[:, 0])
print(grid[1, :])
6.0
[21.  22.8 21.4]
[ 22.8   4.  108. ]
print(mpg_array[mpg_array > 20])
[21.  22.8 21.4]

Part 4: Vectorized Math and Aggregation#

weights_1000lb = np.array([2.62, 2.875, 2.32, 3.215, 3.44])
weights_lb = weights_1000lb * 1000
print(weights_lb)
print(weights_lb + 50)
[2620. 2875. 2320. 3215. 3440.]
[2670. 2925. 2370. 3265. 3490.]
mpg_array = np.array([21.0, 22.8, 21.4, 18.7, 18.1, 14.3, 24.4, 22.8, 32.4])
print(mpg_array.mean())
print(mpg_array.std())
print(mpg_array.min())
print(mpg_array.max())
print(mpg_array.sum())
21.766666666666666
4.731220185580506
14.3
32.4
195.9
print(np.round(mpg_array.mean(), 2))
print(np.sort(mpg_array))
print(np.median(mpg_array))
21.77
[14.3 18.1 18.7 21.  21.4 22.8 22.8 24.4 32.4]
21.4

Part 5: Loading a Real CSV Column with NumPy#

mpg_column = np.genfromtxt('mtcars.csv', delimiter=',', skip_header=1, usecols=1)
print(mpg_column[:5])
print(mpg_column.shape)
print(mpg_column.mean())
[21.  21.  22.8 21.4 18.7]
(32,)
20.090625000000003
hp_column = np.genfromtxt('mtcars.csv', delimiter=',', skip_header=1, usecols=4)
efficient_mask = mpg_column > 20
print(hp_column[efficient_mask].mean())
print(hp_column[~efficient_mask].mean())
88.5
191.94444444444446

Part 6: Reshaping Arrays#

flat = np.arange(12)
print(flat)
grid = flat.reshape(3, 4)
print(grid)
print(grid.reshape(-1))
[ 0  1  2  3  4  5  6  7  8  9 10 11]
[[ 0  1  2  3]
 [ 4  5  6  7]
 [ 8  9 10 11]]
[ 0  1  2  3  4  5  6  7  8  9 10 11]

Wrap-Up: What You Learned#

  • Why NumPy arrays beat plain lists for numeric work: fixed types and vectorized, element-wise math.
  • Creating arrays with array, zeros, ones, arange, and linspace.
  • Indexing and slicing, in one and two dimensions, plus boolean masking.
  • Aggregation functions: mean, std, min, max, sum, sort, and median.
  • Loading a real CSV column directly into NumPy with genfromtxt, and combining masks across two real columns.
  • Reshaping arrays with reshape.
  • NumPy is the engine under the hood of pandas, which is exactly where we're headed. Video seven starts a multi-part pandas block with Series and DataFrames. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.