Lesson 4 · Data visualisation in python
Understanding NumPy Arrays and Random Data Generation for Data Analysis in Python
In this lesson, we will start with basics and move towards using real datasets. NumPy helps us handle collections of numbers, especially for time series…
- CourseData visualisation in python
- Lesson4 of 34
- Video11 min
- FormatJupyter notebook · 23 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWelcome to NumPy Arrays & Random Data for Time Series!#
In this lesson, we will start with basics and move towards using real datasets. NumPy helps us handle collections of numbers, especially for time series data. Let us see how these ideas connect to real-world problems!
import warnings; warnings.filterwarnings("ignore")
# We turn off warnings to keep our notebook clean.
import numpy as np
# NumPy lets us work with powerful math tools in Python.
What is a NumPy array?#
A NumPy array is like a list, but better for math & big datasets. Arrays can store lots of numbers in a single object. We will build and explore some arrays next!
# Let us make our first simple NumPy array
arr = np.array([10, 20, 30, 40, 50])
print(arr)
# NumPy arrays work a bit like lists, but they are more powerful.
print("First value:", arr[0])
print("Middle value:", arr[2])
print("Last value:", arr[-1])
Why not use regular Python lists?#
NumPy arrays are faster and use less memory. You can do math on whole arrays at once without loops. This is very important for time series datasets with lots of data!
# Math on arrays: add 5 to every value
new_arr = arr + 5
print(new_arr)
# What is inside an array? Let us check types and shapes.
print("Array type:", type(arr))
print("Data type:", arr.dtype)
print("Shape:", arr.shape)
# Making arrays of zeros or random numbers
zeroes = np.zeros(5)
print("All zeros:", zeroes)
rand_nums = np.random.rand(5)
print("Random values:", rand_nums)
# Generating random integers between 1 and 100
rand_ints = np.random.randint(1, 101, size=10)
print("Random integers:", rand_ints)
# Reshaping arrays: turning a 1D array into 2D (like a table)
table = np.arange(1, 7).reshape((2,3))
print(table)
# Accessing groups of values (slicing)
print(arr[1:4]) # from 2nd to 4th item (not including the last index)
Handling errors safely#
Sometimes you might try to get data that is not there. Let us see how to safely access array values and catch errors.
# What happens if you use an index too large?
try:
print(arr[10])
except IndexError:
print("That position does not exist! Try a smaller number.")
Data setup#
Now let us work with a real dataset! We will use the Shampoo Sales data, which is monthly sales in thousands. Let us load it and peek at the data.
import pandas as pd
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/shampoo.csv"
shampoo = pd.read_csv(url)
print("Shape:", shampoo.shape)
print(shampoo.head())
# Visualizing the time series using matplotlib
import matplotlib.pyplot as plt
plt.figure(figsize=(10,4))
plt.plot(shampoo["Sales"])
plt.title("Shampoo Sales Over Time")
plt.xlabel("Month")
plt.ylabel("Sales (Thousands)")
plt.show()
# Converting pandas data to a NumPy array for number crunching
sales_arr = shampoo["Sales"].values
print(type(sales_arr))
print(sales_arr[:10])
# Simple stats: understanding your data
print("Mean:", np.mean(sales_arr))
print("Minimum:", np.min(sales_arr))
print("Maximum:", np.max(sales_arr))
# Creating a new array: normalized sales as percentages of max
norm_sales = sales_arr / np.max(sales_arr) * 100
print(norm_sales[:5])
# Find months where sales were above 350 (thousands)
high_months = np.where(sales_arr > 350)[0]
print("Months with highest sales:", high_months)
# Mini-project part 1: Generate synthetic sales data for 24 months
np.random.seed(42)
fake_sales = np.random.normal(loc=300, scale=50, size=24)
print("First 5 fake sales:", fake_sales[:5])
# Mini-project part 2: Plot real vs. synthetic sales
plt.figure(figsize=(10,4))
plt.plot(sales_arr[:24], label="Real Sales")
plt.plot(fake_sales, label="Simulated Sales")
plt.legend()
plt.title("Compare Real and Simulated Sales")
plt.xlabel("Month")
plt.ylabel("Sales (Thousands)")
plt.show()
# Some best practices for working with NumPy arrays
# 1. Always set random seeds for reproducible results.
# 2. Keep your arrays the right shape for operations.
# 3. Label your charts and data for clarity.
# Troubleshooting: When arrays do not match up
try:
arr_a = np.array([1,2,3])
arr_b = np.array([4,5])
arr_sum = arr_a + arr_b
except ValueError as e:
print("Error! Arrays are not the same size:", e)
# Extra tip: You can use input() to get numbers from a user
user_num = input("Type a sales number: ")
user_num = float(user_num)
print("You entered:", user_num)
# Challenge: Find average only for months with sales above 300
high_sales = sales_arr[sales_arr > 300]
avg_high = np.mean(high_sales)
print("Average of high sales:", avg_high)
Recap#
Today you learned about NumPy arrays, random numbers, how to load and plot a real-world time series, and how to make your own synthetic data. You practiced best habits and saw some troubleshooting tips. Keep exploring these basics. They will help in any data science project!
Thank you for learning with us!#
Go try out these codes in your own notebook. Like and subscribe on YouTube for more friendly Python lessons.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



