import numpy as np
prices = np.array([12.0, 9.0, 15.0])
print(prices)
print(prices.shape) # how many, in each dimension
print(prices.dtype) # what kind of number
print(prices.size) # how many in total[12. 9. 15.]
(3,)
float64
3
Programming with Python
The investor wants a numbers deck. Two weeks of orders sit in the system, one row each: minutes, price, zone.
. . .
Tobi’s plan is a 40-tab spreadsheet, one tab per zone, copied by hand. He looks at the pile, then at his list-of-lists, and says the only true thing he’ll say all week:
“We need a bigger boat than a list.”
. . .
Today we get the bigger boat: NumPy, one object that holds a thousand numbers and does arithmetic on all of them at once.
A NumPy array holds many numbers under one name, like a list, but built for math. You make one from a list, then ask it about itself:
import numpy as np
prices = np.array([12.0, 9.0, 15.0])
print(prices)
print(prices.shape) # how many, in each dimension
print(prices.dtype) # what kind of number
print(prices.size) # how many in total[12. 9. 15.]
(3,)
float64
3
. . .
np.array([...]) wraps a list; import numpy as np is the nickname everyone uses. .shape is (3,), .dtype is float64, .size is 3. One dtype for the whole array: every element shares the same type; that’s part of what makes it fast.
Here’s the bigger boat. The deck needs gross prices: 19% VAT on all three. With a list you loop; with an array you just multiply:
# the painful way: a loop, item by item
gross = []
for p in [12.0, 9.0, 15.0]:
gross.append(p * 1.19)import numpy as np
prices = np.array([12.0, 9.0, 15.0])
print(prices * 1.19) # every element, one expression[14.28 10.71 17.85]
. . .
One operation lands on all elements at once: [14.28 10.71 17.85]. No loop, no .append, and on a thousand orders it’s also far faster.
Compare an array to a number and you don’t get one True/False. You get a whole array of them, one per element. That’s a mask, and you can filter with it:
import numpy as np
times = np.array([25, 41, 18, 33])
print(times > 30) # a True/False for every element
print(times[times > 30]) # keep only where the mask is True[False True False True]
[41 33]
. . .
times > 30 is the mask; times[times > 30] reads the array through the mask and returns just the matching values. No loop, no if.
A mask answers two investor questions at once. .sum() counts the Trues (each counts as 1); .mean() gives the share that are True:
import numpy as np
times = np.array([25, 41, 18, 33, 52, 29, 44, 12])
late = times > 30
print(int(late.sum())) # how many were late
print(float(times[late].mean())) # average of just the late ones
print(float(late.mean())) # the SHARE that were late4
42.5
0.5
. . .
4 deliveries over 30 minutes, averaging 42.5, and late.mean() says half the run was late, one line each, straight into the deck.
What does calling .sum() on the mask give?
print((np.array([1, 5, 3]) > 2).sum())a) True b) 2 c) [False, True, True]
. . .
Predict first. Pick a letter, then I reveal the answer.
b) 2: .sum() adds the mask up, and each True is worth 1, each False 0. Two elements clear the bar, so the count is 2:
import numpy as np
print(np.array([1, 5, 3]) > 2) # [False True True]
print((np.array([1, 5, 3]) > 2).sum()) # Trues add up to 2[False True True]
2
. . .
Summing a mask counts; averaging a mask shares. Same two tricks the lab asks for.
The investor wants deliveries that took more than 20 and less than 40 minutes. Tobi writes what he’d write for two numbers:
times = np.array([25, 41, 18, 33])
print(times > 20 and times < 40)a) an error: ambiguous truth value b) [ True False False True] c) True, both sides hold somewhere
. . .
Predict first. Pick a letter, then I reveal the answer.
and wants one truth, a mask has foura) an error. and asks the left side “are you true?”, and a four-element mask has no single answer. Between two masks use &, and wrap each side in parentheses (& binds tighter than >):
import numpy as np
times = np.array([25, 41, 18, 33])
print(times > 20 and times < 40)--------------------------------------------------------------------------- ValueError Traceback (most recent call last) Cell In[9], line 4 1 import numpy as np 3 times = np.array([25, 41, 18, 33]) ----> 4 print(times > 20 and times < 40) ValueError: The truth value of an array with more than one element is ambiguous. Use a.any() or a.all()
print((times > 20) & (times < 40)) # element by element: & combines masks[ True False False True]
Real data isn’t one row. Stack rows and you get a 2D array: here three days (rows) across four zones (columns: Nord, Sued, Hafen, Altstadt):
import numpy as np
deliveries = np.array([[ 9, 14, 11, 6],
[15, 12, 8, 9],
[13, 20, 16, 11]])
print(deliveries.shape) # (rows, columns) → (3, 4)
print(deliveries[0, 2]) # row 0, column 2(3, 4)
11
. . .
.shape is now (3, 4): three days, four zones. One index picks the row, a second the column: deliveries[0, 2] is day 0, zone 2 (Hafen).
Tobi builds a grid with two more days. What does .shape say?
grid = np.array([[ 9, 14, 11, 6],
[15, 12, 8, 9],
[13, 20, 16, 11],
[10, 17, 12, 8],
[ 7, 11, 9, 5]])
print(grid.shape)a) (20,) b) (5, 4) c) (4, 5)
. . .
Predict first. Pick a letter, then I reveal the answer.
b) (5, 4). .shape is always (rows, columns): five inner lists make five rows, each with four numbers. (20,) would be one flat row of twenty; 20 is .size, the total count:
import numpy as np
grid = np.array([[ 9, 14, 11, 6],
[15, 12, 8, 9],
[13, 20, 16, 11],
[10, 17, 12, 8],
[ 7, 11, 9, 5]])
print(grid.shape) # (rows, columns)
print(grid.size) # rows × columns(5, 4)
20
. . .
Same order as indexing: grid[row, column], rows first.
To sum a grid you must say which way to collapse it. The axis tells NumPy which direction disappears:
Nord Sued Hafen Altstadt
day0 [ 9 14 11 6 ]
day1 [ 15 12 8 9 ]
day2 [ 13 20 16 11 ]
axis=0 collapses DOWN the rows:
↓ ↓ ↓ ↓
37 46 35 26 one number per column (zone)
axis=1 collapses ACROSS the columns:
day0 → 40 · day1 → 44 · day2 → 60 one number per row (day)
axis in codeThe same grid, the same two collapses, one keyword each:
import numpy as np
deliveries = np.array([[ 9, 14, 11, 6],
[15, 12, 8, 9],
[13, 20, 16, 11]])
print(deliveries.sum(axis=0)) # DOWN the rows → per zone
print(deliveries.sum(axis=1)) # ACROSS the columns → per day[37 46 35 26]
[40 44 60]
. . .
axis=0 collapses DOWN the rows, one number per column (zone): [37 46 35 26]. axis=1 collapses across, one per day: [40 44 60].
. . .
Download your .py before you leave. Closing the tab without downloading loses your work.
np.array([...]) holds many numbers; arithmetic hits every element at once. No loop. Ask it .shape, .dtype, .size to know what you’re holding.arr > 30 is a True/False array: arr[mask] filters, .sum() counts the Trues, .mean() gives their share.axis=0 collapses DOWN the rows (one number per column), axis=1 across the columns (one per row); argmax tells you where the maximum sits.. . .
Next episode starts with Checkpoint 4: 40 minutes, AI allowed, everything from Episodes 6-7. Then the investor opens a data room, and Tobi lets an AI write his pandas: the arrays get column names, and eighty orders become a table you can query.
. . .
NumPy has excellent free docs: the NumPy absolute beginner’s guide covers everything in this session and a little more.
. . .
For more, see the literature list of this course.