SpiralTrain
Exercises › Block 1 · Exercise 8

Block 1 · Exercise 8

NumPy

Starter notebook
08-numpy-starter
Fabric path
/lakehouse/default/Files/data/solutions/08.numpy/

Open the notebook 08-numpy-starter in your workspace and run its setup cell. It copies the course data next to the notebook, and the Step 1 cell loads server-load.csv from it : twenty-four hourly CPU readings, as a percentage, for web01, web02, web03, db01 and cache01. Each step has a cell with comments where your code goes, and later cells use load and SERVERS from the Step 1 cell.

Do one part at a time, 1a, 1b and so on, and compare with the expected output before you go on. Answer every part with an array expression, not with a loop over an array. The solution is in 08-numpy-solution, under every part on the exercise site, and at the back of the exercises PDF.

Step 1Load the file

The Step 1 cell has read the file into load. The first row of the CSV is a header and the first column is the hour, so neither is in the array.

1a

Print load.shape, load.dtype and load.size, one per line.

Hint 1

The slide Shape and Dimensions shows what shape and size report.

Check your output
(24, 5)
float64
120
Show solutionHide solution
python
print(load.shape)
print(load.dtype)
print(load.size)

1b

Print the first two rows, to see real numbers on the screen.

Hint 1

The slide Indexing and Slicing shows how to take rows from the front of an array.

Check your output
[[13.5 13.4  9.5 31.3  8.3]
 [ 9.3 12.7 13.  31.1  9.8]]
Show solutionHide solution
python
print(load[:2])

Step 2Per server

Each column is one server, so a per-server answer collapses the hours.

2a

Compute the mean, the minimum and the maximum of every server, as means, lows and highs. Print the shape of means.

Hint 1

The slide axis Argument says which dimension disappears for each value of axis.

Hint 2

If the shape is (24,), you collapsed the other axis. One number per server is (5,).

Check your output
(5,)
Show solutionHide solution
python
means = load.mean(axis=0)
lows = load.min(axis=0)
highs = load.max(axis=0)
print(means.shape)

2b

Print the name, mean, minimum and maximum of every server on one line, with one decimal.

Hint 1

zip(SERVERS, means, lows, highs, strict=True) walks the four sequences together.

Check your output
web01        36.8     9.3    65.7
web02        35.8    12.1    67.8
web03        31.7     9.5    55.1
db01         35.6    21.4    93.7
cache01       8.2     5.2    10.4

Your layout may differ. The numbers have to match.

Show solutionHide solution
python
for name, mean, low, high in zip(SERVERS, means, lows, highs, strict=True):
    print(f"{name:<9}{mean:8.1f}{low:8.1f}{high:8.1f}")

2c

Print the busiest server on average, with its mean.

Hint 1

np.argmax gives a position, not a value.

Hint 2

SERVERS[position] turns the position back into a name.

Check your output
web01 at 36.8 %
Show solutionHide solution
python
busiest = int(np.argmax(means))
print(f"{SERVERS[busiest]} at {means[busiest]:.1f} %")

Step 3Per hour

A per-hour answer keeps the rows and loses the servers.

3a

Compute the mean load per hour as per_hour. Print its shape.

Hint 1

One number per hour is twenty-four numbers. If the shape is (5,), use the other axis.

Check your output
(24,)
Show solutionHide solution
python
per_hour = load.mean(axis=1)
print(per_hour.shape)

3b

Print the busiest hour of the day and its mean, then the quietest hour and its mean.

Hint 1

The rows are in order from 00:00, so the position that np.argmax gives is the hour.

Check your output
busiest hour  : 13:00 at 42.9 % mean
quietest hour : 01:00 at 15.2 % mean
Show solutionHide solution
python
busiest = int(np.argmax(per_hour))
quietest = int(np.argmin(per_hour))
print(f"busiest hour  : {busiest:02d}:00 at {per_hour[busiest]:.1f} % mean")
print(f"quietest hour : {quietest:02d}:00 at {per_hour[quietest]:.1f} % mean")

3c

Print the total load per hour, rounded to whole points, to see the shape of the working day.

Hint 1

Total load is the sum of the five servers in one row.

Check your output
[ 76.  76. 144. 148.  94.  97. 108. 123. 142. 157. 176. 188. 198. 214.
 207. 212. 205. 195. 172. 157. 131. 125. 112.  97.]
Show solutionHide solution
python
print(load.sum(axis=1).round(0))

Step 4Alerts

4a

Build a boolean mask busy that is True for every reading above 80 percent. Print how many readings are over the line.

Hint 1

The slide Boolean Masking shows a comparison on a whole array. True counts as 1 when you sum.

Check your output
2
Show solutionHide solution
python
busy = load > 80.0
print(busy.sum())

4b

Print one line per alert, with the server, the hour and the reading. Get the positions from np.nonzero(busy), which returns one array per dimension, and unpack them into hours and servers. This one loop is allowed : it formats text and computes nothing.

Hint 1

zip(hours, servers, strict=True) walks the two arrays together.

Hint 2

load[hour, server] is the reading at that position. SERVERS[server] is the name.

Check your output
db01 at 02:00 -> 93.3 %
db01 at 03:00 -> 93.7 %
Show solutionHide solution
python
hours, servers = np.nonzero(busy)
for hour, server in zip(hours, servers, strict=True):
    print(f"{SERVERS[server]} at {hour:02d}:00 -> {load[hour, server]:.1f} %")

4c

Answer the same question a second way : print the servers that ever go over 80 percent, without a loop.

Hint 1

busy.any(axis=0) gives one True or False per server.

Hint 2

A boolean mask picks elements out of another array, but SERVERS is a list. Make it an array with np.array first.

Check your output
['db01']
Show solutionHide solution
python
print(np.array(SERVERS)[busy.any(axis=0)])

Step 5Baseline

A raw percentage says little. The question is how far each reading is from normal for that server.

5a

Take the median of every server as baseline. Print its shape.

Hint 1

np.median takes an axis argument, like mean.

Check your output
(5,)
Show solutionHide solution
python
baseline = np.median(load, axis=0)
print(baseline.shape)

5b

Subtract each server's median from its column, in one expression, as deviation. Print the shape of deviation.

Hint 1

The slide Broadcasting lines shapes up from the right. (24, 5) and (5,) need no reshaping.

Check your output
(24, 5)
Show solutionHide solution
python
deviation = load - baseline
print(deviation.shape)

5c

Print the largest deviation of every server, rounded to one decimal.

Hint 1

The largest value per column is the maximum along axis=0.

Check your output
[33.4 34.6 24.5 62.   2.2]
Show solutionHide solution
python
print(deviation.max(axis=0).round(1))

5d

Find the single worst deviation in the whole table. Print its server, its hour and its size.

Hint 1

np.argmax on a two-dimensional array gives a position in the flattened array.

Hint 2

np.unravel_index(position, shape) turns that position back into a row and a column.

Check your output
db01 at 03:00, 62.0 points above its median
Show solutionHide solution
python
hour, server = np.unravel_index(np.argmax(deviation), deviation.shape)
print(f"{SERVERS[server]} at {hour:02d}:00, {deviation[hour, server]:.1f} points above its median")

If time permits

  • Use nested np.where to label the whole table in one expression : critical above 90, warning above 60, and fine otherwise. Count how many readings fall into each band. Comparing the label array with a string gives a mask, and a mask sums to a count.
  • Print the number of warnings per server, by summing the warning mask along axis=0.
  • Write the hourly means to a new CSV in DATA with np.savetxt.
  • Cap a copy of load at 90 percent, without disturbing the original. Print the maximum of both arrays to prove the original survived. Then leave the .copy() out and run it again, to see what happens when you forget.

Tried it yourself first?

The solution is a spoiler. Work through the hints first : a wrong attempt teaches more than a solution you only read.