SpiralTrain
Exercises › Block 1 · Appendix A3

Block 1 · Appendix A3

Comprehensions

Starter notebook
22-comprehensions-starter
Fabric path
/lakehouse/default/Files/data/solutions/22.comprehensions/

Open the notebook 22-comprehensions-starter in your workspace and run it once. It has two cells : the puzzles of Steps 1 to 4, and a loop-heavy delay report for Steps 5 to 7. The flights in it are made up, and a comment marks where the code of each part goes.

Do one part at a time, 1a, 1b and so on, and compare with the expected output before you go on. The solution is in 22-comprehensions-solution, under every part on the exercise site, and at the back of the exercises PDF.

Step 1List comprehension

1a

Write one line of Python that makes a new list holding only the even elements of a, and print it.

Hint 1

The slide Filtered List Comprehension shows the if clause at the end of a comprehension.

Check your output
[4, 16, 36, 64, 100]
Show solutionHide solution
python
print([x for x in a if x % 2 == 0])

Step 2Perfect squares

2a

The starter finds the perfect squares from 1 to 100 with a for loop and prints them one per line. Write a list comprehension that produces the same numbers as a list, and print it.

Hint 1

The condition of the if in the loop is the condition of the if clause in the comprehension.

Check your output
[1, 4, 9, 16, 25, 36, 49, 64, 81, 100]
Show solutionHide solution
python
print([i for i in range(1, 101) if int(i ** 0.5) == i ** 0.5])

Step 3Matrix flattening

3a

Run the starter's flatten_for(matrix), which places the rows of the matrix one after the other, and read its output.

Check your output
[0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14]

3b

Write a list comprehension that gives the same list as flatten_for(matrix), and print it. Put the two for clauses in the same order as the nested loops.

Hint 1

The slide Comprehension Features shows two for clauses in one comprehension.

Hint 2

The outer loop of flatten_for goes over the rows, and it comes first.

Check your output
[0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14]
Show solutionHide solution
python
print([x for row in matrix for x in row])

Step 4Dictionary comprehension

4a

Build a dictionary with the countries as keys and the capitals as values, with a dictionary comprehension over zip. Do not index with range(len(...)). Print it.

Hint 1

The slides Zip Function and Dictionary Comprehensions cover this.

Check your output
{'Germany': 'Berlin', 'France': 'Paris', 'Belgium': 'Brussels', 'England': 'London', 'Spain': 'Madrid', 'Italy': 'Rome'}
Show solutionHide solution
python
by_country = {k: v for k, v in zip(country, capital)}
print(by_country)

4b

Build the reverse dictionary, which maps each capital back to its country, in one comprehension. Print it.

Hint 1

by_country.items() gives the pairs of the dictionary from 4a.

Check your output
{'Berlin': 'Germany', 'Paris': 'France', 'Brussels': 'Belgium', 'London': 'England', 'Madrid': 'Spain', 'Rome': 'Italy'}
Show solutionHide solution
python
print({v: k for k, v in by_country.items()})

Step 5Statistics

The second starter cell computes five statistics about a list of flights, each with a loop. Steps 5b to 5f replace one loop each and keep the print line below it, so the output stays the same.

5a

Run the second cell and read the output. It ends with the results of a report function.

Check your output
late flights  : ['DL0045', 'BA0431', 'KL1003', 'KL1004']
origins       : ['AMS', 'CDG', 'JFK', 'LHR']
delay by flight: {'KL1001': 12, 'KL1002': -3, 'DL0045': 47, 'BA0431': 95, 'KL1003': 22, 'AF1240': 0, 'KL1004': 31}
worst by origin: {'AMS': 31, 'CDG': 0, 'JFK': 47, 'LHR': 95}
total delay   : 207
7
['DL0045', 'BA0431', 'KL1003', 'KL1004']
{'AMS': 31, 'CDG': 0, 'JFK': 47, 'LHR': 95}
unknown report: average

5b

Rewrite late, the flight numbers with a dep_delay above 15, as a list comprehension. Run the cell and compare the late flights line with 5a.

Check your output
late flights  : ['DL0045', 'BA0431', 'KL1003', 'KL1004']
Show solutionHide solution
python
late = [f["flight"] for f in flights if f["dep_delay"] > 15]
print("late flights  :", late)

5c

Rewrite origins, the distinct origin airports, as a set comprehension. Compare the origins line with 5a. Why does the loop version need not in, and the set does not?

Hint 1

The slide Set Comprehensions shows the braces without a colon.

Check your output
origins       : ['AMS', 'CDG', 'JFK', 'LHR']
Show solutionHide solution
python
origins = {f["origin"] for f in flights}
print("origins       :", sorted(origins))

5d

Rewrite delay_by_flight, a dictionary from flight number to delay, as a dictionary comprehension. Compare the line with 5a.

Check your output
delay by flight: {'KL1001': 12, 'KL1002': -3, 'DL0045': 47, 'BA0431': 95, 'KL1003': 22, 'AF1240': 0, 'KL1004': 31}
Show solutionHide solution
python
delay_by_flight = {f["flight"]: f["dep_delay"] for f in flights}
print("delay by flight:", delay_by_flight)

5e

Rewrite worst_by_origin as a dictionary comprehension with one entry per origin. The value of an entry is max(...) over a generator expression of the delays of that origin. Compare the line with 5a. Which of the statistics from 5c and 5e does SQL express as SELECT DISTINCT, and which as GROUP BY origin with MAX?

Hint 1

The slide Generator Expression in the Generators module shows a generator expression as the only argument of a call, so max(...) needs no extra parentheses.

Hint 2

Loop over sorted(origins), so the entries come out in the same order as in 5a.

Check your output
worst by origin: {'AMS': 31, 'CDG': 0, 'JFK': 47, 'LHR': 95}
Show solutionHide solution
python
worst_by_origin = {
    origin: max(f["dep_delay"] for f in flights if f["origin"] == origin)
    for origin in sorted(origins)
}
print("worst by origin:", worst_by_origin)

5f

Rewrite total, the sum of the positive delays, as sum over a generator expression with an if clause. Compare the line with 5a.

Check your output
total delay   : 207
Show solutionHide solution
python
print("total delay   :", sum(f["dep_delay"] for f in flights if f["dep_delay"] > 0))

Step 6Pipeline

6a

Compute the total of the positive delays once more, with map, filter and reduce. Take the delays out of flights with map, keep the positive ones with filter, and add them with reduce. Print the result and compare it with 5f.

Hint 1

The slides Map and Filter and Reduce and Lambda show the three functions. reduce comes from functools, and operator.add is the function to hand to it.

Hint 2

The slide From Python Lists to Spark RDDs shows the same pipeline as RDD code, with the same lambdas.

Check your output
207
Show solutionHide solution
python
from functools import reduce
from operator import add

delays = map(lambda f: f["dep_delay"], flights)
positive = filter(lambda d: d > 0, delays)
print(reduce(add, positive))

6b

Answer two questions with any and all, and print both answers : is any flight more than 90 minutes late, and do all flights leave from AMS?

Hint 1

The slide Logical Reductions with any and all shows both functions. A generator expression of True and False values is the argument.

Check your output
True
False
Show solutionHide solution
python
print(any(f["dep_delay"] > 90 for f in flights))
print(all(f["origin"] == "AMS" for f in flights))

Step 7Dispatch table

7a

Replace the report function with a dictionary reports from report name to function, for "count", "late" and "worst". Small lambdas are enough. Look each name up with reports.get(name), and print unknown report when it returns None. Change the loop at the bottom of the cell so that it uses reports.

Hint 1

The slide Dispatch Tables shows the dictionary of functions and the get lookup.

Hint 2

A lambda without parameters, lambda: len(flights), gives a function that is called without arguments.

Check your output
7
['DL0045', 'BA0431', 'KL1003', 'KL1004']
{'AMS': 31, 'CDG': 0, 'JFK': 47, 'LHR': 95}
unknown report: average
Show solutionHide solution
python
reports = {
    "count": lambda: len(flights),
    "late": lambda: late,
    "worst": lambda: worst_by_origin,
}

for name in ["count", "late", "worst", "average"]:
    handler = reports.get(name)
    print(handler() if handler else f"unknown report: {name}")

7b

Add an "average" report, the mean of the delays, by adding one entry to reports and changing nothing else. Run the loop again.

Hint 1

delay_by_flight.values() holds the delays, and len(delay_by_flight) is their number.

Check your output
7
['DL0045', 'BA0431', 'KL1003', 'KL1004']
{'AMS': 31, 'CDG': 0, 'JFK': 47, 'LHR': 95}
29.1
Show solutionHide solution
python
    "average": lambda: round(sum(delay_by_flight.values()) / len(delay_by_flight), 1),

If time permits

  • Find the latest flight with max(flights, key=...) and the three latest with sorted(..., key=..., reverse=True)[:3].
  • Change the brackets of your Step 2 comprehension to parentheses. Feed the result to sum, then to any, and explain why the second call returns False.
  • Write a set comprehension that collects the distinct word lengths in a sentence.

Tried it yourself first?

The solution is a spoiler. Work through the hints first : a wrong attempt teaches more than a solution you only read.