SpiralTrain
Exercises › Setup

Setup

Getting Into the Course Environment

  • The course runs in Microsoft Fabric, in a browser, and you install nothing
  • You have your own account, your own workspace and your own copies of every notebook
  • A notebook runs on a session that Fabric starts when you run its first cell

Sign in, find your workspace, run a notebook and stop the session afterwards. A new session may take a few minutes to start; read the next steps while it starts.

You sign in with a course account on the trainer's Microsoft tenant, not with your own work account. The accounts are named after scientists, they exist only for this course, and they hold nothing of yours. Your card with the email address and password is on the table.

Step 1Sign in

Open https://app.fabric.microsoft.com in a browser, and use the email address and password from your card.

If you are already signed in somewhere with your own work account, that account will try to take over. Open a private or incognito window instead : in Edge and Chrome that is Ctrl+Shift+N. Everything for the next four days happens in that one window.

You should land on the Fabric home page. Microsoft may ask to keep you signed in ; either answer is fine.


Step 2Find your workspace

Click the workspace icon in the bottom left of the sidebar and pick the workspace named after your account, PySpark-<name>.

It is yours alone. Nobody else can see what you do in it, and nothing you do in it can break anyone else's day. There is also a shared workspace, PySpark-Shared, which holds the course data and is where we work together when an exercise asks for it.

Inside your workspace you will find a demo notebook per module, named after the module, such as 03-functions. An exercise with a starting point has a starter and a solution notebook beside it, such as 03-functions-starter and 03-functions-solution. There is also a lakehouse called work whose Files/data folder holds the course data.

The notebooks of modules 1 to 10 are Python notebooks, which run on a small session of their own. From module 11 on they are PySpark notebooks, which run on a Spark session.

The slides and this exercise guide are in Files/data/docs as PDFs. Select one to open it in the explorer's preview, or download it to keep a copy. The same files, with the demos and solutions, are at https://git.feikowielsma.nl/SpiralTrain/PySpark-Student, where a PDF opens straight on the page.


Step 3Start the session now, then read on

Open 01-python-toolchain, the demo notebook of module 1. Click in its first code cell and press Shift+Enter.

Fabric starts a session and attaches the notebook to it. Leave it running and read the next two steps while you wait.

Later cells reuse the active session.


Step 4How a notebook runs

A notebook is a column of cells. A code cell runs when you press Shift+Enter in it, and its output appears underneath.

Cells share one Python process, so a name defined in an earlier cell remains available. Execution order determines the state. After editing a cell, run it again to apply the change. If the state gets confusing, use Restart session and run the cells from the top.

Run all at the top runs every cell from the top down, which is the way to check that a notebook works from a clean start.

Look at the top right of the notebook while a cell runs. It shows the session, whether it is busy, and how long the cell has been going. Below a Spark cell there is a Spark jobs panel that opens up into what the cluster actually did. From day 3 on, that panel is half the course.


Step 5Stop the session

When you are done for the day, or when the trainer says so, open the session menu at the top of the notebook and choose Stop session.

The environment is one pool of machines split between the nine of us. A notebook whose tab you closed keeps its share of that pool for another hour, doing nothing, and that is an hour somebody else waits for a session. Stopping is the polite half of a shared cluster, and it is also what keeps block 2's performance measurements honest.


If something goes wrong

You end up in the wrong tenant, or Fabric says you have no account

Your own work account signed in ahead of the course account. Sign out entirely, or open a private window, and sign in again with the card.

The session will not start, or it stays queued

Tell the trainer rather than clicking again : each click can ask for another session, and the pool is finite. The trainer sees every session in the Monitor hub and can free one.

%run x.py says it cannot read a notebook

In Fabric, %run runs another notebook, not a Python file. Run a file with runpy.run_path("x.py") after import runpy, or import it as a module.

A cell fails on a name that does not exist

Almost always a cell above it was never run. Run from the top.

You closed the notebook and lost your output

The code is saved, the output is not always. Run it again ; that is quick on a warm session.


Try one of these

  • Make a notebook of your own in your workspace, set its language to Python in the notebook's language dropdown, and print something. A new notebook starts as a PySpark notebook, whose Spark session takes 12 vCores of the shared pool where a Python session takes 2. It is your workspace ; nothing in it is precious.
  • Open Monitor in the sidebar and find the session you just started.
  • Look in the work lakehouse under Files/data.

Appendix : Setting This Up on Your Own Machine

None of this is needed for the course, and none of it is possible on an Achmea laptop : the course runs in the browser. This appendix is here because the toolchain of day 1 is worth having, and one day you may sit at a machine where you are allowed to install it. Everything below installs into your own user profile and needs no administrator rights.

It gives you a working Python 3.11, the uv project tool, the ruff linter, an editor that understands Python, and Java 17 for running Spark locally.

Install uv

uv is the modern Python project manager, and it can install Python itself, so it is the only thing you install by hand. On Windows, run this in PowerShell :

powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

On macOS or Linux, run this in a terminal :

bash
curl -LsSf https://astral.sh/uv/install.sh | sh

Close and reopen the terminal afterwards so that uv is on your PATH, then confirm it works with uv --version.

See also: https://docs.astral.sh/uv/

Install Python 3.11

Let uv fetch the interpreter the course uses :

bash
uv python install 3.11

This does not disturb any Python already on the machine. The course uses 3.11 because it is the Python that Spark 3.5 runs on in Microsoft Fabric, where the course itself takes place.

Install an editor

Install Visual Studio Code from https://code.visualstudio.com, then add two extensions from the Extensions panel (the square icon on the left) :

  • Python, published by Microsoft
  • Ruff, published by Astral

Any editor works. See also: https://docs.astral.sh/ruff/

Create a check project

bash
uv init --python 3.11 course-check
cd course-check
uv add numpy pandas pyarrow openpyxl matplotlib
uv add --dev ruff mypy
uv run python -c "import numpy, pandas, pyarrow, openpyxl, matplotlib; print('all imports OK')"

The first uv add downloads the packages and pins Python 3.11 for the project, and the last line should print all imports OK. Those packages are what days 1 and 2 use.

Confirm you can run code

Scripts run with uv run python some_demo.py. Using uv run guarantees the command uses the project's Python and its installed packages, so you never have to activate anything by hand. Verify with :

bash
uv --version
uv python list          # shows 3.11 installed
uv run python --version # prints Python 3.11.x inside a project
uv run ruff --version

Run the course demos

The demos folder of the course repository, https://git.feikowielsma.nl/SpiralTrain/PySpark-Student, is a uv project of its own. From inside it :

bash
uv run 03.functions/demo01_showmessage.py

Install Java 17 and run Spark locally

Spark runs in the Java Virtual Machine. Spark 3.5 supports Java 8, 11 and 17, and the course platform uses 17. Check what you have :

bash
java -version

If that does not print version 17, download the OpenJDK 17 zip for your platform from https://adoptium.net, choosing the JDK rather than the JRE, and unpack it into your user profile, for instance C:\Users\<you>\jdk-17. Then point JAVA_HOME at it and put its bin folder on your PATH, for your own user only. In PowerShell :

$jdk = "$env:USERPROFILE\jdk-17"
[Environment]::SetEnvironmentVariable("JAVA_HOME", $jdk, "User")
[Environment]::SetEnvironmentVariable("Path", "$jdk\bin;" + [Environment]::GetEnvironmentVariable("Path", "User"), "User")

Close and reopen the terminal and run java -version again. Then add PySpark and start it once :

bash
uv add "pyspark==3.5.*"
uv run python -c "from pyspark.sql import SparkSession; s = SparkSession.builder.master('local').getOrCreate(); s.range(3).show(); s.stop()"

The first start takes a while and prints some warnings. You should see a small table with the numbers 0, 1 and 2.

On Windows, Spark also wants a version-matched hadoop.dll and winutils.exe before it can read or write files, which is a fight worth avoiding : run Spark under WSL instead, where it works as it does on any Linux.

If a download is blocked

On a locked-down machine, downloads may be blocked or routed through a proxy. Set the proxy for the session (ask IT for the address) :

bash
# Windows PowerShell
$env:HTTPS_PROXY = "http://proxy.example.com:8080"
# macOS or Linux
export HTTPS_PROXY=http://proxy.example.com:8080

If your company uses its own TLS certificate, try uv with --native-tls, which makes it trust the certificates Windows trusts, or point it at the certificate file with SSL_CERT_FILE=/path/to/company-ca.pem.

Tried it yourself first?

The solution is a spoiler. Work through the hints first : a wrong attempt teaches more than a solution you only read.