DreamLake

SLURM Provider

[ dev ] — you can register a SLURM provider and list its jobs, but the sbatch submit path is still a stub, so nothing you launch reaches the cluster. See Providers for what the marker means.

--launcher SLURM describes a Slurm cluster reached through its login node. The provider row carries the login-node connection; the per-job choices (partition, walltime, CPU/GPU counts) live on the mode.

Submission is not implemented yet

lakeshore providers test <slurm-provider> prints providers.SLURM is not yet implemented in this build (sbatch path stub) and exits 3. The control plane's /providers/:name/launch route supports only EC2, GCE, and Kube, so a SLURM provider gets a 422 there too.

What does work today: registering the provider, and lakeshore providers instances <name>, which SSHes to the login node and reports your running jobs.

Kwargs

The login-node connection reuses the SSH keys — providers instances builds its ssh argv from the same fields the SSH launcher uses.

FieldTypeDefaultMeaning
host (or ip, or login_node)string—Required. Login-node DNS or IP.
user (or username)stringubuntuSSH user on the login node.
portint22SSH port.
pemstring—Local path to the SSH key.

Kwargs are an open bag, so account-scoped extras such as account, qos, or launch_dir round-trip on the row for the submit path to consume once it lands — they are not read by anything today.

Mode fields

Per-job choices belong on the mode, not the provider. The daemon's Slurm runner (the slurm runner inside nymph) reads these keys from run_config.slurm when it renders an sbatch script:

KeyDefaultMaps to
partition—--partition
time01:00:00--time
gres—--gres
nodes—--nodes
ntasks-per-node—--ntasks-per-node
cpus-per-task—--cpus-per-task
mem—--mem
poll_interval_s5How often the daemon polls squeue

That runner submits job.sbatch, parses the job id out of Submitted batch job <n>, polls squeue -j <id> -h -o '%T', and runs scancel on cancellation. It maps CD to success and F / CA / TO to failure.

Example

.dreamrc tab: The same shape in a local .dreamrc:

bash
lakeshore providers add bos14 --launcher SLURM \
  --kwarg host=bos14.mit.edu \
  --kwarg user=gyang \
  --kwarg pem=~/.ssh/id_ed25519 \
  --kwarg account=lakeshore \
  --kwarg launch_dir=/home/gyang/lakeshore

lakeshore modes add bos14 \
  --field provider=bos14 \
  --field runner=slurm \
  --field slurm.partition=xeon-g6-volta \
  --field slurm.time='"04:00:00"' \
  --field slurm.cpus-per-task=8 \
  --field slurm.gres=gpu:1 \
  --field tags='["bos14","gpu"]'

Listing jobs

bash
lakeshore providers instances bos14 --timeout 15

The CLI SSHes to the login node and tries squeue --me --json first, falling back to squeue --me --noheader -o "%i,%T,%P,%V,%j,%a" when the cluster's Slurm has no JSON support. Each job becomes one row: the id, the partition in the type column, the submit time, and the account as a tag. States map RUNNING → running, PENDING → starting, COMPLETING/STOPPING → stopping, everything else unknown.

Without host / ip / login_node in the kwargs the command errors out rather than returning an empty list.

Credentials

SSH to the login node, exactly as for the SSH provider: pem is a local key path, and the agent is the fallback when it is absent. Slurm itself trusts the SSH session, so there is nothing further to configure.

Running a daemon under an allocation

The alternative to per-call sbatch is a long-lived daemon inside one allocation, pulling work for as long as the walltime lasts. Install it with the SSH-alias path and the --slurm flag, which wraps the launch in an sbatch heredoc instead of setsid nohup:

bash
lakeshore daemon install bos14-login --slurm --slurm-partition xeon-g6-volta

The daemon long-polls the control plane for the life of the allocation and disappears when it ends.

Where the nymph sits — the one decision

--slurm is not the only placement, and it is not the default one. A nymph on the login node is a manager: it does not run your code, it submits your code to the scheduler and babysits the job.

Login node (no flag)Inside a job (--slurm)
Rolemanager — submits workworker — runs work locally
Lifetimeunboundeddies at the sbatch wall-clock
Renewalnot neededmanual, on you
hostFacts.backendlocalslurm

Prefer the login node. Under --slurm nothing renews the allocation — when the wall-clock expires the daemon silently vanishes, and the CLI warns about exactly this at install time.

Install it with the plain SSH-alias path and no flag. The nymph probes the host, finds sbatch, and advertises runners: ["process", "slurm"] without being told to:

bash
lakeshore daemon install bos14-login      # no --slurm

A login node reports backend: local on purpose. It is tagged host=slurm and carries the slurm runner because sbatch is on its PATH — but it is not itself inside a job. Conflating "can submit" with "is running under SLURM" would make every exit reason on that row read wrong.

Match a launch to its worker and logs

The unique-launch installer assigns one ULID per lakeshore daemon install, including installs without --slurm. It appends that ID to the worker label and prints the log path. Under --slurm, the same ID appears in the job name:

SurfaceName
Worker label<alias-or-label>-<launchId>
Slurm job namelakeshore-nymph-<launchId>
Log~/.local/share/nymph-<launchId>.log
Binary, config, token~/.local/share/nymph/launches/<launchId>/
bash
# On the cluster; replace this with the ID printed by the installer.
launch_id=01K4G8AX000000000000000001
squeue --name="lakeshore-nymph-$launch_id" -o '%.18i %.48j %.10T'
tail -f "$HOME/.local/share/nymph-$launch_id.log"

# On the machine with CLI auth:
lakeshore daemon list | grep "$launch_id"

Each launch keeps a private config and token, so two queued jobs cannot read the last install's label when they eventually start. Nymph receives the sectioned config with --config; the installer sets its working directory so server.token_file = "token" resolves to that launch's token.

A fresh install starts another daemon. It does not stop an existing one. Launch directories and logs remain for inspection after exit; remove them only after the corresponding process or Slurm job has stopped.

Older installers share a log and config

If your installer prints no launch ID, it still uses the shared nymph.log and fixed lakeshore-nymph job name. Upgrade to a release containing the unique-launch fix before running multiple installs against the same host. This fixes the SSH daemon installer; the provider submission stub described above is unchanged.

Kill

bash
lakeshore daemon kill <worker-id>

# For an inside-allocation daemon, cancel its specific job:
ssh bos14-login scancel <slurm-job-id>

SIGTERM is caught and reported: the worker row records exitReason: sigterm and moves to gone. A scancel on an inside-a-job nymph does the same thing, and SIGKILL leaves exitReason null — the control plane infers stale after 60s of silence rather than inventing a cause.

In-flight sbatch jobs are not cancelled when the manager stops. Cancel those with scancel directly.