SLURM Provider
[ dev ] — you can register a SLURM provider and list its jobs, but the
sbatch submit path is still a stub, so nothing you launch reaches the cluster.
See Providers for what the marker means.
--launcher SLURM describes a Slurm cluster reached through its login
node. The provider row carries the login-node connection; the per-job
choices (partition, walltime, CPU/GPU counts) live on the mode.
lakeshore providers test <slurm-provider> prints
providers.SLURM is not yet implemented in this build (sbatch path stub)
and exits 3. The control plane's /providers/:name/launch route supports
only EC2, GCE, and Kube, so a SLURM provider gets a 422 there too.
What does work today: registering the provider, and
lakeshore providers instances <name>, which SSHes to the login node
and reports your running jobs.
Kwargs
The login-node connection reuses the SSH keys — providers instances
builds its ssh argv from the same fields the SSH launcher uses.
| Field | Type | Default | Meaning |
|---|---|---|---|
host (or ip, or login_node) | string | — | Required. Login-node DNS or IP. |
user (or username) | string | ubuntu | SSH user on the login node. |
port | int | 22 | SSH port. |
pem | string | — | Local path to the SSH key. |
Kwargs are an open bag, so account-scoped extras such as account,
qos, or launch_dir round-trip on the row for the submit path to
consume once it lands — they are not read by anything today.
Mode fields
Per-job choices belong on the mode, not the provider. The daemon's Slurm
runner (the slurm runner inside nymph) reads these keys from
run_config.slurm when it renders an sbatch script:
| Key | Default | Maps to |
|---|---|---|
partition | — | --partition |
time | 01:00:00 | --time |
gres | — | --gres |
nodes | — | --nodes |
ntasks-per-node | — | --ntasks-per-node |
cpus-per-task | — | --cpus-per-task |
mem | — | --mem |
poll_interval_s | 5 | How often the daemon polls squeue |
That runner submits job.sbatch, parses the job id out of
Submitted batch job <n>, polls squeue -j <id> -h -o '%T', and runs
scancel on cancellation. It maps CD to success and F / CA / TO
to failure.
Example
.dreamrc tab: The same shape in a local .dreamrc:
lakeshore providers add bos14 --launcher SLURM \
--kwarg host=bos14.mit.edu \
--kwarg user=gyang \
--kwarg pem=~/.ssh/id_ed25519 \
--kwarg account=lakeshore \
--kwarg launch_dir=/home/gyang/lakeshore
lakeshore modes add bos14 \
--field provider=bos14 \
--field runner=slurm \
--field slurm.partition=xeon-g6-volta \
--field slurm.time='"04:00:00"' \
--field slurm.cpus-per-task=8 \
--field slurm.gres=gpu:1 \
--field tags='["bos14","gpu"]'providers:
bos14: !providers.SLURM
host: bos14.mit.edu
user: gyang
account: lakeshore
launch_dir: /home/gyang/lakeshore
modes:
bos14:
provider: bos14
runner: slurm
slurm:
partition: xeon-g6-volta
time: "04:00:00"
cpus-per-task: 8
gres: gpu:1
tags: [bos14, gpu]Listing jobs
The CLI SSHes to the login node and tries squeue --me --json first,
falling back to squeue --me --noheader -o "%i,%T,%P,%V,%j,%a" when the
cluster's Slurm has no JSON support. Each job becomes one row: the id,
the partition in the type column, the submit time, and the account as a
tag. States map RUNNING → running, PENDING → starting,
COMPLETING/STOPPING → stopping, everything else unknown.
Without host / ip / login_node in the kwargs the command errors
out rather than returning an empty list.
Credentials
SSH to the login node, exactly as for the
SSH provider: pem is a local key path, and the
agent is the fallback when it is absent. Slurm itself trusts the SSH
session, so there is nothing further to configure.
Running a daemon under an allocation
The alternative to per-call sbatch is a long-lived daemon inside one
allocation, pulling work for as long as the walltime lasts. Install it
with the SSH-alias path and the --slurm flag, which wraps the launch in
an sbatch heredoc instead of setsid nohup:
The daemon long-polls the control plane for the life of the allocation and disappears when it ends.
Where the nymph sits — the one decision
--slurm is not the only placement, and it is not the default one. A nymph on
the login node is a manager: it does not run your code, it submits your
code to the scheduler and babysits the job.
| Login node (no flag) | Inside a job (--slurm) | |
|---|---|---|
| Role | manager — submits work | worker — runs work locally |
| Lifetime | unbounded | dies at the sbatch wall-clock |
| Renewal | not needed | manual, on you |
hostFacts.backend | local | slurm |
Prefer the login node. Under --slurm nothing renews the allocation — when the
wall-clock expires the daemon silently vanishes, and the CLI warns about
exactly this at install time.
Install it with the plain SSH-alias path and no flag. The nymph probes the
host, finds sbatch, and advertises runners: ["process", "slurm"] without
being told to:
A login node reports backend: local on purpose. It is tagged host=slurm
and carries the slurm runner because sbatch is on its PATH — but it is not
itself inside a job. Conflating "can submit" with "is running under SLURM"
would make every exit reason on that row read wrong.
Match a launch to its worker and logs
The unique-launch installer assigns one ULID per lakeshore daemon install,
including installs without --slurm. It appends that ID to the worker label
and prints the log path. Under --slurm, the same ID appears in the job name:
| Surface | Name |
|---|---|
| Worker label | <alias-or-label>-<launchId> |
| Slurm job name | lakeshore-nymph-<launchId> |
| Log | ~/.local/share/nymph-<launchId>.log |
| Binary, config, token | ~/.local/share/nymph/launches/<launchId>/ |
Each launch keeps a private config and token, so two queued jobs cannot read
the last install's label when they eventually start. Nymph receives the
sectioned config with --config; the installer sets its working directory so
server.token_file = "token" resolves to that launch's token.
A fresh install starts another daemon. It does not stop an existing one. Launch directories and logs remain for inspection after exit; remove them only after the corresponding process or Slurm job has stopped.
If your installer prints no launch ID, it still uses the shared
nymph.log and fixed lakeshore-nymph job name. Upgrade to a release containing
the unique-launch fix
before running multiple installs against the same host.
This fixes the SSH daemon installer; the provider submission stub described
above is unchanged.
Kill
SIGTERM is caught and reported: the worker row records
exitReason: sigterm and moves to gone. A scancel on an inside-a-job
nymph does the same thing, and SIGKILL leaves exitReason null — the
control plane infers stale after 60s of silence rather than inventing a
cause.
In-flight sbatch jobs are not cancelled when the manager stops. Cancel
those with scancel directly.