# SLURM Provider

**`[ dev ]`** — you can register a SLURM provider and list its jobs, but the
`sbatch` submit path is still a stub, so nothing you launch reaches the cluster.
See [Providers](/lakeshore/providers.md) for what the marker means.

`--launcher SLURM` describes a Slurm cluster reached through its login
node. The provider row carries the login-node connection; the per-job
choices (partition, walltime, CPU/GPU counts) live on the mode.

> **Warning:** `lakeshore providers test <slurm-provider>` prints
> `providers.SLURM is not yet implemented in this build (sbatch path stub)`
> and exits 3. The control plane's `/providers/:name/launch` route supports
> only `EC2`, `GCE`, and `Kube`, so a SLURM provider gets a 422 there too.
>
> What *does* work today: registering the provider, and
> `lakeshore providers instances <name>`, which SSHes to the login node
> and reports your running jobs.

## Kwargs

The login-node connection reuses the SSH keys — `providers instances`
builds its `ssh` argv from the same fields the SSH launcher uses.

| Field | Type | Default | Meaning |
| ----- | ---- | ------- | ------- |
| `host` (or `ip`, or `login_node`) | string | — | **Required.** Login-node DNS or IP. |
| `user` (or `username`) | string | `ubuntu` | SSH user on the login node. |
| `port` | int | 22 | SSH port. |
| `pem` | string | — | Local path to the SSH key. |

Kwargs are an open bag, so account-scoped extras such as `account`,
`qos`, or `launch_dir` round-trip on the row for the submit path to
consume once it lands — they are not read by anything today.

## Mode fields

Per-job choices belong on the mode, not the provider. The daemon's Slurm
runner (the `slurm` runner inside `nymph`) reads these keys from
`run_config.slurm` when it renders an sbatch script:

| Key | Default | Maps to |
| --- | ------- | ------- |
| `partition` | — | `--partition` |
| `time` | `01:00:00` | `--time` |
| `gres` | — | `--gres` |
| `nodes` | — | `--nodes` |
| `ntasks-per-node` | — | `--ntasks-per-node` |
| `cpus-per-task` | — | `--cpus-per-task` |
| `mem` | — | `--mem` |
| `poll_interval_s` | 5 | How often the daemon polls `squeue` |

That runner submits `job.sbatch`, parses the job id out of
`Submitted batch job <n>`, polls `squeue -j <id> -h -o '%T'`, and runs
`scancel` on cancellation. It maps `CD` to success and `F` / `CA` / `TO`
to failure.

## Example

**.dreamrc tab:** The same shape in a local `.dreamrc`:

**CLI**

```bash
lakeshore providers add bos14 --launcher SLURM \
  --kwarg host=bos14.mit.edu \
  --kwarg user=gyang \
  --kwarg pem=~/.ssh/id_ed25519 \
  --kwarg account=lakeshore \
  --kwarg launch_dir=/home/gyang/lakeshore

lakeshore modes add bos14 \
  --field provider=bos14 \
  --field runner=slurm \
  --field slurm.partition=xeon-g6-volta \
  --field slurm.time='"04:00:00"' \
  --field slurm.cpus-per-task=8 \
  --field slurm.gres=gpu:1 \
  --field tags='["bos14","gpu"]'
```

**.dreamrc**

```yaml file=".dreamrc"
providers:
  bos14: !providers.SLURM
    host: bos14.mit.edu
    user: gyang
    account: lakeshore
    launch_dir: /home/gyang/lakeshore

modes:
  bos14:
    provider: bos14
    runner: slurm
    slurm:
      partition: xeon-g6-volta
      time: "04:00:00"
      cpus-per-task: 8
      gres: gpu:1
    tags: [bos14, gpu]
```

## Listing jobs

```bash
lakeshore providers instances bos14 --timeout 15
```

The CLI SSHes to the login node and tries `squeue --me --json` first,
falling back to `squeue --me --noheader -o "%i,%T,%P,%V,%j,%a"` when the
cluster's Slurm has no JSON support. Each job becomes one row: the id,
the partition in the type column, the submit time, and the account as a
tag. States map `RUNNING → running`, `PENDING → starting`,
`COMPLETING`/`STOPPING` → `stopping`, everything else `unknown`.

Without `host` / `ip` / `login_node` in the kwargs the command errors
out rather than returning an empty list.

## Credentials

SSH to the login node, exactly as for the
[SSH provider](/lakeshore/providers/ssh.md): `pem` is a local key path, and the
agent is the fallback when it is absent. Slurm itself trusts the SSH
session, so there is nothing further to configure.

## Running a daemon under an allocation

The alternative to per-call `sbatch` is a long-lived daemon inside one
allocation, pulling work for as long as the walltime lasts. Install it
with the SSH-alias path and the `--slurm` flag, which wraps the launch in
an `sbatch` heredoc instead of `setsid nohup`:

```bash
lakeshore daemon install bos14-login --slurm --slurm-partition xeon-g6-volta
```

The daemon long-polls the control plane for the life of the allocation
and disappears when it ends.

## Where the nymph sits — the one decision

`--slurm` is not the only placement, and it is not the default one. A nymph on
the **login node** is a *manager*: it does not run your code, it submits your
code to the scheduler and babysits the job.

| | Login node (no flag) | Inside a job (`--slurm`) |
| --- | --- | --- |
| Role | manager — submits work | worker — runs work locally |
| Lifetime | **unbounded** | dies at the sbatch wall-clock |
| Renewal | not needed | manual, on you |
| `hostFacts.backend` | `local` | `slurm` |

Prefer the login node. Under `--slurm` nothing renews the allocation — when the
wall-clock expires the daemon silently vanishes, and the CLI warns about
exactly this at install time.

Install it with the plain SSH-alias path and no flag. The nymph probes the
host, finds `sbatch`, and advertises `runners: ["process", "slurm"]` without
being told to:

```bash
lakeshore daemon install bos14-login      # no --slurm
```

A login node reports `backend: local` on purpose. It is tagged `host=slurm`
and carries the `slurm` runner because `sbatch` is on its PATH — but it is not
itself inside a job. Conflating "can submit" with "is running under SLURM"
would make every exit reason on that row read wrong.

## Match a launch to its worker and logs

The unique-launch installer assigns one ULID per `lakeshore daemon install`,
including installs without `--slurm`. It appends that ID to the worker label
and prints the log path. Under `--slurm`, the same ID appears in the job name:

| Surface | Name |
| --- | --- |
| Worker label | `<alias-or-label>-<launchId>` |
| Slurm job name | `lakeshore-nymph-<launchId>` |
| Log | `~/.local/share/nymph-<launchId>.log` |
| Binary, config, token | `~/.local/share/nymph/launches/<launchId>/` |

```bash
# On the cluster; replace this with the ID printed by the installer.
launch_id=01K4G8AX000000000000000001
squeue --name="lakeshore-nymph-$launch_id" -o '%.18i %.48j %.10T'
tail -f "$HOME/.local/share/nymph-$launch_id.log"

# On the machine with CLI auth:
lakeshore daemon list | grep "$launch_id"
```

Each launch keeps a private config and token, so two queued jobs cannot read
the last install's label when they eventually start. Nymph receives the
sectioned config with `--config`; the installer sets its working directory so
`server.token_file = "token"` resolves to that launch's token.

A fresh install starts another daemon. It does not stop an existing one.
Launch directories and logs remain for inspection after exit; remove them
only after the corresponding process or Slurm job has stopped.

> **Warning:** If your installer prints no launch ID, it still uses the shared
> `nymph.log` and fixed `lakeshore-nymph` job name. Upgrade to a release containing
> [the unique-launch fix](https://github.com/dreamlake-ai/lakeshore/pull/25)
> before running multiple installs against the same host.
> This fixes the SSH daemon installer; the provider submission stub described
> above is unchanged.

## Kill

```bash
lakeshore daemon kill <worker-id>

# For an inside-allocation daemon, cancel its specific job:
ssh bos14-login scancel <slurm-job-id>
```

SIGTERM is caught and reported: the worker row records
`exitReason: sigterm` and moves to `gone`. A `scancel` on an inside-a-job
nymph does the same thing, and `SIGKILL` leaves `exitReason` null — the
control plane infers `stale` after 60s of silence rather than inventing a
cause.

In-flight sbatch jobs are **not** cancelled when the manager stops. Cancel
those with `scancel` directly.
