# Providers

  A provider is the account or connection Lakeshore launches compute through — a
  cloud credential, a cluster endpoint, a machine you already own. A queue points
  at one via `providerRef` when it should grow its own workers.

Each provider picks exactly one **launcher**, and the control plane accepts
exactly five, spelled exactly like this. Entries carry a marker when they are not
yet finished: **`[ draft ]`** means designed but nothing built, **`[ dev ]`**
means you can configure it but it will not carry a workload end to end.

| Launcher            | Launches onto                                     | Server-side launch |
| ------------------- | ------------------------------------------------- | ------------------ |
| [EC2](/lakeshore/providers/ec2.md)     | An AWS instance it provisions      | yes                |
| [GCE](/lakeshore/providers/gce.md)     | A Google Cloud VM it provisions    | yes                |
| [Kube](/lakeshore/providers/kube.md)   | A Pod on a cluster you already run | yes                |
| [SSH](/lakeshore/providers/ssh.md) `[ dev ]`     | A host that already exists | no — CLI only |
| [SLURM](/lakeshore/providers/slurm.md) `[ dev ]` | A cluster login node       | no — stub     |

EC2, GCE and Kube are the three a queue can grow workers on. SSH registers and
runs from the CLI but has no server-side launch route; SLURM's `sbatch` submit
path is still a stub. A half-supported provider can exist at all because
`Provider.kind` is a free-form `String` rather than a constrained enum, and the
two lists in front of it disagree — registration accepts all five names, the
launch routes accept only three, and answer 422 for the other two.

Credentials never sit in the row as plaintext: kwargs reference them with a
`{ "$secret": "<name>" }` marker, and the secret is decrypted in memory inside
the launch handler only.

Each launcher's own fields, routes and setup live on its page above. The rest of
this page is the cross-cutting machinery — how providers are registered, stored,
merged and scoped, whichever launcher you picked.

## Dispatch

**Dispatch** picks how work reaches the host:

- `direct` — the launch is initiated per call from outside: an SSH session, an
  `sbatch` submission, a `kubectl apply`. Nothing long-lived runs on the host.
- `daemon` — a `nymph` daemon on the host long-polls the control plane and pulls
  work. The daemon always initiates the connection, so this works behind NAT with
  no inbound port.

`EC2` and `GCE` default to `daemon`; `Kube`, `SSH` and `SLURM` default to
`direct`. Set `dispatch` explicitly on a provider to override its default.

## Registering a provider

Providers live on the control plane. Register them with the CLI:

```bash
lakeshore providers add aws-us-east --launcher EC2 \
  --kwarg region=us-east-1 \
  --kwarg key_name=lakeshore-ops \
  --kwarg security_group=sg-0123abc \
  --kwarg iam_instance_profile_arn='arn:aws:iam::123456789012:instance-profile/lakeshore-worker'
```

Names must match `^[a-z0-9][a-z0-9-]*$`. `--kwarg` is repeatable and dotted keys
nest (`--kwarg aws_credentials.accessKeyId=AKIA…` builds a nested object).
`--kwargs-file <path>` loads a YAML or JSON base that `--kwarg` then layers on
top.

All five launchers register through this one command shape — an SSH or SLURM
provider is accepted and stored exactly like a cloud one. It is the launch side
that is not there yet for those two.

### Wire shape

The API body is `{ name, launcher, kwargs, dispatch? }`. The database column
names differ, which matters when reading the Prisma schema:

| API field  | DB column                  |
| ---------- | -------------------------- |
| `launcher` | `Provider.kind`            |
| `kwargs`   | `Provider.config`          |
| `dispatch` | `Provider.config.dispatch` |

Reads split `dispatch` back out and fill in the launcher default when it was
never set, so a `GET` always reports a concrete value.

Kwargs are an open bag — the control plane validates only that
`{ "$secret": "<name>" }` markers name a secret that exists in the namespace
(422 otherwise). Everything else round-trips verbatim, so a new launcher field
works without a schema change.

### Editing

```bash
lakeshore providers update aws-us-east --kwarg region=us-west-2
lakeshore providers update aws-us-east --dispatch daemon
lakeshore providers update aws-us-east --tunnel wg-lab     # "" clears it
```

`--kwarg` deep-merges (`key=null` deletes a key); `--kwargs-file` replaces kwargs
wholesale. `--launcher` is **immutable** — supplying it is an error, so changing
launcher means remove and re-add.

## The `.dreamrc` fallback

For solo dev the CLI also reads a local `.dreamrc`, resolving providers without a
control-plane round trip. Entries are YAML-tagged mappings and the tag names the
launcher class. Only `!providers.SSH`, `!providers.SLURM`, `!providers.EC2`,
`!providers.GCE`, and `!providers.Kube` are known tags — anything else is a hard
parse error.

`.dreamrc` is resolved first-hit-wins in this order: `--dreamrc <path>`,
`$DREAMRC`, `./.dreamrc` walked up from cwd, `$HOME/.dreamrc`, a server pull
(needs auth or `LAKESHORE_URL`), the on-disk cache of a previous pull, then an
empty config. The server pull composes a `DreamRc` from the namespace's
`providers` and `modes` endpoints, so the two forms converge.

Credentials in a local `.dreamrc` are inline. Treat the file like a `.env` and
never commit it.

## Referencing a provider from a mode

A **mode** is a stored `RunConfig`. It carries the per-machine fields and names
its provider by string:

```bash
lakeshore modes add h100 \
  --field provider=aws-us-east \
  --field instance_type=g5.xlarge \
  --field image_id=ami-0abcdef1234567890 \
  --field runner=docker \
  --field image=ghcr.io/lakeshore-py/cuda12.4-pytorch:2.4 \
  --field resources.gpu=1
```

> **Warning:** The override flag is `--field` on `modes add` / `modes update` and `--kwarg` on
> `providers add` / `providers update`. Both are repeatable and both nest on dotted
> keys, but the names are not interchangeable.

The mode's recognised fields are `backend`, `runner`, `image`, `tags`,
`resources`, `env`, `timeout_s`, and `provider`. Anything else is stashed under
`resources.__extras` and folded back in on read, so unknown keys survive a round
trip. Neither `backend` nor `runner` is validated by the API: the create handler
defaults them to `fabric` and `process` and otherwise accepts any string.

In a local `.dreamrc`, the root of the file *is* the default `RunConfig` and
named alternatives live under `modes:` (alias `mode:`).

## What happens at launch

Merge rules are the same wherever the merge happens:

1. Resolve `provider` on the `RunConfig` to the provider row.
2. **Deep-merge** the provider kwargs with the `RunConfig`'s launcher-shaped
   kwargs. Nested objects merge recursively, arrays concatenate and dedupe, and
   on a scalar conflict the `RunConfig` wins.
3. Runner-side config (`runner`, `image`, `resources`, `env`, `tags`,
   `timeout_s`) stays on the `RunConfig` — the runtime layer consumes it, not the
   launcher.
4. Resolve `{ "$secret": "<name>" }` markers by decrypting the named `Secret` in
   memory. An `aws_keypair` secret is stored as
   `<access_key_id>:<secret_access_key>` and expands to
   `{ accessKeyId, secretAccessKey }`; every other kind substitutes its UTF-8
   plaintext. Plaintext never leaves the handler and is never serialized back to
   the wire. See
   [Secrets](https://lakeshore.dreamlake.ai/api/auth-and-secrets#secrets).
5. Call the cloud API, or open the SSH session.

The control plane holds the SDK clients and the decrypted credentials; the CLI
never talks to AWS, Google, or the apiserver directly. Every launched instance is
tagged or labelled with `lakeshore=true`, the namespace, and the provider name,
so `instances` listings filter cleanly.

## Visibility and lifecycle

`remove` is a **soft delete**. The row keeps its history and gets a tombstone
suffix (`<name> [deleted <iso>]`) appended so the `(namespace, name)` unique
index stays intact and a later `providers add <name>` does not collide.
Responses strip the suffix, so callers always see the original name.

| State       | Set by                                              | Effect on listings |
| ----------- | --------------------------------------------------- | ------------------ |
| `hidden`    | `PATCH .../providers/:name` with `{ hidden: true }` | Excluded by default; `--hidden` opts back in. |
| `deletedAt` | `DELETE .../providers/:name`                        | Excluded by default; `--deleted` opts back in. Restore with `POST .../providers/:name/restore`. |

```bash
lakeshore providers list              # active + visible
lakeshore providers list --hidden     # also hidden
lakeshore providers list --deleted    # also soft-deleted
lakeshore providers list --all        # everything
lakeshore providers list --json
```

The table prints `NAME`, `LAUNCHER`, `DISPATCH`, `KWARGS`. Hidden rows render as
`name  [H]` and soft-deleted rows as `name  [D <iso>]`, with the bare name still
first so `grep` by name keeps working.

### Interactive editor

```bash
lakeshore providers edit
```

Two steps. First a picker over the listing — the first entry is a synthetic
toggle that flips the hidden + deleted filter without leaving the editor, the
second is quit. Then an action menu that adapts to the row: an active row offers
hide and delete, a hidden row offers unhide and delete, a soft-deleted row offers
restore. The list refetches between iterations, so each action shows up in the
next frame. Ctrl-C exits cleanly.

## Namespace and org scoping

Providers belong to a namespace. What you see depends on whether that namespace
has an `orgId`:

- **With an `orgId`** — listings, name lookups, and duplicate-name checks span
  every namespace sharing that `orgId`. Provider names must be unique across the
  org, and a token scoped to one member namespace works against its siblings.
- **Without an `orgId`** — the listing is *unscoped*: the namespace sees every
  provider on the control plane.

Give tenant namespaces an `orgId` if you want provider isolation between teams.
See [Admin setup](https://lakeshore.dreamlake.ai/admin/setup) for how to set one.

## Not built yet

- **SLURM launch.** The launch routes support only `EC2`, `GCE`, and `Kube`, and
  `providers test` on a SLURM provider prints a stub message. Clearing that is
  what clears its `[ dev ]` marker.
- **SSH through the server.** The CLI drives SSH directly, but no server-side
  route does, so a queue cannot grow workers on an SSH provider.
- **Hybrid dispatch** — go direct on the first call, leave a daemon behind for
  later ones. Pick one mode for now.
- **Cross-provider scheduling** — listing, claiming, and leasing capacity across
  providers. Design only.
