DreamLake

Providers

A provider is the account or connection Lakeshore launches compute through — a cloud credential, a cluster endpoint, a machine you already own. A queue points at one via providerRef when it should grow its own workers.

Each provider picks exactly one launcher, and the control plane accepts exactly five, spelled exactly like this. Entries carry a marker when they are not yet finished: [ draft ] means designed but nothing built, [ dev ] means you can configure it but it will not carry a workload end to end.

LauncherLaunches ontoServer-side launch
EC2An AWS instance it provisionsyes
GCEA Google Cloud VM it provisionsyes
KubeA Pod on a cluster you already runyes
SSH [ dev ]A host that already existsno — CLI only
SLURM [ dev ]A cluster login nodeno — stub

EC2, GCE and Kube are the three a queue can grow workers on. SSH registers and runs from the CLI but has no server-side launch route; SLURM's sbatch submit path is still a stub. A half-supported provider can exist at all because Provider.kind is a free-form String rather than a constrained enum, and the two lists in front of it disagree — registration accepts all five names, the launch routes accept only three, and answer 422 for the other two.

Credentials never sit in the row as plaintext: kwargs reference them with a { "$secret": "<name>" } marker, and the secret is decrypted in memory inside the launch handler only.

Each launcher's own fields, routes and setup live on its page above. The rest of this page is the cross-cutting machinery — how providers are registered, stored, merged and scoped, whichever launcher you picked.

Dispatch

Dispatch picks how work reaches the host:

  • direct — the launch is initiated per call from outside: an SSH session, an sbatch submission, a kubectl apply. Nothing long-lived runs on the host.
  • daemon — a nymph daemon on the host long-polls the control plane and pulls work. The daemon always initiates the connection, so this works behind NAT with no inbound port.

EC2 and GCE default to daemon; Kube, SSH and SLURM default to direct. Set dispatch explicitly on a provider to override its default.

Registering a provider

Providers live on the control plane. Register them with the CLI:

bash
lakeshore providers add aws-us-east --launcher EC2 \
  --kwarg region=us-east-1 \
  --kwarg key_name=lakeshore-ops \
  --kwarg security_group=sg-0123abc \
  --kwarg iam_instance_profile_arn='arn:aws:iam::123456789012:instance-profile/lakeshore-worker'

Names must match ^[a-z0-9][a-z0-9-]*$. --kwarg is repeatable and dotted keys nest (--kwarg aws_credentials.accessKeyId=AKIA… builds a nested object). --kwargs-file <path> loads a YAML or JSON base that --kwarg then layers on top.

All five launchers register through this one command shape — an SSH or SLURM provider is accepted and stored exactly like a cloud one. It is the launch side that is not there yet for those two.

Wire shape

The API body is { name, launcher, kwargs, dispatch? }. The database column names differ, which matters when reading the Prisma schema:

API fieldDB column
launcherProvider.kind
kwargsProvider.config
dispatchProvider.config.dispatch

Reads split dispatch back out and fill in the launcher default when it was never set, so a GET always reports a concrete value.

Kwargs are an open bag — the control plane validates only that { "$secret": "<name>" } markers name a secret that exists in the namespace (422 otherwise). Everything else round-trips verbatim, so a new launcher field works without a schema change.

Editing

bash
lakeshore providers update aws-us-east --kwarg region=us-west-2
lakeshore providers update aws-us-east --dispatch daemon
lakeshore providers update aws-us-east --tunnel wg-lab     # "" clears it

--kwarg deep-merges (key=null deletes a key); --kwargs-file replaces kwargs wholesale. --launcher is immutable — supplying it is an error, so changing launcher means remove and re-add.

The .dreamrc fallback

For solo dev the CLI also reads a local .dreamrc, resolving providers without a control-plane round trip. Entries are YAML-tagged mappings and the tag names the launcher class. Only !providers.SSH, !providers.SLURM, !providers.EC2, !providers.GCE, and !providers.Kube are known tags — anything else is a hard parse error.

.dreamrc is resolved first-hit-wins in this order: --dreamrc <path>, $DREAMRC, ./.dreamrc walked up from cwd, $HOME/.dreamrc, a server pull (needs auth or LAKESHORE_URL), the on-disk cache of a previous pull, then an empty config. The server pull composes a DreamRc from the namespace's providers and modes endpoints, so the two forms converge.

Credentials in a local .dreamrc are inline. Treat the file like a .env and never commit it.

Referencing a provider from a mode

A mode is a stored RunConfig. It carries the per-machine fields and names its provider by string:

bash
lakeshore modes add h100 \
  --field provider=aws-us-east \
  --field instance_type=g5.xlarge \
  --field image_id=ami-0abcdef1234567890 \
  --field runner=docker \
  --field image=ghcr.io/lakeshore-py/cuda12.4-pytorch:2.4 \
  --field resources.gpu=1
modes uses --field, providers uses --kwarg

The override flag is --field on modes add / modes update and --kwarg on providers add / providers update. Both are repeatable and both nest on dotted keys, but the names are not interchangeable.

The mode's recognised fields are backend, runner, image, tags, resources, env, timeout_s, and provider. Anything else is stashed under resources.__extras and folded back in on read, so unknown keys survive a round trip. Neither backend nor runner is validated by the API: the create handler defaults them to fabric and process and otherwise accepts any string.

In a local .dreamrc, the root of the file is the default RunConfig and named alternatives live under modes: (alias mode:).

What happens at launch

Merge rules are the same wherever the merge happens:

  1. Resolve provider on the RunConfig to the provider row.
  2. Deep-merge the provider kwargs with the RunConfig's launcher-shaped kwargs. Nested objects merge recursively, arrays concatenate and dedupe, and on a scalar conflict the RunConfig wins.
  3. Runner-side config (runner, image, resources, env, tags, timeout_s) stays on the RunConfig — the runtime layer consumes it, not the launcher.
  4. Resolve { "$secret": "<name>" } markers by decrypting the named Secret in memory. An aws_keypair secret is stored as <access_key_id>:<secret_access_key> and expands to { accessKeyId, secretAccessKey }; every other kind substitutes its UTF-8 plaintext. Plaintext never leaves the handler and is never serialized back to the wire. See Secrets.
  5. Call the cloud API, or open the SSH session.

The control plane holds the SDK clients and the decrypted credentials; the CLI never talks to AWS, Google, or the apiserver directly. Every launched instance is tagged or labelled with lakeshore=true, the namespace, and the provider name, so instances listings filter cleanly.

Visibility and lifecycle

remove is a soft delete. The row keeps its history and gets a tombstone suffix (<name> [deleted <iso>]) appended so the (namespace, name) unique index stays intact and a later providers add <name> does not collide. Responses strip the suffix, so callers always see the original name.

StateSet byEffect on listings
hiddenPATCH .../providers/:name with { hidden: true }Excluded by default; --hidden opts back in.
deletedAtDELETE .../providers/:nameExcluded by default; --deleted opts back in. Restore with POST .../providers/:name/restore.
bash
lakeshore providers list              # active + visible
lakeshore providers list --hidden     # also hidden
lakeshore providers list --deleted    # also soft-deleted
lakeshore providers list --all        # everything
lakeshore providers list --json

The table prints NAME, LAUNCHER, DISPATCH, KWARGS. Hidden rows render as name [H] and soft-deleted rows as name [D <iso>], with the bare name still first so grep by name keeps working.

Interactive editor

bash
lakeshore providers edit

Two steps. First a picker over the listing — the first entry is a synthetic toggle that flips the hidden + deleted filter without leaving the editor, the second is quit. Then an action menu that adapts to the row: an active row offers hide and delete, a hidden row offers unhide and delete, a soft-deleted row offers restore. The list refetches between iterations, so each action shows up in the next frame. Ctrl-C exits cleanly.

Namespace and org scoping

Providers belong to a namespace. What you see depends on whether that namespace has an orgId:

  • With an orgId — listings, name lookups, and duplicate-name checks span every namespace sharing that orgId. Provider names must be unique across the org, and a token scoped to one member namespace works against its siblings.
  • Without an orgId — the listing is unscoped: the namespace sees every provider on the control plane.

Give tenant namespaces an orgId if you want provider isolation between teams. See Admin setup for how to set one.

Not built yet

  • SLURM launch. The launch routes support only EC2, GCE, and Kube, and providers test on a SLURM provider prints a stub message. Clearing that is what clears its [ dev ] marker.
  • SSH through the server. The CLI drives SSH directly, but no server-side route does, so a queue cannot grow workers on an SSH provider.
  • Hybrid dispatch — go direct on the first call, leave a daemon behind for later ones. Pick one mode for now.
  • Cross-provider scheduling — listing, claiming, and leasing capacity across providers. Design only.