DreamLake

Host Setup

For enrolling a machine and starting nymph, see Enroll a host. This page covers runtime environment preparation before or between workloads.

The type

Setup arrives in two places. The first is on the queue's daemonTemplate, read at launch time:

daemonTemplate (setup fields)ts
interface DaemonTemplate {
  bootstrap_script?: string     // runs once, at instance creation
  setup_scripts?:    string[]   // seeds the setup-command FIFO
  // …instance_type, image_id, runner — see the Queue model
}

The second is the command itself, which is what the daemon actually executes:

SetupCommandts
interface SetupCommand {
  script:                string    // required
  sudo?:                 boolean   // default false
  refresh_capabilities?: boolean   // default false — re-report caps after running
  name?:                 string    // optional human label
}

setup_scripts on the template is mapped into setup_commands on the launch body by the elasticity controller. Both routes below append to the same FIFO.

A worker is rarely useful the moment it boots. Host setup is the ordered list of commands it runs before — and, when you append later, between — units of work: installing a driver, warming a cache, mounting something by hand.

Coming from a jaynes config

Jaynes puts setup on the runner, and a runner is constructed with the mount list — mounts is the first parameter on every runner class, "reserved by jaynes to pass in the mount objects". That wiring is what lets the rest of the config point at it:

.jaynes.ymlyaml
runners:
  - !runners.Tmux &tmux_runner
    setup: >                              # on the host, before the session
      source $HOME/.jaynes_bash;
    startup: >                            # inside the session, before your code
      conda activate torch;
    envs: >
      LANG=utf-8 LC_CTYPE=en_US.UTF-8
    pypath: "{mounts[0].host_path}"       # both derived from the mount
    work_dir: "{mounts[0].host_path}"
    entry_script: "python -u -m jaynes.entry"
    post_script: "tmux list-sessions"

Three phases, not one

The field to be careful about is setup versus startup — they are different phases, and only the first is host setup:

jaynesRunsLakeshore
setupOn the host, to prepare the environmenthost_setup (shipped as setup_scripts) / --setup-script / the setup FIFO — this page
startupInside the session, before your code is bootstrappedNothing. run_setup is the proposed name for it; the only lever today is baking it into image_id
entry_scriptThe command that runs your payloadThe worker importing module:qualname, or command on the exec wire
post_scriptAfter the run scriptNothing — nothing runs after work completes

Jaynes has three phases where Lakeshore has two, and they do not line up end to end. Lakeshore's bootstrap_script and setup commands are both host phases — one at instance creation, one at any point after. Everything jaynes does per-session, between the host being ready and the payload starting, has no Lakeshore equivalent at all.

The rest

jaynesWhat it expressesLakeshore
mountsThe runner is handed the mount objects, so other fields can reference themNothing. No queue, mode or RunConfig field names a mount; mounts= on the decorator is the proposal that would
work_dirWorking directory, conventionally {mounts[0].host_path}Fixed at /workspace; not derived from a mount and not settable per run
pypathAdd the mount's path to PYTHONPATHAutomatic for code — the worker inserts the extracted snapshot on ImportError. Nothing for data mounts
envsEnvironment for the sessionRunConfig.env, per invocation rather than per host

The work_dir and pypath rows are the same fact twice: jaynes derives paths from mounts; Lakeshore fixes them by convention. {mounts[0].host_path} interpolation exists because a jaynes mount chooses its own location, so everything downstream has to ask it where that was. /workspace and /mounts/<name> are fixed, so nothing has to be interpolated — and nothing can be, which is the cost.

Naming the phases

Three phases, and the names should say what scope they run at, because that is the only thing anyone needs to know to place a script correctly:

PhaseProposed nameDeclared onRuns
Instance creationbootstrap_scriptdaemonTemplateOnce per machine, before the daemon exists
Hosthost_setupdaemonTemplateOn a live worker, at launch or any time after
Invocationrun_setupRunConfigBefore each body, every time

host_setup replaces setup_scripts, which named the type of the value rather than the phase — and stopped being accurate once the same list could be appended to a running worker.

run_setup is the answer to what jaynes calls startup. Three reasons it beats carrying that name over:

  • startup does not say startup of what. Read cold, it sounds like the machine coming up, which is the phase above it. host_setup / run_setup name their scope, and the pair reads as one axis.
  • The name says where the field lives. RunConfig is already the per-invocation config object — the thing carrying env, image, runner, timeout_s. A per-invocation script belongs there and nowhere else, and run_setup puts it there by name. host_setup sits on daemonTemplate for the same reason.
  • It parallels the phase it is most confused with. The mistake people actually make is running a per-host install on every invocation, or expecting a per-run tweak to persist. Two names differing in one word make the question "per host or per run?" impossible to skip.

There is a sharper argument than any of those three, and it only becomes visible once the host key exists: the pair is exactly "in the key" and "not in the key." host_setup participates in the hash, so declaring different setup means a different host. run_setup cannot, because it runs before every body by definition and a phase that runs every time can never force a different machine. The two names are not a stylistic pairing — they are the two sides of one derived predicate.

Rejected alternatives: warmup collides with preloading, which caches across invocations rather than running before each one — the opposite property. pre_run says when but not what. invocation_setup is exact but long, and run is already the word this SDK uses for the ambient per-invocation context (dls.run.read, run_worker, RunConfig).

Both fields exist; run_setup is not executed yet

RunConfig.host_setup and RunConfig.run_setup have shipped, and daemonTemplate.host_setup is accepted as an alias for setup_scripts. Nothing was renamed — setup_scripts keeps working and stays the written spelling, because daemonTemplate is an untyped Json column with no schema, written by three separate CLI paths and read by the elasticity controller, so a rename would be an unvalidated data migration across independent release trains. Accepting both spellings costs one fallback.

What is declared and what runs are still different things. host_setup on a RunConfig is hashed into the host key but never executed — the daemon runs the host_setup in its own config at boot. run_setup has no execution path at all yet. So both fields today affect placement, not behaviour.

One config, not two — and a derived host key

The obvious move is to split RunConfig into a HostConfig and a RunConfig, because the seam is already in the code: the docstring joins two things with the word "plus" — "per-machine launcher kwargs (instance_type, partition, …) plus the runtime-side knobs" — and _RUNTIME_FIELDS is a hand-maintained set commented "not forwarded to the launcher."

That is the wrong conclusion, and ephemerality is why. If a host that does not match can simply be relaunched, then a host configuration is not a peer of the run configuration. It is a consequence of it. Nobody should be writing one.

So: one declaration, written by the author. The host requirement is derived from it.

Written byWhat it is
Run declarationWhoever writes the function — increasingly an agentimage, resources, env, timeout_s, runner, host_setup, run_setup, source / target / mounts
Host keyNobody — computedA hash of the subset that affects placement. Two runs with the same key can share a worker; a different key means launch
Fleet policySet once, outside the code being writtenproviderRef, elasticity, min / max, what instance types are permitted at all. Lives on the queue

The middle row is the whole idea. host_setup does not need to live on a separate object to be shared — it needs to be part of the key. Two functions declaring the same setup hash the same and land on the same warm worker; one that declares different setup gets its own. The boundary between "host" and "run" stops being a typing decision and becomes a caching decision, which is what it always was.

That also explains the warm/cold behaviour better than a split does. Preloaded functions survive because the worker process survives; the host key is exactly the predicate for "may this invocation land on that process."

Agent-written code argues for one object, not two

The classic case for splitting was that operators and authors are different people with different review. When an agent writes the function, they are the same hand — so the split buys no separation and costs a second object to keep consistent.

That cost is the real one. daemonTemplate duplicating part of RunConfig is a sync burden today; asking an agent to maintain both, in step, across every generated function, is a class of drift with no upside. A derived key cannot drift from the declaration it is computed from.

The key function

"A hash of the subset that affects placement" is only useful if the subset is written down, because three languages have to compute the same string. It is:

host_key = "hk1_" + base32-lower-nopad( sha256( canonical_json(subset) )[0..10] )

The subset is exactly seven keys, always all seven, in every case:

KeyTypeValue when absent
vintegerAlways 1
runnerstring | nullnull
imagestring | nullnull
resourcesobject, string → int | string{}
host_setuparray of string[]
providerstring | nullnull
extrasobject{}

Canonical JSON means UTF-8, no whitespace, and object keys sorted by code point recursively. Absent, empty and null are the same input — there is exactly one representation of "not set," which is what makes the hash stable across three implementations rather than merely deterministic within one. Floats are a hard error rather than a silent coercion: resources values are int or str, because IEEE-754 round-tripping through msgpack and three JSON serializers is not something to bet a cache key on.

Two things are asymmetric on purpose. host_setup order is significant — the array is not sorted, because running pip install torch before apt-get install ffmpeg is a different host preparation, and a worker prepared one way must not serve a job that declared the other. resources key order is not — it is a mapping, so it is normalised.

The exclusions each carry a reason:

ExcludedWhy
envApplied by the runner at exec time, per invocation. It does not change what the machine is — and if it did, every distinct variable would be a cold start
run_setupRuns per invocation, before each body, by definition. A phase that runs every time cannot force a different host
timeout_sA longer deadline does not need a different machine
tagsAn operator concern, already matched separately by the pool spec. Folding it in would double-count

Ten bytes is eighty bits, which is exactly sixteen base32 characters with no padding — fixed-width, lowercase, URL-safe, and short enough to sit in a table without wrapping. hk1_ is a version prefix: changing the subset or the encoding bumps it to hk2_, so an old key never silently means something new.

Two worked values, so an implementation can be checked against something:

{}                      →  hk1_ljwqko7vjw2cnyny
{"runner": "process"}   →  hk1_jjml3vzthmtzu76a

Python computes the key at submit — that is the only place the declaration exists, so it is authoritative. The daemon recomputes it from the config it was launched with and advertises it on hello and every poll. The control plane computes nothing; it stores what the SDK sent, receives what the daemon advertised, and compares two strings. A hash comparison the control plane does not participate in cannot drift from the hash function.

Two independent implementations of the same function are the obvious risk, and the only real mitigation is a shared fixture: one golden-vector file, copied verbatim into every repo that implements the key, with a test in each that reads it. An edited copy fails immediately, which turns a certainty of drift into a visible event.

What the split got right, and where it still applies

Two objections survive, and both are about hosts that are not ephemeral:

Not every worker can be relaunched. providerRef: null is the bring-your-own-worker case — a lab machine, a SLURM allocation you already hold, fixed elasticity. There the host predates the run and cannot be re-derived, so a declaration that does not match is an error rather than a launch. The derived key still works: it just fails to find a match instead of provisioning one.

Cost needs a bound. If the run declares everything, it can ask for a p4d.24xlarge and nothing says no. Splitting into a HostConfig does not fix that, because the same hand writes both — and when that hand is an agent, it writes both in the same edit.

A bound only works if it lives outside the file being edited. That is the queue: it already carries providerRef and elasticity, it is set once rather than per function, and a UDF cannot widen it from the inside. Permitted instance types belong there, next to the spend limits. The distinction that matters is not who writes it but what can rewrite it — a guardrail inside the generated artifact is not a guardrail.

What this means for `_RUNTIME_FIELDS`

The filter does not become a type; it becomes the key function. Today it answers "which fields do I hide from the launcher," maintained by hand, with anything unlisted silently forwarded — so a typo becomes a launcher kwarg. As a key function it answers "which fields force a new host," which is a question with an observable right answer: get it wrong and you either over-launch or land a job on a worker that cannot serve it.

Proposal

RunConfig is one class today, daemonTemplate is an untyped Json column that duplicates part of it, and there is no derived host key — no host_key() in the SDK, nothing on the wire, and nothing in the claim query. The elasticity controller builds a launch body from daemonTemplate directly. Neither host_setup nor run_setup exists, and _RUNTIME_FIELDS is still the hand-maintained filter described above, with no readers at all.

The key function, the subset and the two worked values in The key function are the specification the three implementations will be written against, not a description of any of them.

Two mechanisms, different lifetimes

bootstrap_scriptSetup commands
WhenOnce, at instance creationAny time, including long after join
Where declareddaemonTemplate.bootstrap_scriptdaemonTemplate.setup_scripts, --setup-script, or the HTTP route
Re-runnableNo — new instance onlyYes, append whenever
Failure blast radiusThe instance never joinsOne command; the worker stays

bootstrap_script is the cloud-init-shaped one: it belongs to the machine, not to the queue's ongoing life. Setup commands are the opposite — a durable queue the daemon drains, which is what makes them appendable.

Bake what you can into the image

Setup runs per worker, not once per queue. An elastic queue that scales to eight workers runs the same install eight times, and pays it again after every scale-to-zero. Anything stable belongs in image_id; setup is for what genuinely varies per launch.

How the FIFO drains

The queue is fire-and-forget, and the daemon pulls rather than being pushed:

POST …/workers/:id/setup   →  append to the worker's pending-setup FIFO
                              worker flips active → setting_up
POST /v1/daemon/poll       →  pops the head of the FIFO, one per poll
                              daemon runs it, reports back via
                              updated_capabilities on the next poll

The state flip is conservative. A worker in active moves to setting_up; pending and joining are left alone because they reach active naturally and will pick the commands up then, and the terminal-ish states — draining, stale, gone — are never flipped silently.

refresh_capabilities: true is the flag that matters when the script changes what the worker can do. Installing CUDA without it leaves the control plane still believing the worker is CPU-only.

The setting_up flip is the only readiness signal a worker has today, and it only covers setup the control plane itself queued — setup the daemon runs on its own initiative is invisible to it. Worker lifecycle is where that gap, and the phase field that closes it, are worked out.

The interface

Python

There is none. Setup is operator-side — declared on the queue's daemonTemplate or pushed at a worker — and a @udf cannot request it.

CLI

Seed setup at launch:

terminalbash
lakeshore daemon launch --queue gpu \
    --setup-script ./install-cuda.sh \
    --setup-script ./warm-cache.sh

lakeshore daemon launch --queue gpu --no-setup

--setup-script is repeatable and takes a path; the file's contents become the command. Scripts declared in the project config under daemon.setup_scripts are read first and the CLI flags append after them, so config supplies the baseline and the flag adds to it. --no-setup drops both — the config-declared scripts and the flags — which makes it the way to launch a deliberately bare worker.

HTTP

POST /v1/namespaces/:ns/workers/:id/setupjson
{
  "script": "apt-get install -y ffmpeg",
  "sudo": true,
  "refresh_capabilities": false,
  "name": "install-ffmpeg"
}
json
202 { "commandId": "cmd_01JD2K…" }

A malformed body is a 400; an unknown worker id is a 404, and so is a worker in another namespace — the same posture the daemon GET handler takes, so a probe cannot distinguish "does not exist" from "not yours".

The companion path is the launch route, which seeds the same FIFO at creation:

POST /v1/namespaces/:ns/daemons/launch    { …, setup_commands: SetupCommand[] }
Elastic queues →

Where daemonTemplate lives, and why setup cost is paid per worker.

Mounting Storage →

The other half of preparing a host — and what a jaynes mount maps onto.

Worker lifecycle →

When a worker is willing to be given work, and how it says so — the phase the setup FIFO is really gating.