Host Setup
For enrolling a machine and starting nymph, see Enroll a host. This page covers runtime environment preparation before or between workloads.
The type
Setup arrives in two places. The first is on the queue's daemonTemplate, read
at launch time:
The second is the command itself, which is what the daemon actually executes:
setup_scripts on the template is mapped into setup_commands on the launch
body by the elasticity controller. Both routes below append to the same FIFO.
A worker is rarely useful the moment it boots. Host setup is the ordered list of commands it runs before — and, when you append later, between — units of work: installing a driver, warming a cache, mounting something by hand.
Coming from a jaynes config
Jaynes puts setup on the runner, and a runner is constructed with the mount
list — mounts is the first parameter on every runner class, "reserved by
jaynes to pass in the mount objects". That wiring is what lets the rest of the
config point at it:
Three phases, not one
The field to be careful about is setup versus startup — they are different
phases, and only the first is host setup:
| jaynes | Runs | Lakeshore |
|---|---|---|
setup | On the host, to prepare the environment | host_setup (shipped as setup_scripts) / --setup-script / the setup FIFO — this page |
startup | Inside the session, before your code is bootstrapped | Nothing. run_setup is the proposed name for it; the only lever today is baking it into image_id |
entry_script | The command that runs your payload | The worker importing module:qualname, or command on the exec wire |
post_script | After the run script | Nothing — nothing runs after work completes |
Jaynes has three phases where Lakeshore has two, and they do not line up end to
end. Lakeshore's bootstrap_script and setup commands are both host phases —
one at instance creation, one at any point after. Everything jaynes does
per-session, between the host being ready and the payload starting, has no
Lakeshore equivalent at all.
The rest
| jaynes | What it expresses | Lakeshore |
|---|---|---|
mounts | The runner is handed the mount objects, so other fields can reference them | Nothing. No queue, mode or RunConfig field names a mount; mounts= on the decorator is the proposal that would |
work_dir | Working directory, conventionally {mounts[0].host_path} | Fixed at /workspace; not derived from a mount and not settable per run |
pypath | Add the mount's path to PYTHONPATH | Automatic for code — the worker inserts the extracted snapshot on ImportError. Nothing for data mounts |
envs | Environment for the session | RunConfig.env, per invocation rather than per host |
The work_dir and pypath rows are the same fact twice: jaynes derives paths
from mounts; Lakeshore fixes them by convention. {mounts[0].host_path}
interpolation exists because a jaynes mount chooses its own location, so
everything downstream has to ask it where that was. /workspace and
/mounts/<name> are fixed, so nothing has to be interpolated — and nothing
can be, which is the cost.
Naming the phases
Three phases, and the names should say what scope they run at, because that is the only thing anyone needs to know to place a script correctly:
| Phase | Proposed name | Declared on | Runs |
|---|---|---|---|
| Instance creation | bootstrap_script | daemonTemplate | Once per machine, before the daemon exists |
| Host | host_setup | daemonTemplate | On a live worker, at launch or any time after |
| Invocation | run_setup | RunConfig | Before each body, every time |
host_setup replaces setup_scripts, which named the type of the value
rather than the phase — and stopped being accurate once the same list could be
appended to a running worker.
run_setup is the answer to what jaynes calls startup. Three reasons it
beats carrying that name over:
startupdoes not say startup of what. Read cold, it sounds like the machine coming up, which is the phase above it.host_setup/run_setupname their scope, and the pair reads as one axis.- The name says where the field lives.
RunConfigis already the per-invocation config object — the thing carryingenv,image,runner,timeout_s. A per-invocation script belongs there and nowhere else, andrun_setupputs it there by name.host_setupsits ondaemonTemplatefor the same reason. - It parallels the phase it is most confused with. The mistake people actually make is running a per-host install on every invocation, or expecting a per-run tweak to persist. Two names differing in one word make the question "per host or per run?" impossible to skip.
There is a sharper argument than any of those three, and it only becomes visible
once the host key exists: the pair is exactly "in the key"
and "not in the key." host_setup participates in the hash, so declaring
different setup means a different host. run_setup cannot, because it runs
before every body by definition and a phase that runs every time can never force
a different machine. The two names are not a stylistic pairing — they are the
two sides of one derived predicate.
Rejected alternatives: warmup collides with
preloading, which caches across invocations rather
than running before each one — the opposite property. pre_run says when but
not what. invocation_setup is exact but long, and run is already the word
this SDK uses for the ambient per-invocation context (dls.run.read,
run_worker, RunConfig).
RunConfig.host_setup and RunConfig.run_setup have shipped, and
daemonTemplate.host_setup is accepted as an alias for setup_scripts.
Nothing was renamed — setup_scripts keeps working and stays the written
spelling, because daemonTemplate is an untyped Json column with no schema,
written by three separate CLI paths and read by the elasticity controller, so
a rename would be an unvalidated data migration across independent release
trains. Accepting both spellings costs one fallback.
What is declared and what runs are still different things. host_setup on a
RunConfig is hashed into the host key
but never executed — the daemon runs the host_setup in its own config at
boot. run_setup has no execution path at all yet. So both fields today
affect placement, not behaviour.
One config, not two — and a derived host key
The obvious move is to split RunConfig into a HostConfig and a RunConfig,
because the seam is already in the code: the docstring joins two things with the
word "plus" — "per-machine launcher kwargs (instance_type, partition, …)
plus the runtime-side knobs" — and _RUNTIME_FIELDS is a hand-maintained
set commented "not forwarded to the launcher."
That is the wrong conclusion, and ephemerality is why. If a host that does not match can simply be relaunched, then a host configuration is not a peer of the run configuration. It is a consequence of it. Nobody should be writing one.
So: one declaration, written by the author. The host requirement is derived from it.
| Written by | What it is | |
|---|---|---|
| Run declaration | Whoever writes the function — increasingly an agent | image, resources, env, timeout_s, runner, host_setup, run_setup, source / target / mounts |
| Host key | Nobody — computed | A hash of the subset that affects placement. Two runs with the same key can share a worker; a different key means launch |
| Fleet policy | Set once, outside the code being written | providerRef, elasticity, min / max, what instance types are permitted at all. Lives on the queue |
The middle row is the whole idea. host_setup does not need to live on a
separate object to be shared — it needs to be part of the key. Two functions
declaring the same setup hash the same and land on the same warm worker; one
that declares different setup gets its own. The boundary between "host" and
"run" stops being a typing decision and becomes a caching decision, which is
what it always was.
That also explains the warm/cold behaviour better than a split does. Preloaded functions survive because the worker process survives; the host key is exactly the predicate for "may this invocation land on that process."
The classic case for splitting was that operators and authors are different people with different review. When an agent writes the function, they are the same hand — so the split buys no separation and costs a second object to keep consistent.
That cost is the real one. daemonTemplate duplicating part of RunConfig
is a sync burden today; asking an agent to maintain both, in step, across
every generated function, is a class of drift with no upside. A derived key
cannot drift from the declaration it is computed from.
The key function
"A hash of the subset that affects placement" is only useful if the subset is written down, because three languages have to compute the same string. It is:
The subset is exactly seven keys, always all seven, in every case:
| Key | Type | Value when absent |
|---|---|---|
v | integer | Always 1 |
runner | string | null | null |
image | string | null | null |
resources | object, string → int | string | {} |
host_setup | array of string | [] |
provider | string | null | null |
extras | object | {} |
Canonical JSON means UTF-8, no whitespace, and object keys sorted by code point
recursively. Absent, empty and null are the same input — there is exactly
one representation of "not set," which is what makes the hash stable across three
implementations rather than merely deterministic within one. Floats are a hard
error rather than a silent coercion: resources values are int or str,
because IEEE-754 round-tripping through msgpack and three JSON serializers is not
something to bet a cache key on.
Two things are asymmetric on purpose. host_setup order is significant — the
array is not sorted, because running pip install torch before apt-get install ffmpeg is a different host preparation, and a worker prepared one way must not
serve a job that declared the other. resources key order is not — it is a
mapping, so it is normalised.
The exclusions each carry a reason:
| Excluded | Why |
|---|---|
env | Applied by the runner at exec time, per invocation. It does not change what the machine is — and if it did, every distinct variable would be a cold start |
run_setup | Runs per invocation, before each body, by definition. A phase that runs every time cannot force a different host |
timeout_s | A longer deadline does not need a different machine |
tags | An operator concern, already matched separately by the pool spec. Folding it in would double-count |
Ten bytes is eighty bits, which is exactly sixteen base32 characters with no
padding — fixed-width, lowercase, URL-safe, and short enough to sit in a table
without wrapping. hk1_ is a version prefix: changing the subset or the encoding
bumps it to hk2_, so an old key never silently means something new.
Two worked values, so an implementation can be checked against something:
Python computes the key at submit — that is the only place the declaration exists, so it is authoritative. The daemon recomputes it from the config it was launched with and advertises it on hello and every poll. The control plane computes nothing; it stores what the SDK sent, receives what the daemon advertised, and compares two strings. A hash comparison the control plane does not participate in cannot drift from the hash function.
Two independent implementations of the same function are the obvious risk, and the only real mitigation is a shared fixture: one golden-vector file, copied verbatim into every repo that implements the key, with a test in each that reads it. An edited copy fails immediately, which turns a certainty of drift into a visible event.
What the split got right, and where it still applies
Two objections survive, and both are about hosts that are not ephemeral:
Not every worker can be relaunched. providerRef: null is the
bring-your-own-worker case — a lab machine, a SLURM allocation you already hold,
fixed elasticity. There the host predates the run and cannot be re-derived, so
a declaration that does not match is an error rather than a launch. The derived
key still works: it just fails to find a match instead of provisioning one.
Cost needs a bound. If the run declares everything, it can ask for a
p4d.24xlarge and nothing says no. Splitting into a HostConfig does not fix
that, because the same hand writes both — and when that hand is an agent, it
writes both in the same edit.
A bound only works if it lives outside the file being edited. That is the
queue: it already carries providerRef and elasticity, it is set once rather
than per function, and a UDF cannot widen it from the inside. Permitted instance
types belong there, next to the spend limits. The distinction that matters is
not who writes it but what can rewrite it — a guardrail inside the generated
artifact is not a guardrail.
The filter does not become a type; it becomes the key function. Today it answers "which fields do I hide from the launcher," maintained by hand, with anything unlisted silently forwarded — so a typo becomes a launcher kwarg. As a key function it answers "which fields force a new host," which is a question with an observable right answer: get it wrong and you either over-launch or land a job on a worker that cannot serve it.
RunConfig is one class today, daemonTemplate is an untyped Json column
that duplicates part of it, and there is no derived host key — no
host_key() in the SDK, nothing on the wire, and nothing in the claim query.
The elasticity controller builds a launch body from daemonTemplate directly.
Neither host_setup nor run_setup exists, and _RUNTIME_FIELDS is still the
hand-maintained filter described above, with no readers at all.
The key function, the subset and the two worked values in The key function are the specification the three implementations will be written against, not a description of any of them.
Two mechanisms, different lifetimes
bootstrap_script | Setup commands | |
|---|---|---|
| When | Once, at instance creation | Any time, including long after join |
| Where declared | daemonTemplate.bootstrap_script | daemonTemplate.setup_scripts, --setup-script, or the HTTP route |
| Re-runnable | No — new instance only | Yes, append whenever |
| Failure blast radius | The instance never joins | One command; the worker stays |
bootstrap_script is the cloud-init-shaped one: it belongs to the machine, not
to the queue's ongoing life. Setup commands are the opposite — a durable queue
the daemon drains, which is what makes them appendable.
Setup runs per worker, not once per queue. An elastic queue that scales to
eight workers runs the same install eight times, and pays it again after every
scale-to-zero. Anything stable belongs in image_id; setup is for what
genuinely varies per launch.
How the FIFO drains
The queue is fire-and-forget, and the daemon pulls rather than being pushed:
The state flip is conservative. A worker in active moves to setting_up;
pending and joining are left alone because they reach active naturally and
will pick the commands up then, and the terminal-ish states — draining,
stale, gone — are never flipped silently.
refresh_capabilities: true is the flag that matters when the script changes
what the worker can do. Installing CUDA without it leaves the control plane
still believing the worker is CPU-only.
The setting_up flip is the only readiness signal a worker has today, and it
only covers setup the control plane itself queued — setup the daemon runs on
its own initiative is invisible to it. Worker
lifecycle is where
that gap, and the phase field that closes it, are worked out.
The interface
Python
There is none. Setup is operator-side — declared on the queue's
daemonTemplate or pushed at a worker — and a @udf cannot request it.
CLI
Seed setup at launch:
--setup-script is repeatable and takes a path; the file's contents become the
command. Scripts declared in the project config under daemon.setup_scripts are
read first and the CLI flags append after them, so config supplies the baseline
and the flag adds to it. --no-setup drops both — the config-declared scripts
and the flags — which makes it the way to launch a deliberately bare worker.
HTTP
A malformed body is a 400; an unknown worker id is a 404, and so is a worker in another namespace — the same posture the daemon GET handler takes, so a probe cannot distinguish "does not exist" from "not yours".
The companion path is the launch route, which seeds the same FIFO at creation: