# Host Setup

For enrolling a machine and starting nymph, see [Enroll a host](/lakeshore/hosts/enroll.md).
This page covers runtime environment preparation before or between workloads.

## The type

Setup arrives in two places. The first is on the queue's `daemonTemplate`, read
at launch time:

```ts file="daemonTemplate (setup fields)"
interface DaemonTemplate {
  bootstrap_script?: string     // runs once, at instance creation
  setup_scripts?:    string[]   // seeds the setup-command FIFO
  // …instance_type, image_id, runner — see the Queue model
}
```

The second is the command itself, which is what the daemon actually executes:

```ts file="SetupCommand"
interface SetupCommand {
  script:                string    // required
  sudo?:                 boolean   // default false
  refresh_capabilities?: boolean   // default false — re-report caps after running
  name?:                 string    // optional human label
}
```

`setup_scripts` on the template is mapped into `setup_commands` on the launch
body by the elasticity controller. Both routes below append to the same FIFO.

  A worker is rarely useful the moment it boots. **Host setup** is the ordered
  list of commands it runs before — and, when you append later, between — units
  of work: installing a driver, warming a cache, mounting something by hand.

## Coming from a jaynes config

Jaynes puts setup on the **runner**, and a runner is constructed with the mount
list — `mounts` is the first parameter on every runner class, "reserved by
jaynes to pass in the mount objects". That wiring is what lets the rest of the
config point at it:

```yaml file=".jaynes.yml"
runners:
  - !runners.Tmux &tmux_runner
    setup: >                              # on the host, before the session
      source $HOME/.jaynes_bash;
    startup: >                            # inside the session, before your code
      conda activate torch;
    envs: >
      LANG=utf-8 LC_CTYPE=en_US.UTF-8
    pypath: "{mounts[0].host_path}"       # both derived from the mount
    work_dir: "{mounts[0].host_path}"
    entry_script: "python -u -m jaynes.entry"
    post_script: "tmux list-sessions"
```

### Three phases, not one

The field to be careful about is `setup` versus `startup` — they are different
phases, and only the first is host setup:

| jaynes | Runs | Lakeshore |
| --- | --- | --- |
| `setup` | **On the host**, to prepare the environment | `host_setup` (shipped as `setup_scripts`) / `--setup-script` / the setup FIFO — **this page** |
| `startup` | **Inside the session**, before your code is bootstrapped | **Nothing.** [`run_setup`](#naming-the-phases) is the proposed name for it; the only lever today is baking it into `image_id` |
| `entry_script` | The command that runs your payload | The worker importing `module:qualname`, or `command` on the [exec wire](/lakeshore/command-programs.md) |
| `post_script` | After the run script | **Nothing** — nothing runs after work completes |

Jaynes has three phases where Lakeshore has two, and they do not line up end to
end. Lakeshore's `bootstrap_script` and setup commands are both *host* phases —
one at instance creation, one at any point after. Everything jaynes does
per-session, between the host being ready and the payload starting, has no
Lakeshore equivalent at all.

### The rest

| jaynes | What it expresses | Lakeshore |
| --- | --- | --- |
| `mounts` | The runner is handed the mount objects, so other fields can reference them | **Nothing.** No queue, mode or `RunConfig` field names a mount; [`mounts=` on the decorator](/lakeshore/access.md#mounts) is the proposal that would |
| `work_dir` | Working directory, conventionally `{mounts[0].host_path}` | Fixed at **`/workspace`**; not derived from a mount and not settable per run |
| `pypath` | Add the mount's path to `PYTHONPATH` | Automatic for code — the worker inserts the extracted snapshot on `ImportError`. Nothing for data mounts |
| `envs` | Environment for the session | `RunConfig.env`, per invocation rather than per host |

The `work_dir` and `pypath` rows are the same fact twice: **jaynes derives paths
from mounts; Lakeshore fixes them by convention.** `{mounts[0].host_path}`
interpolation exists because a jaynes mount chooses its own location, so
everything downstream has to ask it where that was. `/workspace` and
`/mounts/<name>` are fixed, so nothing has to be interpolated — and nothing
*can* be, which is the cost.

## Naming the phases

Three phases, and the names should say **what scope they run at**, because that
is the only thing anyone needs to know to place a script correctly:

| Phase | Proposed name | Declared on | Runs |
| --- | --- | --- | --- |
| Instance creation | `bootstrap_script` | `daemonTemplate` | Once per machine, before the daemon exists |
| Host | **`host_setup`** | `daemonTemplate` | On a live worker, at launch or any time after |
| Invocation | **`run_setup`** | `RunConfig` | Before each body, every time |

`host_setup` replaces `setup_scripts`, which named the *type* of the value
rather than the phase — and stopped being accurate once the same list could be
appended to a running worker.

**`run_setup` is the answer to what jaynes calls `startup`.** Three reasons it
beats carrying that name over:

- **`startup` does not say startup of what.** Read cold, it sounds like the
  machine coming up, which is the phase *above* it. `host_setup` / `run_setup`
  name their scope, and the pair reads as one axis.
- **The name says where the field lives.** `RunConfig` is already the
  per-invocation config object — the thing carrying `env`, `image`, `runner`,
  `timeout_s`. A per-invocation script belongs there and nowhere else, and
  `run_setup` puts it there by name. `host_setup` sits on `daemonTemplate` for
  the same reason.
- **It parallels the phase it is most confused with.** The mistake people
  actually make is running a per-host install on every invocation, or expecting a
  per-run tweak to persist. Two names differing in one word make the question
  "per host or per run?" impossible to skip.

There is a sharper argument than any of those three, and it only becomes visible
once the [host key](#the-key-function) exists: **the pair is exactly "in the key"
and "not in the key."** `host_setup` participates in the hash, so declaring
different setup means a different host. `run_setup` cannot, because it runs
before every body by definition and a phase that runs every time can never force
a *different* machine. The two names are not a stylistic pairing — they are the
two sides of one derived predicate.

Rejected alternatives: `warmup` collides with
[preloading](/lakeshore/preloaded.md), which caches *across* invocations rather
than running before each one — the opposite property. `pre_run` says when but
not what. `invocation_setup` is exact but long, and `run` is already the word
this SDK uses for the ambient per-invocation context (`dls.run.read`,
`run_worker`, `RunConfig`).

> **Warning:** `RunConfig.host_setup` and `RunConfig.run_setup` have shipped, and
>   `daemonTemplate.host_setup` is accepted as an *alias* for `setup_scripts`.
>   Nothing was renamed — `setup_scripts` keeps working and stays the written
>   spelling, because `daemonTemplate` is an untyped `Json` column with no schema,
>   written by three separate CLI paths and read by the elasticity controller, so
>   a rename would be an unvalidated data migration across independent release
>   trains. Accepting both spellings costs one fallback.
>
>   What is declared and what runs are still different things. `host_setup` on a
>   `RunConfig` is **hashed into the [host key](#one-config-not-two--and-a-derived-host-key)
>   but never executed** — the daemon runs the `host_setup` in its own config at
>   boot. `run_setup` has no execution path at all yet. So both fields today
>   affect *placement*, not behaviour.

## One config, not two — and a derived host key

The obvious move is to split `RunConfig` into a `HostConfig` and a `RunConfig`,
because the seam is already in the code: the docstring joins two things with the
word "plus" — *"per-machine launcher kwargs (`instance_type`, `partition`, …)
**plus** the runtime-side knobs"* — and `_RUNTIME_FIELDS` is a hand-maintained
set commented "not forwarded to the launcher."

**That is the wrong conclusion, and ephemerality is why.** If a host that does
not match can simply be relaunched, then a host configuration is not a peer of
the run configuration. It is a *consequence* of it. Nobody should be writing one.

So: **one declaration, written by the author. The host requirement is derived
from it.**

| | Written by | What it is |
| --- | --- | --- |
| **Run declaration** | Whoever writes the function — increasingly an agent | `image`, `resources`, `env`, `timeout_s`, `runner`, `host_setup`, `run_setup`, `source` / `target` / `mounts` |
| **Host key** | Nobody — computed | A hash of the subset that affects placement. Two runs with the same key can share a worker; a different key means launch |
| **Fleet policy** | Set once, outside the code being written | `providerRef`, `elasticity`, `min` / `max`, what instance types are permitted at all. Lives on the [queue](/lakeshore/queues/elastic.md) |

The middle row is the whole idea. `host_setup` does not need to live on a
separate object to be shared — it needs to be *part of the key*. Two functions
declaring the same setup hash the same and land on the same warm worker; one
that declares different setup gets its own. The boundary between "host" and
"run" stops being a typing decision and becomes a **caching decision**, which is
what it always was.

That also explains the warm/cold behaviour better than a split does.
[Preloaded functions](/lakeshore/preloaded.md) survive because the worker process
survives; the host key is exactly the predicate for "may this invocation land on
that process."

> **Note:** The classic case for splitting was that operators and authors are different
>   people with different review. When an agent writes the function, they are the
>   same hand — so the split buys no separation and costs a second object to keep
>   consistent.
>
>   That cost is the real one. `daemonTemplate` duplicating part of `RunConfig`
>   is a sync burden today; asking an agent to maintain both, in step, across
>   every generated function, is a class of drift with no upside. A derived key
>   cannot drift from the declaration it is computed from.

### The key function

"A hash of the subset that affects placement" is only useful if the subset is
written down, because three languages have to compute the same string. It is:

```
host_key = "hk1_" + base32-lower-nopad( sha256( canonical_json(subset) )[0..10] )
```

The subset is exactly seven keys, always all seven, in every case:

| Key | Type | Value when absent |
| --- | --- | --- |
| `v` | integer | Always `1` |
| `runner` | string \| null | `null` |
| `image` | string \| null | `null` |
| `resources` | object, string → int \| string | `{}` |
| `host_setup` | array of string | `[]` |
| `provider` | string \| null | `null` |
| `extras` | object | `{}` |

Canonical JSON means UTF-8, no whitespace, and object keys sorted by code point
**recursively**. **Absent, empty and null are the same input** — there is exactly
one representation of "not set," which is what makes the hash stable across three
implementations rather than merely deterministic within one. Floats are a hard
error rather than a silent coercion: `resources` values are `int` or `str`,
because IEEE-754 round-tripping through msgpack and three JSON serializers is not
something to bet a cache key on.

Two things are asymmetric on purpose. **`host_setup` order is significant** — the
array is not sorted, because running `pip install torch` before `apt-get install
ffmpeg` is a different host preparation, and a worker prepared one way must not
serve a job that declared the other. **`resources` key order is not** — it is a
mapping, so it is normalised.

The exclusions each carry a reason:

| Excluded | Why |
| --- | --- |
| `env` | Applied by the runner at exec time, per invocation. It does not change what the machine *is* — and if it did, every distinct variable would be a cold start |
| `run_setup` | Runs per invocation, before each body, by definition. A phase that runs every time cannot force a *different* host |
| `timeout_s` | A longer deadline does not need a different machine |
| `tags` | An operator concern, already matched separately by the pool spec. Folding it in would double-count |

Ten bytes is eighty bits, which is exactly sixteen base32 characters with no
padding — fixed-width, lowercase, URL-safe, and short enough to sit in a table
without wrapping. `hk1_` is a version prefix: changing the subset or the encoding
bumps it to `hk2_`, so an old key never silently means something new.

Two worked values, so an implementation can be checked against something:

```
{}                      →  hk1_ljwqko7vjw2cnyny
{"runner": "process"}   →  hk1_jjml3vzthmtzu76a
```

Python computes the key at submit — that is the only place the declaration
exists, so it is authoritative. The daemon recomputes it from the config it was
launched with and advertises it on hello and every poll. **The control plane
computes nothing**; it stores what the SDK sent, receives what the daemon
advertised, and compares two strings. A hash comparison the control plane does
not participate in cannot drift from the hash function.

Two independent implementations of the same function are the obvious risk, and the
only real mitigation is a shared fixture: one golden-vector file, copied verbatim
into every repo that implements the key, with a test in each that reads it. An
edited copy fails immediately, which turns a certainty of drift into a visible
event.

### What the split got right, and where it still applies

Two objections survive, and both are about hosts that are *not* ephemeral:

**Not every worker can be relaunched.** `providerRef: null` is the
bring-your-own-worker case — a lab machine, a SLURM allocation you already hold,
`fixed` elasticity. There the host predates the run and cannot be re-derived, so
a declaration that does not match is an error rather than a launch. The derived
key still works: it just fails to find a match instead of provisioning one.

**Cost needs a bound.** If the run declares everything, it can ask for a
`p4d.24xlarge` and nothing says no. Splitting into a `HostConfig` does not fix
that, because the same hand writes both — and when that hand is an agent, it
writes both in the same edit.

A bound only works if it lives **outside the file being edited**. That is the
queue: it already carries `providerRef` and `elasticity`, it is set once rather
than per function, and a UDF cannot widen it from the inside. Permitted instance
types belong there, next to the spend limits. The distinction that matters is
not who writes it but *what can rewrite it* — a guardrail inside the generated
artifact is not a guardrail.

> **Note:** The filter does not become a type; it becomes the **key function**. Today it
>   answers "which fields do I hide from the launcher," maintained by hand, with
>   anything unlisted silently forwarded — so a typo becomes a launcher kwarg. As
>   a key function it answers "which fields force a new host," which is a question
>   with an observable right answer: get it wrong and you either over-launch or
>   land a job on a worker that cannot serve it.

> **Warning:** `RunConfig` is one class today, `daemonTemplate` is an untyped `Json` column
>   that duplicates part of it, and there is **no derived host key** — no
>   `host_key()` in the SDK, nothing on the wire, and nothing in the claim query.
>   The elasticity controller builds a launch body from `daemonTemplate` directly.
>   Neither `host_setup` nor `run_setup` exists, and `_RUNTIME_FIELDS` is still the
>   hand-maintained filter described above, with no readers at all.
>
>   The key function, the subset and the two worked values in [The key
>   function](#the-key-function) are the specification the three implementations
>   will be written against, not a description of any of them.

## Two mechanisms, different lifetimes

| | `bootstrap_script` | Setup commands |
| --- | --- | --- |
| When | Once, at instance creation | Any time, including long after join |
| Where declared | `daemonTemplate.bootstrap_script` | `daemonTemplate.setup_scripts`, `--setup-script`, or the HTTP route |
| Re-runnable | No — new instance only | Yes, append whenever |
| Failure blast radius | The instance never joins | One command; the worker stays |

`bootstrap_script` is the cloud-init-shaped one: it belongs to the machine, not
to the queue's ongoing life. Setup commands are the opposite — a durable queue
the daemon drains, which is what makes them appendable.

> **Note:** Setup runs **per worker**, not once per queue. An elastic queue that scales to
>   eight workers runs the same install eight times, and pays it again after every
>   scale-to-zero. Anything stable belongs in `image_id`; setup is for what
>   genuinely varies per launch.

## How the FIFO drains

The queue is fire-and-forget, and the daemon pulls rather than being pushed:

```
POST …/workers/:id/setup   →  append to the worker's pending-setup FIFO
                              worker flips active → setting_up
POST /v1/daemon/poll       →  pops the head of the FIFO, one per poll
                              daemon runs it, reports back via
                              updated_capabilities on the next poll
```

The state flip is conservative. A worker in `active` moves to `setting_up`;
`pending` and `joining` are left alone because they reach `active` naturally and
will pick the commands up then, and the terminal-ish states — `draining`,
`stale`, `gone` — are never flipped silently.

`refresh_capabilities: true` is the flag that matters when the script changes
what the worker can do. Installing CUDA without it leaves the control plane
still believing the worker is CPU-only.

The `setting_up` flip is the only readiness signal a worker has today, and it
only covers setup the **control plane itself queued** — setup the daemon runs on
its own initiative is invisible to it. [Worker
lifecycle](/lakeshore/lifecycle.md#readiness-is-a-claim-not-an-inference) is where
that gap, and the phase field that closes it, are worked out.

## The interface

### Python

There is none. Setup is operator-side — declared on the queue's
`daemonTemplate` or pushed at a worker — and a `@udf` cannot request it.

### CLI

Seed setup at launch:

```bash file="terminal"
lakeshore daemon launch --queue gpu \
    --setup-script ./install-cuda.sh \
    --setup-script ./warm-cache.sh

lakeshore daemon launch --queue gpu --no-setup
```

`--setup-script` is repeatable and takes a path; the file's contents become the
command. Scripts declared in the project config under `daemon.setup_scripts` are
read first and the CLI flags append after them, so config supplies the baseline
and the flag adds to it. `--no-setup` drops both — the config-declared scripts
*and* the flags — which makes it the way to launch a deliberately bare worker.

### HTTP

```json file="POST /v1/namespaces/:ns/workers/:id/setup"
{
  "script": "apt-get install -y ffmpeg",
  "sudo": true,
  "refresh_capabilities": false,
  "name": "install-ffmpeg"
}
```

```json
202 { "commandId": "cmd_01JD2K…" }
```

A malformed body is a 400; an unknown worker id is a 404, and so is a worker in
another namespace — the same posture the daemon GET handler takes, so a probe
cannot distinguish "does not exist" from "not yours".

The companion path is the launch route, which seeds the same FIFO at creation:

```
POST /v1/namespaces/:ns/daemons/launch    { …, setup_commands: SetupCommand[] }
```

    Where `daemonTemplate` lives, and why setup cost is paid per worker.

    The other half of preparing a host — and what a jaynes mount maps onto.

    When a worker is willing to be given work, and how it says so — the phase
    the setup FIFO is really gating.
