# Worker lifecycle

The `phase` field, the `bye` route, the reported in-flight count, `host_key`
matching and setup-result reporting are all **built** — this page describes the
daemon you can run today. Four narrower things are still proposal; the
[Proposal callout](#what-does-not-ship-yet) at the end is the exact list.

  This page is about **workers**, not invocations — an invocation's `queued →
  running → succeeded` machine lives on the [task
  queue](/lakeshore/queues/task-queue.md#the-lifecycle) page and is a different
  thing entirely. Here: how a `nymph` daemon comes up, when it is willing to be
  given work, and what it says on the way out.

## The three machines that already exist

There is no shortage of state. A running fleet already carries three separate
machines, each answering a different question:

| Machine | Where it lives | Values | Who writes it |
| --- | --- | --- | --- |
| `Worker.state` | Control plane | `pending`, `joining`, `active`, `setting_up`, `draining`, `stale`, `gone` | The control plane, from hello, poll and terminate |
| `Worker.hostStatus` | Control plane | `pending`, `running`, `stopped`, `terminated` | The launcher — EC2, GCE, Kube — out of band |
| The daemon's run state | Inside `nymph` | `Starting`, `Polling`, `Backoff`, `Hibernating`, `Draining` | The daemon; visible only on the loopback introspect socket |

`Worker.state` is **computed on read**, not stored for its derived values: a
terminal state is never relabelled, a `hostStatus` of `stopped` or `terminated`
becomes `gone`, and a worker unheard-from for longer than the 60-second stale
window becomes `stale`. The daemon's own run state **never crosses the wire** —
`/status` binds loopback only and refuses a routable address by design, so it can
never be a remote probe.

The problem is not a missing machine. It is that the only thing crossing the wire
is **the absence of a poll**. So this design adds **one field**, and makes the
control-plane machine a projection of it. It does not add a fourth machine.

## The phase field

```ts file="DaemonPhase"
type DaemonPhase =
  | 'booting'      // pre-hello. Never observed by the control plane, by definition
  | 'setup'        // running host_setup or a setup command. Heartbeat yes, claim no
  | 'ready'        // polling and accepting work
  | 'draining'     // shutdown started. No new claims; inflight finishing
  | 'terminating'  // grace expired. About to exit

interface PollRequest {
  // …existing fields…
  phase?:    DaemonPhase   // absent on an old daemon — the CP falls back to inference
  host_key?: string        // see Host Setup → The key function
}
```

Both fields are optional in both directions on purpose. An old daemon sends
neither and behaves exactly as it does today; an old control plane ignores both.

### Projection onto `Worker.state`

| `phase` | `Worker.state` the poll handler writes |
| --- | --- |
| *absent* | Today's behaviour — `joining → active` on the first poll |
| `setup` | `setting_up` |
| `ready` | `active` |
| `draining` | `draining` |
| `terminating` | `draining` — the `bye` that follows writes `gone` |

The projection never widens the conservative state flips documented under [How
the FIFO drains](/lakeshore/host-setup.md#how-the-fifo-drains) — a reported `phase`
can move a worker the existing rules already allow to move, and nothing more.

## Readiness is a claim, not an inference

Today "ready" means "a poll arrived." Locally the daemon flips to `Polling` on
its first successful poll; on the control plane the worker flips `joining →
active` on the first poll received. Both are **liveness**, not readiness — the
daemon reaches that point before it has verified its workdir, before any declared
setup has run, and, under the derived host key, before it even knows which key it
will advertise.

The daemon has no way to say *I am up, keep counting me alive, but do not give me
work.* It cannot withhold the poll, because withholding the poll is how it dies.

`phase: setup` is exactly that sentence: **heartbeat yes, claim no.** The poll
handler skips the claim and returns an empty command list.

This matters the moment `host_setup` is [derived from a
`RunConfig`](/lakeshore/host-setup.md#one-config-not-two--and-a-derived-host-key)
rather than pushed by the control plane. The existing guard short-circuits the
claim loop only when the control plane's *own* pending-setup FIFO is non-empty —
it protects setup the control plane queued. A `host_setup` the daemon runs on its
own initiative is invisible to that guard, so without `phase` the worker is
handed work mid-install.

It is deliberately **not** an authorization boundary. A buggy or hostile daemon
can assert a `ready` it has not earned. That is still strictly better than the
status quo, where nothing is asserted at all and readiness is inferred from a TCP
round-trip.

## The state machine

```
                     (process start)
                            │
                     ┌──────▼──────┐
                     │   booting   │  config load, host detection,
                     └──────┬──────┘  introspect bind — no wire contact
                            │  POST /v1/daemon/hello  (bounded retry)
             ┌──────────────┴──────────────┐
             │ rejected, or retries        │ 200 → worker_id
             ▼ exhausted                   ▼
         (exit 1)                    ┌───────────┐
                                     │   setup   │  host_setup, and the
                     ┌──────────────▶│           │  control-plane setup FIFO
      a setup command│               └─────┬─────┘  CP: setting_up, no claims
      arrives        │                     │  FIFO empty && workdir writable
                     │                     ▼
                     │               ┌───────────┐
                     └───────────────┤   ready   │  claims and runs work
                                     │           │  CP: active
                                     └─────┬─────┘
                                           │
        ┌──────────────────────────────────┼──────────────────────────────┐
        │ SIGTERM / SIGINT                 │ keep_alive_s elapsed         │
        │                                  │            CP drain command  │
        └──────────────────────────────────┼──────────────────────────────┘
                                           ▼
                                     ┌───────────┐
                                     │  draining │  poll loop stopped; inflight
                                     │           │  finishing inside grace_s
                                     └─────┬─────┘  CP: draining
                                           │  grace expired, work still running
                                           ▼
                                   ┌──────────────┐
                                   │ terminating  │  SIGTERM to children, wait
                                   └──────┬───────┘  kill_after_s, then SIGKILL
                                          │  POST /v1/daemon/bye  (best effort)
                                          ▼
                                      (exit 0)       CP: gone
```

`booting` is never observed remotely — by definition, the daemon has not spoken
yet. Everything else is a phase the control plane can read off a poll.

## Termination — four cases, one rule

**The daemon reports every exit it chooses; the control plane infers every exit
that is done to it.**

That one sentence decides the whole matrix. An idle exit, a signal and a
commanded drain all leave the daemon alive enough to speak, so it speaks. A spot
reclaim, a crash and a partition do not, so the control plane is left to notice.

| Case | Who knows first | Daemon reports | Control plane mechanism | Worst-case detection |
| --- | --- | --- | --- | --- |
| Idle exit (`keep_alive_s` elapsed) | The daemon | `bye{reason: "idle"}` | — | Immediate |
| SIGTERM / SIGINT | The OS | `bye{reason: "signal"}` | Falls back to `stale` if the POST fails | Immediate, else 60 s |
| Control-plane drain | The control plane | `phase: draining`, then `bye{reason: "drain"}` | It issued the command | Immediate |
| Launcher terminate / spot reclaim | The launcher | **Nothing.** The power is gone | `hostStatus → terminated` ⇒ `gone` | Next host poll, else 60 s |
| Crash, OOM, kernel panic | Nobody | Nothing | `stale` | 60 s |
| Network partition | Nobody | Nothing — the POSTs fail too | `stale` | 60 s |

The last three rows are why `bye` is **not a reliability mechanism**. The
60-second stale timeout stays exactly as it is and remains the backstop for every
case. `bye` is a **latency** mechanism: it turns a 60-second ambiguity into a
zero-second certainty for the three cases where the daemon is still able to
speak. Nothing here is allowed to depend on the message arriving.

## `POST /v1/daemon/bye`

```json file="POST /v1/daemon/bye"
{
  "worker_id": "wrk_01JD2K…",
  "reason": "signal",
  "inflight_abandoned": 0
}
```

```
204 No Content
```

`reason` is one of `idle`, `signal`, `drain`, `hello_failed`. The control plane
sets `state: "gone"` and `lastSeenAt: now`, and does nothing else.

The constraints on the client side are the interesting part, and they are
non-negotiable:

- **One attempt, two-second timeout.** A failure is logged at `debug` and
  ignored.
- **It never delays exit.** A shutdown that hangs because it was being polite is
  worse than the ambiguity it was trying to remove.
- **It is sent after the drain returns**, so `inflight_abandoned` is a real
  number rather than a guess.

## What crosses the wire

| Call | When | Carries | What it changes |
| --- | --- | --- | --- |
| `POST /v1/daemon/hello` | Once, at boot | `machine_id`, `label`, `tags`, `lanes`, `capabilities`, `versions`, **`host_key`** | Creates or updates the `Worker` row; `state = joining` |
| `POST /v1/daemon/poll` | Every cycle | `worker_id`, `queue`, `timeout_s`, `invocations`, **`phase`**, **`host_key`**, `updated_capabilities` including **`current_invocations`**, **`setup_results`** | Bumps `lastSeenAt`; sets `state` from `phase`; updates capabilities; gates the claim on `phase` and `host_key` |
| `POST /v1/daemon/ack` | Once per run | `invocation_id`, `state`, `result_blob`, `error`, `worker_id` | The invocation reaches a terminal state |
| `POST /v1/daemon/bye` | Once, at exit | `worker_id`, `reason`, `inflight_abandoned` | `state = gone`, `lastSeenAt = now` |

`current_invocations` is the row with a consequence outside this page: it is what
closes the structurally-zero utilization gap described under [Scale to
zero](/lakeshore/queues/elastic.md#scale-to-zero-and-the-cold-start-it-buys).
Sending the count on every poll closes it without a control-plane change,
because the five call sites already read the key.

The same count fixes a second bug for free. The daemon tracks in-flight work in
four different places and only one of them covers `exec` as well as `run`, so a
worker busy with a long `exec` can look idle and time itself out mid-job. One
counter, incremented at both dispatch sites, is what the idle gate should have
been reading all along.

> **Note:** The alternative was a new sibling field on the poll. Reusing
>   `updated_capabilities` needs **zero** control-plane change — that key is already
>   consumed and already written through to the worker row. The price is one
>   redundant field write per poll, on a document that is already written every poll
>   to bump `lastSeenAt`.

## The interface

### Python

The hooks run **on the worker**, not on the submitter — they are the daemon-side
half of the [run declaration](/lakeshore/host-setup.md).

```python file="lifecycle_upload.py"
import dreamlake.lakeshore as dls

MODEL: dict[str, object] = {}

@dls.lifecycle.on_host_setup          # once per worker process, before the first body
def warm() -> None:
    MODEL["weights"] = load("checkpoint.pt")

@dls.lifecycle.on_run_setup           # before every body, in the invocation workdir
def prep() -> None:
    scratch.mkdir(exist_ok=True)

@dls.lifecycle.on_shutdown            # inside the grace window — flush, then return
def flush() -> None:
    dls.run.publish("./out", name="checkpoint")

dls.lifecycle.phase()      # 'setup' | 'ready' | 'draining'
dls.lifecycle.host_key()   # 'hk1_…' — the key this worker was configured with, or None
```

The hooks take **no arguments**. There is no context object: a hook reads and
writes ordinary module state — a global, an open handle — which is exactly the
thing a shell `host_setup` cannot hand to the next body. Environment and working
directory come from the `RunConfig`, not from a parameter.

`phase()` returns a **three-value subset** of the five-value `DaemonPhase` on the
wire. `booting` and `terminating` have no Python spelling by construction: the
interpreter is not running yet during `booting`, and by `terminating` the grace
window has already expired, so no hook is still entitled to run.

`on_host_setup` and `on_run_setup` are the two phases named on the [host
setup](/lakeshore/host-setup.md#naming-the-phases) page, reachable from Python
rather than from a queue template. `on_shutdown` runs inside the grace window and
must return before it expires — a hook that outlives the grace is killed with
everything else, which is the whole reason the escalation ends in SIGKILL. The
runnable version of all three, including the upload, is
`examples/lifecycle_upload.py` in `lakeshore-py`.

## What does not ship yet

> **Warning:** `phase`, `host_key`, `current_invocations`, `setup_results` and
>   `POST /v1/daemon/bye` are all on the wire, and `claimOne` filters on the host
>   key. Four things on or adjacent to this page are still design only:
>
>   - **Nothing launches on an unmatched host key.** Matching ships: a worker
>     advertising a different key is not handed the invocation. But an invocation
>     whose key matches *no* live worker simply stays queued — provisioning one is
>     the next slice.
>   - **A `RunConfig`'s `host_setup` is hashed but never executed.** The daemon
>     runs the `host_setup` in *its own* config at boot, and hashes the one in an
>     incoming run config into the key it matches against — so a declared
>     `host_setup` selects a host that already has it, and never installs it. Only
>     `run_setup` is executed from a run config, per invocation, before the body.
>   - **`ExecBody` carries no `run_config`**, so the `exec` path — what the CLI and
>     SDK drive today — has no host key and no placement story.
>   - **Setup failures are logged, not tabulated.** A non-success is written to the
>     control plane's log and the most recent one is kept on the worker row; there
>     is no per-command status table to read a history out of.

    What runs before a worker takes work, and the derived host key that decides
    which worker it may be.

    Where the utilization number is read, and why it currently understates load.

    Run a daemon on your own machine and watch these phases happen.
