DreamLake

Worker lifecycle

The phase field, the bye route, the reported in-flight count, host_key matching and setup-result reporting are all built — this page describes the daemon you can run today. Four narrower things are still proposal; the Proposal callout at the end is the exact list.

This page is about workers, not invocations — an invocation's queued → running → succeeded machine lives on the task queue page and is a different thing entirely. Here: how a nymph daemon comes up, when it is willing to be given work, and what it says on the way out.

The three machines that already exist

There is no shortage of state. A running fleet already carries three separate machines, each answering a different question:

MachineWhere it livesValuesWho writes it
Worker.stateControl planepending, joining, active, setting_up, draining, stale, goneThe control plane, from hello, poll and terminate
Worker.hostStatusControl planepending, running, stopped, terminatedThe launcher — EC2, GCE, Kube — out of band
The daemon's run stateInside nymphStarting, Polling, Backoff, Hibernating, DrainingThe daemon; visible only on the loopback introspect socket

Worker.state is computed on read, not stored for its derived values: a terminal state is never relabelled, a hostStatus of stopped or terminated becomes gone, and a worker unheard-from for longer than the 60-second stale window becomes stale. The daemon's own run state never crosses the wire — /status binds loopback only and refuses a routable address by design, so it can never be a remote probe.

The problem is not a missing machine. It is that the only thing crossing the wire is the absence of a poll. So this design adds one field, and makes the control-plane machine a projection of it. It does not add a fourth machine.

The phase field

DaemonPhasets
type DaemonPhase =
  | 'booting'      // pre-hello. Never observed by the control plane, by definition
  | 'setup'        // running host_setup or a setup command. Heartbeat yes, claim no
  | 'ready'        // polling and accepting work
  | 'draining'     // shutdown started. No new claims; inflight finishing
  | 'terminating'  // grace expired. About to exit

interface PollRequest {
  // …existing fields…
  phase?:    DaemonPhase   // absent on an old daemon — the CP falls back to inference
  host_key?: string        // see Host Setup → The key function
}

Both fields are optional in both directions on purpose. An old daemon sends neither and behaves exactly as it does today; an old control plane ignores both.

Projection onto Worker.state

phaseWorker.state the poll handler writes
absentToday's behaviour — joining → active on the first poll
setupsetting_up
readyactive
drainingdraining
terminatingdraining — the bye that follows writes gone

The projection never widens the conservative state flips documented under How the FIFO drains — a reported phase can move a worker the existing rules already allow to move, and nothing more.

Readiness is a claim, not an inference

Today "ready" means "a poll arrived." Locally the daemon flips to Polling on its first successful poll; on the control plane the worker flips joining → active on the first poll received. Both are liveness, not readiness — the daemon reaches that point before it has verified its workdir, before any declared setup has run, and, under the derived host key, before it even knows which key it will advertise.

The daemon has no way to say I am up, keep counting me alive, but do not give me work. It cannot withhold the poll, because withholding the poll is how it dies.

phase: setup is exactly that sentence: heartbeat yes, claim no. The poll handler skips the claim and returns an empty command list.

This matters the moment host_setup is derived from a RunConfig rather than pushed by the control plane. The existing guard short-circuits the claim loop only when the control plane's own pending-setup FIFO is non-empty — it protects setup the control plane queued. A host_setup the daemon runs on its own initiative is invisible to that guard, so without phase the worker is handed work mid-install.

It is deliberately not an authorization boundary. A buggy or hostile daemon can assert a ready it has not earned. That is still strictly better than the status quo, where nothing is asserted at all and readiness is inferred from a TCP round-trip.

The state machine

                     (process start)
                            │
                     ┌──────▼──────┐
                     │   booting   │  config load, host detection,
                     └──────┬──────┘  introspect bind — no wire contact
                            │  POST /v1/daemon/hello  (bounded retry)
             ┌──────────────┴──────────────┐
             │ rejected, or retries        │ 200 → worker_id
             ▼ exhausted                   ▼
         (exit 1)                    ┌───────────┐
                                     │   setup   │  host_setup, and the
                     ┌──────────────▶│           │  control-plane setup FIFO
      a setup command│               └─────┬─────┘  CP: setting_up, no claims
      arrives        │                     │  FIFO empty && workdir writable
                     │                     ▼
                     │               ┌───────────┐
                     └───────────────┤   ready   │  claims and runs work
                                     │           │  CP: active
                                     └─────┬─────┘
                                           │
        ┌──────────────────────────────────┼──────────────────────────────┐
        │ SIGTERM / SIGINT                 │ keep_alive_s elapsed         │
        │                                  │            CP drain command  │
        └──────────────────────────────────┼──────────────────────────────┘
                                           ▼
                                     ┌───────────┐
                                     │  draining │  poll loop stopped; inflight
                                     │           │  finishing inside grace_s
                                     └─────┬─────┘  CP: draining
                                           │  grace expired, work still running
                                           ▼
                                   ┌──────────────┐
                                   │ terminating  │  SIGTERM to children, wait
                                   └──────┬───────┘  kill_after_s, then SIGKILL
                                          │  POST /v1/daemon/bye  (best effort)
                                          ▼
                                      (exit 0)       CP: gone

booting is never observed remotely — by definition, the daemon has not spoken yet. Everything else is a phase the control plane can read off a poll.

Termination — four cases, one rule

The daemon reports every exit it chooses; the control plane infers every exit that is done to it.

That one sentence decides the whole matrix. An idle exit, a signal and a commanded drain all leave the daemon alive enough to speak, so it speaks. A spot reclaim, a crash and a partition do not, so the control plane is left to notice.

CaseWho knows firstDaemon reportsControl plane mechanismWorst-case detection
Idle exit (keep_alive_s elapsed)The daemonbye{reason: "idle"}—Immediate
SIGTERM / SIGINTThe OSbye{reason: "signal"}Falls back to stale if the POST failsImmediate, else 60 s
Control-plane drainThe control planephase: draining, then bye{reason: "drain"}It issued the commandImmediate
Launcher terminate / spot reclaimThe launcherNothing. The power is gonehostStatus → terminated ⇒ goneNext host poll, else 60 s
Crash, OOM, kernel panicNobodyNothingstale60 s
Network partitionNobodyNothing — the POSTs fail toostale60 s

The last three rows are why bye is not a reliability mechanism. The 60-second stale timeout stays exactly as it is and remains the backstop for every case. bye is a latency mechanism: it turns a 60-second ambiguity into a zero-second certainty for the three cases where the daemon is still able to speak. Nothing here is allowed to depend on the message arriving.

POST /v1/daemon/bye

POST /v1/daemon/byejson
{
  "worker_id": "wrk_01JD2K…",
  "reason": "signal",
  "inflight_abandoned": 0
}
204 No Content

reason is one of idle, signal, drain, hello_failed. The control plane sets state: "gone" and lastSeenAt: now, and does nothing else.

The constraints on the client side are the interesting part, and they are non-negotiable:

  • One attempt, two-second timeout. A failure is logged at debug and ignored.
  • It never delays exit. A shutdown that hangs because it was being polite is worse than the ambiguity it was trying to remove.
  • It is sent after the drain returns, so inflight_abandoned is a real number rather than a guess.

What crosses the wire

CallWhenCarriesWhat it changes
POST /v1/daemon/helloOnce, at bootmachine_id, label, tags, lanes, capabilities, versions, host_keyCreates or updates the Worker row; state = joining
POST /v1/daemon/pollEvery cycleworker_id, queue, timeout_s, invocations, phase, host_key, updated_capabilities including current_invocations, setup_resultsBumps lastSeenAt; sets state from phase; updates capabilities; gates the claim on phase and host_key
POST /v1/daemon/ackOnce per runinvocation_id, state, result_blob, error, worker_idThe invocation reaches a terminal state
POST /v1/daemon/byeOnce, at exitworker_id, reason, inflight_abandonedstate = gone, lastSeenAt = now

current_invocations is the row with a consequence outside this page: it is what closes the structurally-zero utilization gap described under Scale to zero. Sending the count on every poll closes it without a control-plane change, because the five call sites already read the key.

The same count fixes a second bug for free. The daemon tracks in-flight work in four different places and only one of them covers exec as well as run, so a worker busy with a long exec can look idle and time itself out mid-job. One counter, incremented at both dispatch sites, is what the idle gate should have been reading all along.

Why capabilities on every poll

The alternative was a new sibling field on the poll. Reusing updated_capabilities needs zero control-plane change — that key is already consumed and already written through to the worker row. The price is one redundant field write per poll, on a document that is already written every poll to bump lastSeenAt.

The interface

Python

The hooks run on the worker, not on the submitter — they are the daemon-side half of the run declaration.

lifecycle_upload.pypython
import dreamlake.lakeshore as dls

MODEL: dict[str, object] = {}

@dls.lifecycle.on_host_setup          # once per worker process, before the first body
def warm() -> None:
    MODEL["weights"] = load("checkpoint.pt")

@dls.lifecycle.on_run_setup           # before every body, in the invocation workdir
def prep() -> None:
    scratch.mkdir(exist_ok=True)

@dls.lifecycle.on_shutdown            # inside the grace window — flush, then return
def flush() -> None:
    dls.run.publish("./out", name="checkpoint")

dls.lifecycle.phase()      # 'setup' | 'ready' | 'draining'
dls.lifecycle.host_key()   # 'hk1_…' — the key this worker was configured with, or None

The hooks take no arguments. There is no context object: a hook reads and writes ordinary module state — a global, an open handle — which is exactly the thing a shell host_setup cannot hand to the next body. Environment and working directory come from the RunConfig, not from a parameter.

phase() returns a three-value subset of the five-value DaemonPhase on the wire. booting and terminating have no Python spelling by construction: the interpreter is not running yet during booting, and by terminating the grace window has already expired, so no hook is still entitled to run.

on_host_setup and on_run_setup are the two phases named on the host setup page, reachable from Python rather than from a queue template. on_shutdown runs inside the grace window and must return before it expires — a hook that outlives the grace is killed with everything else, which is the whole reason the escalation ends in SIGKILL. The runnable version of all three, including the upload, is examples/lifecycle_upload.py in lakeshore-py.

What does not ship yet

Proposal

phase, host_key, current_invocations, setup_results and POST /v1/daemon/bye are all on the wire, and claimOne filters on the host key. Four things on or adjacent to this page are still design only:

  • Nothing launches on an unmatched host key. Matching ships: a worker advertising a different key is not handed the invocation. But an invocation whose key matches no live worker simply stays queued — provisioning one is the next slice.
  • A RunConfig's host_setup is hashed but never executed. The daemon runs the host_setup in its own config at boot, and hashes the one in an incoming run config into the key it matches against — so a declared host_setup selects a host that already has it, and never installs it. Only run_setup is executed from a run config, per invocation, before the body.
  • ExecBody carries no run_config, so the exec path — what the CLI and SDK drive today — has no host key and no placement story.
  • Setup failures are logged, not tabulated. A non-success is written to the control plane's log and the most recent one is kept on the worker row; there is no per-command status table to read a history out of.
Host Setup →

What runs before a worker takes work, and the derived host key that decides which worker it may be.

Elastic queues →

Where the utilization number is read, and why it currently understates load.

Local dev loop →

Run a daemon on your own machine and watch these phases happen.