Worker lifecycle
The phase field, the bye route, the reported in-flight count, host_key
matching and setup-result reporting are all built — this page describes the
daemon you can run today. Four narrower things are still proposal; the
Proposal callout at the end is the exact list.
This page is about workers, not invocations — an invocation's queued → running → succeeded machine lives on the task
queue page and is a different
thing entirely. Here: how a nymph daemon comes up, when it is willing to be
given work, and what it says on the way out.
The three machines that already exist
There is no shortage of state. A running fleet already carries three separate machines, each answering a different question:
| Machine | Where it lives | Values | Who writes it |
|---|---|---|---|
Worker.state | Control plane | pending, joining, active, setting_up, draining, stale, gone | The control plane, from hello, poll and terminate |
Worker.hostStatus | Control plane | pending, running, stopped, terminated | The launcher — EC2, GCE, Kube — out of band |
| The daemon's run state | Inside nymph | Starting, Polling, Backoff, Hibernating, Draining | The daemon; visible only on the loopback introspect socket |
Worker.state is computed on read, not stored for its derived values: a
terminal state is never relabelled, a hostStatus of stopped or terminated
becomes gone, and a worker unheard-from for longer than the 60-second stale
window becomes stale. The daemon's own run state never crosses the wire —
/status binds loopback only and refuses a routable address by design, so it can
never be a remote probe.
The problem is not a missing machine. It is that the only thing crossing the wire is the absence of a poll. So this design adds one field, and makes the control-plane machine a projection of it. It does not add a fourth machine.
The phase field
Both fields are optional in both directions on purpose. An old daemon sends neither and behaves exactly as it does today; an old control plane ignores both.
Projection onto Worker.state
phase | Worker.state the poll handler writes |
|---|---|
| absent | Today's behaviour — joining → active on the first poll |
setup | setting_up |
ready | active |
draining | draining |
terminating | draining — the bye that follows writes gone |
The projection never widens the conservative state flips documented under How
the FIFO drains — a reported phase
can move a worker the existing rules already allow to move, and nothing more.
Readiness is a claim, not an inference
Today "ready" means "a poll arrived." Locally the daemon flips to Polling on
its first successful poll; on the control plane the worker flips joining → active on the first poll received. Both are liveness, not readiness — the
daemon reaches that point before it has verified its workdir, before any declared
setup has run, and, under the derived host key, before it even knows which key it
will advertise.
The daemon has no way to say I am up, keep counting me alive, but do not give me work. It cannot withhold the poll, because withholding the poll is how it dies.
phase: setup is exactly that sentence: heartbeat yes, claim no. The poll
handler skips the claim and returns an empty command list.
This matters the moment host_setup is derived from a
RunConfig
rather than pushed by the control plane. The existing guard short-circuits the
claim loop only when the control plane's own pending-setup FIFO is non-empty —
it protects setup the control plane queued. A host_setup the daemon runs on its
own initiative is invisible to that guard, so without phase the worker is
handed work mid-install.
It is deliberately not an authorization boundary. A buggy or hostile daemon
can assert a ready it has not earned. That is still strictly better than the
status quo, where nothing is asserted at all and readiness is inferred from a TCP
round-trip.
The state machine
booting is never observed remotely — by definition, the daemon has not spoken
yet. Everything else is a phase the control plane can read off a poll.
Termination — four cases, one rule
The daemon reports every exit it chooses; the control plane infers every exit that is done to it.
That one sentence decides the whole matrix. An idle exit, a signal and a commanded drain all leave the daemon alive enough to speak, so it speaks. A spot reclaim, a crash and a partition do not, so the control plane is left to notice.
| Case | Who knows first | Daemon reports | Control plane mechanism | Worst-case detection |
|---|---|---|---|---|
Idle exit (keep_alive_s elapsed) | The daemon | bye{reason: "idle"} | — | Immediate |
| SIGTERM / SIGINT | The OS | bye{reason: "signal"} | Falls back to stale if the POST fails | Immediate, else 60 s |
| Control-plane drain | The control plane | phase: draining, then bye{reason: "drain"} | It issued the command | Immediate |
| Launcher terminate / spot reclaim | The launcher | Nothing. The power is gone | hostStatus → terminated ⇒ gone | Next host poll, else 60 s |
| Crash, OOM, kernel panic | Nobody | Nothing | stale | 60 s |
| Network partition | Nobody | Nothing — the POSTs fail too | stale | 60 s |
The last three rows are why bye is not a reliability mechanism. The
60-second stale timeout stays exactly as it is and remains the backstop for every
case. bye is a latency mechanism: it turns a 60-second ambiguity into a
zero-second certainty for the three cases where the daemon is still able to
speak. Nothing here is allowed to depend on the message arriving.
POST /v1/daemon/bye
reason is one of idle, signal, drain, hello_failed. The control plane
sets state: "gone" and lastSeenAt: now, and does nothing else.
The constraints on the client side are the interesting part, and they are non-negotiable:
- One attempt, two-second timeout. A failure is logged at
debugand ignored. - It never delays exit. A shutdown that hangs because it was being polite is worse than the ambiguity it was trying to remove.
- It is sent after the drain returns, so
inflight_abandonedis a real number rather than a guess.
What crosses the wire
| Call | When | Carries | What it changes |
|---|---|---|---|
POST /v1/daemon/hello | Once, at boot | machine_id, label, tags, lanes, capabilities, versions, host_key | Creates or updates the Worker row; state = joining |
POST /v1/daemon/poll | Every cycle | worker_id, queue, timeout_s, invocations, phase, host_key, updated_capabilities including current_invocations, setup_results | Bumps lastSeenAt; sets state from phase; updates capabilities; gates the claim on phase and host_key |
POST /v1/daemon/ack | Once per run | invocation_id, state, result_blob, error, worker_id | The invocation reaches a terminal state |
POST /v1/daemon/bye | Once, at exit | worker_id, reason, inflight_abandoned | state = gone, lastSeenAt = now |
current_invocations is the row with a consequence outside this page: it is what
closes the structurally-zero utilization gap described under Scale to
zero.
Sending the count on every poll closes it without a control-plane change,
because the five call sites already read the key.
The same count fixes a second bug for free. The daemon tracks in-flight work in
four different places and only one of them covers exec as well as run, so a
worker busy with a long exec can look idle and time itself out mid-job. One
counter, incremented at both dispatch sites, is what the idle gate should have
been reading all along.
The alternative was a new sibling field on the poll. Reusing
updated_capabilities needs zero control-plane change — that key is already
consumed and already written through to the worker row. The price is one
redundant field write per poll, on a document that is already written every poll
to bump lastSeenAt.
The interface
Python
The hooks run on the worker, not on the submitter — they are the daemon-side half of the run declaration.
The hooks take no arguments. There is no context object: a hook reads and
writes ordinary module state — a global, an open handle — which is exactly the
thing a shell host_setup cannot hand to the next body. Environment and working
directory come from the RunConfig, not from a parameter.
phase() returns a three-value subset of the five-value DaemonPhase on the
wire. booting and terminating have no Python spelling by construction: the
interpreter is not running yet during booting, and by terminating the grace
window has already expired, so no hook is still entitled to run.
on_host_setup and on_run_setup are the two phases named on the host
setup page, reachable from Python
rather than from a queue template. on_shutdown runs inside the grace window and
must return before it expires — a hook that outlives the grace is killed with
everything else, which is the whole reason the escalation ends in SIGKILL. The
runnable version of all three, including the upload, is
examples/lifecycle_upload.py in lakeshore-py.
What does not ship yet
phase, host_key, current_invocations, setup_results and
POST /v1/daemon/bye are all on the wire, and claimOne filters on the host
key. Four things on or adjacent to this page are still design only:
- Nothing launches on an unmatched host key. Matching ships: a worker advertising a different key is not handed the invocation. But an invocation whose key matches no live worker simply stays queued — provisioning one is the next slice.
- A
RunConfig'shost_setupis hashed but never executed. The daemon runs thehost_setupin its own config at boot, and hashes the one in an incoming run config into the key it matches against — so a declaredhost_setupselects a host that already has it, and never installs it. Onlyrun_setupis executed from a run config, per invocation, before the body. ExecBodycarries norun_config, so theexecpath — what the CLI and SDK drive today — has no host key and no placement story.- Setup failures are logged, not tabulated. A non-success is written to the control plane's log and the most recent one is kept on the worker row; there is no per-command status table to read a history out of.