DreamLake

Elastic queues

The deployment entries below give current status. Earlier entries retain the validation and deployment status recorded when each change merged.

2026-09-15 — CLI release validation

Lakeshore #28 tracks a release workflow defect: the publish job used a fresh checkout without building the npm package and suppressed publication errors. PR #29 merged at 1d0e794; it builds and packs the selected source, checks a clean installation against real HTTP responses, and publishes that tested tarball. A fresh registry installation must pass before R2 latest advances. Requested tags must match the package version.

The local 0.2.1 candidate passed 653 TypeScript/Node tests, one skipped, and nine HTTP checks using a clean tarball installation. Linux CI passed the package gate, but the first native build exposed an optional DevTools dependency. The revised candidate pins the build dependency and passes nine local native macOS HTTP checks. Updated-head CI passed package verification and all three native builds. Linux x64 and macOS arm64 passed the same command checks; Linux arm64 is cross-compiled and checksum-verified without runtime acceptance.

GitHub lists no repository or organization release secrets and no release environment. The existing credential setup has been requested from Ge; no secrets were copied or new access granted. Version 0.2.1 is prepared, not published. The cancellation implementation remains merged in #27; runtime activation and Ge acceptance remain separate.

2026-09-15 — CLI cancellation intent

Lakeshore CLI #27 merged at a018025. jobs kill now prints cancellation requested when the server records intent on a still-running or queued job. A terminal response keeps its actual state, including success or failure when completion wins the race. List and show expose cancellation intent separately. JSON responses and legacy missing fields remain compatible.

The new regressions failed seven cases on parent eec3335. Candidate a6daef8 passed all 26 jobs tests and the full TypeScript/Node suite: 653 passed, one skipped. The repository has no pull-request test workflow; these are local test results. The README documents the commands and response semantics.

Package release and installed-client acceptance remain open in CLI #26. This merge does not deploy the control plane or enable worker cancellation capability. Recovery after losing the child handle, runtime rollout, UI acceptance and Ge testing remain separate tasks in the DreamLake task tree.

2026-09-14 — Cancellation API merged with current signed capability

Control-plane #44 merged at c1c72b7. Running cancellation records the dispatched attempt and returns 202 only for a worker advertising support in its latest signed poll. State remains running until a terminal worker acknowledgement. Unsupported workers retain 409; queued cancellation is unchanged.

Review reproduced a stale-advertisement bug: after a replacement worker omitted support, stored inventory still allowed cancellation. Candidate 519ba07 clears support on omission, false/malformed values, unsigned polls and setup-only metadata. Ordinary inventory is preserved. The local suite passed 716 tests, thirteen skipped, and CI passed. The separate real HTTP/Mongo fixture passed withdrawal and re-advertisement, nonzero-attempt dispatch, intent retry, terminal-result preservation and legacy string assignment. Its processes and database were removed. Receipt · Reproduced failure.

The server merge preserves the private tracked API. It does not deploy a control plane or enable a worker capability. Nymph #38 remains draft; replacement-worker cancellation, uncertainty resolution, client/UI acceptance and schema/runtime rollout remain open. Ge acceptance is unconfirmed.

2026-09-14 — Ordinary-process termination refusal merged

Nymph #44 merged at 28a65b9. When ownership inspection refuses termination, invocation bodies and setup scripts retain the original child and keep their task pending. They retry verification once per second; terminal acknowledgement and capacity release wait for verified termination. Persistent refusal keeps shutdown pending.

Candidate 2dff3cb passed 389 local tests, 395 Linux default tests and 397 Linux all-feature tests. The Linux fixtures verify actual ownership refusal through runner cancellation and daemon accounting. The submission-timeout fixture now allows normal preflight time and proves that submission began before recovery. Private-runner main, 0.1.6 metadata and optional CA configuration are preserved. Validation.

Streaming #38, 3e75f99, remains draft. Its combined branch passed 419 local tests, seven paired scenarios, and both Linux workflows (425 default and 427 all-feature tests). Cancellation after losing the child handle and uncertainty resolution remain in #43.

This merge does not change frozen 0.1.6 artifacts, release a package, deploy a worker or establish Ge acceptance. Cancellation capability remains disabled.

2026-09-14 — Linux streaming termination refusal

Draft Nymph #38, b5a3ec8, passed all eight paired scenarios in a non-root Linux container on the shared development host. During ownership refusal, the worker kept capacity occupied and withheld terminal results and acknowledgements. After the fixture restored visibility, cancellation completed without rerunning the producer. A separate supervisor's heartbeat continued; fixture, container and temporary-directory cleanup passed. Receipt.

The two tests passed in 156.05 seconds. The collector then failed to parse JSON prefixed by a Rust test name. Evidence was reconciled from the saved original log; the receipt retains the collector failure. The run was not repeated.

This verifies isolated Linux process behavior with a simulated control plane, not Slurm, host PID-namespace behavior, private credentials or deployment. Integration with the newly merged private runner is being validated separately. Uncertainty resolution and cancellation after losing the child handle remain in #43. Cancellation capability remains disabled; runtime release and Ge acceptance are still pending.

2026-09-14 — Retain uncertainty when terminal evidence is missing

Draft Nymph #38, 8be4503, retains the reservation and occupied slot when recovery has no outcome checkpoint or acknowledged terminal frame. It keeps reporting the invocation and preserves the producer fence. Missing evidence no longer creates a terminal failure.

The unit and actual-daemon regressions failed on the previous revision. The candidate passed with three worker starts, one producer execution and zero terminal acknowledgements. After the original worker crashes and its producer exits without a terminal record, two replacements each report the occupied slot and reject duplicate execution. Cleanup passed. The local Rust suite passed 403 tests, fourteen ignored. Pinned receipt · Test guide. The CP is simulated; this case does not test surviving descendants or Linux ownership refusal. Full paired and Linux CI results are tracked in #38.

Resolving uncertainty, cancelling after loss of the child handle, and Linux streaming-refusal acceptance remain in #43. Previously acknowledged failures are not undone. Cancellation capability remains disabled; no runtime merge, release, deployment or Ge acceptance is claimed.

2026-09-14 — Full-worker heartbeats and blocked-frame cancellation

Draft Nymph #38, 7d83430, includes the tested #44 termination-refusal candidate. Full workers now periodically poll with their occupied capacity; a freed slot wakes them immediately. Fresh and recovered executions use the same heartbeat path.

The local suite passed 403 tests, thirteen ignored. All six paired daemon/Python cases passed: success, failure, replacement, cancellation, legacy dispatch and cancellation during a blocked frame request. The blocked case sent 63 full-capacity reports, stopped one producer, and withheld terminal acknowledgement until pending frame delivery resumed. Frame/result retries and cleanup passed. The earlier regression failed because the full worker stopped polling. Receipt · Test guide. The paired CP checks signature presence, not cryptography; Linux CI is tracked in #38.

Recovery after cancellation or loss of the child handle, missing terminal evidence, and Linux streaming-refusal acceptance remain in #43. Automatic cancellation capability stays disabled. No runtime merge, release, deployment or Ge acceptance is claimed.

2026-09-14 — Termination refusal blocks cancellation acceptance

Both Linux workflows reproduced Nymph #43: when ownership inspection is denied, the runner returns a terminal failure while the child remains alive. Green CI on streaming #38 did not cover this path.

Draft Nymph #44, 7559378, retains the original child and retries verified termination once per second. It covers invocation bodies and setup scripts. The regression restores ownership visibility and requires cancellation to finish; a daemon test also checks that no terminal acknowledgement is sent during refusal. The local macOS suite passed 372 tests, eight ignored; the Linux baseline failed as expected. Candidate Linux results are tracked in the PR. Test guide.

Persistent refusal keeps shutdown pending. Streaming checkpoints, full-capacity heartbeats, blocked-frame cancellation and recovery after worker replacement still need integration and acceptance. #43 remains open; no cancellation capability, runtime release or deployment is enabled.

2026-09-14 — Fresh streaming executions consume cancellation

Draft Nymph #38, f984118, now includes merged #42 and polls signed control for fresh executions with a known attempt. Matching running intent cancels through retained group ownership before the outcome is checkpointed for durable delivery. Wrong attempts, authentication errors and false intent leave execution running. Legacy dispatches without an attempt make no control requests; already completed outcomes retain their state.

The full local suite passed 403 tests, thirteen ignored. Five actual-daemon/Python cases passed: success, failure, worker replacement, fresh cancellation and legacy dispatch. The cancellation case survived wrong-attempt/401/false responses, stopped on matching intent, and retried its identical killed result after response loss. One producer ran and cleanup passed in every case. SDK dafc96c; the CP fixture checks IDs and signature headers, not cryptography. Receipt.

Current-head CI is pending. Cancellation after worker replacement remains unwired; blocked-frame cancellation and termination-refusal/uncertainty accounting need validation before activation. Failed ownership must not become a claim that work stopped or its capacity is free. No cancellation capability, runtime release or deployment is enabled. Hosted, CLI/UI and escaped-group acceptance remain open.

2026-09-14 — Process-group fix merged after Linux acceptance

Nymph #42 merged after both Linux CI jobs passed at 1317e8c. An isolated non-root container on the shared development host passed all ten ownership tests and five repetitions of both cancellation tests. A separately marked supervisor process continued throughout all six commands. The foreign fixture was reaped; the container and uploaded directory were removed. The cleanup verifier's lowercase Docker-message mismatch was corrected against the same completed run, without rerunning the tests.

Receipt and exact fixture. The full local suite passed 372 tests, eight ignored. This verifies descendants within the spawned process group; existing-worker rollout, escaped groups, recovery without a retained child handle and actual shared-cluster/Slurm acceptance remain separate. Nymph #38 and CP #44 still need cancellation integration; no package release, runtime deployment or cancellation capability is enabled.

The preceding source notes through #477 are live at 862c0a1, Netlify 6aa8d4c69be3dbd1e6bb8786: nine public pages/assets matched the build. Publication preserved the already-live Vault baseline 652a3fa; the older publication branch was not used to overwrite that newer deployment.

2026-09-14 — Unix exit-window follow-up

Draft Nymph #42 is now at 1317e8c. The Linux all-features run at prior 9cf8465 passed group termination but failed in the non-group fixture's signal-based cleanup. That fixture now exits through stdin. A local check also exposed macOS EPERM before exit became waitable; bounded ownership rechecks now apply on both Unix platforms.

The final full local suite passed 372 tests, eight ignored. Current Linux CI and shared-machine acceptance remain pending. Receipt. The runtime PR remains draft; no release, deployment or queue cancellation activation.

2026-09-14 — Completion race and Linux ownership retry

Draft Nymph #42, 9cf8465, adds a deterministic completion race: the fixture child exits before the runner resumes, then cancellation is made ready. Success and failure retain their original outcomes. The descendant regression also passes. The full local suite passed 372 tests, eight ignored; both integration tests passed again after the Linux-only adjustment.

Both Linux CI jobs at prior 2b5e221 failed during immediate escalation: the ownership read was denied or the marker was unavailable. The candidate retries failed checks for at most 500 ms while the original owned group and retained child remain valid. Each attempt verifies ownership again; errors do not permit signals. A Linux-only test holds process inspection disabled and checks that persistent denial leaves the child alive before the fixture exits through stdin.

Receipt. Current-head Linux CI and shared-machine acceptance remain pending. Escaped groups and recovery without the original child handle require separate work. The runtime PR remains draft; no release, deployment or queue cancellation activation is claimed.

2026-09-14 — Retained group ownership during cancellation

Draft Nymph #42, 2b5e221, fixes the reproduced descendant case locally. Cancellation keeps the original child waitable through the final group signal and reaps it afterward. The spawn-time group capability and retained child establish ownership; foreign owners, non-group spawns and already-reaped handles are rejected.

On macOS, a group containing only its retained exited leader returns EPERM. The candidate accepts this only when kernel group membership is exactly that leader; other permission or membership failures remain errors. Earlier candidate failures exposed the macOS group-lookup and EPERM cases.

The descendant regression and full local suite passed: 371 tests, eight ignored. Reproduction and receipt. Linux CI, deterministic completion-race coverage and shared-machine acceptance remain pending. This covers descendants in the spawned process group; escaped groups and recovery without the original child handle require separate work. Nymph #38 and CP #44 remain draft, with cancellation disabled. No runtime release or deployment is included.

2026-09-14 — Process-tree cancellation gap

Nymph #41 blocks worker cancellation activation. The process runner can return killed after its leader exits on SIGTERM while a descendant that ignores SIGTERM continues running. A bounded fixture reproduced this on streaming draft 8e5155a and independently on main 9f9cda5: the descendant wrote a marker after cancellation returned. The cleanup assertion failed in both runs.

Draft Nymph #42 contains the failing regression and receipt; no fix is included yet. The termination helper is unchanged from v0.1.5. Supervisor identity remains released, but it does not establish process-tree cleanup for this case. The fix must preserve verifiable ownership through termination; a cached PID or process-group ID after leader exit is not permission to signal.

Nymph #38 and CP #44 remain draft. Automatic cancellation, recovery, hosted acceptance and Ge testing remain open. Both CI jobs passed for the earlier #38 head; those checks did not cover this newly reproduced case.

2026-09-14 — Signed worker cancellation-control read

Draft Nymph #38, 8e5155a, adds Client::invocation_cancel_requested. It signs the worker and invocation IDs and requires the caller's known accepted attempt. Only matching intent for a running invocation returns true. Terminal states return false; another attempt, malformed response or transport/authentication failure is an error, not permission to cancel. The request has a ten-second timeout.

Two focused tests and the actual Rust client against CP #44 65919ff passed real signature verification, matching/incorrect attempts, terminal intent, missing invocation and invalid-key checks. Reads left the invocation unchanged; fixture cleanup passed. The full worker suite passed 399 tests, thirteen ignored. Reproduction and receipt.

This is a read API only. It does not poll automatically, signal processes or report a terminal result. Worker control consumption, owned termination and surviving- producer cancellation remain open. No cancellation capability is advertised; runtime release, hosted acceptance and Ge testing remain separate.

2026-09-14 — Cancellation attempt and assignment

Draft CP #44 now dispatches the stored attempt and records which attempt was cancelled. An old request does not cancel a later attempt. New claims store the worker ID as an ObjectId; cancellation accepts both that form and legacy strings while preserving its atomic namespace, worker, state and attempt checks.

The first real-dispatch check exposed a 409: raw claim stored a string worker ID, but the cancellation predicate expected an ObjectId. The earlier fixture seeded assignment through Prisma and missed it. At 65919ff, 709 tests and real signed HTTP/Mongo dispatch/cancellation checks passed, including attempt 7, retries, terminal acknowledgment, a later attempt and legacy assignment. Cleanup passed. Receipt.

Draft Nymph #38 accepts an optional attempt and retains it before Python starts. Reopening the scoped journal preserves it. Legacy requests/records return unknown, not zero; they cannot authorize API cancellation. 397 tests passed, twelve ignored. Older worker binaries cannot recover the new attempt-bearing records; do not remove or rewrite retained evidence to downgrade them.

Worker control consumption, owned termination, surviving-producer cancellation, client/UI display, paired cancellation acceptance and rollout remain open. No cancellation capability, runtime release, deployment or Ge testing is claimed.

2026-09-14 — Signed frame worker identity

Nymph #39 fixes frame uploads that signed the request but omitted worker_id from its MessagePack body. The real control-plane authentication hook requires that field. Earlier paired daemon checks used a simulated CP and did not establish compatibility with real CP auth.

The wire test now verifies the signature and worker ID in the exact request bytes. It fails against the old encoder and passes with the fix. The paired daemon fixture also checks the field. Three uploader tests passed on the main-based fix; integration #38 passed 395 tests, twelve ignored, and all three daemon/Python success, failure and restart cases with cleanup. The first full run failed an existing Slurm timeout test with a missing submission marker; isolated and full reruns passed without changing that test. Its cause is not established.

Nymph #39 merged at af8a065 after both CI jobs passed. The actual Rust uploader also passed against full local CP routes, real signature verification and isolated Mongo: duplicate delivery stored one exact frame; conflicting bytes, a missing invocation and an unregistered key were rejected. Owned cleanup passed. Reusable fixture and receipt #40. The CP checkout was d4d9ba1; its result-frame route is unchanged from development 7d8504f. This is local protocol acceptance, not hosted streaming or a runtime release/deployment. Running cancellation remains open.

2026-09-14 — Running cancellation intent (draft)

Control-plane #44 records running cancellation separately from the terminal state. A capable worker's job returns HTTP 202 with cancelRequested: true; it remains running until the worker reports an outcome. Older workers still return 409. Completion, assignment and attempt changes are protected against a concurrent cancellation write.

The assigned worker reads the request through a signature-required control route. The full suite passed 708 tests, twelve skipped. Real HTTP routes, signature verification and isolated Mongo passed repeated cancellation, intent reads and a successful terminal acknowledgment after cancellation intent; cleanup passed. Behavior and reproduction.

No shipped worker advertises this capability. Worker consumption and owned-process termination, surviving-producer recovery, client/UI display, paired acceptance and schema rollout remain open. The control-plane PR remains draft and is not deployed.

2026-09-14 — Worker capacity and daemon streaming recovery

Control-plane #43 merged at 7d8504f. Workers reporting full capacity continue to send heartbeats and receive control messages. Queued invocations and exec requests wait until the worker reports a free slot. A zero limit remains unbounded; older workers that omit capacity retain their existing behavior. This checks each poll, not a shared quota across simultaneous polls using the same worker identity.

Development now runs 7d8504f. The built image passed 695 tests, twelve skipped. Signed requests through the public HTTPS endpoint verified full/over-capacity backpressure, control delivery, capacity release and legacy compatibility. Both original workers polled the sole new control-plane pod. All fixture records and the uploaded build directory were removed; no queued/running work remained. There is no schema change. Production rollout remains pending. Deployment receipt.

Nymph #38 now wires streaming recovery into daemon startup behind runtime.streaming_frame_limit_bytes. Recovered producers occupy worker capacity until their outcomes are recovered. The default remains disabled. The daemon does not rerun a reserved execution.

Runtime ad95f42 passed 395 tests, eleven ignored. Follow-up 9df17d9 passed three actual-daemon/Python checks with SDK dafc96c: generator success, generator failure, and daemon replacement while the original producer survives. Each check lost one frame response and one terminal-result response, verified identical-byte retries and a single producer execution, and cleaned up its owned processes. The replacement daemon reported the surviving producer against capacity and left it unfenced while its lock was held. CI passed at both Nymph revisions. Reproduction and receipt.

The daemon fixture uses a simulated control plane; its signing-header checks do not verify signatures. The separate capacity fixture verifies actual signatures. Hosted streaming execution, running API cancellation, runtime release and Ge review/testing remain open. Nymph #38 and SDK #23 remain drafts.

2026-09-14 — Streaming execution recovery

Nymph #38 adds explicit recovery without restarting Python. A busy producer stays running. A stopped producer restores its recorded outcome or acknowledged terminal output; missing terminal evidence becomes InterruptedExecution. This does not establish that user code had no side effects. Invalid evidence is retained for inspection.

Acknowledged results retain an exact-result receipt before their pending records are removed. Recovery does not recreate delivered results. Lost responses retry identical bytes; a conflicting result is rejected. Reservation discovery lists executions within the worker journal without taking consumer locks or touching output, and rejects invalid directory scope.

The full suite passed 390 tests, eleven ignored, at b5b5399. Discovery follow-up d49331c passed fourteen handoff unit tests and all four paired Python/Rust checks with SDK dafc96c. The HTTP fixture verifies lost frame and terminal-result responses, recovery with its server stopped, and receipt/pending-record coexistence. It does not simulate a daemon crash during Python execution. Acceptance artifact.

Daemon startup wiring, surviving-producer capacity accounting, API cancellation, hosted acceptance and Ge review/testing remain pending. The runtime is still a draft; publishing these notes does not release or deploy it.

2026-09-14 — Worker-side producer fencing

Nymph #36 6a39cb4 and integration #38 40498f7 add worker-side producer fencing. New handoffs advertise producer_fence_version: 1; SDK #23 dafc96c validates it. Legacy handoffs remain readable for frame delivery but cannot be fenced or used for a new process launch through this API.

Nymph attempts the producer's lock without waiting. A busy producer stays untouched. Otherwise Nymph persists a scoped fence and holds the lock until recovery releases its guard. The fence remains afterward. This neither signals a process nor decides its terminal job status.

The SDK suite passed 279 tests, seven skipped. Four paired checks passed on the consumer branch, including a live producer remaining unfenced and a late child refusing the worker fence before user code. The combined integration suite passed 385 tests, eleven ignored; all four paired checks also passed on the integration branch. Reproduction. Recovery-result classification, daemon activation, cancellation ownership and hosted acceptance remain open. No package/runtime release or deployment.

2026-09-14 — Producer recovery fence

SDK #23, at 743a2cb, records producer startup durably while holding its producer lock. A previous start or recovery fence blocks another producer even if no frame exists. Refused handoff setup preserves prior result/error files. The SDK holds the lock through user-code execution and terminal delivery, releasing it in finally.

The unit suite passed 273 tests, seven skipped. Three regressions failed before the fix and passed afterward; additional checks cover process exit before first yield and start-marker sync failure. Three paired Rust/SDK checks passed with Nymph a4155b6. Worker-side lock/fence recovery, cancellation ownership, daemon activation and hosted acceptance remain unfinished. No package/runtime release or deployment.

Nymph Linux CI at a4155b6 failed on an extra immediate lock reacquisition in a retained-output test; its all-features job passed. The test now validates the saved reservation directly. Integration 585ab40 passed 244 local library tests, three ignored; fresh CI is pending. Runtime locking is unchanged, and the exact source of the contention has not been established.

2026-09-14 — Streaming terminal-result journal integration

Nymph #38, at b535de0, combines the artifact and streaming drafts. The explicit process API now requires invocation result delivery and journals the completed outcome before returning. Replay holds the handoff lock, delivers any pending frame, then sends the terminal result. This preserves the last queued value when cancellation coincides with an outage.

Focused tests and three paired SDK/Rust checks passed. A signed HTTP fixture lost the first frame response, reopened the client/journal and verified identical frame retry before the killed result. The real Python runner check verified its outcome was journaled. The combined suite passed 382 tests, ten ignored; the three paired checks ran separately. Reproduction.

Input review at consumer f98b0d0 and integration b80d1ad added rejection of path-like invocation IDs before staging. The regression checks empty IDs, dot segments, absolute/nested paths, backslashes and NUL, asserting no staging or handoff directory is created. Six integrated handoff checks and three paired SDK/Rust checks passed; prior full-suite evidence remains pinned to b535de0. Current-head CI is pending.

At a4155b6, the actual daemon startup replay check passed. It seeds a pending value frame and journaled killed result, loses the first frame response, kills and reaps the first daemon, then starts its replacement with the same journal. The replacement retries identical frame bytes before terminal acknowledgment. Pending records and owned cleanup are checked. The CP is simulated and no workload is dispatched; recovery of a job interrupted mid-execution remains open. Run the daemon check.

This branch contains drafts #34 and #36; it does not merge them. Daemon activation, execution-reservation recovery, API cancellation, actual worker-process restart and hosted acceptance remain open. Artifact deployment and retention-policy gates are unchanged. No runtime release, deployment or Ge acceptance is claimed.

2026-09-14 — Explicit process streaming execution

Nymph #36, at 74d6934, adds Client::run_frame_process. It provisions a private handoff, records execution intent before spawning Python, replaces the caller's handoff path and consumes frames concurrently. Output stays staged for durable terminal-result handling. A prior reservation requires recovery; an empty handoff cannot trigger another execution. Successful exit without terminal-frame acknowledgment returns an error.

Three paired Python/Rust checks passed: automatic delivery with response-loss retry and duplicate-execution refusal, cancellation while waiting for acknowledgment, and generator success/failure with consumer reopening. The full Rust suite passed 372 tests, ten ignored; the three paired tests were invoked separately. Reproduction.

Daemon polling does not call this API yet. It supports Python process/subprocess execution under the worker Unix user. Restart recovery, terminal-result persistence, API-driven cancellation and deployed CP acceptance remain open. No package/runtime release, deployment or Ge acceptance is claimed.

2026-09-14 — Retain output until delivery is durable

Nymph #37 extracts shared output-retention hooks from artifact draft #34. Callers can keep successful runner output until they persist the delivery result, then explicitly clean it up. Existing dispatch calls retain their automatic cleanup behavior. Failure and args-bearing script diagnostics remain on disk.

At ce82b94, focused process/subprocess tests passed for retained result bytes, explicit cleanup, existing automatic cleanup and script/failure diagnostics. The full suite passed 368 tests, seven ignored. Container and scheduler paths have no new live acceptance. The hooks have no active caller yet; streaming runner integration, release, deployment and Ge acceptance remain open.

2026-09-14 — Explicit generator entrypoint streaming

SDK #23 now connects the invocation entrypoint to the handoff when DREAMLAKE_FRAME_HANDOFF_V1 is explicitly provided. Sync and async generator yields become VALUE frames. The entrypoint persists and syncs its result or error file before sending END or ERROR. Without that configuration, yields still collect into the final result.

At SDK 4046904, 268 unit tests passed, seven skipped. The paired check at Nymph 246d6e7 passed generator success and failure after a yield using the real SDK entrypoint in a Python process. The HTTP fixture observed terminal frames only after the corresponding output file existed. Lost-response retry, consumer reopening, producer pause, terminal-state reopening and owned cleanup passed. Reproduction.

At Nymph 92bec30 with SDK 4046904, the paired checks now use Nymph's actual process runner to launch the Python entrypoint. Success and failure after a yield passed, with staged output retained. Cancelling while the first frame awaited acknowledgment returned a killed outcome, preserved pending bytes and arguments, and prevented producer advancement. Owned directory cleanup passed. Run the checks.

Nymph's runner does not provision or enable the handoff yet. API-driven cancellation, deployed CP and actual worker-process restart acceptance remain open. SDK #23 and Nymph #36 stay draft; neither is package/runtime released, deployed or Ge accepted.

2026-09-14 — Python/Rust handoff acceptance

The paired handoff check passed at Nymph 17d0e62 and SDK 88a1643. A real Python process persisted a value and stayed paused after the first upload response was lost. Reopening the Rust consumer retried the same bytes; only its acknowledgment and removal of the pending file let Python advance. Python then sent a terminal frame. Its acknowledged terminal state survived another reopen. The test asserts Python exit and removal of its owned directory.

Reproduce the check. The HTTP endpoint is a control-plane fixture. This does not establish deployed CP, generator entrypoint, cancellation or actual worker-process restart acceptance. SDK #23 and Nymph #36 remain inactive drafts. No release, deployment or Ge acceptance is claimed.

2026-09-14 — Frame handoff consumer draft

Nymph #36 adds the inactive consumer paired with SDK #23. It uploads one pending frame, persists its acknowledgment, then removes the frame. A lost response retains the bytes. Reopening after an accepted upload and failed cleanup removes the same frame without another upload. The acknowledgment includes sequence, digest and terminal state on both sides.

Nymph's local suite passed 368 tests, seven ignored; Python passed 266, seven skipped. The HTTP/filesystem fixture exercises recovery and preserves opaque MessagePack values. Real Python/Rust interoperability, runner wiring, cancellation and terminal-result coordination remain unfinished. Both PRs stay draft and inactive; no release, deployment or Ge acceptance is claimed.

2026-09-14 — Python frame handoff draft

SDK #23 adds an inactive Unix handoff writer. A worker-provisioned private manifest binds its invocation and size limit. Python persists one frame and waits for a matching sequence/digest acknowledgment and removal of the pending file before advancing. Timeout, acknowledgment conflict and directory-sync failure retain pending bytes.

The unit suite passed 266 tests, with seven skipped. This is producer-side testing; Nymph's consumer, child entrypoint wiring, terminal ordering and process-restart acceptance remain unfinished. No entrypoint enables the writer. The draft is not released or deployed, and Ge acceptance is not claimed. Track integration in #406.

2026-09-14 — Frame transport and truncated reads

Nymph #35 merged as 1c07b0f after both CI jobs passed at 8ccfdc5. The local suite passed 366 tests, with seven ignored. The upload method checks the existing 1 MiB encoded-request limit, requires worker signing and accepts only a matching sequence acknowledgment.

Python SDK #22 merged as 9c9dd5b. Previously, HTTP EOF inside a MessagePack frame silently discarded unfinished bytes. The reader now raises DispatchError. Both partial first-frame and partial-after-complete-frame regressions failed before the fix; 260 unit tests passed afterward, with seven skipped. Empty terminal streams keep their existing behavior.

These changes are merged but not package/runtime released. They do not enable generator streaming: child frame delivery, bounded storage and terminal ordering remain open in #406. No deployment or Ge acceptance is claimed.

2026-09-14 — Signed frame transport draft

Nymph #35 adds a signed client for the existing result-frame upload endpoint. It requires an enrolled identity, preserves the caller's sequence and bytes on retry, and accepts only a matching positive acknowledgment. Each upload attempt has a ten-second timeout.

Focused HTTP tests passed for signature verification, a lost response followed by an identical retry, conflict rejection and invalid acknowledgments. The draft does not connect Python yields to uploads or enable streaming. The child spool, buffering limits, restart delivery and terminal ordering remain in the proposal. No release, deployment or Ge acceptance is claimed.

2026-09-14 — Published SDK remote acceptance

A fresh remote installation of Python SDK 0.4.0 from PyPI passed against development control plane 50d68bb. Two temporary signed workers enrolled in an isolated namespace. Eight concurrent submissions and eight terminal retries executed the Python function once; the client received its expected exception and the accepted terminal record stayed unchanged.

Both workers stopped. Fixture records and the remote environment were removed; the original workers still had fresh polls afterward. Acceptance receipt. This checks published-client failure/retry behavior, not artifact runtime activation, running cancellation, streaming, Slurm or Ge acceptance. Those remain in #406.

2026-09-14 — Python SDK 0.4.0 released

SDK #21 merged as 00e3cf4. 0.4.0 is published: 258 unit tests passed, seven skipped; wheel and source distribution built. PyPI hashes match both local distributions. A fresh uncached PyPI installation verified version metadata, imports and default/explicit runner selection.

This release includes the already-merged run declaration and lifecycle helpers, namespace authentication and artifact reader. RunConfig() defaults to process; use RunConfig(runner="docker", image="python:3.11-slim") to select Docker. The worker must support the selected runner.

Artifact runtime support remains draft and undeployed. Hosted storage identity, retention and provider acceptance remain open in #406. Ge review/testing is not claimed by package validation.

2026-09-14 — Large results and daemon restart

The Python reader now supports result artifacts. SDK #20 merged as 0d2aaa4; 258 unit tests passed, with seven skipped. Downloads verify the byte count and SHA-256 before returning a result. They use a separate HTTP client without namespace credentials. Inline results keep their existing behavior. The package has not been released.

CP #42 and Nymph #34 remain drafts. The worker saves the original result before cleanup and retains it until the control plane acknowledges it. Results over 256 KiB use a signed upload grant, with a 100 MiB artifact limit. The control plane checks the object before accepting the reference. Older readers receive an upgrade error for artifact results. Worker invocation journaling remains disabled by default.

Real S3 checks passed for corrupt-byte rejection, signed size, create-only upload retries, stored metadata and downloads. A configured Nymph daemon then ran one Python fixture that produced a serialized 2 MiB value. After its acknowledgment response was dropped, abrupt restart replayed the saved artifact without another execution. The accepted MongoDB row and completion time stayed unchanged. Python result() and get() verified the downloaded value. Hello and one-job dispatch were simulated; artifact routes, signature verification, MongoDB and S3 were real. Both daemon processes and all owned objects, database, journal and identity files were cleaned up. The check passed again on a fresh ordinary daemon build. Daemon receipt · S3 receipt.

Nymph CI passed at 1fe23b0; the local suite passed 372 tests with seven ignored using two test threads. The initial default-concurrency run failed an existing Slurm timeout test; its isolated rerun passed. The cause remains unconfirmed and no Slurm source changed. CP runtime tests passed: 700 tests, 12 skipped.

Development still runs CP 50d68bb. A read-only preflight found two original workers polling and no queued/running work, but no artifact bucket or AWS region configured in the pod. Storage configuration and its credential source must be verified before hosted acceptance. Completed-result retention, abandoned-upload cleanup, hosted/provider acceptance and Ge review/testing remain open in #406. No bucket retention rule or runtime deployment changed.

2026-09-14 — Invocation result retries

CP #41 prevents an acknowledgment retry from changing a job's completion timestamp or replacing its result. An identical outcome returns success without a write. A conflicting outcome returns 409 and preserves the accepted result. The first result write requires the same running job, assigned worker and namespace.

Local Mongo and HTTP checks use the real signature hook and simulated workers. They cover binary results, both stored worker-ID formats, retries after reconnect, invalid and foreign requests, and 20 competing-result races. Each race accepted one outcome and rejected the other. Unsigned, tampered, revoked-key and repeated nonce requests were rejected; a freshly signed identical retry succeeded. The temporary Mongo process and database directory were removed. Acceptance script and receipt.

CI passed and the merged fix is deployed to development at 50d68bb. The image passed 693 tests with 12 skipped. Public HTTPS checks verified eight identical retries, conflicting outcomes, auth and namespace rejection, and twelve competing result races. Both original workers polled the sole new pod; public health passed and no active work remained. Fixture records and the uploaded build directory were removed. Image f5b5f156c212 remains available for rollback. Hosted receipt.

The first hosted probe failed because decoding a saved Buffer added metadata used by a later equality assertion. Its fixtures were cleaned up; decoding a copy fixed the probe without changing the runtime. These checks use simulated workers, not launched workloads. Worker result persistence, cancellation delivery, production deployment and Ge review/testing remain open in #406.

2026-09-14 — Queue-filtered jobs and demand counts

CP #40 fixes queue-filtered job listing and bulk cancellation returning 500 on Mongo. Both used JSON filters unsupported by the Prisma Mongo connector. Listing now filters the nested queue field before pagination and preserves the existing response format. Bulk cancellation changes only queued, unassigned jobs in the selected namespace and queue; a concurrent claim remains intact.

Queue stats and scaler demand counts had a fallback that loaded all queued invocations in the namespace. They now count matching documents in Mongo. The test stub rejects the unsupported filter forms instead of accepting them.

Typecheck and 690 tests passed, with 12 skipped. The real-Mongo acceptance script checks filtering, pagination, literal queue names, namespace isolation, counts without loading invocation payloads, and cancellation retries. Twenty bulk-cancel and claim races preserved the winner: nine cancellations and eleven claims. Its temporary Mongo process and database were removed. Local receipt.

CI passed and development runs the merged fix at f5b5f15. The built image passed 690 tests with 12 skipped. Hosted checks verified filtered pagination, worker/state filters, stats and scaler counts, namespace isolation and bulk-cancel retries. Running jobs and other queues were preserved. All 16 concurrent bulk-cancel and signed-poll races were won by cancellation; separate claim-first checks verified that assigned work remained intact. Hosted receipt.

The checks used signed simulated workers, without launching a workload or a provider. Fixture records and the uploaded build directory were removed. Both original workers polled after only the new pod remained; public health passed and no active jobs remained. Image 70c2438f1d71 is retained for rollback. Running cancellation, restart recovery, streaming, job UI integration and Ge acceptance remain open in #406.

2026-09-14 — Queued cancellation and worker claims

CP #39 fixes a race in single-invocation cancellation. A worker could claim or finish a job after the handler read its queued state, and cancellation would overwrite that new state with killed. The write now requires the invocation to remain queued and unassigned in the same namespace. A competing update returns 409 and retains the worker's state and result.

Five regression cases failed before the fix. Afterward, all 21 focused tests and 690 full-suite tests passed, with 12 skipped; typechecking passed. Local real Mongo and HTTP checks covered competing running/succeeded/failed writes, null or absent worker IDs, and a claim attempted after successful cancellation. The competing writer was injected; this does not establish hosted worker acceptance. Local receipt.

Forty concurrent cancellation/claim races preserved the winner: 15 cancellations succeeded and 25 returned 409. Mongo profiling confirmed the conditional update runs in a transaction; transaction write conflicts also return 409.

The fix merged after CI passed and is deployed to development at 70c2438. The built image passed 690 tests with 12 skipped. Public HTTPS checks verified cancellation before claim, unchanged running/completed jobs after rejected cancellation, namespace isolation and same-ID retries of cancelled jobs. All 16 concurrent cancellation/poll races were won by cancellation; a separate check claimed the job first and verified that cancellation returned 409. Hosted receipt.

These checks used signed simulated workers, without launching a workload. Fixture records and the uploaded build directory were removed. Both original workers polled after only the new pod remained; public health passed and no active jobs remained. Image 963c34180927 is retained for rollback. Running-job cancellation, restart recovery, streaming, job UI integration and Ge acceptance remain open in #406.

2026-09-14 — Python failure and retry acceptance

Two isolated Nymph workers enrolled before eight concurrent submissions of one invocation. The Python function recorded one execution and raised a known ValueError. The producer received a DispatchError containing that message; the stored invocation was failed. Eight more submissions after failure returned the same invocation, preserved its error and did not execute the function again. Failure receipt.

The check used SDK source a5db72f, development CP 963c341 and the same observed Nymph binary identified in the earlier receipts. Both fixture workers stopped; their credentials, database records and Python environment were removed. Both original workers remained active with recent polls, and no active jobs remained.

This verifies Python exception propagation and retries of a failed invocation. Cancellation, restart recovery, generator streaming, job UI integration, Slurm/GPU, package release and Ge acceptance remain open in #406.

2026-09-14 — Real Python execution and concurrent retries

An isolated Nymph enrolled through development HTTPS, claimed a queued function, executed it through the Python daemon.entry contract and returned 42. The stored invocation was succeeded, assigned to the worker named in the consumed enrollment grant. Single-worker receipt.

A second check enrolled two Nymph workers before submitting one invocation eight times concurrently. The function wrote an execution counter and returned 42. The counter contained one entry. Eight more submissions after completion returned the same invocation and left the counter unchanged. Two-worker receipt.

Both checks used SDK source a5db72f and development CP 963c341. The receipts identify the existing Nymph binary by version and checksum; its build provenance was not independently verified. Test daemons, credentials, database records and Python environments were removed. The original workers remained active.

These checks establish scalar execution and duplicate-submission behavior with concurrent workers. Generator streaming, failure/cancel/restart behavior, Slurm/GPU, tracked-workload integration, package release and Ge acceptance remain open in #406.

2026-09-14 — Public signed result uploads

Development runs 963c341 from CP #38. POST /v1/daemon/result-chunk is available through development HTTPS after ingress #419. Uploads require the assigned worker's signature and namespace. An identical retry returns success; different bytes or a different final marker at the same sequence return 409 and preserve the accepted frame.

The result journal saves frames in Mongo before sending Redis notifications. Its separate Redis connection uses bounded waits and no offline command queue. Readers check Mongo when notifications are unavailable. Typechecking and 684 tests passed, with 12 skipped. A silent-connection test verified acknowledgment and Mongo catch-up during an outage; a real Redis check verified frame notifications and exact two-frame reads.

Public HTTPS acceptance verified owner uploads, retries and result reads, plus unsigned, foreign-worker and conflicting-write rejection. Legacy producer routes remain denied. Temporary credentials and records were removed; both existing workers freshly polled the sole new pod. The previous bd6a32e image is retained for rollback. Hosted receipt.

Python PR #19, merged as a5db72f, uses the selected namespace for HttpDispatch submission and result reads. Its 238 unit tests passed, with seven skipped; the Python client passed public HTTPS checks using a completed fixture. This change is not yet released as a package.

Actual worker execution, worker integration, tracked workloads, production rollout and Ge acceptance remain open in #406. The HTTP probes used fixture records; they did not execute a user function.

2026-09-14 — Namespaced producer development rollout

Development runs bd6a32e, including CP #37. The built image passed typechecking and 679 tests, with 12 skipped. The rollout confirmed no active workloads and rechecked deployment identity before changing the application image.

Public HTTPS checks verified token rejection, same-org target namespace scoping, eight identical retries, conflicting payload rejection and archived-queue admission rejection. Authorized await returned the saved result; the result stream returned exact seeded frame bytes. Reads through the wrong target returned 404. Legacy producer routes remain denied at public ingress.

The probe used a completed invocation and seeded result frames. It created no runnable workload and removed all fixture namespaces, tokens, queues, functions, invocations and frames. Both original workers freshly polled the sole new pod after the old pod was removed. TLS health passed; the previous 4d1cc71 image remains available for rollback.

Hosted acceptance receipt. SDK/CLI adoption, authenticated public result-frame upload, actual worker execution, tracked workloads, production deployment and Ge acceptance remain pending in queue integration #406.

2026-09-14 — Namespaced producer API

Control-plane PR #37 adds authenticated operations under /v1/namespaces/:ns/producer:

  • PUT /queues creates or finds a queue by name.
  • POST /submit persists a waiting invocation and returns its receipt.
  • POST /await reads the invocation result or returns pending, waiting at most 60 seconds.
  • GET /invocations/:id/result/stream reads the invocation's result frames.

Existing namespace tokens authorize the requested target, including same-org siblings. Queues, functions, parent references and result reads stay in that target. Workers resolve queue prefixes, pools and wake notifications in their own namespace. Inactive queues reject new namespaced submissions; retries of accepted invocations remain available.

At 40ec8ca, typecheck and 679 tests passed, with 12 skipped. A real local Mongo/API probe with required worker signatures verified token rejection, sibling targeting, concurrent retries, restart survival, one claim, results/stream bytes, pool scope and inactive admission. The private legacy round trip also passed. Acceptance receipt.

CI passed; the fix merged as bd6a32e. No hosted rollout, SDK release, actual worker execution or Ge acceptance is claimed. Legacy public producer routes remain denied by Caddy. Public result-frame upload, SDK/CLI adoption, tracked workloads and atomic drain coordination remain open in #406; the registration design is unchanged.

2026-09-14 — Invocation ownership development rollout

Development runs 4d1cc71, including CP #36. The built image passed typechecking and 674 tests, with 12 skipped. The rollout rechecked deployment identity and confirmed no active jobs before changing only the application image.

On the running API, signed callbacks from another worker or namespace returned HTTP 404 without changing the invocation or adding events. The assigned owner's signed progress and result were accepted. All temporary namespaces, workers, keys, functions, invocations and events were removed. No workload was launched.

Both original workers polled after the old pod was removed. Public TLS health passed; the previous 9ff0a0d image remains available for rollback. Namespace-filtered claims were verified in the local signed Mongo/API probe; this hosted check covers callback ownership. Production deployment, producer authentication, tracked-workload integration and Ge acceptance remain open. Hosted acceptance receipt.

2026-09-14 — Invocation worker ownership

Control-plane PR #36 checks that a signed acknowledgment or progress event belongs to the invocation's assigned worker and namespace. A valid signature alone previously allowed changes to another worker's invocation. Rejected requests now return HTTP 404 before the handler runs. Claims also filter by worker namespace in the atomic database query.

At dd0d56b, typechecking and 674 tests passed, with 12 skipped. Local API/Mongo checks with required Ed25519 signatures reproduced both gaps before the fix. Afterward, foreign callbacks were rejected without mutation, the assigned owner could report progress/results, and a worker claimed its own namespace's work while older work in another namespace remained queued. Local acceptance receipt.

CI passed and the fix merged as 4d1cc71. These checks use simulated workers; no hosted rollout, actual worker execution or Ge acceptance is claimed. Producer authentication, queue configuration scoping, result-journal access and tracked-workload integration remain open in #406. Legacy unsigned callback behavior is unchanged.

2026-09-14 — Producer retry development rollout

Development runs 9ff0a0d, the merged producer-retry fix from CP #35. The built image passed typechecking and 672 tests, with 12 skipped. A fresh preflight found no queued/running invocations or exec jobs; only the application image changed.

On the running API, eight identical retries of a completed fixture returned HTTP 200 and the original receipt. A conflicting payload returned HTTP 409. The completed state, result and timestamps stayed unchanged, and the fixture never became runnable. Its invocation, function and archived queue were removed.

Both original workers polled after the old pod was removed. The sole new pod is ready and public TLS health passed. The previous image remains available for rollback. This checks hosted retry behavior, not actual worker execution or full recovery. Production deployment, producer authentication, tracked-workload integration and Ge review/testing remain open. Acceptance receipt.

2026-09-14 — Producer submission retries

Control-plane PR #35 returns the existing invocation ID and queue when a producer repeats an identical submission. Reusing the ID with different work returns HTTP 409. Retries preserve worker assignment, execution state and results. Previously an identical retry returned HTTP 500.

At 2267d10, typechecking and 672 local tests passed, with 12 skipped. A fresh local Mongo replica set with the Prisma schema verified workerless persistence, API restart, one command from two concurrent worker polls, result delivery from a simulated worker, and retries after completion. Eight concurrent submissions for a new queue/function created one invocation and all returned 200. An earlier fixture omitted unique indexes; that fixture was corrected before this check.

CI passed and the PR merged as 9ff0a0d. Runtime rollout and Ge review/testing remain pending. This probe uses simulated workers and legacy unsigned polls; it does not verify Python/Slurm execution, signed-worker ownership or full recovery.

The existing Invocation queue already stores waiting work. Integration #406 will reuse it. Producer namespace authentication, cancellation/deadline/drain coordination and tracked-workload integration remain open. The legacy producer routes accepted unauthenticated requests in the isolated API test even with an admin token configured; public ingress exposure was not tested.

2026-09-14 — Dashboard confirmation deployed

Staging #251 deployed ba90c59 after all CI passed: 124 browser tests passed, 17 skipped, plus component tests, build/typecheck and formatting. It includes the Notes membership and navigation repair from #254 and the controlled hover-test timing from #256. The prior failures remain in CI history; the repair issues are closed.

Production #252 passed its own CI and deployed 7bc9962. Its only change is the confirmation wording from #250: running jobs continue, and machine shutdown is managed separately. It does not promote the newer staging Notes implementation or change an API request.

Published-site pointers and served confirmation modules were verified for both environments. Staging's authenticated Lakeshores page reloads successfully under ge-yang-33ce5a, but has no connected Lakeshore. No live queue action or Ge review/testing was performed. Queue admission, scaling and provider termination remain separate work; deployment of this text does not verify those behaviors.

Deploy receipts: staging 6aa7dbd16df47500083e6b26; production 6aa7dd562775ae0008c9f693. Durable waiting admission is tracked in #406.

2026-09-14 — Development control-plane rollout

Development now runs aba4a1e, containing CP #31–#34. The built image passed typecheck and 668 tests, with 12 skipped. A fresh preflight found no active jobs; the rollout changed only the application image, with no schema change.

Authenticated HTTP checks against the running API rejected draining/archived queue exec requests and an unsupported tracked manifest. Reactivation restored admission eligibility. The isolated namespace had no workers, so its active queue returned HTTP 503; no workload was submitted. All fixture queues and the namespace were removed, with zero fixture jobs/workers confirmed.

Both existing workers polled after the old pod was removed. The single new pod is ready, public TLS health passed, and the uploaded build directory was removed. Receipt and reusable API check. Production rollout, live scaling, durable tracked admission and provider shutdown remain open. The draft Slurm runtime was not deployed; Ge testing is unconfirmed.

2026-09-14 — Queue drain confirmation

Dashboard PR #250 corrects the drain confirmation: running jobs continue and machine shutdown is managed separately. The previous text promised worker scale-down. The endpoint marks the queue as draining; the controller skips inactive queues. Neither confirms that provider resources have stopped.

The change is text-only; API requests are unchanged. CI passed and #250 merged as 364dd19. Staging and production deployment are recorded above; Ge testing remains open. The admission correction in CP #34 passed CI and merged as aba4a1e; its documentation was published in #401. Runtime deployment, atomic drain coordination and provider shutdown remain open.

2026-09-14 — Reject exec admission to inactive queues

Control-plane PR #34 returns HTTP 409 when a queue-routed exec selects a draining or archived queue. Previously the route checked existence but ignored state, so these queues could still receive new work. Rejection creates no exec record or worker FIFO entry. Reactivating the queue permits submission again; an absent implicit default queue retains its existing behavior. Direct worker-pinned requests do not select a queue.

Four route regressions cover named/default queues in both inactive states, including reactivation. All four reproduced HTTP 202 before the fix and pass afterward. At 30302f4, typecheck and 668 tests passed, with 12 skipped. CI and runtime deployment remain pending; Ge review/testing is unconfirmed.

This state check does not serialize a concurrent drain with admission. Atomic drain/claim coordination and durable tracked waiting admission remain open. The queue controller skips non-active queues, and its default removal hook deletes worker records; draining does not establish provider-resource shutdown. The UI shutdown wording correction is tracked in dashboard PR #250 above.

2026-09-14 — Reject unsupported tracked queue requests

Control-plane PR #33 rejects an explicit tracked_run field on the legacy namespace exec endpoint. Previously, the endpoint ignored the manifest and queued the accompanying shell command. Requests now receive HTTP 400 before creating an exec record or modifying the worker FIFO, whether they select a queue or pin a worker. Legacy shell requests keep their existing behavior.

Four route regressions reproduced HTTP 202 before the fix and pass with HTTP 400 afterward, including explicit null manifests. At cb3db2c, typecheck and 664 tests passed, with 12 skipped. CI passed and the fix merged as 06c9001; runtime deployment remains pending.

Queue admission still needs implementation

The current namespace exec route selects an eligible worker before storing a job. If none exists, it returns HTTP 503. The elasticity controller counts queued legacy invocations. These paths do not provide durable waiting admission for tracked runs.

The next queue implementation must persist accepted demand independently of worker availability, then assign compatible capacity with durable ownership and replay. Counting existing exec records alone would not enable scale-from-zero submission. Providers #243 remain first; queues #253 tracks admission, capacity limits, atomic drain and provider termination.

2026-09-14 — Exclude unfinished exec jobs from scale-down

Control-plane PR #32 fixes a scale-down candidate check. Workers reporting zero invocations could still own pending or running exec jobs. Standalone queues and shared pools now exclude those workers, including jobs assigned through another queue and workers without invocation-count capabilities. Terminal exec records do not prevent scale-down.

At 4c41634, eight active-job regression cases failed before the fix. All 12 new cases pass afterward; the full suite passed typecheck and 660 tests, with 12 skipped. CI passed and the fix merged as e6ba5bf. Deployment and Ge review/testing remain pending.

This check does not reserve a worker against new assignments. Atomic draining, provider termination, compose ownership and live acceptance remain in queues #253. Scaling remains behind the existing explicit enablement setting.

Client and UI inventory

SurfaceReusable implementationRemaining integration
CLI 546d028Queue-definition CRUD and launch --queueRuntime depth, scaling, events and drain operations
Python e9ab8b3Queue placement in workflow schemaRuntime queue operations and examples
UI 0700a76Lakeshore queue list, create, drain and delete controlsDepth, capacity, scaling limits, events and linked jobs

Source inventory. The UI's drain confirmation promises worker scale-down, while the controller's default removal hook only deletes a worker record. Provider cleanup must be implemented and verified before these controls count as safe idle release. Control-plane registration #274 retains its design hold. Providers #243 remain first in the delivery order in master #247.

2026-09-14 — Controller inventory and launch expiry

Queues #253 follows providers #243 in master #247. The control plane already has queue routes, scaling policies, shared pools, worker matching, hysteresis and decision events. It does not need a replacement controller before these gaps are addressed.

AreaExisting behaviorRemaining work
Scaling policyFixed, fully elastic, threshold pool and maximum-count policiesVerify limits against real provider capacity and failure recovery
Launch accountingIn-process holds with a 15-second default expiryDurable reservations, restart recovery and coordination across controllers
Launch expiryScoped keys could retain expired holds indefinitelyFocused fix in CP #31; no live scaling claim
Scale-downChecks invocation counts, then calls a kill hook; the default hook deletes the worker rowDrain and verify provider termination; retain ownership after failure
Provider integrationExisting provider references and daemon templatesIntegrate owned associations and tracked executions; readiness is not capacity
Fleet ownershipQueue membership selects scale-down candidatesPreserve compose-owned fleets and revalidate ownership before retirement

Source: controller and bootstrap. The bootstrap enables elasticity only when ELASTICITY_ENABLED=true and uses the default kill hook. This inventory does not establish the setting on any live host.

Control-plane PR #31 fixes expiry for namespace-scoped queue keys. The old code split a combined key at the first delimiter and failed to remove the matching launch. The fix stores queue keys and token deadlines separately. A deterministic regression failed before the fix and passes afterward; namespace, shared-pool and token-delimiter cases are covered. pnpm test passed typecheck and 648 tests, with 12 skipped.

The fix merged as 5b0735c after CI passed. Deployment, real scaling, CLI/Python/UI inventory and Ge review/testing remain separate. Do not infer safe provider shutdown or distributed launch limits from these tests.