Provider execution
2026-09-15 — GPU allocation configuration recheck
Bounded diagnostic job 147184 requested one GPU on bos14-node-031.
While running, Slurm again reported gpu:rtxpro6000:0(IDX:) and no
CUDA_VISIBLE_DEVICES. It completed 0:0 after 45 seconds; its owned stage
was removed after checking job ownership and terminal state.
The controller enables configless operation, but the worker's live command is
/usr/sbin/slurmd --systemd. A local slurm.conf and GPU device mapping exist;
/run/slurm/conf and /var/spool/slurmd/conf-cache are absent. This supports
local configuration use; the daemon environment and actual opened file were
not traced. A missing configless cache is not proof of a missing local mapping.
No cluster settings changed. GPU allocation and device isolation remain
unverified in #354.
Follow-up job 147185 confirmed configuration drift: the worker and controller
have different configuration hashes, and the worker's local node definition uses
an older address. Both still advertise one GPU. Installed worker tools report
Slurm 23.11.4. This does not prove the drift causes the allocation failure.
Both default and explicit-local slurmd -G diagnostics exited 1 during protected
cgroup initialization, before producing merged GRES output. The collector itself
completed 0:0 after 45 seconds; that is not a passing GRES diagnostic. Its stage
was removed after ownership and terminal checks. Configuration changes require
coordination with the cluster owner; no reconfigure or daemon restart was run.
Receipt
· Slurm configuration precedence.
2026-09-15 — Current-source bos14 acceptance
API draft #389 now incorporates main's private-setup rejection, queue-mount execution boundaries and owner-scoped recovery warnings alongside Slurm placement. Typecheck and 2,621 unit tests passed. Fresh npm CLI 0.21.1 and PyPI Python 0.16.2 passed 81 focused HTTP/Mongo checks; one optional browser test was skipped.
On bos14, main eee37a89, CP 8531b60 (runtime 8a3787b) and Nymph 18fc21e
completed four CPU-only probe/workload jobs through the public API. Retries kept
the same receipt IDs; returned UID, selected input and exact arguments matched.
Four signed results arrived. The isolated worker, remote staging, local services
and databases were cleaned up. Host identity was seeded for this test.
Live receipt.
The deadline check at main 7de53132, with the same CP and worker, delayed a
one-second request past expiry and dropped the first CP response. Seven identical
delivery attempts recovered one timeout/-2. The expired request had no scheduler
observation or worker stage, and every recorded deadline matched its original API
receipt. Four normal jobs also succeeded; owned runtime cleanup passed.
Deadline receipt.
Slurm briefly reported InvalidAccount while a job was pending; it cleared
without an account change. The user queue was empty after cleanup. These checks
do not cover production enrollment, GPU isolation or hosted browser acceptance.
Missing scheduler-history policy, claim-v1 Slurm integration, ordered rollout
and Ge review/testing remain open. Runtime PRs remain draft.
2026-09-15 — Paired Slurm control-plane reconciliation
Control-plane draft #29
is reconciled with current main at runtime source 8a3787b; head 8531b60 adds
its receipt and documentation. Private authorization, exact receipt recovery and
capacity checks remain in place. Registered claim workers cannot acquire or
recover Slurm jobs through stale capabilities or a missing claim envelope;
ordinary claim-based jobs still proceed. An existing incompatible Slurm claim
returns HTTP 409 without changing the job or claim. This matches Nymph 18fc21e.
Typecheck and 728 local tests passed (17 skipped). All 19 isolated real Mongo/HTTP checks passed. The paired process fixture recovered one scheduler submission and one signed result after lost poll/submission responses and token-free restart; the accepted request and diagnostic survived. Owned processes and data were removed. The scheduler commands were local fixtures, not a live cluster.
Real Mongo rejected the initial JSON-filter queries; the final implementation keeps the indexed scalar query and checks claim metadata on the returned row. The failed attempts are retained in the integration receipt. Fresh CI executed zero steps because GitHub account payments/spending availability blocked startup. Claim-v1 Slurm admission/recovery, missing-history policy, current-source live acceptance and release remain open. No runtime was deployed.
2026-09-15 — Slurm and current worker protocols
Nymph draft #31 is being reconciled with current main's private-run and claim-admission protocols. Slurm requests now unwrap the ordinary tracked-request variant explicitly. Private progress carries no scheduler observation; private requests retain their separate dispatch path.
The combined source initially bypassed claim admission for Slurm. A regression reproduced a claimed request producing a result without admission. The integration now rejects claim-v1 Slurm envelopes until admission and reserved-submission recovery are paired. Claim-configured workers advertise tracked Slurm as disabled in hello, poll and refreshed capabilities. Envelope-free requests also pass the sink's capability check, so a claim-configured worker cannot silently accept the older protocol. Rejection does not produce a terminal result or fall back to local execution.
This is an integration gate, not completed claim-v1 Slurm support. The unresolved
scheduler-history policy, paired control-plane integration, current-source live
acceptance, release and Ge review/testing remain open. At 18fc21e, 436 local tests passed with 11 ignored; formatting passed. The first
combined-suite run timed out in a cancellation/deadline test, which passed
unchanged in focused and subsequent full runs. Fresh CI and current-head live
acceptance remain separate. No runtime was deployed by this change.
Integration receipt.
2026-09-14 — Kubernetes pod ownership check
Nymph PR #33 adds ownership checks for the tracked Kubernetes adapter. A saved receipt binds the accepted execution, request hash, namespace and pod name to the Kubernetes UID. Recovery rejects a replacement with the same name; deletion requests include a UID precondition.
Formatting and the full Rust suite passed at a1fa777: 363 tests, seven ignored.
The JSON probe built at 03662a2 exercised the same ownership implementation
against the existing dreamlake-dev k3s API. The Rust helper rejected a replacement,
and the API rejected deletion with the stale UID. The replacement remained intact
and unscheduled. Both pod generations and the isolated namespace were removed,
with absence confirmed. Scheduling gates prevented workload execution.
Receipt and reusable probe.
The current PR head 6aee9e6 adds the probe and evidence; fresh CI is pending.
This code is not connected to dispatch. Durable submission, source/result transfer,
readiness, original deadlines and confirmed cancellation remain under
providers #243.
No runtime release, deployment or Ge review/testing is claimed.
2026-09-13 — Slurm discovery verified on bos14
Worker 17075ce uses a user-scoped scheduler query for discovery. The previous
cluster-wide query exceeded its 1 MiB response cap and could not submit a job.
Identity checks and the response cap remain unchanged; an oversized history for
the same Unix user can still leave discovery unresolved.
Live job 146919 completed with exit 0 on bos14 using CP fa326f4 (runtime
0c33cd8). Signed RUNNING/job-ID observations were readable through the CP API.
Both output streams exceeded 3 MiB; each capture retained 2 MiB with a truncation
notice and no cancellation. One owned scheduler job and one signed result were
observed. The worker, tunnel, remote stage, test control plane and database were
removed after completion. This was a CPU workload; no GPU was allocated.
383 worker tests passed, seven were ignored, and current-head CI passed for both runtime PRs. The signed restart fixture also passed with one submission/result. Acceptance record retains the three preceding failed probes and their cleanup evidence.
Worker #31 and CP #29 remain draft. Missing-history recovery, public provider placement, hosted rollout and Ge review/testing remain open. This test does not establish those outcomes.
2026-09-13 — Signed observations reach the DreamLake API
The candidate now passes the complete observation read path: Nymph 1b53d4b
sends signed progress to the real control plane, and DreamLake returns those
facts through its authenticated run API. Three fresh DreamLake server processes
verified reads before worker restart, after token-free worker restart, and after
successful completion. Each denied an unrelated user. The final response kept
execution success authoritative while retaining the last scheduler observation
as history. One scheduler submission and one signed result were observed; all
owned processes, both test databases and runtime files were removed.
Exact evidence
records the component revisions and probe hash. The main server typecheck passed;
full CI at 74d6f74
also passed before adding this probe and evidence. Later heads require their own
CI result.
The fixture seeds an owned run receipt and uses scheduler command fixtures. It verifies observation integration, not public Slurm submission, provider readiness, placement or host enrollment. Those paths, job UI, missing-history recovery, hosted deployment and Ge review/testing remain open.
2026-09-13 — DreamLake run observation candidate
The run API candidate retains optional schedulerObservation metadata with
observedAt, raw states, jobId and diagnostic. Existing owner checks apply.
Worker-claim status remains separate from scheduler facts; older or malformed
reports cannot replace newer facts or prevent a valid terminal result. Retained
facts after completion are historical.
Typecheck and 42 tests passed against real HTTP/auth and isolated MongoDB with a control-plane protocol stub. Checks include outage/HTTP restart, stale reports, terminal precedence and a concurrent write with an identical row-update timestamp. The first concurrency guard used a JSON filter rejected by Prisma's Mongo runtime; a separate observation timestamp now provides that guard. Owned database and files were removed. Acceptance record.
This is read-side integration. Public Slurm submission/placement, the complete signed-worker-to-DreamLake acceptance path, job UI, missing-history recovery and hosted schema/server deployment remain open. Ge review/testing is unconfirmed.
2026-09-13 — Submission diagnostic survives restart
Draft Nymph #31 at 1b53d4b
saves the first submission/receipt error as one private 1 KiB escaped-text record.
Later monitors include it as a previous error beside the latest scheduler
observation. It cannot authorize submission, cancellation or completion.
Unreadable diagnostics do not stop reconciliation.
The worker suite passed 382 tests with seven ignored. A real worker/control-plane/ Mongo check then simulated an accepted submission whose response failed. The worker was killed and restarted without its enrollment token while the fixture job remained RUNNING. A newer signed observation retained the same original diagnostic and was readable through the API. One scheduler submission produced one successful signed result; no cancellation occurred. Owned processes, database and runtime files were removed.
Exact restart evidence and reusable check. This uses scheduler command fixtures, not a live Slurm cluster. Both runtime PRs remain draft. Permanent missing-history resolution, public DreamLake projection, job UI, live-cluster acceptance, runtime rollout and Ge review/testing remain open.
2026-09-13 — Scheduler observations and recovery review
Draft Nymph #31 at d2beefb
and control-plane #29
at 0c33cd8 carry optional timestamped Slurm states, job identity and diagnostics
through signed progress. The control plane retains newer observations and returns
them with the execution record. A worker claim still has status running while
Slurm may report PENDING; consumers must show observation age. Missing records
remain unresolved. Terminal results take precedence over historical observations.
The worker suite passed 380 tests with seven ignored. Control-plane checks passed 655 tests with 15 skipped; four native Mongo/TCP checks also passed, including stale-report protection. A real worker/control-plane/Mongo fixture verified signed observation delivery, API reads and token-free worker restart, with one scheduler submission and one result. The scheduler used command fixtures. Owned processes, databases and staged files were removed. Acceptance record.
Lifecycle review
keeps two recovery gaps explicit: retaining the original submission diagnostic
through reconciliation and resolving permanently missing scheduler history.
The current worker logs an sbatch error, but later progress may only report a
missing record. A missing record cannot authorize resubmission or establish a
terminal outcome. No operator-resolution endpoint was found in the tracked
execution routes reviewed at this revision.
Both runtime PRs remain draft. Public DreamLake projection, job UI, live-cluster acceptance at these revisions, runtime delivery and Ge review/testing remain open. Earlier live results below apply to their pinned revisions.
2026-09-13 — Deadline while queued
Slurm job 146764 stayed PENDING / JobHeldUser until its 45-second accepted
deadline. A test-only wrapper added sbatch --hold for the disposable worker;
no shared scheduler configuration changed. Nymph 4dc7df1 recorded deadline
intent and delivered one signed timeout result with exit code -2 after Slurm
confirmed cancellation. No API cancellation was requested. The terminal record
had no allocated nodes, TRES or batch host, and the workload produced no output.
One scheduler submission and one signed result were observed. Accepted parameters were unchanged. Owned workers, tunnels, databases and staged files were removed; remote removal was rechecked over a fresh SSH connection. Queued-deadline evidence.
The first probe, job 146763, failed a test-only timestamp assertion and was
cleaned up. Slurm 23.11.4 sets both start and end timestamps when cancelling a
pending job; a terminal start timestamp alone does not prove execution.
Slurm cancellation source.
The corrected check uses the held state, allocation and output evidence above.
This verifies deadline expiry under an explicit hold, not natural queue contention
or resource rejection. Control-plane running marks a worker claim, so provider
status must distinguish that from scheduler start. Public placement/UI,
unresolved-outcome recovery and GPU isolation remain open. Both runtime PRs stay
draft; no runtime deployment or Ge review/testing is claimed.
2026-09-13 — Slurm log capture
Nymph #31 at 4dc7df1 captures
the first 2 MiB of stdout and stderr without cancelling the Slurm workload when
that limit is reached. The final result preserves the scheduler exit status and
includes a truncation notice. Full log files remain in the stage; durable artifact
upload remains separate work. File-read errors and changes to the captured prefix
still request cancellation. The tracked local process runner is unchanged.
The full local suite passed 378 tests with seven ignored. The overflow regression covers both streams, successful and nonzero exits, unchanged full staged logs, no cancellation request, and identical results after terminal-checkpoint recovery.
Live bos14 job 146703 completed with exit 0 after writing more than 3 MiB to
each stream. The signed result captured exactly 2 MiB per stream and included the
truncation notice. Full staged stdout and stderr were 3,146,010 and 3,145,741 bytes;
no cancellation intent was created. One scheduler job and one signed result were
observed. The owned worker, tunnel, database and files were removed, with remote
removal rechecked over a fresh SSH connection.
Live log-capture evidence.
Both runtime PRs remain draft. Public provider placement/UI, unresolved-outcome recovery, queue-only expiry and GPU isolation remain open. No hosted runtime deployment, installed-worker rollout or Ge review/testing is claimed.
2026-09-13 — Accepted Slurm deadline
Slurm job 146702 ran a CPU workload with a 45-second accepted deadline through
Nymph 29e7244 and control plane 4ffc678. No API cancellation was requested.
The worker persisted deadline intent; Slurm confirmed CANCELLED about one
second after the deadline, and the signed result reported timeout with exit
code -2. Selected input/output and accepted parameters matched. One scheduler
job and one final signed result were observed.
The full worker suite passed 377 tests with seven ignored, and both CI checks
passed at 29e7244. The live check used a disposable local control plane and a
private callback tunnel. The owned worker, tunnel, database and staged files were
removed; remote removal was rechecked over a fresh SSH connection.
Deadline evidence.
This covers deadline expiry after a workload starts. Queue-only expiry, rejected placement, public provider dispatch/UI and GPU isolation remain open. Both runtime PRs remain draft; no runtime rollout or Ge review/testing is claimed.
2026-09-13 — Slurm review findings
Focused review identified two user-visible behaviors that green CI does not resolve: missing scheduler records can leave a submission unresolved indefinitely, and captured output above 2 MiB causes workload cancellation.
Nymph #31 now retains a failed scheduler command's exit status and the first 4 KiB of stderr in the worker diagnostic, with terminal controls escaped. A nonzero client exit still does not prove that no job was submitted. Recovery preserves the reservation and never submits it again merely because the scheduler returned no record.
Exposing diagnostics and a recovery action in the job API/UI remains open. The Slurm log-capture change above replaces cancellation on overflow; the tracked local process runner retains that behavior. No runtime release, deployment or Ge review/testing is implied.
2026-09-13 — Slurm worker dispatch and recovery
Nymph #31 at f76410f and
control-plane #29
at 4ffc678 remain draft. The worker dispatches explicit Slurm requests, monitors
scheduler state and delivers results through the existing durable journal. The
control plane validates new requests, assigns a deadline that includes queue time,
and redelivers missing running jobs to the same signed worker.
After submission is reserved, recovery reconciles the existing job without submitting again. Missing scheduler evidence leaves the outcome unresolved. A stage lock excludes simultaneous local owners, and active-ID polls identify which claimed jobs still have a worker task. Unix workers configured with the Slurm runner advertise protocol version 1; this does not establish cluster or shared-filesystem readiness. The worker now accepts exact integer timestamps from the control plane's JavaScript MessagePack encoder.
Validation:
- Worker: 375 local tests passed, seven ignored. Remote formatting, clippy and
test checks passed at
f76410f. - Control-plane runtime
fe4345a: 652 tests passed, 14 skipped; three native MongoDB/TCP checks passed. The later4ffc678change adds the paired harness, instructions and evidence without changing runtime code. - Paired recovery: the real control plane, MongoDB and worker passed a lost claimed response and restart without the enrollment token. One scheduler submission produced one signed result with unchanged accepted timestamps. Scheduler commands were fixtures; no live Slurm workload ran.
- Launcher failure: an invalid fixture executable exited 1. Both successful and failed checks removed their owned processes, database and runtime directories.
The bos14 submission-filesystem probe passed on /export/work (ext4): a second
process could not acquire the held lock, and a new process acquired the unchanged
lock file after the owned holder was killed. The private test directory was
removed. This checks competing processes on the submission host, not cross-host
migration or NFS-client locking. No Slurm job or installed worker was changed.
Lock evidence.
The live CPU worker check also passed: Slurm job 146649 ran on
bos14-node-031 through the candidate worker and a disposable control plane.
After signed running output arrived, only the test worker was killed. It restarted
with the same key after its enrollment token was removed, then delivered one
signed result with exit 0 and unchanged accepted parameters. Selected input,
quoted/Unicode arguments, stdout and stderr matched. One scheduler job was observed.
All owned processes, runtime directories and the temporary database were removed;
remote removal was verified over a new SSH connection.
Slurm allocated the whole 32-CPU node for a one-CPU request. No GPU was requested; this check makes no GPU isolation claim. The control plane ran locally behind a private SSH callback tunnel, not as a hosted deployment. Live worker evidence.
Additional live CPU checks passed through the same candidate worker and control plane:
- Job
146651ended in SlurmFAILEDand preserved exit code 7. The control-plane wire status isdonefor a finished process, including nonzero exits; DreamLake's existing run service maps this result tofailed. No runtime change was needed. - Job
146652was cancelled after running output arrived. Cancellation intent was recorded, then Slurm confirmedCANCELLEDand the signed result returnedcancelledwith exit code-3. - Job
146653completed after its callback tunnel was interrupted for about nine seconds. No signed polls arrived during a five-second quiet window. The same worker process stayed alive, Slurm was observed running during the outage, and polling resumed on the restored endpoint. One scheduler job and one final signed result were observed.
Selected input/output and accepted parameters matched in all three cases. Owned processes, runtime directories, temporary databases and generated server credentials were removed; remote removal was rechecked over fresh SSH connections. The callback test launcher disappeared after the workload checks passed; its remaining owned database was stopped and removed manually. The closure cause is unconfirmed. Failure and cancellation evidence · Callback interruption evidence.
Paired evidence · Run the check.
Remaining: lifecycle/complexity review, DreamLake provider placement, queue/resource rejection and deadline cases. Locking on other submission filesystems requires separate acceptance. The job UI must expose an unresolved outcome when scheduler evidence is lost; provider UI acceptance tracks this requirement. GPU allocation/isolation #354 is separate. These drafts have no runtime release, deployment or worker rollout. Ge review/testing is unconfirmed.
2026-09-13 — Proposed shared contract
Issue #243 · Master #247 · Contract draft
The draft reuses DreamLake's provider collection and namespace permissions, adds versioned host/runner associations, and extends tracked runs with explicit placement. Missing placement keeps existing process execution; invalid placement must fail without a local fallback. Configuration belongs in DreamLake and secrets in Vault. Legacy provider writes need revision guards before associations can rely on that configuration.
Current DreamLake main 47393e7 dispatches tracked runs to an exact enrolled
worker. Nymph main 317f989 runs their uv commands locally. Its Slurm/Kubernetes
invocation runners use a separate protocol. Kubernetes assumes shared hostPath
storage; Slurm still lacks streaming logs on the invocation path.
The draft covers mutation recovery, scheduler identity, cancellation confirmation, readiness observations, paired proposed client operations and real bos14/AWS acceptance. Concrete cluster storage/output transfer and the final contract remain to be agreed. #274 registration/Host UI and #301 Notes holds are unchanged.
This is design documentation. No API, client, scheduler or deployment changed. The full docs build passed (62 indexed pages), and the new hidden Dev Note rendered with noindex. Source inspection establishes the integration gaps; it does not establish live provider execution. Ge review/testing is unconfirmed.
2026-09-13 — bos14 prerequisites
A one-CPU batch probe (146455) briefly showed InvalidAccount, then started on
bos14-node-031 and failed immediately with signal 53. No output files were
created. The signal's cause remains unresolved.
A separate one-CPU srun probe (146456) completed on the same node as geyang
(UID 107400011), exit 0. It found the controller's private staging directory absent
and no uv on PATH. Controller uv also failed its version check. This disproves
readiness of that staging/PATH combination; it does not rule out other shared
storage or installed binaries elsewhere.
The contract now requires explicit source delivery or a verified shared location, plus a verified node execution environment. Direct node SSH stopped at host-key verification; no trust settings changed. Both jobs are confirmed terminal and absent from the queue, and the exact owned staging directory was removed. No DreamLake scheduler API, GPU workload or durable output transfer was tested.
2026-09-13 — Verified Slurm path mapping and uv workload
The node mounts /work over NFS. Its /home/geyang symlink points to the missing
/export/work/geyang path. Controller /export/work/geyang maps to node
/work/geyang; the contract now records separate submission and compute roots.
Job 146458 completed a plain Python script through uv on bos14-node-031, as
geyang (UID 107400011), exit 0. It verified source/input hashes, preserved the
argument vector (including an empty argument and Unicode), and produced matching
stdout/result JSON plus the expected stderr marker. Output was retrieved through
a new SSH connection after the submitting connection exited.
The test used a private uv 0.8.15 release binary,
verified against the official archive checksum and checked again by binary hash
on the node. Python used /usr/bin/python3; cache paths were inside the owned
directory, without changing HOME. Both this job and the filesystem probe 146457
are confirmed completed and absent from the queue. The owned directory, binary,
cache and test outputs were removed after evidence capture.
This verifies one mapped-storage CPU workload through Slurm. It does not deliver DreamLake provider dispatch, durable output archival, cancellation/restart recovery or GPU execution. No persistent node environment was installed, and the earlier signal-53 failure is not diagnosed by this successful follow-up.
2026-09-13 — Registration API merged
The implementation reuses the provider collection for typed, immutable declarations. Registration and its recovery receipt commit in one MongoDB transaction. Equivalent retries preserve the provider/operation IDs, and lookup by request ID recovers a lost response. Receipts store metadata only and do not expire in this slice.
Slurm associations require the caller's own host enrollment, an exact declaration revision, permitted partition/account settings and explicit submission/compute paths. Creation serializes with retirement. Namespace provider permissions do not grant access to another user's Unix enrollment. Vault references are stored without reading values or granting authorization.
Legacy provider CRUD remains available for old rows; it cannot mutate versioned
declarations. Status returns not_checked, not inferred connectivity or capacity.
Startup verifies the required uniqueness indexes and rejects expiring receipts.
The existing database schema deployment step must run before startup.
Validation uses real TCP HTTP, JWT/namespace permissions and an isolated MongoDB replica set. Typecheck and 38 tests passed across the new API, existing REST routes and GraphQL reads, including concurrent retries, post-commit response loss, server restart, enrollment isolation, retirement races and missing/TTL index gates. The final startup test also checks that the port stays closed when the guard fails.
PR #321 is merged at 95710f2. Full CI
passed 3,815 product tests (two skipped), 47 migration tests, typecheck and index
verification. The merged tree matches the tree recorded by CI. The API has not
been deployed. Kubernetes associations, readiness probes, conditional configuration
updates/legacy upgrades and tracked scheduler dispatch remain open.
The #274 registration/Host UI and #301 Notes design holds are unchanged. Ge
review/testing is unconfirmed.
2026-09-13 — CLI and Python registration clients
CLI PR #45 · Python PR #32 · Acceptance evidence
Both clients support declarations, owned Slurm associations, status, retirement and historical receipt lookup. Writes require an explicit request ID. Transport failures retain that ID; clients never retry writes automatically. Receipt lookup checks the returned identity. Paired guides cover path mapping, recovery and retirement.
Six checks against real local HTTP routes and MongoDB passed with subprocess CLI and Python consumers. They covered cross-client replay, association/status, retirement/history, conflicts, two deliberately dropped post-commit responses, and denial of another authenticated user. Cleanup removed four providers, one association and seven receipts; the owned MongoDB stopped.
CLI validation passed all 911 tests plus package and docs builds (26 indexed pages). Python validation passed 488 tests with 60 skipped; its wheel includes the provider module. Sphinx built with 38 warnings in other pages. Python PR #32 and CLI PR #45 are merged; CLI CI also passed. Package publication and hosted acceptance remain pending. This is metadata API acceptance, not scheduler execution or readiness evidence. Ge review/testing is unconfirmed.
2026-09-13 — Staging deployment and release candidates
Staging evidence · Reproduce the checks
Staging runs 73cd00f, task revision 148. The image build, database index steps,
canary rollout and final ECS health checks passed. Native CLI 0.14.0 and an isolated
Python 0.11.0 wheel passed hosted registration through both clients, replay,
receipt lookup, changed-payload conflicts, association metadata and unsupported
field rejection. Unauthenticated receipt lookup was denied.
Cleanup retired the test association and both providers, then verified retired status and active-detail absence. Historical receipts and retired metadata remain under the API retention contract. No existing host service or job changed; the association did not verify its paths, executables or scheduler readiness.
CLI release #46 and Python release #33 are merged. CLI release CI passed; all eight native targets built and the macOS arm64 binary passed a detached smoke test. Python wheel/sdist and fresh wheel installation passed. Package publication and production acceptance remain pending. Ge review/testing is unconfirmed.
2026-09-13 — Production API acceptance
Production evidence · Deployment
Production runs 73cd00f, task revision 110. The database index steps, canary
rollout and final ECS image/health verification passed. Native CLI 0.14.0 and an
isolated Python 0.11.0 wheel passed registration through both clients, replay,
receipt lookup, changed-payload conflicts and rejection of unsupported fields.
Unauthenticated receipt lookup was denied. Both providers were retired and their
active-detail absence verified; historical receipts and retired metadata remain.
There is no production enrollment fixture, so association acceptance remains staging-only. These checks do not test provisioning, readiness or scheduler execution. Package publication and fresh public-package acceptance are in progress. Ge review/testing is unconfirmed.
2026-09-13 — Public native and Python clients
Fresh-install evidence · CLI and Python examples
Python 0.11.0 is published on PyPI; both distribution hashes match the release
build. All eight CLI 0.14.0 native downloads passed byte verification. A fresh
macOS native download and uncached PyPI installation passed provider metadata
checks against staging and production at 73cd00f. Staging also passed owned
Slurm association checks. Every fixture was retired and verified, and both
temporary client installations were removed.
npm publication is incomplete. The Linux x64 package returned a publication
conflict before its metadata became visible; its public tarball now matches the
release bytes. Remaining platform packages and the root launcher are being
published. Native latest is still 0.13.0; fresh npm acceptance remains pending.
Readiness, scheduler execution and Ge review/testing remain open.
2026-09-13 — CLI 0.14.0 and Python 0.11.0 released
Publication and fresh-install evidence · CLI/Python guide · Python release
All nine npm packages match the generated release tarballs. All eight native
CLI downloads are verified, and native latest is 0.14.0. Python's PyPI and
GitHub wheel/sdist downloads match the release build. The earlier npm conflicts
and delayed visibility are resolved; no version was replaced or unpublished.
Fresh npm CLI 0.14.0 and uncached PyPI 0.11.0 installations passed metadata
acceptance on staging and production at 73cd00f. Staging also passed owned
Slurm association checks. Every fixture was retired and verified, and temporary
installations were removed. Production association checks remain absent because
there is no production enrollment fixture.
The CLI guide is live at 337e2b.
This release delivers provider metadata operations. Readiness, tracked scheduler
execution, Kubernetes associations and Ge review/testing remain open.
2026-09-13 — Tracked Slurm integration draft
Nymph draft #29 · Implementation and remaining work
The draft prepares selected files under an explicit submission root and generates compute-side paths for the batch script and outputs. Staging survives the preparation call. Existing execution directories cannot be overwritten, and the source manifest uses the same validation as local tracked runs.
A durable intent reserves submission before sbatch may run. Concurrent reservations and interrupted writes cannot silently create a second submission. Scheduler receipts preserve cluster, Unix UID and ownership marker; conflicting identities cannot replace an earlier receipt.
Eleven focused tests passed, including execution of the generated script with mapped paths and empty, quoted and Unicode arguments. The full Nymph suite passed 340 tests with six ignored; Clippy completed with warnings in unchanged code. These are local tests. The module is not wired to dispatch and no live Slurm job was submitted. Readiness, placement, submission/reconciliation, logs, confirmed cancellation, real bos14 acceptance and Ge review/testing remain open.
2026-09-13 — Controller recovery and direct adapter acceptance
Nymph draft #29 · Adapter evidence
The draft now submits jobs, recovers lost responses by ownership marker and records controller observations. Cancellation remains a durable request until a fresh terminal observation confirms the result. Empty job lists, unknown states and incomplete exit details remain unresolved. Environment options cannot override the accepted placement, automatic requeue is disabled, and run limits are explicit.
The Linux adapter at f870667 submitted bos14 job 146465. Separate monitor
processes recovered its receipt and terminal state: completed, exit 0. Source hash,
selected input, empty/quoted/Unicode arguments, stdout, stderr and result JSON
matched. The request used one CPU, 512 MiB, no GPUs and a two-minute walltime;
Slurm reported 32 allocated CPUs. The private staging directory, uv cache, binaries
and local test assets were removed. No installed worker or host service changed.
bos14 runs Slurm 23.11.4 with accounting and job-completion storage disabled and
MinJobAge=300. A completed job can disappear before a restarted monitor sees it.
Persisted observations survive that expiry, but a missed terminal observation
remains unknown. Explicit --clusters=bos14 also fails without accounting storage;
the adapter verifies and uses its local controller. Schema probe 146464 completed
with exit 0 and supplied the reduced JSON parser fixture.
Nineteen focused tests and the full local suite passed (348 tests, six ignored). Both Linux CI jobs passed. Local tests cover lost responses, command timeouts, foreign identities, cancellation during submission and later terminal confirmation. The live check covers direct adapter execution, not public CLI/Python run submission or daemon dispatch. Those connections, live fault injection, GPU execution and Ge review/testing remain open. The runtime PR stays draft.
2026-09-13 — Terminal checkpoint recovery
Nymph change 5058c7d ·
Linux check ·
Linux tests
After observing a confirmed terminal state, the adapter saves an immutable checkpoint. A restarted monitor validates the stage, Unix owner and scheduler receipt, then returns that checkpoint without querying the controller. This preserves a recorded result when the controller is unavailable or its job record has expired. A job whose terminal state was never observed remains unresolved.
The recovery test removes controller access after saving completion and verifies
that a new monitor still returns the terminal result. Nineteen focused tests and
the full local suite passed (348 tests, six ignored); both Linux CI jobs passed at
5058c7d. The earlier bos14 acceptance used f870667 and does not test this change
on the cluster.
Upstream result acknowledgment, durable output delivery, public run placement and daemon dispatch remain open. The runtime PR remains draft; no worker was upgraded and Ge review/testing is unconfirmed.
2026-09-13 — Slurm review and fingerprint correction
Nymph draft #29 · Review guide
The submission fingerprint omitted CPU, memory, GPU and time limits. It now includes all four. A regression test changes each limit while keeping the execution ID, paths and workload fixed; it fails with the previous calculation. An additional check rejects receipt recording with an altered stage descriptor. Stage/intent validation now uses one shared function instead of three copies.
The review guide identifies four guarantees: preserve the accepted request, never blindly resubmit, act only on the recorded job identity, and report only observed completion. It separates the adapter from public dispatch and output delivery, and records the journal retention work still needed for long-running monitoring. The PR remains draft; passing tests do not imply Ge approval.
Focused validation passed 21 tests. A first full-suite run hit two ordinary
fixture command timeouts. Those fixtures now use the production command limit;
the dedicated timeout test retains its deliberately short limit. The corrected full suite passed 350 tests with six ignored at db3540f;
Linux CI for that revision is pending. No cluster job or installed worker changed.
2026-09-13 — Exact result acknowledgment
Control-plane PR #27 · Local acceptance evidence
A tracked result callback could report success without storing its result, or when its payload disagreed with a stored terminal result. The callback now checks the committed exit code, duration, error and both output streams before returning success. Identical replay is accepted; conflicts and unclaimed jobs return 409. A final write also compares the log prefixes it validated, preserving progress accepted while the callback was in flight.
The old route fails all three new regressions. Typecheck and the full local suite passed (646 tests, 12 skipped). A separate real MongoDB/TCP check passed competing results, replay after server/client reconnect and the progress/final-write race. The disposable Mongo process stopped and its directory was removed. Signature checks use the existing authentication suite; the Mongo harness isolates storage and HTTP behavior. Hosted rollout and acceptance remain open.
The Nymph db3540f Linux Check rerun passed, as did its all-features job. The first
Check attempt timed out in the unchanged local-uv descendant test; that failure
remains recorded in CI. Nymph #29 stays draft. This separate callback fix does
not connect Slurm dispatch or make worker result storage durable.
2026-09-13 — Development result acknowledgment deployed
Control-plane #27 · Hosted acceptance evidence
Development now runs 826e976. The built image passed 646 tests (12 skipped).
Before rollout, no queued or running jobs were present. The update changed the
application image with a deployment-version guard; configuration, secrets and
schema were unchanged. Both existing workers sent fresh polls after the old pod
was removed. The new pod is the sole ready replica; public TLS health and the
authenticated API passed.
Temporary signed callback fixtures passed exact replay, conflicts in all supported result fields, unclaimed-result rejection and unsigned/foreign/revoked signer denial. The deployed route source hash matches the reviewed source. Fixture jobs, keys and namespace were removed and verified absent; the uploaded build directory was removed. No workload was dispatched to either existing worker.
This verifies the callback fix in development. Production rollout, worker-side durable result delivery and acknowledgment handling, already-claimed run recovery, and public Slurm dispatch remain open. Nymph #29 remains draft; no worker binary was upgraded and Ge review/testing is unconfirmed.
2026-09-13 — Durable tracked result delivery draft
Nymph draft #30 · Behavior and limits
The draft saves a completed tracked result before sending it and removes the pending record only after an explicit acknowledgment. Startup replays pending records independently of new work. The journal is private, beside the identity file, and scoped to the control-plane address and worker ID. Key rotation keeps that worker scope; another address or worker cannot replay it. A process lock prevents concurrent journal owners.
The local executor retains a result while alive and retries rather than giving up after eight callbacks. Records are immutable; conflicting or corrupt records remain on disk for reconciliation. Each record is bounded to 6 MiB, and replay loads one result at a time. Disk loss or process loss before successful persistence can still lose a result. Running-work and lost-claim recovery are separate work.
At 13d0e89, the full local suite passed 338 tests with six ignored. Journal tests
cover reopen, conflicting records/acknowledgments, locking, scope isolation,
corruption and symlinks. A real HTTP test loses a response, recreates the client,
returns negative/missing acknowledgments, and verifies identical payload bytes
through final acceptance. Both Linux jobs passed for 13d0e89.
Subsequent acceptance is recorded below. No worker binary was upgraded. This PR is separate from Slurm #29; both remain draft, and Ge review/testing is unconfirmed.
2026-09-13 — Result delivery acceptance and persistence retry
Nymph draft #30 · Hosted acceptance record
The actual daemon-process test passed: a saved result survived process termination, and the restarted daemon replayed identical bytes before removing the record after acknowledgment. This local fixture did not dispatch a workload.
Hosted delivery passed against development control plane 826e976. A proxy
forwarded the Rust client's signed callback over HTTPS and dropped the first
successful response. A recreated client replayed the saved result; the server
acknowledged the identical retry. Stored result fields matched. Disposable database
records and the temporary private key were removed.
Review found that a failed directory sync could leave a visible result record that
subsequent delivery treated as durable. Both foreground retry and background replay
now require a successful directory sync before sending. The injected-failure test
rejects the old behavior and passes with the fix at cc6a526. The full local
suite passed: 340 tests, seven ignored (including opt-in hosted acceptance, run
separately).
The descendant-capture test now separates Python startup from its capture-drain deadline. It passed locally and failed when the production drain deadline was deliberately extended; production behavior was restored before full validation.
New-head CI and review remain pending. These checks do not prove that an executing workload survives a hosted daemon restart. Running-claim recovery, Slurm dispatch, runtime release, worker rollout and Ge review/testing remain separate.
2026-09-13 — Merged result delivery and simpler Slurm checkpoints
Result delivery #30 merged at
94cd852 after both Linux CI jobs passed for e90500b. The full local suite
passed 341 tests with seven ignored. Review corrected directory-sync retries and
force-drain behavior: failed callbacks cannot keep retry backoff alive indefinitely
after the drain deadline. Saved results remain available for startup replay.
Slurm draft #29 at c13b082 removes
unused per-poll observation history. Live and missing states are returned without
adding files; only the first confirmed terminal result is checkpointed. Repeated
polling, controller-unavailable recovery and duplicate/conflicting checkpoint
checks passed in all 16 focused tracked-Slurm tests. The combined branch at
983e8ac includes merged result delivery and passed its full local suite (357
tests, seven ignored). New-head CI and runtime integration remain pending.
The result-delivery merge does not release a runtime version or upgrade a worker. Slurm public dispatch, stage/output cleanup after acknowledgment, live fault/GPU acceptance and running-claim recovery remain open. Ge review/testing is unconfirmed.
2026-09-13 — Live Slurm cancellation and checkpoint recovery
Slurm draft #29 · Acceptance record
The Linux adapter at 983e8ac passed three CPU-only checks on bos14:
| Job | Check | Observed result |
|---|---|---|
146484 | Normal completion | Exit 0; selected input, argv, stdout/stderr and mapped working directory matched. |
146485 | Cancellation while pending | The adapter requested cancellation and later observed CANCELLED. |
146486 | Cancellation after RUNNING | Identity-checked cancellation was followed by confirmed CANCELLED. |
After Slurm's 300-second retention period, controller queries confirmed all three records were absent. New adapter processes recovered the exact saved terminal observations; checkpoint hashes were unchanged. Each stage had one terminal checkpoint, no polling-history files and no leftover checkpoint temporary files. Both private staging directories and local test assets were removed.
The running jobs used bos14-node-031. One CPU and 512 MiB were requested; Slurm
allocated 32 CPUs and no GPU resources. This validates the selected shared-path
mapping on that node, not every node or GPU execution.
Both Linux CI jobs passed at 983e8ac. A subsequent review correction at bf6120b
retries directory sync for an existing cancellation intent before calling the
scheduler; its full local suite passed 357 tests with seven ignored. New-head CI
remains pending. The live runs above used the earlier 983e8ac binary.
The adapter remains draft. Public placement/daemon dispatch, output archival and cleanup after acknowledgment, overall queue-plus-runtime deadlines, running-claim recovery and GPU acceptance remain open. No installed worker changed, and Ge review/testing is unconfirmed.
2026-09-13 — GPU allocation acceptance remains open
Nymph #29 at 7e06502 requests
single-node GPUs with --gres=gpu:N, which supports bos14's select/linear.
The full local suite passed: 358 tests, seven ignored.
Jobs 146488 and 146489 completed a CUDA kernel with the expected result,
41 + 1 = 42. During RUNNING, job 146489 reported
gpu:rtxpro6000:0(IDX:) despite requesting one GPU. GPU allocation and isolation
are therefore unverified. Inspecting only a completed job's empty GRES record
was insufficient; the evidence now includes a running observation.
New adapter processes recovered the terminal checkpoints. All owned test stages
were removed. Acceptance evidence
includes the earlier --gpus run 146487 without treating it as GPU allocation
acceptance. No cluster configuration, installed worker or runtime release changed.
The PR remains draft; new-head CI and the remaining integration gates are separate.
2026-09-13 — Reject unsupported tracked-run fields
Control-plane #28 rejects unknown fields in tracked requests, execution options and source records before storing a job. Previously a request containing provider placement returned HTTP 202, while the current worker ignored the placement field and ran locally. The corrected route returns HTTP 400. Provider placement itself remains unimplemented.
At cd6fdbe, typecheck and 647 local tests passed (12 skipped). The HTTP regression
fails against the previous validator with 202 instead of 400. The current
DreamLake run service sends the supported fields; the legacy exec route is unchanged.
Control-plane #28 merged at 93c295a after CI passed. Hosted deployment and
worker upgrade remain separate.
2026-09-13 — Hosted tracked-request validation
Control-plane #28 is deployed in development at 93c295a. The built image passed
647 tests (12 skipped); rollout started after verifying no queued/running work.
Public HTTPS checks returned 400 for five unsupported-field requests and stored
no jobs. A valid pre-cancelled request returned 202, then replayed the same receipt
with 200. No workload was dispatched.
Both established workers polled the new control plane; public health returned 200. The disposable namespace, worker and execution records were removed, as was the uploaded build directory. The first verifier attempt omitted a required fixture field; cleanup was verified before correcting it and rerunning. Hosted evidence.
This is development acceptance. Production rollout, provider placement, runtime worker upgrades and Ge review/testing remain separate.