Providers: register infrastructure and run workloads on Slurm and Kubernetes
Status: Proposed. Tracking: Providers #243. Depends on host enrollment #218; credential storage follows vault #241. Commands below are interface sketches, not implemented guarantees.
Proposed provider/run contract covers stable associations, placement, compatibility and recovery. It is not frozen or implemented.
Outcome and boundary
Register infrastructure, optionally provision it, associate an enrolled submission host, and run the same Python workload on bos14/Slurm and AWS Kubernetes with durable results. A provider describes infrastructure and access scope; a host is a machine; its runner submits the work. Existing hosts enroll without providers. Provider registration, nymph connectivity, and Terraform validation alone do not prove execution.
Keep host identity/bootstrap in #218 and vault authorization/secret lifecycle in #241. Store provider configuration and credential references, not copied secret payloads. An infrastructure account or host namespace grants no implicit vault access. Source/runnable registration remains separate from infrastructure registration.
Shared lifecycle
| Operation | Required behavior |
|---|---|
| Register/adopt | Validate an existing cluster/account declaration and return a stable provider ID; no infrastructure creation. |
| Provision | Preview/apply an explicit Terraform resource plan when new infrastructure is requested. Preserve independent operation IDs and partial outcomes. |
| Associate | Bind the provider to authorized enrolled host IDs and runner configuration; maintain per-Unix-user Slurm execution identities. |
| Status | Report infrastructure access, nymph connection, runner readiness, capacity, and queued/running work separately. |
| Run | Resolve placement unambiguously, stage code/data, submit with an authorized runner, and return a durable run ID plus scheduler/pod identifiers. |
| Logs/results | Stream and retrieve output after client exit; preserve success, failure, cancellation, and lost-connectivity states. |
| Cleanup | Cancel owned workloads and remove owned temporary artifacts. Deregistration does not destroy adopted resources, stop shared hosts, or delete shared credentials. Provisioned infrastructure teardown is a separate explicit operation. |
Freeze one HTTP contract with equivalent CLI/Python operations, typed results/errors, authentication, idempotency, and status lookup. CLI commands require Python counterparts and paired documentation examples; do not invent supported SDK methods before implementation. Current singular provider create scaffolds configuration; reconcile it deliberately with the proposed public plural group.
Guide 1 — Adopt bos14 and execute through Slurm
Start with enrolled fortyfive/bos14/bos14-ctrl and the operator's existing Slurm account. Register the existing cluster, associate the host, and validate partition/account/shared workdir and uv. Nymph submits under that Unix user; compute nodes execute the workload. No Terraform provisioning or daemon on every node is required.
Prove a real sbatch job starts on an allocated compute node, reads staged code and required data, streams output, and leaves retrievable results. Cover nonzero exit, cancellation, queue delay, insufficient resources, and connection loss. Cluster status must agree with scheduler evidence. A signed Slurm heartbeat is connectivity evidence only.
Related docs:
- Host enrollment — Required host identity, bootstrap, and readiness contract.
- Slurm reference — Existing submission-manager and scheduler constraints.
- Slurm adoption template — Partition, account, host, and shared-workdir inputs.
Guide 2 — Adopt or provision AWS Kubernetes and execute a pod
Choose the actual target explicitly: existing EC2/k3s and managed EKS are different environments. An existing cluster needs no Terraform apply. For new EKS, review the resource example separately; private API access, IAM/Kubernetes permissions, node pools, and an operator/submission host must be established before association. Provisioning a cluster does not install a ready runner.
Prove a real workload pod is scheduled with the expected image/resources, receives staged code/data, and produces durable output and terminal state. Cover image pull/scheduling failures, cancellation, client disconnect, and node/pod loss. Cleanup removes only the run's owned resources. Keep Kubernetes API authentication, SSH to the submission host, and secrets consumed by workloads separate.
Related docs:
- Resource setup examples — AWS SSH-host/EKS Terraform examples, adoption boundaries, and network prerequisites.
- AWS EKS Terraform — Reviewed example variables and mocked resource tests; no live apply evidence.
- Kubernetes reference — Existing cluster submission behavior and limitations.
- Managing host credentials — #241 credential references and independent lifecycle; no implicit workload-secret injection.
Implementation and acceptance
- Freeze provider declarations, stable IDs, association, placement/runner selection, and access checks jointly across HTTP, CLI, and Python.
- Implement registration/adoption and optional provision-plan/apply with independent failure/retry tracking; never provision during registration implicitly.
- Implement cluster status with infrastructure, host, scheduler/capacity, and workload evidence.
- Deliver durable run/log/result/cancel APIs and code/data staging that work from a plain non-Git script.
- Complete real bos14 execution and real AWS Kubernetes execution with Python and CLI consumers, including negative paths and owned-resource cleanup.
- Test namespace/user isolation, credential authorization, interrupted operations, duplicate retries, and references surviving host/provider renames or retirement.
- Publish concise paired CLI/Python guides and API/error references, retaining both code variants in generated artifacts.
Acceptance uses real HTTP routes, an isolated persistent database, and subprocess CLI/Python consumers. Mocked scheduler/Kubernetes tests are useful but do not replace the two real target runs. Record exact cluster, revision, run IDs, scheduler/pod IDs, outputs, tests, skips, and blockers. Existing 844 CLI tests and mocked Terraform checks are baseline evidence only; no provider lifecycle or end-to-end workload completion is claimed by this plan.