DreamLake

Declarations vs. execution

Proposal — in the tree, not yet in the database

This page is the design of record as of 2026-08-06 — argued in plans/2026-08-06_dreamlake-declaration-collections/README.md, with the Prisma blocks verbatim in that directory's schema.md. The five models, the twelve permission keys, the host-key helper and the collection routes are now in the dreamlake-server tree, and the left-nav rows are wired.

They are not live on any shared server. prisma db push is a deliberate, out-of-band, human step against a shared Atlas cluster, and this work did not run it — so the collections exist in source and not yet as documents. Read the shapes below as settled and the availability as pending. Parts that are further out still carry their own Proposal callout.

The rule

DreamLake owns declarations. The lakeshore controlplane owns execution. A declaration is a thing an author writes down once and then names — a provider, a storage connection, a UDF, an agent, the environment a UDF runs inside, the git tree it imports from. It is durable, editable, and it belongs to a namespace. Execution is everything that happens because somebody submitted work — a queue, an invocation, a worker, an event, a live channel. It is ephemeral, it is high-churn, and it belongs to the controlplane that is actually running it. Every placement question below resolves against that one sentence.

Ge, 2026-08-06:

those configs live on dreamlake server. lakeshore cp is responsible for the runtime / ephemeral states and the queue.

DreamLake (declarations)Lakeshore controlplane (execution)
Tenancy — Namespace, users, orgs, teamsQueues — including fleet policy and elasticity
Credentials — the AES-256-GCM secret blobsInvocations — one row per submitted run
LakeshoreProvider — launcher credentials + placementWorkers / daemons — the live fleet
Source — data and storage connectorsEvents — the run log
Repo / RepoSnapshot — registered git treesSessions and live channels
Runnable / RunnableVersion — UDFs, sessions, agentsInvocation.runConfigRef + Invocation.resolvedConfig
RunConfig — the named environment template

The last row is the one worth staring at. RunConfig is the template: a named, reusable, namespace-scoped record that a human edits. The instance — the merged snapshot of what a particular run actually got — is Invocation.resolvedConfig, and an invocation is execution, so it stays on the controlplane. Splitting the two is what keeps RunConfig from naming two different things.

Fleet policy stays with the queue for the same reason it always has: a guardrail that lives inside a generated artifact is not a guardrail. The permitted instance types, the elasticity bounds and providerRef are the controlplane's to enforce.


The four collections

Ge named the set as provider / storage / UDF / agents. They cohere as a set, but not because they share a landing page — there is deliberately no /:ns/declarations index. They cohere in two concrete ways.

One URL grammar and one authorization root. Every collection sits at /namespaces/:slug/<collection> and behaves identically: the same response envelope, the same { items, page, pageSize, total, totalPages } list shape, the same id-or-name resolution (24-hex reads as an id, anything else as a name), the same soft delete via deletedAt, the same 409 on a duplicate name, and the same authorize(userId, '<x>:<verb>', { kind: 'lakeshore', namespaceId }, prisma) call. Reusing ctx.kind: 'lakeshore' is a choice, not an accident — AuthContext is a closed union, and a new arm would need a role mapper byte-identical to resolveLakeshoreRole. The consequence is that these collections inherit exactly the ADMIN/READ derivation that lakeshore connections already have.

One left-nav group. The lakeshore section of DreamLakeLeftPanel is the index. It already ships dead rows for "UDFs" and "Providers" with no onClick; wiring them is one line each.

CollectionRouteModelStatus
Providers/namespaces/:slug/lakeshore-providersLakeshoreProviderroutes exist, no UI
Storage/namespaces/:slug/lakeshore-storagesLakeshoreStorageroutes exist, collapsing into Source
Sources/namespaces/:slug/sourcesSourceships today, UI at /source/:ns
Runnables/namespaces/:slug/runnablesRunnable + RunnableVersionnew
RunConfigs/namespaces/:slug/runconfigsRunConfignew
Repos/namespaces/:slug/reposRepo + RepoSnapshotnew

Two sub-resources hang off the two-layer collections: /namespaces/:slug/runnables/:idOrName/versions and /namespaces/:slug/repos/:idOrName/snapshots. Both are list / get / register. Registration is idempotent by content: re-registering an existing (runnableId, version) or (repoId, commit) returns 200 with the existing row and rewrites nothing, because there the key is the content.

The left-nav rows the set produces:

RowRoute
Overview/:ns/lakeshore
Sources/source/:ns
Providers/:ns/providers
RunConfigs/:ns/runconfigs
UDFs & Agents/:ns/runnables
Repos/:ns/repos

Runnable — one model, three lifetimes

Runnable is a single model with a free-string kind ∈ udf | session | agent. The kind distinguishes lifetime, not substance — all three are named code you invoke. The ladder, plainly:

  • udf — one-shot. You submit, you get a result back, it is over.
  • session — long-lived, with a bidirectional pipe attached.
  • agent — a session, plus a policy driving it.

Each rung adds to the one below it; nothing is taken away. kind is a plain String rather than a Prisma enum, matching Source.kind and LakeshoreProvider.kind, so a fourth lifetime never needs a schema change.

An Agent is a UDF plus exactly two fields

Ge, 2026-08-06:

For the agents, start the definition with stuff from the UDF, plus a channel and a markdown. channel does not need to be configed at real time I think but we can include it.

Checked field by field against the controlplane's Function, that holds up. An agent is decorated code, so it has a module and a qualname; it has a source and therefore a version; its signature is the entrypoint's parameters, which is degenerate but not meaningless. Nothing on the UDF side is empty for an agent. Only the reverse holds, and only twice — markdown and channel. A definition stated as "X plus two fields" is a discriminator, not a second shape, so there is no Agent model.

markdown is read as behaviour — the agent's instructions, authored as Markdown and rendered on the detail page. It is not a description card; description String? already covers documentation.

Both fields live on the mutable Runnable head, not on RunnableVersion, because RunnableVersion.version is sha256(source) and must stay byte-identical to the controlplane's Function.version. Folding prose into that hash forks the identity for nothing. The accepted cost is that an agent's prose has no version history, so "which prompt was this invocation run with" is currently unanswerable. That is the design's largest known compromise, and it is recorded rather than hidden.

channel is a small typed object in a Json? column, documented shape { kind, target?, config? } — a stored declaration of where the agent's pipe attaches, never a live connection:

Runnable.channeljson
{
  "kind": "slack",
  "target": "#eng-alerts",
  "config": { "sourceRef": "slack-eng" }
}

kind is a free string (stdio, http-sse, websocket, slack, …), following the same discriminator-plus-open-blob shape Source.config and LakeshoreProvider.config already use. A bare string like "slack:#eng" was rejected: it forces a parser and has nowhere to put an SSE path, a Slack channel id, or a reference to the credential that opens the pipe.

No credentials in channel.config

channel.config may hold a reference to a credential and never the credential itself. A Slack token belongs on a Source or a LakeshoreProvider, both of which already carry the AES-256-GCM blob and the never-return-ciphertext machinery. Point at one by name — { "sourceRef": "slack-eng" } — and let the server resolve it. The route enforces this: any key under channel.config matching /token|secret|password|api_?key/i is a 400, with the response naming the model the field should have gone to instead.

The split tripwire

Written down so nobody re-argues it. Split Agent into its own model — with a foreign key back to a shared Runnable, never by duplication — the first time any one of these holds:

  1. Agent-only fields exceed three, or any agent-only field becomes required. A column that is required for one kind and forbidden for another is a lie the schema cannot express.
  2. markdown or channel must be version-addressed — i.e. the answer to "which prompt ran" becomes load-bearing. This is the likeliest trigger, and it is exactly the cost accepted above.
  3. Agents acquire a relation UDFs must not have — the moment Runnable.sessions would be meaningless for kind="udf".

Until one of those fires: one model, one kind.


Two-layer identity

Two of the three new collections split into a mutable head plus an immutable, content-addressed child. The third does not.

ModelKeyMutability
Runnable@@unique([namespaceId, name, deletedAt]), name = "<module>.<qualname>"mutable head
RunnableVersion@@unique([runnableId, version]), version = sha256(source)immutable, append-only
Repo@@unique([namespaceId, name, deletedAt])mutable head
RepoSnapshot@@unique([repoId, commit])immutable, append-only
RunConfig@@unique([namespaceId, name, deletedAt])mutable, no versions

Why Runnable and Repo split. The controlplane's Function is flat, keyed (namespaceId, module, qualname, version), so it accumulates one row per source edit with no head — and "list my UDFs" degrades into a distinct plus a max(createdAt) group-by. The two-layer form hands the UI a name list and a history for free, and it is the house's own answer twice over (Pipeline/PipelineVersion, Workflow/WorkflowVersion).

Why RunConfig does not. It is a hand-edited operator declaration, and every reference to it is by name on purpose. Content-addressing it would mean editing gpu-a100 mints gpu-a100@v2 and every reference has to be repointed. Where content-addressing is genuinely needed — placement — the derived hostKey supplies it without freezing the record.

Why deletedAt is in the unique key. All three heads use the Pipeline/Workflow form @@unique([namespaceId, name, deletedAt]), not the Source form @@unique([namespaceId, name]). The latter covers tombstones, so deleting a provider named ec2 burns that name forever. These rows target the production server; delete-and-recreate has to work.

Append-only rows carry createdAt and deliberately have no updatedAt and no deletedAt. Versions and snapshots are never deleted by any endpoint; they become invisible when their parent head is soft-deleted.


RunConfig serves two audiences

A RunConfig is read by two consumers that do not talk to each other, so it stores them in two columns rather than one bag:

  • runtime — read by the runtime layer: backend, server, runner, image, resources, env, timeout_s, host_setup, run_setup.
  • extras — forwarded verbatim to the launcher: instance_type, partition, time_limit, n_cpu, n_gpu, and whatever the next launcher wants.

Alongside them sit two typed columns. tags String[] is typed and separate because it is the one field forwarded to both halves — it cannot honestly live in either bag. provider String? is a name, not a foreign key: the controlplane resolves providers by name, late, and an FK across a service boundary breaks the moment the CP resolves against its own copy.

One config bag was rejected, and not on taste. That form has already failed three times, always the same way — by making code reconstruct a partition that was never stored. _RUNTIME_FIELDS in lakeshore-py is a hand-maintained list of the runtime half with zero readers, while the real partition happens twice more in two lists that disagree with it. The controlplane's TYPED_KEYS omits server, so a runtime field round-trips through resources.__extras and works by accident. And the two sides default backend and runner differently, so a partial config changes meaning when stored and re-read. Two columns make the boundary a schema fact, so there is nothing to reconstruct and nothing to drift.

A GPU config with a non-trivial host setup, which is the acceptance shape:

RunConfig — gpu-a10gjson
{
  "runtime": {
    "backend": "fabric",
    "runner": "docker",
    "image": "nvcr.io/nvidia/pytorch:24.10-py3",
    "resources": { "gpu": 1, "cpu": 8, "memory": "32Gi" },
    "env": { "HF_HOME": "/mnt/cache/hf", "TORCH_CUDA_ARCH_LIST": "8.6" },
    "timeout_s": 3600,
    "host_setup": [
      "sudo nvidia-ctk runtime configure --runtime=docker",
      "sudo systemctl restart docker",
      "docker pull nvcr.io/nvidia/pytorch:24.10-py3"
    ],
    "run_setup": ["pip install -e ."]
  },
  "extras": { "instance_type": "g5.2xlarge", "n_gpu": 1, "n_cpu": 8 },
  "tags": ["gpu", "a10g"],
  "provider": "ec2-us-east-1"
}

The GPU appears in both halves, and it means two different things. runtime.resources.gpu is what the runner asks of the host it lands on — a requirement, checked against a machine that already exists. extras.n_gpu and extras.instance_type are what the launcher provisions — an order, placed before any host exists. They are usually consistent and they are not the same statement, and a single bag cannot keep them apart.

One more gap, recorded rather than solved: image is a single string, so a multi-container host setup is not expressible. Nothing in the acceptance shape needs it.

Proposal — host_setup / run_setup do not exist on the controlplane's Mode

Both keys are introduced by the host-key design. The controlplane's Mode route has no such fields: its TYPED_KEYS set is exactly {backend, runner, image, tags, resources, env, timeout_s, provider}, and anything outside it is swept into resources.__extras. So a host_setup sent to POST /v1/namespaces/:ns/modes today does not round-trip as a runtime field — it survives, buried, in the launcher-extras bucket. (server has the same problem and it is a runtime field, which is the bug that made the two-column split worth doing.)

RunConfig.runtime on dreamlake-server carries both keys as first-class runtime fields, and the Python SDK already emits them on the wire. Making the CP's Mode agree — or deleting Mode in favour of RunConfig, which is the staged plan — is a separate slice.


The derived host key

There is one declaration, written by the author: the RunConfig. The host requirement is derived from it — a hash over the placement-affecting subset — and never written down beside it. That is the whole point: a second, hand-written host field is a second source of truth that can disagree with the first.

The algorithm is not restated here. It is normative in the Host Setup design page, which owns the key's definition, its canonicalisation rules and its golden vectors.

Where it is computed, and who is allowed to be right:

SideWhenAuthority
Python SDKat submit, RunConfig.host_key()authoritative — stamps run_config["host_key"]
nymphat boot and after a setup refreshadvertises it on hello / poll
TypeScript CLIon lakeshore config showdisplay only
lakeshore controlplaneon every daemon poll / claimcomputes nothing; matches two strings (see below)
dreamlake-serveron every RunConfig writedisplay only, non-authoritative

dreamlake-server keeps a fourth copy in RunConfig.hostKey, recomputed on every write from { runner, image, resources, host_setup } out of runtime, plus the provider column and the whole extras column. It is null when it cannot be computed — a float in resources, for instance, is a hard error rather than a coercion.

It does not travel, and nothing matches on it. No scheduler, no worker and no queue ever reads dreamlake's copy. It exists for exactly two reasons: the dashboard can show which RunConfigs share a host, and a divergence between the SDK's stamp and the stored value becomes a diffable string instead of a silent mismatch. It exists at all only because a RunConfig is now authored in the UI and the SDK never writes a RunConfig row — without a server-side computation the column would be permanently null for precisely the rows this work exists to show. The mitigation is mandatory: the golden vector fixture is copied verbatim into the server and every vector is asserted, so an edited copy fails immediately.

Proposal — the controlplane's half is landing in a concurrent slice

Read the authority table above as a moving target. As of this writing the controlplane does consume the host key: claimOne composes a hostKeyFilter clause onto its findAndModify query, so a worker may only claim an invocation whose declaration hashes to the same hk1_…. That landed with the nymph lifecycle work and is not yet exercised end to end by a rolled-forward fleet.

The compatibility rule is the part worth knowing, because it is what stops a half-upgraded fleet from deadlocking: an absent key on either end degrades to exactly the previous matching. A worker that advertises no host key gets no clause at all; an invocation with no runConfig.host_key is claimable by any worker. Only when both sides carry a key does the key narrow anything. That is why the SDK omits the field rather than nulling it.

None of this changes dreamlake-server's position. Its RunConfig.hostKey is still the fourth copy, still display-only, and still matches nothing — the string the controlplane compares came from the SDK's stamp on the invocation, never from this column.


What is not here

Everything on the execution side of the rule. None of it moves, and none of it gains a mirror on dreamlake-server. The dashboard still shows all of it — by reading the controlplane through the proxy, with no model in between. That is Execution views.

  • Queues are not here. Including daemonTemplate, elasticity policy and the permitted instance types. Ge reaffirmed it on the same day the rest of this was settled, and a queue is the clearest case the rule has: it is a piece of running machinery with a depth that changes by the second.
  • Invocations — runConfigRef (a name) and resolvedConfig (the merged per-run snapshot) both stay on the controlplane.
  • Workers and daemons — the live fleet, its capabilities and its heartbeats.
  • Events — the run log.
  • Sessions and live channels — a Runnable declares where a channel attaches; opening one is execution. There is no session model, no socket and no reconfiguration endpoint in this slice.
Proposal — LakeshoreQueueDef removal

LakeshoreQueueDef still exists on dreamlake-server and is not deleted by this work. It is deprecated and frozen: its routes and GraphQL fields keep serving existing callers, no field changes, and no new UI offers it. Dropping a model orphans documents in a shared Mongo database and there is no migration to undo that. Full removal — the model, four permission keys, the CollectionSpec entry, the GraphQL type and its two Query fields, the test-teardown line, and re-pointing two tests that use queue-defs as their sole coverage vehicle — is staged for Ge's approval, not executed.

Proposal — LakeshoreStorage collapses into Source

LakeshoreStorage is a third copy of what Source already is: a connector with non-secret config and an AES-256-GCM secret blob. The plan is to collapse it into Source, carrying over the immutable-target constraint @@unique([namespaceId, bucket, prefix, endpoint]) — for which this schema has no precedent, since every existing namespace-scoped unique key is over name. It needs a data migration against the live database, so it is argued and staged, not executed. This is why no Storage page is built: Source already has a full UI at /source/:ns, and building a page with a known expiry date is the definition of throwaway work.


Where to read further

  • Execution views — the other half of the line: runs, elastic queues and cluster topology, all read through the proxy, none of them a collection.
  • plans/2026-08-06_dreamlake-declaration-collections/README.md — the full argument, including every rejected alternative.
  • plans/2026-08-06_dreamlake-declaration-collections/schema.md — the five Prisma models verbatim, with their comment blocks.
  • DECISIONS.md, 2026-08-06 — the two decision entries this page implements.
  • Host Setup — the host-key derivation and why the requirement is derived rather than declared.