DreamLake

Execution views

The companion to Declarations vs. execution. That page covers the half DreamLake owns. This one covers the half DreamLake only looks at: the runs that are in flight, the queues that are scaling, and the machines they are scaling onto.

The rule, first

Topology, elastic queues and invocations add no Prisma model and no dreamlake-server collection. Not one. Queues, invocations, elasticity events, workers and everything topology is drawn from are execution, and execution stays on the controlplane — Ge, 2026-08-06:

those configs live on dreamlake server. lakeshore cp is responsible for the runtime / ephemeral states and the queue.

All three views are DreamLake-side read surfaces over a proxy that already ships. There is nothing to sync, nothing to invalidate, and no second copy that can be stale, because there is no copy at all.


Everything goes through one proxy

GET /namespaces/:slug/lakeshores/:id/proxy/<controlplane-path>[?query]

The wildcard is forwarded verbatim to ${baseUrl}/<controlplane-path> with the Lakeshore's decrypted bearer token injected. Four properties of that route are load-bearing for anything built on top of it.

It is deliberately dumb, and single-target. The proxy does not know what a queue is. The client builds the controlplane path itself — including the lakeshoreNs segment, which it reads off the Lakeshore record — and the proxy resolves baseUrl from that same record and nothing else. That is what keeps it from being pointed at any host but this Lakeshore's controlplane. The cost is that the browser has to know the controlplane's URL grammar; the benefit is that a new controlplane endpoint needs zero dreamlake-server work.

GET is lakeshore:read; every mutating method is lakeshore:use, which is ADMIN-only. All three views are read-only, so an org member with READ sees all three. Anything that would need lakeshore:use — killing an invocation, forcing a scale action, draining a worker — is out of scope here precisely because it changes the permission story.

Only content-type and cache-control come back. The proxy mirrors the upstream status and exactly those two headers, then streams the body. Everything else is dropped.

Pagination cannot ride in a header

Because the proxy forwards only content-type and cache-control, a Link: rel="next" header from the controlplane is silently discarded — the client sees a 200 with a complete body and no hint that more exists. Any paginated controlplane endpoint these views consume must carry its cursor in the response body or in the query string, never in a header. The existing feeds do this correctly (?limit&before on events, ?limit&offset on invocations); the constraint is written down here so the next endpoint does not learn it the hard way.

An unreachable controlplane is a 502. The fetch is wrapped, and a transport failure returns 502 { error: "controlplane unreachable", message } — not a 503, not a 504, not a hung request. Views distinguish three states from that: a 502 is "the Lakeshore is down", a 4xx passed through is "the controlplane answered and said no", and a network error before the proxy is "dreamlake-server is down".

Resolving which Lakeshore

Every proxy URL needs a lakeshoreId, which comes from GET /namespaces/:slug/lakeshores on dreamlake-server. The views pick the first connected record — the literal reading of the requirement, "in a single connected lakeshore."

Zero connected lakeshores is a first-class state, not an error state: each of the three pages renders a "connect a lakeshore" prompt linking to /:ns/lakeshore. Never a spinner that never resolves, and never a red error card — an unconfigured namespace has not failed at anything.


Topology — one spine, four levels

The requirement is across providers, in a single connected lakeshore. That rules out two provider-shaped panels side by side, and it rules out mirroring Kubernetes objects.

Lakeshore                the one connected controlplane
  └── Provider           Worker.providerName, else "byo"
        └── Pool         derived — see below
              └── Worker machineId · state · hostStatus · runners · queues · gpus

The whole design is the third level. An EC2 auto-scaling group and a Karpenter NodePool are the same object at this altitude: a named group of interchangeable machines that a provider grows and shrinks. Once you say that out loud, both live clusters fall out of one spine with no special case:

ClusterProviderPoolsWorkers
aws-workspace-ge (EC2 + ASG)one aws-ec2 providerASG pools cpu, gpuone nymph per EC2 instance
dreamlake-k8s (EKS + Karpenter)one k8s providerNodePools cpu, gpuone nymph per pod
Nothing below Worker is drawn

No pods, no nodeclaims, no ASG scaling activities, no container statuses. k9s already exists and is better at every one of those than a dashboard panel will be. The view answers "how many GPU nodes are up, on which provider, in what state" — a fleet question — and stops there. Rendering is an indented grouped list, not a graph: a fleet is a containment tree, not a dataflow DAG, and a force layout makes the count harder to read, not easier.

Pool-key derivation

Computed client-side from the GET v1/namespaces/:ns/workers response. First hit wins:

  1. worker.capabilities.pool, when the daemon reports it.
  2. A tag matching /^pool[:=](.+)$/i — else a bare tag in {cpu, gpu}.
  3. worker.capabilities.gpu present and non-zero → "gpu", otherwise "cpu".
  4. "default".

Rule 3 is why the view works today, before any daemon cooperation: a g5.2xlarge reports GPUs and a spot CPU node does not, so the two pools separate correctly on both live clusters with no change to the nymph. Rule 1 is what makes it authoritative later; when a daemon starts reporting capabilities.pool, nothing in this view changes except that a marker stops appearing.

Pools derived by rule 3 render with a visible inferred badge. That badge is not decoration. Rules 1 and 2 report something an operator declared; rule 3 is the dashboard guessing from a GPU count. A guess and a fact must not look identical, because the moment they do, "the topology says these are two pools" stops being evidence of anything.

Provider comes from Worker.providerName, which the controlplane stamps on workers it launched; SSH-installed workers have none and group under byo. GET v1/admin/providers is best-effort enrichment only — it wants an admin-capable token, so a 401 or 403 there must degrade to provider names without kind/region metadata. It must never produce an error page: losing the region label is not losing the topology.


Queues — the events feed is primary

Two controlplane endpoints, both already shipped, both with zero UI consumers before this work:

EndpointReturns
GET v1/namespaces/:ns/queues/:name/events?limit&before&verbosethe ElasticityEvent feed
GET v1/namespaces/:ns/queues/:name/stats{ queueName, daemon_count, utilization, pending_invocations, recent_scale_actions[] }

The feed is the primary panel and the stats strip is secondary, which is the inverse of how a stats endpoint usually gets treated. The reason is one row type: the controller logs tick_noop with a reason every time it decides not to act, which makes the feed the only thing in the entire stack that answers "why didn't my cluster scale up" — precisely the question a live demo produces.

/stats deliberately filters those out (action: { not: "tick_noop" }), so recent_scale_actions is the highlight reel and the feed is the transcript. Both are useful; only one of them explains a non-event. Default the feed to verbose=false with a toggle, because noops are the point and they are also roughly 95% of the rows.

Queue policy — elasticity kind, min/max/threshold, admission.max_depth, providerRef — comes from GET v1/namespaces/:ns/queues and renders beside the live signals. "max: 4" is the answer to a good half of the "why didn't it scale" questions, and it is useless on a different screen from the depth it is capping.

daemon_count and pending_invocations are computed server-side from real rows. They are trustworthy today; render them plainly.

Utilization, and why it renders as a dash

Do not trust /stats utilization on its own

utilization is sumCurrent / sumMax, where sumCurrent sums capabilities.current_invocations ?? 0 across the queue's workers. The ?? 0 is the problem: a daemon that has never heard of the key contributes exactly what a fully idle daemon contributes. /stats therefore exposes no way to tell "0% busy" from "nobody reported", and a dial pinned at 0% while a GPU node is melting is a confident lie.

The dashboard must not guess from the number. It checks the source: the topology view already fetches GET v1/namespaces/:ns/workers, so the queue panel reuses that response and asks whether any worker in this queue carries a capabilities.current_invocations key at all.

  • some worker reports it → render the percentage
  • daemon_count > 0 and none reports it → render —, footnoted "utilization not reported by these daemons"
  • daemon_count === 0 → render no daemons, which is a real answer and not a missing one

A concurrent workstream is landing the daemon-side field (the nymph now stamps current_invocations from its in-flight counter on every poll). The dash is the correct rendering until every daemon in a fleet has been rolled forward — and it stays correct afterwards, because a mixed-version fleet is the normal state of a fleet.


Runs — invocations, and the join between the halves

GET v1/namespaces/:ns/invocations?limit&offset&state&queue and GET …/invocations/:id.

List columns: id, function, state, queue, worker, submitted, duration. The state and queue filters are server-side query parameters, not client-side array filters — the list is unbounded, and filtering a page of 50 client-side gives an answer about the page rather than about the namespace.

The detail page shows what a run was: the function ref, runConfigRef, the resolvedConfig snapshot, the queue, the worker, and the repo commit from the envelope's code stamp.

That page is the join between the two halves of the system. An invocation names a RunConfig by name and a Runnable by module/qualname, and the detail view links both back into the declaration collections — /:ns/runconfigs/<name> and /:ns/runnables/<name> — by name, not by id. This is exactly what the collections API's id-or-name resolution exists for: a 24-hex segment reads as an id, anything else as a name. Without it, the link would need a lookup table across a service boundary, which is the thing the whole declaration/execution split exists to avoid.

Note what this does not claim: runConfigRef is a name that was resolved on the controlplane at submit time, so the RunConfig it points at may since have been edited. resolvedConfig is the snapshot of what actually ran and is authoritative for the run; the link is navigation, not provenance.

Out of scope, stated plainly

  • The log SSE stream. streamUrl and a working SSE reader already exist in the locked lib/lakeshore/; the proxy streams SSE correctly. Wiring it into these pages is a separate slice.
Proposal — invocation lineage

parentInvocationId and rootInvocationId exist on the controlplane's Invocation model, and the submit handler now walks the parent to stamp a root. No lineage view is built here, and the fields are not rendered on the run detail page. A tree view needs a fan-out read the list endpoint does not offer, and it needs every SDK path to populate the parent link before the tree means anything — today only the Python HttpDispatch sends parent_invocation_id, and only when a caller passes one. Treat lineage as designed and partially wired, not as available.


The three routes

RouteReads (all via /proxy/*)Poll
/:ns/runsv1/namespaces/:ns/invocations, …/invocations/:id6 s
/:ns/queuesv1/namespaces/:ns/queues, …/:name/stats, …/:name/events8 s
/:ns/topologyv1/namespaces/:ns/workers, v1/admin/providers10 s

Every poll pauses while the tab is hidden. Each tick is a proxied round-trip from the browser through dreamlake-server to the controlplane, and the existing overview panel already mounts five pollers at 4–6 s. Three more unconditional fast pollers would multiply proxy load for a dashboard nobody is looking at. Pause on document.hidden, resume on visibility, and refetch once immediately on resume so the first frame after a tab switch is not stale.


Where to read further

  • Declarations vs. execution — the other half of the line, and the collections these views link into.
  • plans/2026-08-06_dreamlake-declaration-collections/README.md — the full argument, including every rejected alternative.
  • DECISIONS.md, 2026-08-06 — the decision this page implements.