Execution views
The companion to Declarations vs. execution. That page covers the half DreamLake owns. This one covers the half DreamLake only looks at: the runs that are in flight, the queues that are scaling, and the machines they are scaling onto.
The rule, first
Topology, elastic queues and invocations add no Prisma model and no dreamlake-server collection. Not one. Queues, invocations, elasticity events, workers and everything topology is drawn from are execution, and execution stays on the controlplane — Ge, 2026-08-06:
those configs live on dreamlake server. lakeshore cp is responsible for the runtime / ephemeral states and the queue.
All three views are DreamLake-side read surfaces over a proxy that already ships. There is nothing to sync, nothing to invalidate, and no second copy that can be stale, because there is no copy at all.
Everything goes through one proxy
The wildcard is forwarded verbatim to ${baseUrl}/<controlplane-path> with
the Lakeshore's decrypted bearer token injected. Four properties of that route
are load-bearing for anything built on top of it.
It is deliberately dumb, and single-target. The proxy does not know what a
queue is. The client builds the controlplane path itself — including the
lakeshoreNs segment, which it reads off the Lakeshore record — and the proxy
resolves baseUrl from that same record and nothing else. That is what keeps it
from being pointed at any host but this Lakeshore's controlplane. The cost is
that the browser has to know the controlplane's URL grammar; the benefit is that
a new controlplane endpoint needs zero dreamlake-server work.
GET is lakeshore:read; every mutating method is lakeshore:use, which is
ADMIN-only. All three views are read-only, so an org member with READ sees all
three. Anything that would need lakeshore:use — killing an invocation, forcing
a scale action, draining a worker — is out of scope here precisely because it
changes the permission story.
Only content-type and cache-control come back. The proxy mirrors the
upstream status and exactly those two headers, then streams the body. Everything
else is dropped.
Because the proxy forwards only content-type and cache-control, a Link: rel="next" header from the controlplane is silently discarded — the client
sees a 200 with a complete body and no hint that more exists. Any paginated
controlplane endpoint these views consume must carry its cursor in the response
body or in the query string, never in a header. The existing feeds do this
correctly (?limit&before on events, ?limit&offset on invocations); the
constraint is written down here so the next endpoint does not learn it the hard
way.
An unreachable controlplane is a 502. The fetch is wrapped, and a transport
failure returns 502 { error: "controlplane unreachable", message } — not a 503,
not a 504, not a hung request. Views distinguish three states from that: a 502 is
"the Lakeshore is down", a 4xx passed through is "the controlplane answered and
said no", and a network error before the proxy is "dreamlake-server is down".
Resolving which Lakeshore
Every proxy URL needs a lakeshoreId, which comes from
GET /namespaces/:slug/lakeshores on dreamlake-server. The views pick the first
connected record — the literal reading of the requirement, "in a single
connected lakeshore."
Zero connected lakeshores is a first-class state, not an error state: each of the
three pages renders a "connect a lakeshore" prompt linking to /:ns/lakeshore.
Never a spinner that never resolves, and never a red error card — an unconfigured
namespace has not failed at anything.
Topology — one spine, four levels
The requirement is across providers, in a single connected lakeshore. That rules out two provider-shaped panels side by side, and it rules out mirroring Kubernetes objects.
The whole design is the third level. An EC2 auto-scaling group and a Karpenter NodePool are the same object at this altitude: a named group of interchangeable machines that a provider grows and shrinks. Once you say that out loud, both live clusters fall out of one spine with no special case:
| Cluster | Provider | Pools | Workers |
|---|---|---|---|
aws-workspace-ge (EC2 + ASG) | one aws-ec2 provider | ASG pools cpu, gpu | one nymph per EC2 instance |
dreamlake-k8s (EKS + Karpenter) | one k8s provider | NodePools cpu, gpu | one nymph per pod |
No pods, no nodeclaims, no ASG scaling activities, no container statuses. k9s already exists and is better at every one of those than a dashboard panel will be. The view answers "how many GPU nodes are up, on which provider, in what state" — a fleet question — and stops there. Rendering is an indented grouped list, not a graph: a fleet is a containment tree, not a dataflow DAG, and a force layout makes the count harder to read, not easier.
Pool-key derivation
Computed client-side from the GET v1/namespaces/:ns/workers response.
First hit wins:
worker.capabilities.pool, when the daemon reports it.- A tag matching
/^pool[:=](.+)$/i— else a bare tag in{cpu, gpu}. worker.capabilities.gpupresent and non-zero →"gpu", otherwise"cpu"."default".
Rule 3 is why the view works today, before any daemon cooperation: a
g5.2xlarge reports GPUs and a spot CPU node does not, so the two pools separate
correctly on both live clusters with no change to the nymph. Rule 1 is what makes
it authoritative later; when a daemon starts reporting capabilities.pool,
nothing in this view changes except that a marker stops appearing.
Pools derived by rule 3 render with a visible inferred badge. That badge is
not decoration. Rules 1 and 2 report something an operator declared; rule 3 is the
dashboard guessing from a GPU count. A guess and a fact must not look identical,
because the moment they do, "the topology says these are two pools" stops being
evidence of anything.
Provider comes from Worker.providerName, which the controlplane stamps on
workers it launched; SSH-installed workers have none and group under byo.
GET v1/admin/providers is best-effort enrichment only — it wants an
admin-capable token, so a 401 or 403 there must degrade to provider names without
kind/region metadata. It must never produce an error page: losing the region
label is not losing the topology.
Queues — the events feed is primary
Two controlplane endpoints, both already shipped, both with zero UI consumers before this work:
| Endpoint | Returns |
|---|---|
GET v1/namespaces/:ns/queues/:name/events?limit&before&verbose | the ElasticityEvent feed |
GET v1/namespaces/:ns/queues/:name/stats | { queueName, daemon_count, utilization, pending_invocations, recent_scale_actions[] } |
The feed is the primary panel and the stats strip is secondary, which is the
inverse of how a stats endpoint usually gets treated. The reason is one row type:
the controller logs tick_noop with a reason every time it decides not to
act, which makes the feed the only thing in the entire stack that answers "why
didn't my cluster scale up" — precisely the question a live demo produces.
/stats deliberately filters those out (action: { not: "tick_noop" }), so
recent_scale_actions is the highlight reel and the feed is the
transcript. Both are useful; only one of them explains a non-event. Default
the feed to verbose=false with a toggle, because noops are the point and they
are also roughly 95% of the rows.
Queue policy — elasticity kind, min/max/threshold, admission.max_depth,
providerRef — comes from GET v1/namespaces/:ns/queues and renders beside the
live signals. "max: 4" is the answer to a good half of the "why didn't it scale"
questions, and it is useless on a different screen from the depth it is capping.
daemon_count and pending_invocations are computed server-side from real rows.
They are trustworthy today; render them plainly.
Utilization, and why it renders as a dash
utilization is sumCurrent / sumMax, where sumCurrent sums
capabilities.current_invocations ?? 0 across the queue's workers. The ?? 0
is the problem: a daemon that has never heard of the key contributes exactly
what a fully idle daemon contributes. /stats therefore exposes no way to tell
"0% busy" from "nobody reported", and a dial pinned at 0% while a GPU node is
melting is a confident lie.
The dashboard must not guess from the number. It checks the source: the topology
view already fetches GET v1/namespaces/:ns/workers, so the queue panel reuses
that response and asks whether any worker in this queue carries a
capabilities.current_invocations key at all.
- some worker reports it → render the percentage
daemon_count > 0and none reports it → render—, footnoted "utilization not reported by these daemons"daemon_count === 0→ renderno daemons, which is a real answer and not a missing one
A concurrent workstream is landing the daemon-side field (the nymph now stamps
current_invocations from its in-flight counter on every poll). The dash is the
correct rendering until every daemon in a fleet has been rolled forward — and it
stays correct afterwards, because a mixed-version fleet is the normal state of a
fleet.
Runs — invocations, and the join between the halves
GET v1/namespaces/:ns/invocations?limit&offset&state&queue and
GET …/invocations/:id.
List columns: id, function, state, queue, worker, submitted, duration. The
state and queue filters are server-side query parameters, not client-side
array filters — the list is unbounded, and filtering a page of 50 client-side
gives an answer about the page rather than about the namespace.
The detail page shows what a run was: the function ref, runConfigRef, the
resolvedConfig snapshot, the queue, the worker, and the repo commit from the
envelope's code stamp.
That page is the join between the two halves of the system. An invocation
names a RunConfig by name and a Runnable by module/qualname, and the detail
view links both back into the declaration collections — /:ns/runconfigs/<name>
and /:ns/runnables/<name> — by name, not by id. This is exactly what the
collections API's id-or-name resolution exists for: a 24-hex segment reads as an
id, anything else as a name. Without it, the link would need a lookup table
across a service boundary, which is the thing the whole declaration/execution
split exists to avoid.
Note what this does not claim: runConfigRef is a name that was resolved on the
controlplane at submit time, so the RunConfig it points at may since have been
edited. resolvedConfig is the snapshot of what actually ran and is authoritative
for the run; the link is navigation, not provenance.
Out of scope, stated plainly
- The log SSE stream.
streamUrland a working SSE reader already exist in the lockedlib/lakeshore/; the proxy streams SSE correctly. Wiring it into these pages is a separate slice.
parentInvocationId and rootInvocationId exist on the controlplane's
Invocation model, and the submit handler now walks the parent to stamp a root.
No lineage view is built here, and the fields are not rendered on the run
detail page. A tree view needs a fan-out read the list endpoint does not offer,
and it needs every SDK path to populate the parent link before the tree means
anything — today only the Python HttpDispatch sends parent_invocation_id, and
only when a caller passes one. Treat lineage as designed and partially wired, not
as available.
The three routes
| Route | Reads (all via /proxy/*) | Poll |
|---|---|---|
/:ns/runs | v1/namespaces/:ns/invocations, …/invocations/:id | 6 s |
/:ns/queues | v1/namespaces/:ns/queues, …/:name/stats, …/:name/events | 8 s |
/:ns/topology | v1/namespaces/:ns/workers, v1/admin/providers | 10 s |
Every poll pauses while the tab is hidden. Each tick is a proxied round-trip
from the browser through dreamlake-server to the controlplane, and the existing
overview panel already mounts five pollers at 4–6 s. Three more unconditional
fast pollers would multiply proxy load for a dashboard nobody is looking at.
Pause on document.hidden, resume on visibility, and refetch once immediately on
resume so the first frame after a tab switch is not stale.
Where to read further
- Declarations vs. execution — the other half of the line, and the collections these views link into.
plans/2026-08-06_dreamlake-declaration-collections/README.md— the full argument, including every rejected alternative.DECISIONS.md, 2026-08-06 — the decision this page implements.