# Execution views

The companion to [Declarations vs. execution](/dev/declaration-collections.md). That
page covers the half DreamLake **owns**. This one covers the half DreamLake only
**looks at**: the runs that are in flight, the queues that are scaling, and the
machines they are scaling onto.

## The rule, first

**Topology, elastic queues and invocations add no Prisma model and no
dreamlake-server collection.** Not one. Queues, invocations, elasticity events,
workers and everything topology is drawn from are execution, and execution stays
on the controlplane — Ge, 2026-08-06:

> those configs live on dreamlake server. lakeshore cp is responsible for the
> runtime / ephemeral states and the queue.

All three views are DreamLake-side **read surfaces** over a proxy that already
ships. There is nothing to sync, nothing to invalidate, and no second copy that
can be stale, because there is no copy at all.

---

## Everything goes through one proxy

```
GET /namespaces/:slug/lakeshores/:id/proxy/<controlplane-path>[?query]
```

The wildcard is forwarded **verbatim** to `${baseUrl}/<controlplane-path>` with
the Lakeshore's decrypted bearer token injected. Four properties of that route
are load-bearing for anything built on top of it.

**It is deliberately dumb, and single-target.** The proxy does not know what a
queue is. The *client* builds the controlplane path itself — including the
`lakeshoreNs` segment, which it reads off the `Lakeshore` record — and the proxy
resolves `baseUrl` from that same record and nothing else. That is what keeps it
from being pointed at any host but this Lakeshore's controlplane. The cost is
that the browser has to know the controlplane's URL grammar; the benefit is that
a new controlplane endpoint needs zero dreamlake-server work.

**GET is `lakeshore:read`; every mutating method is `lakeshore:use`, which is
ADMIN-only.** All three views are read-only, so an org member with READ sees all
three. Anything that would need `lakeshore:use` — killing an invocation, forcing
a scale action, draining a worker — is out of scope here precisely because it
changes the permission story.

**Only `content-type` and `cache-control` come back.** The proxy mirrors the
upstream status and exactly those two headers, then streams the body. Everything
else is dropped.

> **Warning:** Because the proxy forwards only `content-type` and `cache-control`, a `Link:
> rel="next"` header from the controlplane is **silently discarded** — the client
> sees a 200 with a complete body and no hint that more exists. Any paginated
> controlplane endpoint these views consume must carry its cursor **in the response
> body or in the query string**, never in a header. The existing feeds do this
> correctly (`?limit&before` on events, `?limit&offset` on invocations); the
> constraint is written down here so the next endpoint does not learn it the hard
> way.

**An unreachable controlplane is a 502.** The `fetch` is wrapped, and a transport
failure returns `502 { error: "controlplane unreachable", message }` — not a 503,
not a 504, not a hung request. Views distinguish three states from that: a 502 is
"the Lakeshore is down", a 4xx passed through is "the controlplane answered and
said no", and a network error before the proxy is "dreamlake-server is down".

### Resolving *which* Lakeshore

Every proxy URL needs a `lakeshoreId`, which comes from
`GET /namespaces/:slug/lakeshores` on dreamlake-server. The views pick the first
**connected** record — the literal reading of the requirement, *"in a single
connected lakeshore."*

Zero connected lakeshores is a first-class state, not an error state: each of the
three pages renders a "connect a lakeshore" prompt linking to `/:ns/lakeshore`.
Never a spinner that never resolves, and never a red error card — an unconfigured
namespace has not failed at anything.

---

## Topology — one spine, four levels

The requirement is *across providers, in a single connected lakeshore*. That
rules out two provider-shaped panels side by side, and it rules out mirroring
Kubernetes objects.

```
Lakeshore                the one connected controlplane
  └── Provider           Worker.providerName, else "byo"
        └── Pool         derived — see below
              └── Worker machineId · state · hostStatus · runners · queues · gpus
```

The whole design is the third level. **An EC2 auto-scaling group and a Karpenter
NodePool are the same object at this altitude**: a named group of interchangeable
machines that a provider grows and shrinks. Once you say that out loud, both live
clusters fall out of one spine with no special case:

| Cluster | Provider | Pools | Workers |
|---|---|---|---|
| `aws-workspace-ge` (EC2 + ASG) | one `aws-ec2` provider | ASG pools `cpu`, `gpu` | one nymph per EC2 instance |
| `dreamlake-k8s` (EKS + Karpenter) | one `k8s` provider | NodePools `cpu`, `gpu` | one nymph per pod |

> **Note:** No pods, no nodeclaims, no ASG scaling activities, no container statuses. **k9s
> already exists** and is better at every one of those than a dashboard panel will
> be. The view answers "how many GPU nodes are up, on which provider, in what
> state" — a fleet question — and stops there. Rendering is an indented grouped
> list, not a graph: a fleet is a containment tree, not a dataflow DAG, and a force
> layout makes the count harder to read, not easier.

### Pool-key derivation

Computed client-side from the `GET v1/namespaces/:ns/workers` response.
**First hit wins:**

1. `worker.capabilities.pool`, when the daemon reports it.
2. A tag matching `/^pool[:=](.+)$/i` — else a bare tag in `{cpu, gpu}`.
3. `worker.capabilities.gpu` present and non-zero → `"gpu"`, otherwise `"cpu"`.
4. `"default"`.

Rule 3 is why the view works **today**, before any daemon cooperation: a
`g5.2xlarge` reports GPUs and a spot CPU node does not, so the two pools separate
correctly on both live clusters with no change to the nymph. Rule 1 is what makes
it authoritative later; when a daemon starts reporting `capabilities.pool`,
nothing in this view changes except that a marker stops appearing.

**Pools derived by rule 3 render with a visible `inferred` badge.** That badge is
not decoration. Rules 1 and 2 report something an operator declared; rule 3 is the
dashboard guessing from a GPU count. A guess and a fact must not look identical,
because the moment they do, "the topology says these are two pools" stops being
evidence of anything.

`Provider` comes from `Worker.providerName`, which the controlplane stamps on
workers it launched; SSH-installed workers have none and group under `byo`.
`GET v1/admin/providers` is **best-effort enrichment only** — it wants an
admin-capable token, so a 401 or 403 there must degrade to provider names without
kind/region metadata. It must never produce an error page: losing the region
label is not losing the topology.

---

## Queues — the events feed is primary

Two controlplane endpoints, both already shipped, both with zero UI consumers
before this work:

| Endpoint | Returns |
|---|---|
| `GET v1/namespaces/:ns/queues/:name/events?limit&before&verbose` | the `ElasticityEvent` feed |
| `GET v1/namespaces/:ns/queues/:name/stats` | `{ queueName, daemon_count, utilization, pending_invocations, recent_scale_actions[] }` |

**The feed is the primary panel and the stats strip is secondary**, which is the
inverse of how a stats endpoint usually gets treated. The reason is one row type:
the controller logs `tick_noop` with a `reason` every time it decides *not* to
act, which makes the feed the only thing in the entire stack that answers **"why
didn't my cluster scale up"** — precisely the question a live demo produces.

`/stats` deliberately filters those out (`action: { not: "tick_noop" }`), so
`recent_scale_actions` is the **highlight reel** and the feed is the
**transcript**. Both are useful; only one of them explains a non-event. Default
the feed to `verbose=false` with a toggle, because noops are the point and they
are also roughly 95% of the rows.

Queue *policy* — elasticity kind, min/max/threshold, `admission.max_depth`,
`providerRef` — comes from `GET v1/namespaces/:ns/queues` and renders beside the
live signals. "max: 4" is the answer to a good half of the "why didn't it scale"
questions, and it is useless on a different screen from the depth it is capping.

`daemon_count` and `pending_invocations` are computed server-side from real rows.
They are trustworthy today; render them plainly.

### Utilization, and why it renders as a dash

> **Warning:** `utilization` is `sumCurrent / sumMax`, where `sumCurrent` sums
> `capabilities.current_invocations ?? 0` across the queue's workers. The `?? 0`
> is the problem: **a daemon that has never heard of the key contributes exactly
> what a fully idle daemon contributes.** `/stats` therefore exposes no way to tell
> "0% busy" from "nobody reported", and a dial pinned at 0% while a GPU node is
> melting is a confident lie.
>
> The dashboard must not guess from the number. It checks the source: the topology
> view already fetches `GET v1/namespaces/:ns/workers`, so the queue panel reuses
> that response and asks whether **any** worker in this queue carries a
> `capabilities.current_invocations` key at all.
>
> - some worker reports it → render the percentage
> - `daemon_count > 0` and none reports it → render **`—`**, footnoted
>   *"utilization not reported by these daemons"*
> - `daemon_count === 0` → render **`no daemons`**, which is a real answer and not
>   a missing one
>
> A concurrent workstream is landing the daemon-side field (the nymph now stamps
> `current_invocations` from its in-flight counter on every poll). The dash is the
> correct rendering until every daemon in a fleet has been rolled forward — and it
> stays correct afterwards, because a mixed-version fleet is the normal state of a
> fleet.

---

## Runs — invocations, and the join between the halves

`GET v1/namespaces/:ns/invocations?limit&offset&state&queue` and
`GET …/invocations/:id`.

List columns: id, function, state, queue, worker, submitted, duration. **The
`state` and `queue` filters are server-side query parameters, not client-side
array filters** — the list is unbounded, and filtering a page of 50 client-side
gives an answer about the page rather than about the namespace.

The detail page shows what a run *was*: the function ref, `runConfigRef`, the
`resolvedConfig` snapshot, the queue, the worker, and the repo commit from the
envelope's code stamp.

**That page is the join between the two halves of the system.** An invocation
names a `RunConfig` by name and a `Runnable` by module/qualname, and the detail
view links both back into the declaration collections — `/:ns/runconfigs/<name>`
and `/:ns/runnables/<name>` — **by name, not by id**. This is exactly what the
collections API's id-or-name resolution exists for: a 24-hex segment reads as an
id, anything else as a name. Without it, the link would need a lookup table
across a service boundary, which is the thing the whole declaration/execution
split exists to avoid.

Note what this does *not* claim: `runConfigRef` is a name that was resolved on the
controlplane at submit time, so the `RunConfig` it points at may since have been
edited. `resolvedConfig` is the snapshot of what actually ran and is authoritative
for the run; the link is navigation, not provenance.

### Out of scope, stated plainly

- **The log SSE stream.** `streamUrl` and a working SSE reader already exist in
  the locked `lib/lakeshore/`; the proxy streams SSE correctly. Wiring it into
  these pages is a separate slice.

> **Warning:** `parentInvocationId` and `rootInvocationId` exist on the controlplane's
> `Invocation` model, and the submit handler now walks the parent to stamp a root.
> **No lineage view is built here**, and the fields are not rendered on the run
> detail page. A tree view needs a fan-out read the list endpoint does not offer,
> and it needs every SDK path to populate the parent link before the tree means
> anything — today only the Python `HttpDispatch` sends `parent_invocation_id`, and
> only when a caller passes one. Treat lineage as designed and partially wired, not
> as available.

---

## The three routes

| Route | Reads (all via `/proxy/*`) | Poll |
|---|---|---|
| `/:ns/runs` | `v1/namespaces/:ns/invocations`, `…/invocations/:id` | **6 s** |
| `/:ns/queues` | `v1/namespaces/:ns/queues`, `…/:name/stats`, `…/:name/events` | **8 s** |
| `/:ns/topology` | `v1/namespaces/:ns/workers`, `v1/admin/providers` | **10 s** |

**Every poll pauses while the tab is hidden.** Each tick is a proxied round-trip
from the browser through dreamlake-server to the controlplane, and the existing
overview panel already mounts five pollers at 4–6 s. Three more unconditional
fast pollers would multiply proxy load for a dashboard nobody is looking at.
Pause on `document.hidden`, resume on visibility, and refetch once immediately on
resume so the first frame after a tab switch is not stale.

---

## Where to read further

- [Declarations vs. execution](/dev/declaration-collections.md) — the other half of
  the line, and the collections these views link into.
- `plans/2026-08-06_dreamlake-declaration-collections/README.md` — the full
  argument, including every rejected alternative.
- `DECISIONS.md`, 2026-08-06 — the decision this page implements.
