Agent API
The agent declaration is settled and in the tree: Runnable with
kind="agent", plus markdown and channel. Everything on this page is the
execution half, and none of it exists — there is no agent route on the
controlplane, no Agent export in the SDK, and no agent entry in the CP
schema. This page proposes the surface and answers the six questions left open
on Agent lifecycle so
that they stop blocking.
An agent is a Session with a policy driving it. That decision is already taken. What it does not settle is the wire: how a message reaches a running agent, how a reply comes back, what the SDK calls it, and what a run records about itself afterwards. This page is that.
Why now, and why this doc has a second audience
The forcing function is not agents. It is that a data labeler is the same shape — long-lived, holds context across many items, receives work over a channel, writes results back — and if we design its API separately we will build a second dispatch plane and a second channel that do the same job slightly differently. The labeler coupling section at the bottom names exactly which primitives are shared and which are genuinely different. Read it before ratifying anything above it.
What is already decided, and is not re-opened here
Four things are settled. This page builds on them rather than re-arguing them.
| Decision | Where | What it fixes |
|---|---|---|
| An Agent is a Session plus a loop | DECISIONS.md, 2026-08-06 | Agents get no separate lifecycle. UDF = one-shot; Session = long-lived + pipe; Agent = long-lived + pipe + policy. |
| DreamLake owns declarations, the CP owns execution | The four collections | The agent record is DreamLake's. Everything live — the channel, the invocation, the worker — is the CP's. |
| Placement is derived, never declared | Agents | An agent inherits host_key from its RunConfig unchanged. There is no agent-specific placement field, and this page does not add one. |
| Provenance is recorded, not reconstructed | same | The rendered prompt is written to the invocation so a drifting template cannot rewrite history. |
Three surfaces, one of which exists
| Surface | Owner | Status |
|---|---|---|
| Declaration — create, edit, version, list agents | DreamLake server | Routes exist under /namespaces/:slug/runnables. |
| Control — start, stop, drain, inspect a running agent | Lakeshore CP | Proposed below. Nothing today. |
| Channel — send a message, read replies | Lakeshore CP | Proposed below. The channel field is stored and read by nothing. |
The split matters because the control surface is where the lifecycle states live,
and the channel surface is where the throughput is. Conflating them produces one
endpoint that is both a state machine and a byte pipe, which is the thing
dev/sessions already rejected.
Control surface
Following the CP's existing grammar exactly — /v1/namespaces/:ns/<resource>,
named resources on :name, sub-verbs as POST sub-paths, the same shape
queues and providers already use.
:name addresses the declaration; the run is identified by the
agent_run_id a start returns. That asymmetry is deliberate and matches
providers/:name/instances/:id: you ask for an agent by the name a human wrote,
and the platform tells you which instance answered.
start is ensure-shaped, not create-shaped — starting an already-running
agent returns the existing run rather than a 409. This is the idempotency rule
from Agent lifecycle applied to its
own API: the caller most likely to hit this endpoint is an agent that lost
context and re-issued.
State, and where it rides
The states are declared → starting → ready → (draining → stopped | failed | lost), unchanged from the lifecycle doc. The proposal here is only about
transport: they ride the existing per-entity monotonic status_seq with
push-with-outbox for terminal transitions, and lost is detected by lease
expiry — the same mechanism that already marks a daemon-less invocation lost. No
new machinery, and specifically no agent-specific poller.
Channel surface
This is the load-bearing part, and there is a shipped primitive that fits.
Topic already has exactly the right semantics — post / listen / stream
over the frame journal, with a cursor so a reconnecting subscriber passes the
last seq it saw (dreamlake/lakeshore/queues.py:199). It is not usable over the
wire today for one specific reason: a topic append becomes
POST /v1/daemon/result-chunk with invocation_id: "<topic-name>", and a topic
name is not an invocation row, so it 404s.
Proposal: make the journal addressable by a channel id that is not an
invocation id, and the agent channel is Topic with no further work.
The SSE half already has a working precedent in /exec/:id/log/stream —
log-cache.ts runs a Redis Stream per exec with Last-Event-ID backfill, and
result-journal.ts has both the write and the cursor-read halves over a
ResultChunk journal whose @@unique([invocationId, seq]) is simultaneously its
dedup key and its replay index. The agent channel needs the same two halves over
a differently-keyed journal, not a new design.
Alternative considered — a WebSocket. Rejected for v1, not on the merits of
the protocol but because SSE is what the deploy target is known to survive
(dev/sessions flags the Heroku H12 cap as a thing to check for WS, and nobody
has checked). SSE plus a POST for the inbound half is one round-trip worse per
message and zero new operational unknowns. Revisit when message rate makes it
hurt.
SDK surface
Agent slots in beside UDF and Topic in dreamlake.lakeshore.__all__. The
shape mirrors Queue, deliberately — same verbs, async default, sync twin.
Three things this shape commits to:
sendtakes the declared arguments as kwargs, not a formatted string. The substitution stays on the platform, which is the only reason the rendered prompt can be recorded on the invocation.stream(since=…)is the same cursor contract asTopic.stream. A reconnecting caller passes the last seq it saw. No separate reconnect verb.- No
Agent(...)decorator. An agent is authored as markdown, not as a decorated Python function — that is what makes it not a UDF. A decorator here would quietly re-introduce the conflation the declaration deliberately avoids.
The six open questions, answered
Each of these is open on Agent lifecycle. Recommendations, with the reason and the cost.
1. Workdir on restart — fresh, with an explicit escape
A restart is a fresh start. The workdir survives invocations, not restarts.
Why. Preserving is friendlier and strictly riskier: a poisoned workdir outlives the fix that was supposed to clear it, and the failure it produces is indistinguishable from the original bug. The weaker promise is also the one the lifecycle doc already writes down — "treat it as a warm cache, not a system of record" — so this ratifies the text rather than contradicting it.
Cost. A restart re-clones and re-installs. Mitigated by the fact that mounts
are the durable layer: a git mount re-resolves to the same pinned commit
cheaply, so the expensive part is the virtualenv, not the source.
Escape. --preserve-workdir on restart, for the debugging case where the
whole point is to inspect what the agent left behind.
2. Concurrency within one agent — serialized by default, declared opt-in
max_concurrency: 1 on the agent record. Authors may raise it.
Why. The persistent workdir is shared mutable state, and two invocations
writing it concurrently is a data race the platform cannot see, cannot detect,
and cannot attribute. Serializing also gives the channel a single writer, which
is what makes seq monotonic without a merge step.
Honest limitation. max_concurrency: N is the author asserting their agent
does not write the workdir. We cannot verify that. It is a declaration, not a
guarantee, and the docs should say so rather than implying enforcement.
3. Channel durability — replayable, bounded the way sessions are
Replayable, with a cursor, two-tier retention.
Why. The machinery already exists and already does the harder version. Choosing
best-effort would mean building a new, weaker path alongside a stronger one
that ships today. Tier 1 is Redis Streams with MAXLEN while the agent is live;
Tier 2 flushes to the storage prefix on terminal transition. This is the log
storage design applied unchanged.
Bound. Replay is "resume from your cursor", not "scroll back through the agent's entire history". Anything the caller needs permanently is a write to storage, not a message it expects to re-read next week.
4. Where the definition lives — DreamLake, ratified
This is already true in the tree; the question is only whether it is written down. Write it down.
The consequence worth stating plainly is the one that is already biting:
agent prose has no version history, because RunnableVersion.version is
sha256(source) and must stay byte-identical to the CP's Function.version.
Folding the prompt into that hash forks the identity of the row.
Proposal. Do not touch the hash. Add an append-only RunnablePrompt
side-table keyed (runnableId, seq), holding one immutable row per edit to
markdown. Cheap, additive, no effect on placement or identity — and it is what
makes question 6 answerable, so the two land together.
5. timeout_s and run_setup under managed — re-read per message
| Field | external | managed |
|---|---|---|
timeout_s | The process. Yours. | One message. A deadline per unit of work the platform dispatched. |
run_setup | Before the body. | Once per agent process, not per message. |
on_host_setup | Per host, host-key-scoped | unchanged |
on_shutdown | At exit | At drain |
Plus one new field: idle_timeout_s, the deadline for the process under
managed — because once timeout_s means "one message", nothing else answers
"how long may this agent sit doing nothing". Without it, a managed agent with no
traffic is immortal.
Deliberately not added: an on_message_setup hook. Per-message setup is what
the prompt is for. A fourth lifecycle hook would need a fourth place to reason
about failure, and dls.lifecycle already exports three that nothing calls.
6. Is the rendered prompt enough of an audit record — no; pair it
Not on its own. Record four things per invocation:
- the rendered prompt (already designed),
- the resolved argument values, separately, so you can re-render,
- the pinned mount commits (
ref→commitat dispatch, already designed), - the prompt version from question 4's side-table.
What is still not reconstructable, stated rather than papered over: the accumulated workdir. An agent that has been running for a week is not derivable from its calls. Question 1 bounds this usefully — since a restart clears the workdir, the state is bounded by one agent lifetime, and if the invocation sequence within that lifetime is recorded in order, a replay is well-defined: re-run that lifetime's invocations from the start.
So the honest claim is reproducible within a lifetime, not across restarts — and that is worth writing on the user-facing page, because the alternative is readers assuming the stronger thing.
What the labeler inherits
A data labeler — human or model — is structurally a managed agent. It holds context across many items, it is dispatched work rather than called, and it writes results back. Four of the decisions above transfer without modification:
| Primitive | Agent | Labeler |
|---|---|---|
| Long-lived worker with a persistent workdir (§1) | checkout, virtualenv, cache | decoded frames, model weights, a video cache |
| Serialized by default (§2) | one message at a time | one item at a time — and here it is also correctness, since two labelers on one item is a conflict, not a speedup |
| Replayable channel with a cursor (§3) | messages | the item queue; a labeler that reconnects must not lose or re-do items |
| Provenance triple (§6) | rendered prompt + prompt version + pinned commit | rendered guidelines + schema version + item revision |
The one real difference, which the API must accommodate rather than route
around: an agent's result is a value; a labeler's result is a write — an
Annotation layer, a Track value, an Episode.revise — and the thing worth
returning on the channel is a reference to that write, not the payload.
So the channel result type cannot be "bytes" or "the return value". It has to be able to carry a side-effect reference. Designing that in now costs one field. Discovering it later costs a second channel.
Do not build dreamlake labelers as a parallel dispatch plane. If the surface
above is right, a labeler is an agent with a different kind and an
annotation-typed result — and if the surface above is not right, we should
find that out now, while there is one caller instead of two.
What this does not decide
- Fleet policy. Unchanged and deliberately out of scope: how many agents to keep warm, which instance types are permitted, scale-down delay. Queue-level, operator-owned.
- Convergence with workflow
udanodes. They now share thetoolsvocabulary and nothing else. Still two declarations; still future work. - Human-in-the-loop auth for labelers. A human labeler is an identity with a session, not a token on a daemon. Genuinely unaddressed.
- The mount activation seam. Everything above that involves a mount inherits the fact that no runner activates mounts today, for any kind — see the status callout on Mounts.
Implementation order
Mirrors the order dev/sessions argues for, because it is the same substrate.
- Journal addressable by non-invocation id. Unblocks
Topicover HTTP, which is the channel. Smallest change here and the highest leverage. RunnablePromptside-table on dreamlake-server. Additive; unblocks questions 4 and 6 together.- Agent process manager in nymph. The genuinely absent piece — no PTY, no
child registry, no
Sessionsstruct today. - CP control routes + state transport on the existing
status_seqwire. - SSE channel proxy.
- SDK
Agent, then the CLI. - Labeler
kind— only after 1–6 have one working caller.
Steps 1 and 2 are independent of everything else and of each other. They are the right thing to start on if this proposal is accepted in outline but not in detail.