DreamLake

Agent API

Proposal — nothing here is implemented

The agent declaration is settled and in the tree: Runnable with kind="agent", plus markdown and channel. Everything on this page is the execution half, and none of it exists — there is no agent route on the controlplane, no Agent export in the SDK, and no agent entry in the CP schema. This page proposes the surface and answers the six questions left open on Agent lifecycle so that they stop blocking.

An agent is a Session with a policy driving it. That decision is already taken. What it does not settle is the wire: how a message reaches a running agent, how a reply comes back, what the SDK calls it, and what a run records about itself afterwards. This page is that.

Why now, and why this doc has a second audience

The forcing function is not agents. It is that a data labeler is the same shape — long-lived, holds context across many items, receives work over a channel, writes results back — and if we design its API separately we will build a second dispatch plane and a second channel that do the same job slightly differently. The labeler coupling section at the bottom names exactly which primitives are shared and which are genuinely different. Read it before ratifying anything above it.

What is already decided, and is not re-opened here

Four things are settled. This page builds on them rather than re-arguing them.

DecisionWhereWhat it fixes
An Agent is a Session plus a loopDECISIONS.md, 2026-08-06Agents get no separate lifecycle. UDF = one-shot; Session = long-lived + pipe; Agent = long-lived + pipe + policy.
DreamLake owns declarations, the CP owns executionThe four collectionsThe agent record is DreamLake's. Everything live — the channel, the invocation, the worker — is the CP's.
Placement is derived, never declaredAgentsAn agent inherits host_key from its RunConfig unchanged. There is no agent-specific placement field, and this page does not add one.
Provenance is recorded, not reconstructedsameThe rendered prompt is written to the invocation so a drifting template cannot rewrite history.

Three surfaces, one of which exists

SurfaceOwnerStatus
Declaration — create, edit, version, list agentsDreamLake serverRoutes exist under /namespaces/:slug/runnables.
Control — start, stop, drain, inspect a running agentLakeshore CPProposed below. Nothing today.
Channel — send a message, read repliesLakeshore CPProposed below. The channel field is stored and read by nothing.

The split matters because the control surface is where the lifecycle states live, and the channel surface is where the throughput is. Conflating them produces one endpoint that is both a state machine and a byte pipe, which is the thing dev/sessions already rejected.

Control surface

Following the CP's existing grammar exactly — /v1/namespaces/:ns/<resource>, named resources on :name, sub-verbs as POST sub-paths, the same shape queues and providers already use.

POST   /v1/namespaces/:ns/agents/:name/start     → { agent_run_id, state }
GET    /v1/namespaces/:ns/agents/:name           → current run + state + worker
GET    /v1/namespaces/:ns/agents                 → list, with state filter
POST   /v1/namespaces/:ns/agents/:name/drain     → finish in-flight, then stop
DELETE /v1/namespaces/:ns/agents/:name           → SIGTERM now

:name addresses the declaration; the run is identified by the agent_run_id a start returns. That asymmetry is deliberate and matches providers/:name/instances/:id: you ask for an agent by the name a human wrote, and the platform tells you which instance answered.

start is ensure-shaped, not create-shaped — starting an already-running agent returns the existing run rather than a 409. This is the idempotency rule from Agent lifecycle applied to its own API: the caller most likely to hit this endpoint is an agent that lost context and re-issued.

State, and where it rides

The states are declared → starting → ready → (draining → stopped | failed | lost), unchanged from the lifecycle doc. The proposal here is only about transport: they ride the existing per-entity monotonic status_seq with push-with-outbox for terminal transitions, and lost is detected by lease expiry — the same mechanism that already marks a daemon-less invocation lost. No new machinery, and specifically no agent-specific poller.

Channel surface

This is the load-bearing part, and there is a shipped primitive that fits.

Topic already has exactly the right semantics — post / listen / stream over the frame journal, with a cursor so a reconnecting subscriber passes the last seq it saw (dreamlake/lakeshore/queues.py:199). It is not usable over the wire today for one specific reason: a topic append becomes POST /v1/daemon/result-chunk with invocation_id: "<topic-name>", and a topic name is not an invocation row, so it 404s.

Proposal: make the journal addressable by a channel id that is not an invocation id, and the agent channel is Topic with no further work.

POST   /v1/namespaces/:ns/agents/:name/messages      → { seq }
GET    /v1/namespaces/:ns/agents/:name/stream        → SSE, honours Last-Event-ID

The SSE half already has a working precedent in /exec/:id/log/stream — log-cache.ts runs a Redis Stream per exec with Last-Event-ID backfill, and result-journal.ts has both the write and the cursor-read halves over a ResultChunk journal whose @@unique([invocationId, seq]) is simultaneously its dedup key and its replay index. The agent channel needs the same two halves over a differently-keyed journal, not a new design.

Alternative considered — a WebSocket. Rejected for v1, not on the merits of the protocol but because SSE is what the deploy target is known to survive (dev/sessions flags the Heroku H12 cap as a thing to check for WS, and nobody has checked). SSE plus a POST for the inbound half is one round-trip worse per message and zero new operational unknowns. Revisit when message rate makes it hurt.

SDK surface

Agent slots in beside UDF and Topic in dreamlake.lakeshore.__all__. The shape mirrors Queue, deliberately — same verbs, async default, sync twin.

python
from dreamlake.lakeshore import Agent

triage = Agent("run-triage")          # resolves the declaration by name

run = await triage.start()            # ensure-shaped; returns the live run
await run.send(invocation_id="inv_01H…", max_log_lines=200)
async for msg in run.stream(since=-1):
    print(msg)
await run.drain()

Three things this shape commits to:

  1. send takes the declared arguments as kwargs, not a formatted string. The substitution stays on the platform, which is the only reason the rendered prompt can be recorded on the invocation.
  2. stream(since=…) is the same cursor contract as Topic.stream. A reconnecting caller passes the last seq it saw. No separate reconnect verb.
  3. No Agent(...) decorator. An agent is authored as markdown, not as a decorated Python function — that is what makes it not a UDF. A decorator here would quietly re-introduce the conflation the declaration deliberately avoids.

The six open questions, answered

Each of these is open on Agent lifecycle. Recommendations, with the reason and the cost.

1. Workdir on restart — fresh, with an explicit escape

A restart is a fresh start. The workdir survives invocations, not restarts.

Why. Preserving is friendlier and strictly riskier: a poisoned workdir outlives the fix that was supposed to clear it, and the failure it produces is indistinguishable from the original bug. The weaker promise is also the one the lifecycle doc already writes down — "treat it as a warm cache, not a system of record" — so this ratifies the text rather than contradicting it.

Cost. A restart re-clones and re-installs. Mitigated by the fact that mounts are the durable layer: a git mount re-resolves to the same pinned commit cheaply, so the expensive part is the virtualenv, not the source.

Escape. --preserve-workdir on restart, for the debugging case where the whole point is to inspect what the agent left behind.

2. Concurrency within one agent — serialized by default, declared opt-in

max_concurrency: 1 on the agent record. Authors may raise it.

Why. The persistent workdir is shared mutable state, and two invocations writing it concurrently is a data race the platform cannot see, cannot detect, and cannot attribute. Serializing also gives the channel a single writer, which is what makes seq monotonic without a merge step.

Honest limitation. max_concurrency: N is the author asserting their agent does not write the workdir. We cannot verify that. It is a declaration, not a guarantee, and the docs should say so rather than implying enforcement.

3. Channel durability — replayable, bounded the way sessions are

Replayable, with a cursor, two-tier retention.

Why. The machinery already exists and already does the harder version. Choosing best-effort would mean building a new, weaker path alongside a stronger one that ships today. Tier 1 is Redis Streams with MAXLEN while the agent is live; Tier 2 flushes to the storage prefix on terminal transition. This is the log storage design applied unchanged.

Bound. Replay is "resume from your cursor", not "scroll back through the agent's entire history". Anything the caller needs permanently is a write to storage, not a message it expects to re-read next week.

4. Where the definition lives — DreamLake, ratified

This is already true in the tree; the question is only whether it is written down. Write it down.

The consequence worth stating plainly is the one that is already biting: agent prose has no version history, because RunnableVersion.version is sha256(source) and must stay byte-identical to the CP's Function.version. Folding the prompt into that hash forks the identity of the row.

Proposal. Do not touch the hash. Add an append-only RunnablePrompt side-table keyed (runnableId, seq), holding one immutable row per edit to markdown. Cheap, additive, no effect on placement or identity — and it is what makes question 6 answerable, so the two land together.

5. timeout_s and run_setup under managed — re-read per message

Fieldexternalmanaged
timeout_sThe process. Yours.One message. A deadline per unit of work the platform dispatched.
run_setupBefore the body.Once per agent process, not per message.
on_host_setupPer host, host-key-scopedunchanged
on_shutdownAt exitAt drain

Plus one new field: idle_timeout_s, the deadline for the process under managed — because once timeout_s means "one message", nothing else answers "how long may this agent sit doing nothing". Without it, a managed agent with no traffic is immortal.

Deliberately not added: an on_message_setup hook. Per-message setup is what the prompt is for. A fourth lifecycle hook would need a fourth place to reason about failure, and dls.lifecycle already exports three that nothing calls.

6. Is the rendered prompt enough of an audit record — no; pair it

Not on its own. Record four things per invocation:

  • the rendered prompt (already designed),
  • the resolved argument values, separately, so you can re-render,
  • the pinned mount commits (ref → commit at dispatch, already designed),
  • the prompt version from question 4's side-table.

What is still not reconstructable, stated rather than papered over: the accumulated workdir. An agent that has been running for a week is not derivable from its calls. Question 1 bounds this usefully — since a restart clears the workdir, the state is bounded by one agent lifetime, and if the invocation sequence within that lifetime is recorded in order, a replay is well-defined: re-run that lifetime's invocations from the start.

So the honest claim is reproducible within a lifetime, not across restarts — and that is worth writing on the user-facing page, because the alternative is readers assuming the stronger thing.

What the labeler inherits

A data labeler — human or model — is structurally a managed agent. It holds context across many items, it is dispatched work rather than called, and it writes results back. Four of the decisions above transfer without modification:

PrimitiveAgentLabeler
Long-lived worker with a persistent workdir (§1)checkout, virtualenv, cachedecoded frames, model weights, a video cache
Serialized by default (§2)one message at a timeone item at a time — and here it is also correctness, since two labelers on one item is a conflict, not a speedup
Replayable channel with a cursor (§3)messagesthe item queue; a labeler that reconnects must not lose or re-do items
Provenance triple (§6)rendered prompt + prompt version + pinned commitrendered guidelines + schema version + item revision

The one real difference, which the API must accommodate rather than route around: an agent's result is a value; a labeler's result is a write — an Annotation layer, a Track value, an Episode.revise — and the thing worth returning on the channel is a reference to that write, not the payload.

So the channel result type cannot be "bytes" or "the return value". It has to be able to carry a side-effect reference. Designing that in now costs one field. Discovering it later costs a second channel.

The thing to not do

Do not build dreamlake labelers as a parallel dispatch plane. If the surface above is right, a labeler is an agent with a different kind and an annotation-typed result — and if the surface above is not right, we should find that out now, while there is one caller instead of two.

What this does not decide

  • Fleet policy. Unchanged and deliberately out of scope: how many agents to keep warm, which instance types are permitted, scale-down delay. Queue-level, operator-owned.
  • Convergence with workflow uda nodes. They now share the tools vocabulary and nothing else. Still two declarations; still future work.
  • Human-in-the-loop auth for labelers. A human labeler is an identity with a session, not a token on a daemon. Genuinely unaddressed.
  • The mount activation seam. Everything above that involves a mount inherits the fact that no runner activates mounts today, for any kind — see the status callout on Mounts.

Implementation order

Mirrors the order dev/sessions argues for, because it is the same substrate.

  1. Journal addressable by non-invocation id. Unblocks Topic over HTTP, which is the channel. Smallest change here and the highest leverage.
  2. RunnablePrompt side-table on dreamlake-server. Additive; unblocks questions 4 and 6 together.
  3. Agent process manager in nymph. The genuinely absent piece — no PTY, no child registry, no Sessions struct today.
  4. CP control routes + state transport on the existing status_seq wire.
  5. SSE channel proxy.
  6. SDK Agent, then the CLI.
  7. Labeler kind — only after 1–6 have one working caller.

Steps 1 and 2 are independent of everything else and of each other. They are the right thing to start on if this proposal is accepted in outline but not in detail.