Living design doc. Captures decisions taken so far; edit freely as the design evolves. Companion to DreamLake → Lakeshore Federation.
The declaration it hangs off is settled and documented at
Agents: Claude's agent spec — name, description,
tools, model, permissions, prompt body — plus run_config and typed
arguments. This doc adds the one field that spec has no opinion about,
because Claude Code runs on a machine that is already up: lifecycle.
Problem
An agent is a long-lived caller: it holds a session, issues many invocations over its life, and receives a streamed channel of messages rather than one terminal result. That is a different shape from a UDF invocation, which is request/response and dies at the end.
Today nothing models this. Agents are implicit — a process somewhere holding a token. That leaves three things undefined: who starts and restarts them, where their working state lives between invocations, and how a caller discovers what agents exist.
The external | managed split
The same discriminator providers use, for the same reason: it lets the new behavior land without changing what already works.
| Who runs the process | Lifecycle owner | Workdir | |
|---|---|---|---|
external | You do | You do | Yours |
managed | Lakeshore | Lakeshore | Persistent, platform-owned |
external is today's implicit behavior made explicit: a process you
start, which authenticates and opens a channel. Nothing about it
changes.
managed is the new capability — Lakeshore starts the agent, health-
checks it, restarts it on failure, stops it on request, and gives it a
workdir that survives across invocations.
Lifecycle states
Monotonic, and deliberately the same shape as the run and provider state machines:
Transitions ride the same transport as run status — a per-entity
monotonic status_seq, push-with-outbox for terminal transitions,
and lease expiry as the detector for lost. No new machinery: an
agent whose lease lapses without a heartbeat is lost, exactly as a
daemon-less invocation is.
draining matters for agents in a way it does not for invocations: a
managed agent asked to stop should finish its in-flight invocations and
close its channel cleanly rather than dropping them.
The persistent workdir
This is the substantive difference between a managed agent and a loop that calls UDFs.
An invocation's workdir is scratch — created at dispatch, discarded at exit. A managed agent's workdir persists for the life of the agent, so it accumulates: a checkout, a virtualenv, a cache, intermediate artifacts. The agent stops re-deriving context on every call.
Durability expectations, stated plainly:
- The workdir survives invocations, not necessarily restarts. Treat it as a warm cache, not a system of record. Anything that must outlive the agent goes to the storage prefix named in the launch command.
- It is scoped to one agent. Two agents never share a workdir; use a mount for that.
- It counts against the project's capacity grant like any other resource.
Why git mounts and managed agents belong together
A git mount attached to a managed agent gives it a
durable checkout it can work in across many invocations — which is the
difference between an agent that re-clones on every call and one that
holds a workspace.
The composition is what makes it useful:
The composition is written in the agent file's frontmatter — Claude's fields on top, the platform's underneath:
Because the git mount resolves ref → commit at dispatch and
records the commit on the invocation, every action the agent takes is
attributable to an exact tree state. For an autonomous caller that is
not a nicety — it is the only way to reconstruct why it did what it
did.
This is also the path by which a GitOps desired-state repo reaches a reconciler: mount the declarations repo read-only at a pinned commit, converge, write observed state out to the storage prefix.
AI-native affordances
Agents are the caller these surfaces are designed for, which changes what the API owes them:
- Every verb idempotent. An agent that lost context re-issues.
ensure-shaped, notcreate-shaped. - Typed failure with a next action.
INSUFFICIENT_QUOTA / retry_after,MOUNT_FAILED / fix_and_retry,AGENT_LOST / restart— an agent has to branch on failure, and a stack trace collapses three different behaviors into one. - Plan as a verb. Before a committing action, return structured blast radius: resources touched, estimated cost, reversibility.
- Soft delete is the undo buffer for a caller that errs faster than a human can intervene.
Open questions
- Restart semantics for the workdir — preserve across restart, or treat restart as a fresh start? Preserving is friendlier and riskier (a poisoned workdir survives the fix).
- Concurrency within one agent — may a managed agent have more than one invocation in flight, or is it serialized? Serialized is simpler and probably right first.
- Channel durability — is the message channel replayable after a reconnect, or best-effort? Replay needs a cursor and retention.
- Where the agent definition lives — DreamLake as project-scoped definitional schema, consistent with UDFs and providers. Assumed, not yet ratified.
- Which
lifecyclevalues a RunConfig'stimeout_sandrun_setupapply to. Underexternalthey are unambiguous — the process is yours. Undermanagedthey need re-reading: a deadline for the agent or for one message, setup before the process or before every message. The Agents page states the question; nothing answers it yet. - Whether the rendered prompt is enough of an audit record. The
argument-substitution design records the rendered prompt per
invocation, which fixes reproducibility for a call. A managed agent
accumulating state in a persistent workdir is not reconstructable from
its calls alone, and a
gitmount's pinned commit only covers what it read, not what it wrote.
Status
Design only. Nothing here is implemented. Depends on the mount activation seam, which is unbuilt for every kind — see the status callout on Mounts.