DreamLake

Living design doc. Captures decisions taken so far; edit freely as the design evolves. Companion to DreamLake → Lakeshore Federation.

The declaration it hangs off is settled and documented at Agents: Claude's agent spec — name, description, tools, model, permissions, prompt body — plus run_config and typed arguments. This doc adds the one field that spec has no opinion about, because Claude Code runs on a machine that is already up: lifecycle.

Problem

An agent is a long-lived caller: it holds a session, issues many invocations over its life, and receives a streamed channel of messages rather than one terminal result. That is a different shape from a UDF invocation, which is request/response and dies at the end.

Today nothing models this. Agents are implicit — a process somewhere holding a token. That leaves three things undefined: who starts and restarts them, where their working state lives between invocations, and how a caller discovers what agents exist.

The external | managed split

The same discriminator providers use, for the same reason: it lets the new behavior land without changing what already works.

Who runs the processLifecycle ownerWorkdir
externalYou doYou doYours
managedLakeshoreLakeshorePersistent, platform-owned

external is today's implicit behavior made explicit: a process you start, which authenticates and opens a channel. Nothing about it changes.

managed is the new capability — Lakeshore starts the agent, health- checks it, restarts it on failure, stops it on request, and gives it a workdir that survives across invocations.

Lifecycle states

Monotonic, and deliberately the same shape as the run and provider state machines:

declared → starting → ready → (draining → stopped | failed | lost)
                        ↑                              │
                        └──────── restart ─────────────┘

Transitions ride the same transport as run status — a per-entity monotonic status_seq, push-with-outbox for terminal transitions, and lease expiry as the detector for lost. No new machinery: an agent whose lease lapses without a heartbeat is lost, exactly as a daemon-less invocation is.

draining matters for agents in a way it does not for invocations: a managed agent asked to stop should finish its in-flight invocations and close its channel cleanly rather than dropping them.

The persistent workdir

This is the substantive difference between a managed agent and a loop that calls UDFs.

An invocation's workdir is scratch — created at dispatch, discarded at exit. A managed agent's workdir persists for the life of the agent, so it accumulates: a checkout, a virtualenv, a cache, intermediate artifacts. The agent stops re-deriving context on every call.

Durability expectations, stated plainly:

  • The workdir survives invocations, not necessarily restarts. Treat it as a warm cache, not a system of record. Anything that must outlive the agent goes to the storage prefix named in the launch command.
  • It is scoped to one agent. Two agents never share a workdir; use a mount for that.
  • It counts against the project's capacity grant like any other resource.

Why git mounts and managed agents belong together

A git mount attached to a managed agent gives it a durable checkout it can work in across many invocations — which is the difference between an agent that re-clones on every call and one that holds a workspace.

The composition is what makes it useful:

The composition is written in the agent file's frontmatter — Claude's fields on top, the platform's underneath:

markdown
---
name: infra-operator
description: Reconciles declared provider state against what is actually running.
tools: [Read, Edit, Bash]
model: claude-opus-5
permissions:
  allow: [Read(./providers/**), Bash(lakeshore provider:*)]
  ask:   [Edit(./providers/**)]
  deny:  [Bash(rm:*)]

run_config: k8s-gpu
lifecycle: managed          # ← the field this doc proposes
channel: { kind: stream }
mounts:
  - { kind: git, repo: …/lakeshore-workspace, ref: main, path: providers/ }

arguments:
  - { name: provider_name, type: string, required: true }
---

Reconcile `{{ provider_name }}` against the declarations mounted at
`providers/`. Report drift; do not apply it unless asked.

Because the git mount resolves ref → commit at dispatch and records the commit on the invocation, every action the agent takes is attributable to an exact tree state. For an autonomous caller that is not a nicety — it is the only way to reconstruct why it did what it did.

This is also the path by which a GitOps desired-state repo reaches a reconciler: mount the declarations repo read-only at a pinned commit, converge, write observed state out to the storage prefix.

AI-native affordances

Agents are the caller these surfaces are designed for, which changes what the API owes them:

  • Every verb idempotent. An agent that lost context re-issues. ensure-shaped, not create-shaped.
  • Typed failure with a next action. INSUFFICIENT_QUOTA / retry_after, MOUNT_FAILED / fix_and_retry, AGENT_LOST / restart — an agent has to branch on failure, and a stack trace collapses three different behaviors into one.
  • Plan as a verb. Before a committing action, return structured blast radius: resources touched, estimated cost, reversibility.
  • Soft delete is the undo buffer for a caller that errs faster than a human can intervene.

Open questions

  1. Restart semantics for the workdir — preserve across restart, or treat restart as a fresh start? Preserving is friendlier and riskier (a poisoned workdir survives the fix).
  2. Concurrency within one agent — may a managed agent have more than one invocation in flight, or is it serialized? Serialized is simpler and probably right first.
  3. Channel durability — is the message channel replayable after a reconnect, or best-effort? Replay needs a cursor and retention.
  4. Where the agent definition lives — DreamLake as project-scoped definitional schema, consistent with UDFs and providers. Assumed, not yet ratified.
  5. Which lifecycle values a RunConfig's timeout_s and run_setup apply to. Under external they are unambiguous — the process is yours. Under managed they need re-reading: a deadline for the agent or for one message, setup before the process or before every message. The Agents page states the question; nothing answers it yet.
  6. Whether the rendered prompt is enough of an audit record. The argument-substitution design records the rendered prompt per invocation, which fixes reproducibility for a call. A managed agent accumulating state in a persistent workdir is not reconstructable from its calls alone, and a git mount's pinned commit only covers what it read, not what it wrote.

Status

Design only. Nothing here is implemented. Depends on the mount activation seam, which is unbuilt for every kind — see the status callout on Mounts.