DreamLake

Agent execution overview

An agent is a name and a prompt:

bash
dreamlake agents create triage --prompt 'You triage failed runs. Read the
terminal event first, classify the failure, quote the line that decided it.'

That is a complete agent. Everything else is optional and layered on when the agent earns it — Claude's tools, model and permissions when it should be constrained, an attached RunConfig when it needs a machine, and arguments when its prompt has holes in it.

The configured form is Claude's agent spec — the same frontmatter file you would otherwise drop in .claude/agents/ — with two additions this platform needs:

  • an optional run configuration, because unlike Claude Code an agent here does not run on the machine you are sitting at, and
  • arguments, values substituted into its prompt at dispatch so that a run is reconstructable from what was recorded rather than from a template that may have changed since.

An agent is not a UDF. They share a storage table and nothing else worth knowing: a UDF is addressed as <module>.<qualname> because it is code, an agent has a name the way a person does, and a UDF needs a RunConfig to be placed where an agent usually has none.

Status: declaration only

The agent record is real and DreamLake stores it. Execution is not. There is no agent API in the SDK, no agent route on the control plane, and no agent entry in the CP schema. dls.lifecycle exports on_host_setup / on_run_setup / on_shutdown and nothing calls them. Everything on this page that describes an agent running describes the intended shape. Functions is what runs today.

Where the pieces live

The split follows the platform's general rule — DreamLake owns declarations, Lakeshore owns execution:

OwnerWhat it holds
The agent fileDreamLakename, description, tools, model, permissions, the prompt body, arguments
The RunConfig it namesDreamLakeimage, resources, host_setup, run_setup, env, timeout_s, mounts, provider
The queue, invocations, workers, eventsLakeshore CPeverything ephemeral

The reference for the declaration — Agent declaration reference has the field-by-field tables, the permission rule syntax, and the argument grammar. This page is the execution view: what a run configuration means for something that is supposed to outlive a single call, and where the design is still open.

The file, in brief

markdown
---
name: run-triage
description: Triages failed invocations and proposes the smallest RunConfig fix.
tools: [Read, Grep, Glob, Bash]
model: claude-opus-5

permissions:
  allow: [Read(./runs/**), Bash(lakeshore logs:*)]
  ask:   [Edit(./runconfigs/**)]
  deny:  [Read(./.env), Bash(curl:*)]

run_config: gpu-a10g
queue: gpu-a10g

arguments:
  - { name: invocation_id, type: string, required: true }
  - { name: max_log_lines, type: integer, default: 200 }
---

You triage failed invocations. The invocation is `{{ invocation_id }}`;
read at most `{{ max_log_lines }}` lines from the end of its log.

name, description, tools, model, permissions and the markdown body are Claude's spec verbatim. run_config, queue and arguments are the extensions, and every field above the body is optional — a bare markdown file with no frontmatter at all is a valid agent.

Declare it with the DreamLake CLI, which takes the file, a heredoc or $EDITOR:

bash
dreamlake agents create --file agents/run-triage.md
dreamlake agents create run-triage --edit
dreamlake agents create triage --file - <<'AGENT'
You triage failed invocations. Read the terminal event before the logs.
AGENT

Every frontmatter field also has a flag — --tools, --model, --allow, --ask, --deny, --run-config, --queue, --arg — so a one-property agent does not need a file. The full per-property reference is on Agent declaration reference.

Attaching a run configuration

Optional. A prompt-only agent needs no machine and most do not have one. Attach a RunConfig when the agent has to run somewhere in particular — a GPU, a specific image, a prepared host:

bash
dreamlake agents create colmap-driver --file agent.md \
  --run-config colmap-docker-gpu --queue colmap-gpu

run_config names an existing RunConfig; it does not copy one. That is the point of the indirection — the machine can be re-tuned, or shared with a UDF that needs the same host, without touching the agent. It is the same declaration a queue-bound @udf uses, and the same one whose fields you can read off lakeshore runconfig get.

FieldEffect on the agent
runtime.image, runtime.resources, runtime.host_setup, provider, extrasIn the host key. These decide which worker can serve it.
runtime.env, runtime.run_setup, runtime.timeout_s, tags, runtime.dockerNot in the host key. Applied at exec time.
mountsAttached into the workdir before dispatch. Today metadata only — no runner activates them yet, for any kind of job.
queueWhich queue the work is submitted on, and therefore which fleet policy applies.

The host key is what actually places it

There is no second field where an author says "and this needs a GPU host". The placement requirement is derived from the RunConfig by hashing the subset above:

subset   = { v, runner, image, resources, host_setup, provider, extras }
host_key = "hk1_" + base32(sha256(canonical_json(subset))[:10])

Two agents with the same key can share a warm worker. Two that differ in image, resources or the order of host_setup steps never will, however similar they look — and that is the mechanism working, not a bug. A worker may claim only invocations whose host_key is absent or equal to the one it advertises, so a job whose key nobody advertises waits rather than running somewhere wrong.

This part is not aspirational: the key is computed on every write and the filter is enforced at claim time today, for UDFs. An agent inherits it unchanged.

Where the lifetime mismatch bites

A RunConfig describes a host for one invocation. An agent is supposed to outlive individual invocations, and three of the RunConfig's fields stop being obviously meaningful once it does:

FieldThe question it stops answering
timeout_sA deadline for what — the agent process, or one message it handles?
run_setupIt runs before every body by definition. Before every message?
the workdirScratch, created at dispatch and discarded at exit — or the thing that makes a long-lived agent worth having?

The workdir is the substantive one. An invocation's workdir is scratch; a managed agent's workdir is what stops it re-deriving context on every call — a checkout, a virtualenv, a cache. That changes the durability contract, and the open questions (does it survive a restart, is it per-agent or shared, does it count against the project's capacity grant) are tracked on Agent lifecycle. None of it is built.

Fleet policy is not in the agent file

Nothing in an agent declaration says how many machines to keep warm, which instance types are permitted, or what the scale-down delay is. That is queue-level fleet policy and it belongs to the operator — see Scaling rules. extras.instance_type in a RunConfig is a request; whether it is permitted is not the author's call.

Arguments at dispatch

arguments declares the holes in the prompt; a dispatch fills them.

yaml
arguments:
  - name: invocation_id
    type: string          # string | integer | number | boolean | enum
    required: true
    description: The failed invocation to triage.
  - name: max_log_lines
    type: integer
    default: 200

A {{ name }} placeholder in the body is replaced with the value. The rules, in the order they bite:

  1. A placeholder with no matching arguments entry fails at declare time — a typo is caught before anything runs, not rendered as a blank.
  2. A missing required argument fails at dispatch, before the agent starts.
  3. Values are inserted as text, once, on the literal placeholder. A value containing {{ … }} is not re-scanned. This is not a trust boundary — an argument that reaches a prompt is still untrusted input a model reads — but it does mean an argument cannot manufacture new placeholders.
  4. Placeholders are body-only. {{ … }} in frontmatter is not substituted; an agent does not get to choose its own run_config at dispatch.

Claude Code's positional forms ($1 … $9, $ARGUMENTS) are accepted and lower onto the same list, so a file pasted from .claude/commands/ keeps working. Prefer named placeholders for anything declared — $3 does not survive someone reordering the list.

Why this is a platform feature and not string formatting

You could format the prompt at the call site and send finished text. The reason not to is that the agent's instructions live on the mutable head and have no version history. Pinning an agent to a version pins the code it runs; the prose can be rewritten by anyone with write access and every pinned caller silently gets the new text.

Because the platform does the substitution, it can record the rendered prompt on the invocation. The template can drift afterwards and the run still says what it actually asked. That is the whole argument for the feature, and it is the same reason invocation lineage is recorded rather than reconstructed.

The same logic makes a git mount and a long-lived agent natural partners: a git mount resolves ref → commit at dispatch and records the commit on the invocation, so every action the agent takes is attributable to an exact tree state. Rendered prompt plus pinned commit is the pair that makes an autonomous run reconstructable.

What is missing, by name

Two pieces, both undesigned — not partially built, not behind a flag:

What it would beWhere it stands
LifecycleWhat starts an agent, keeps it alive, restarts it when its worker dies, and stops itdeclared → starting → ready → (draining → stopped | failed | lost) is sketched on Agent lifecycle, reusing the run-status transport. Nothing implements it.
ChannelHow a message reaches a running agent and how a reply returnsThe channel field — { kind, target?, config? } — is stored and validated for credentials. It is read by nothing.

The nearest shipped thing to a channel is Topic (post / listen with cursor semantics), but it is local-Dispatch-only in practice: over HTTP a topic append becomes POST /v1/daemon/result-chunk with invocation_id: "<topic-name>", and a topic name is not an invocation row, so it 404s. Making Topic work over the control plane is the cheapest real path to a bidirectional agent channel, and it is a small, well-shaped change.

Not the same as a workflow uda node

DreamLake workflows have their own agent concept — uda nodes: a typed member of a static workflow graph. That is a node in a graph, not a long-lived worker. The two now share the tools vocabulary, which is worth flagging rather than leaving to be tripped over: they are separate declarations today, and converging them is future work.