# Agent execution overview

An agent is a **name and a prompt**:

```bash
dreamlake agents create triage --prompt 'You triage failed runs. Read the
terminal event first, classify the failure, quote the line that decided it.'
```

That is a complete agent. Everything else is optional and layered on when the
agent earns it — Claude's `tools`, `model` and `permissions` when it should be
constrained, an **attached RunConfig** when it needs a machine, and
**arguments** when its prompt has holes in it.

The configured form is **Claude's agent spec** — the same frontmatter file you
would otherwise drop in `.claude/agents/` — with two additions this platform
needs:

- an optional **run configuration**, because unlike Claude Code an agent here
  does not run on the machine you are sitting at, and
- **arguments**, values substituted into its prompt at dispatch so that a run is
  reconstructable from what was recorded rather than from a template that may
  have changed since.

An agent is **not a UDF**. They share a storage table and nothing else worth
knowing: a UDF is addressed as `<module>.<qualname>` because it is code, an
agent has a name the way a person does, and a UDF needs a RunConfig to be placed
where an agent usually has none.

> **Warning:** The agent *record* is real and DreamLake stores it. **Execution is not.** There
>   is no agent API in the SDK, no agent route on the control plane, and no agent
>   entry in the CP schema. `dls.lifecycle` exports `on_host_setup` /
>   `on_run_setup` / `on_shutdown` and **nothing calls them**. Everything on this
>   page that describes an agent *running* describes the intended shape.
>   [Functions](https://lakeshore.dreamlake.ai/get-started/functions) is what runs today.

## Where the pieces live

The split follows the platform's general rule — **DreamLake owns declarations,
Lakeshore owns execution**:

| | Owner | What it holds |
| --- | --- | --- |
| The agent file | DreamLake | `name`, `description`, `tools`, `model`, `permissions`, the prompt body, `arguments` |
| The RunConfig it names | DreamLake | image, resources, `host_setup`, `run_setup`, `env`, `timeout_s`, mounts, provider |
| The queue, invocations, workers, events | Lakeshore CP | everything ephemeral |

The **reference for the declaration** —
[Agent declaration reference](/lakeshore/agents.md) has the
field-by-field tables, the permission rule syntax, and the argument grammar.
This page is the execution view: what a run configuration means for something
that is supposed to outlive a single call, and where the design is still open.

## The file, in brief

```markdown file=agents/run-triage.md
---
name: run-triage
description: Triages failed invocations and proposes the smallest RunConfig fix.
tools: [Read, Grep, Glob, Bash]
model: claude-opus-5

permissions:
  allow: [Read(./runs/**), Bash(lakeshore logs:*)]
  ask:   [Edit(./runconfigs/**)]
  deny:  [Read(./.env), Bash(curl:*)]

run_config: gpu-a10g
queue: gpu-a10g

arguments:
  - { name: invocation_id, type: string, required: true }
  - { name: max_log_lines, type: integer, default: 200 }
---

You triage failed invocations. The invocation is `{{ invocation_id }}`;
read at most `{{ max_log_lines }}` lines from the end of its log.
```

`name`, `description`, `tools`, `model`, `permissions` and the markdown body are
Claude's spec verbatim. `run_config`, `queue` and `arguments` are the extensions,
and every field above the body is optional — a bare markdown file with no
frontmatter at all is a valid agent.

Declare it with the DreamLake CLI, which takes the file, a heredoc or `$EDITOR`:

```bash
dreamlake agents create --file agents/run-triage.md
dreamlake agents create run-triage --edit
dreamlake agents create triage --file - <<'AGENT'
You triage failed invocations. Read the terminal event before the logs.
AGENT
```

Every frontmatter field also has a flag — `--tools`, `--model`, `--allow`,
`--ask`, `--deny`, `--run-config`, `--queue`, `--arg` — so a one-property agent
does not need a file. The full per-property reference is on
[Agent declaration reference](/lakeshore/agents.md#creating-one).

## Attaching a run configuration

**Optional.** A prompt-only agent needs no machine and most do not have one.
Attach a RunConfig when the agent has to run somewhere in particular — a GPU, a
specific image, a prepared host:

```bash
dreamlake agents create colmap-driver --file agent.md \
  --run-config colmap-docker-gpu --queue colmap-gpu
```

`run_config` **names** an existing
[RunConfig](/lakeshore/registry.md); it does not copy one.
That is the point of the indirection — the machine can be re-tuned, or shared
with a UDF that needs the same host, without touching the agent. It is the same
declaration a queue-bound `@udf` uses, and the same one whose fields you can
read off `lakeshore runconfig get`.

| Field | Effect on the agent |
| --- | --- |
| `runtime.image`, `runtime.resources`, `runtime.host_setup`, `provider`, `extras` | **In the host key.** These decide which worker can serve it. |
| `runtime.env`, `runtime.run_setup`, `runtime.timeout_s`, `tags`, `runtime.docker` | Not in the host key. Applied at exec time. |
| `mounts` | Attached into the workdir before dispatch. Today [metadata only](/lakeshore/mounts.md) — no runner activates them yet, for any kind of job. |
| `queue` | Which queue the work is submitted on, and therefore which fleet policy applies. |

### The host key is what actually places it

There is no second field where an author says "and this needs a GPU host". The
placement requirement is **derived** from the RunConfig by hashing the subset
above:

```
subset   = { v, runner, image, resources, host_setup, provider, extras }
host_key = "hk1_" + base32(sha256(canonical_json(subset))[:10])
```

Two agents with the same key can share a warm worker. Two that differ in
`image`, `resources` or the *order* of `host_setup` steps never will, however
similar they look — and that is the mechanism working, not a bug. A worker may
claim only invocations whose `host_key` is absent or equal to the one it
advertises, so a job whose key nobody advertises **waits rather than running
somewhere wrong**.

This part is not aspirational: the key is computed on every write and the filter
is enforced at claim time today, for UDFs. An agent inherits it unchanged.

### Where the lifetime mismatch bites

A RunConfig describes a **host for one invocation**. An agent is supposed to
outlive individual invocations, and three of the RunConfig's fields stop being
obviously meaningful once it does:

| Field | The question it stops answering |
| --- | --- |
| `timeout_s` | A deadline for *what* — the agent process, or one message it handles? |
| `run_setup` | It runs before every body by definition. Before every *message*? |
| the workdir | Scratch, created at dispatch and discarded at exit — or the thing that makes a long-lived agent worth having? |

The workdir is the substantive one. An invocation's workdir is scratch; a
managed agent's workdir is what stops it re-deriving context on every call — a
checkout, a virtualenv, a cache. That changes the durability contract, and the
open questions (does it survive a restart, is it per-agent or shared, does it
count against the project's capacity grant) are tracked on
[Agent lifecycle](/dev/agent-lifecycle). None of it is built.

> **Note:** Nothing in an agent declaration says how many machines to keep warm, which
>   instance types are permitted, or what the scale-down delay is. That is
>   queue-level fleet policy and it belongs to the operator —
>   see [Scaling rules](https://lakeshore.dreamlake.ai/get-started/scaling-rules). `extras.instance_type` in a
>   RunConfig is a *request*; whether it is permitted is not the author's call.

## Arguments at dispatch

`arguments` declares the holes in the prompt; a dispatch fills them.

```yaml
arguments:
  - name: invocation_id
    type: string          # string | integer | number | boolean | enum
    required: true
    description: The failed invocation to triage.
  - name: max_log_lines
    type: integer
    default: 200
```

A `{{ name }}` placeholder in the body is replaced with the value. The rules,
in the order they bite:

1. A placeholder with **no matching `arguments` entry** fails at declare time —
   a typo is caught before anything runs, not rendered as a blank.
2. A **missing required argument** fails at dispatch, before the agent starts.
3. Values are inserted **as text**, once, on the literal placeholder. A value
   containing `{{ … }}` is not re-scanned. This is not a trust boundary — an
   argument that reaches a prompt is still untrusted input a model reads — but
   it does mean an argument cannot manufacture new placeholders.
4. Placeholders are **body-only**. `{{ … }}` in frontmatter is not substituted;
   an agent does not get to choose its own `run_config` at dispatch.

Claude Code's positional forms (`$1` … `$9`, `$ARGUMENTS`) are accepted and
lower onto the same list, so a file pasted from `.claude/commands/` keeps
working. Prefer named placeholders for anything declared — `$3` does not survive
someone reordering the list.

### Why this is a platform feature and not string formatting

You could format the prompt at the call site and send finished text. The reason
not to is that **the agent's instructions live on the mutable head and have no
version history.** Pinning an agent to a version pins the *code* it runs; the
prose can be rewritten by anyone with write access and every pinned caller
silently gets the new text.

Because the platform does the substitution, it can **record the rendered prompt
on the invocation**. The template can drift afterwards and the run still says
what it actually asked. That is the whole argument for the feature, and it is
the same reason invocation lineage is recorded rather than reconstructed.

The same logic makes a [`git` mount](/lakeshore/mounts.md) and a long-lived agent
natural partners: a `git` mount resolves `ref` → `commit` at dispatch and records
the commit on the invocation, so every action the agent takes is attributable to
an exact tree state. Rendered prompt plus pinned commit is the pair that makes an
autonomous run reconstructable.

## What is missing, by name

Two pieces, both undesigned — not partially built, not behind a flag:

| | What it would be | Where it stands |
| --- | --- | --- |
| **Lifecycle** | What starts an agent, keeps it alive, restarts it when its worker dies, and stops it | `declared → starting → ready → (draining → stopped \| failed \| lost)` is sketched on [Agent lifecycle](/dev/agent-lifecycle), reusing the run-status transport. Nothing implements it. |
| **Channel** | How a message reaches a running agent and how a reply returns | The `channel` field — `{ kind, target?, config? }` — is stored and validated for credentials. It is **read by nothing**. |

The nearest shipped thing to a channel is `Topic` (post / listen with cursor
semantics), but it is local-`Dispatch`-only in practice: over HTTP a topic append
becomes `POST /v1/daemon/result-chunk` with `invocation_id: "<topic-name>"`, and
a topic name is not an invocation row, so it 404s. Making `Topic` work over the
control plane is the cheapest real path to a bidirectional agent channel, and it
is a small, well-shaped change.

## Not the same as a workflow `uda` node

DreamLake workflows have their own agent concept — `uda` nodes: a typed member
of a static workflow graph. That is a node in a graph, not a long-lived worker.
The two now share the `tools` vocabulary, which is worth flagging rather than
leaving to be tripped over: they are separate declarations today, and converging
them is future work.
