# Agents

  An agent is a <strong>name and a prompt</strong>. Everything else is optional
  and layered on when the agent earns it: Claude's <code>tools</code>,
  <code>model</code> and <code>permissions</code> when it should be constrained,
  an <strong>attached RunConfig</strong> when it needs a machine, and
  <strong>arguments</strong> when its prompt has holes in it.

```bash
dreamlake agents create triage --prompt 'You triage failed runs. Read the
terminal event first, classify the failure, quote the line that decided it.'
```

That is a complete agent. No tools list, no permission block, no run config, no
arguments — and it is the normal case, not a degraded one.

> **Status.** The declaration is real: DreamLake stores, validates and serves
> it. **Execution is not.** There is no agent API in the SDK and no agent route
> on the control plane. The [What is stored today](#what-is-stored-today) table
> says field-by-field which half you are in.

> **Warning:** Everything that currently *runs* is a [UDF](/lakeshore/patterns.md) — invoked,
>   runs to completion, returns. If you need remote compute now, that is the
>   surface. This page declares an agent; nothing on it drives one.

## An agent is not a UDF

They share a storage table and nothing else worth knowing, and the difference
shows up first in how they are named. A UDF is addressed as
`<module>.<qualname>` because it is *code*, and that is where code lives. An
agent has a **name**, the way a person does — `triage`, `run-triage`,
`colmap-driver` — scoped to the repository it was created in, so two repos can
each hold a `triage` without colliding.

Nothing about authoring an agent asks you to think in module paths. The CLI
derives the storage segments and prints them once; after that the name is the
name.

| | UDF | Agent |
| --- | --- | --- |
| Identity | `<module>.<qualname>` — a code path | a name, scoped to a repo |
| Body | a Python function | a markdown prompt |
| A RunConfig is | required to place it | **optional**, attached when it needs one |
| Versioned | yes — `version` is `sha256(source)` | the prompt is not; see [below](#why-arguments-and-not-string-formatting) |

## Why Claude's spec

The alternative was to invent a fifth agent format. Claude Code, the Claude
Agent SDK, and Claude Managed Agents already agree on one — `name`,
`description`, `tools`, `model`, a system prompt in the body, and an
allow/ask/deny permission block — and an agent author almost certainly already
writes it. Adopting it means an agent file is portable in both directions: the
same file drives a local `.claude/agents/` subagent and a declared platform
agent, and the fields a reader has to learn are the two that are genuinely ours.

So the rule for this page: **where Claude's spec has an answer, it is the
answer.** The Lakeshore extensions are marked as extensions everywhere they
appear, and they are additive — a plain Claude agent file is a valid
declaration.

## The agent file

The prompt can be a bare markdown file with no frontmatter at all. Add
frontmatter when you want to constrain the agent or give it a machine — this is
the fully-loaded version, and every field in it is optional:

```markdown file=agents/run-triage.md
---
# ── Claude's agent spec — unchanged ────────────────────────────────
name: run-triage
description: >-
  Triages failed invocations: reads a run's logs and events, classifies the
  failure, and proposes the smallest RunConfig change that would fix it.
tools: [Read, Grep, Glob, Bash]
model: claude-opus-5

permissions:
  allow:
    - Read(./runs/**)
    - Bash(lakeshore logs:*)
    - Bash(lakeshore events:*)
    - WebFetch(domain:docs.dreamlake.ai)
  ask:
    - Edit(./runconfigs/**)
  deny:
    - Read(./.env)
    - Read(./secrets/**)
    - Bash(curl:*)
  additionalDirectories: ["../runconfigs/"]
  defaultMode: default

# ── Lakeshore extensions ───────────────────────────────────────────
run_config: gpu-a10g
queue: gpu-a10g
channel:
  kind: http-sse
  target: /v1/agents/triage/stream
  config: { heartbeatSeconds: 15, maxIdleSeconds: 900 }

grants:
  - dreamlake.nodes.read
  - lakeshore.queues.submit:gpu-a10g

arguments:
  - name: invocation_id
    type: string
    required: true
    description: The failed invocation to triage.
  - name: max_log_lines
    type: integer
    default: 200
    description: How much of the tail to read before giving up on a quote.
---

# Run triage

You triage failed invocations on a Lakeshore control plane. You are given one
invocation id and read-only access to its logs, its events, and the RunConfig
it resolved against.

The invocation is `{{ invocation_id }}`. Read at most `{{ max_log_lines }}`
lines from the end of its log.

## Procedure

1. Read the terminal event first…
```

The body **is** the system prompt, exactly as in Claude Code. Everything above
the second `---` is configuration; everything below it is behaviour.

## Creating one

```
dreamlake agents create [name] [options]
```

Three ways in, and they compose — frontmatter supplies the defaults, flags win.

### Writing the prompt in the terminal

A prompt longer than one line wants a **quoted heredoc**. Quote the delimiter —
`<<'AGENT'`, not `<<AGENT` — or the shell expands `$VARIABLES` and backticks
inside your prose before the CLI ever sees it:

```bash
dreamlake agents create triage --file - <<'AGENT'
# Run triage

You triage failed invocations. Read the terminal event before the logs — it
says whether the worker died, the container exited, or the body raised, and
those three send you to different places.

Quote the single log line that decided your classification. One line.
AGENT
```

No frontmatter, and none needed. To open `$EDITOR` on a filled-in scaffold and
declare whatever you save:

```bash
dreamlake agents create run-triage --edit
```

To use a file you already have — the same file that works in `.claude/agents/`:

```bash
dreamlake agents create --file agents/run-triage.md
```

> **Note:** `cat agents/run-triage.md | dreamlake agents create` reads stdin without
>   `--file -`. The explicit form is worth typing anyway: it makes the command
>   read the same whether or not something is piped into it.

### Specifying properties

Every field in the frontmatter has a flag. Reach for the flags when you are
adding one property to an otherwise simple agent, and for the file when the
agent has enough configuration that a diff of it is worth reading.

| Property | Flag | Example |
| --- | --- | --- |
| Name | positional, or `--name` | `dreamlake agents create run-triage` |
| Description | `--description` | `--description 'Triages a failed invocation.'` |
| Tools | `--tools` (CSV) | `--tools Read,Grep,Glob,Bash` |
| Model | `--model` | `--model claude-opus-5` |
| Allowed | `--allow` (repeatable) | `--allow 'Read(./runs/**)'` |
| Ask first | `--ask` (repeatable) | `--ask 'Edit(./runconfigs/**)'` |
| Denied | `--deny` (repeatable) | `--deny 'Bash(rm:*)'` |
| Extra roots | `--add-dir` (repeatable) | `--add-dir ../runconfigs/` |
| Unmatched calls | `--permission-mode` | `--permission-mode acceptEdits` |
| Machine | `--run-config` | `--run-config gpu-a10g` |
| Queue | `--queue` | `--queue gpu-a10g` |
| Resource grants | `--grant` (repeatable) | `--grant dreamlake.nodes.read` |
| Prompt arguments | `--arg` (repeatable) | `--arg 'invocation_id:string!'` |
| Channel | `--channel` | `--channel http-sse:/v1/agents/triage/stream` |

Everything on one line:

```bash
dreamlake agents create run-triage \
  --description 'Triages a failed invocation and proposes the smallest fix.' \
  --tools Read,Grep,Bash \
  --model claude-opus-5 \
  --allow 'Read(./runs/**)' \
  --allow 'Bash(lakeshore logs:*)' \
  --deny  'Read(./.env)' \
  --deny  'Bash(rm:*)' \
  --run-config gpu-a10g --queue gpu-a10g \
  --arg 'invocation_id:string!' \
  --arg 'max_log_lines:integer=200' \
  --arg 'verdict_detail:enum(terse|normal|forensic)=normal' \
  --prompt 'Triage {{ invocation_id }} at {{ verdict_detail }} detail.
            Read at most {{ max_log_lines }} lines from the end of its log.'
```

> **Warning:** `Bash(npm run test:*)` contains a glob and parentheses; `enum(a|b|c)` contains
>   a pipe. Unquoted, the shell gets there first — `|` becomes a pipeline and `*`
>   becomes whatever is in the current directory. Single-quote every `--allow`,
>   `--ask`, `--deny` and `--arg`.

### The `--arg` mini-syntax

Designed so the common cases fit in one shell word:

| Written | Means |
| --- | --- |
| `invocation_id:string!` | required — `!` |
| `max_log_lines:integer=200` | defaulted — `=` |
| `verdict_detail:enum(terse\|normal\|forensic)=normal` | a fixed set, with a default |
| `dry_run:boolean=false` | a boolean default |
| `scene:string! "the scene id"` | a trailing description |

`!` and `=` are mutually exclusive: required plus a default are two statements
that contradict each other, and either reading silently disables the other.

### Naming

An agent's name is scoped to the **repository you are standing in**, so
`dreamlake agents create triage` inside `dreamlake-starter-kit` stores a `triage`
that cannot collide with the `triage` in another repo. Both the name and the
scope are printed before anything is sent:

```
agent    triage
  remote:      https://api.dreamlake.ai
  namespace:   chengdu
  scope:       dreamlake_starter_kit   (from repo /Users/…/dreamlake-starter-kit)
  prompt:      412 characters, from stdin
  tools:       (omitted — Claude reads that as inherit all)
  permissions: 0 allow / 0 ask / 0 deny   defaultMode=default
  run_config:  —   queue: —
  arguments:   0 declared, 0 substituted
```

With no name at all one is **generated** — from the description if there is one,
otherwise as `agent-<8 hex>`. The generated form is deliberately ugly so it
invites being renamed rather than being left in place forever. Override any of
it with `--name`, `--scope` or `--qualname`.

### Checking before writing

`--dry-run` resolves everything, runs every validation, prints the result and
sends nothing:

```bash
dreamlake agents create --file agents/run-triage.md --dry-run
```

That is the cheapest way to see what a placeholder typo or an unquoted rule
actually does, because the checks that matter are the ones that only fire once
the prompt and the spec are read together:

```
✗ the prompt uses {{ invocaton_id }} but declares no such argument.
✗ argument 'x' is required AND carries a default. One of those is a lie.
✗ permissions has unknown key(s) allowed — expected only allow, ask, deny, …
✗ frontmatter.auth.api_key looks like a credential.
```

## Fields taken from Claude, unchanged

| Field | Type | Meaning |
| --- | --- | --- |
| `name` | string | The agent's identifier. Kebab-case, unique in the namespace. |
| `description` | string | *When to use this agent* — written for whatever is choosing between agents, not for a human scanning a table. |
| `tools` | list \| omitted | Which tools the agent may call. **Omit to inherit everything available.** A list is exhaustive — there is no "all except". |
| `model` | string | The model to run it on. Claude Code's aliases (`opus`, `sonnet`, `haiku`, `inherit`) are accepted; so is a full model id such as `claude-opus-5`. |

`tools` names *capabilities*; it is not a permission list. `tools: [Bash]` says
the agent may call Bash at all; `permissions` says which commands. Both are
required for a Bash call to happen, and they are validated independently.

## Permissions — what it may touch

Lifted verbatim from Claude Code's `settings.json` `permissions` block, so the
rule strings are the ones you already write.

| Key | Meaning |
| --- | --- |
| `allow` | Rules that run without asking. |
| `ask` | Rules that pause for confirmation. |
| `deny` | Rules that are refused outright. |
| `additionalDirectories` | Extra roots the agent may read and write outside its workdir. |
| `defaultMode` | `default` \| `acceptEdits` \| `plan` \| `bypassPermissions` — what happens to a call no rule matches. |

**Precedence is `deny` > `ask` > `allow`**, and it is not overridable. A path in
both `deny` and `additionalDirectories` is denied; that is the point of having
`deny`.

### Rule syntax

A rule is `Tool` or `Tool(specifier)`. A bare tool name matches every use of it.

| Rule | Matches |
| --- | --- |
| `Bash(npm run test:*)` | Any Bash command whose text starts with `npm run test`. `:*` is a **prefix** wildcard, not a glob. |
| `Read(./data/**)` | Reads under `data/`, relative to the agent file. |
| `Edit(src/**/*.ts)` | Edits to TypeScript under `src/`. Gitignore-style patterns. |
| `Read(//var/log/**)` | An absolute path — note the leading `//`. |
| `Read(~/.config/**)` | A path under the invoking user's home. |
| `WebFetch(domain:docs.dreamlake.ai)` | Fetches of that host only. |
| `Read` | Every read, anywhere the agent can reach. |

Three traps worth naming, because each of them is a rule that looks like it
works and does not:

- **`Bash` rules match the command string, not the effect.** `Bash(git:*)`
  does not stop `sh -c "git push"`, and it never will. Treat Bash rules as a
  convenience over an already-trusted command set, not as a sandbox boundary.
- **A `deny` on a path does not stop a tool that reaches it another way.**
  `Read(./.env)` denies the Read tool; `Bash(cat:*)` still prints the file.
  Deny the *route*, not only the destination.
- **Relative paths resolve against the agent file**, not the invocation's
  workdir. Under a `run_config` that mounts a repo somewhere else, an
  unanchored `Read(src/**)` is not the path you pictured. Use
  `additionalDirectories` and absolute rules when the layout is set by a mount.

### Tool rules vs. resource grants

Two lists, deliberately, because they answer different questions:

| | Governs | Vocabulary |
| --- | --- | --- |
| `permissions` | Tools and files — what the *process* may do | `Tool(specifier)`, from Claude |
| `grants` | DreamLake and Lakeshore resources — what the *identity* may reach | `domain.resource.verb[:scope]`, IAM-style |

`Read(./data/**)` says nothing about whether the agent may publish a dataset
version; `dreamlake.datasets.release` says nothing about whether it may run
`rm`. The grant registry, its closed verb set, and the scope rules are on
[Agent permissions](/workflows/agent-permissions.md).

## Attaching a run configuration

The first of the two extensions, and the one Claude's spec has no opinion about
because Claude Code runs on the machine you are already sitting at.

**It is optional.** A prompt-only agent needs no machine and most do not have
one. Attach a RunConfig when the agent has to run somewhere in particular — a
GPU, a specific image, a prepared host — and leave it off otherwise.

| Field | Meaning |
| --- | --- |
| `run_config` | The name of a [RunConfig](/lakeshore/registry.md) — image, resources, `host_setup`, `run_setup`, `env`, `timeout_s`, mounts, provider. |
| `queue` | Which queue to submit on. The queue is the operator's surface; the RunConfig is the author's. |

`run_config` **names an existing RunConfig; it does not copy one.** That is the
point of the indirection: the machine can be re-tuned, or shared with a UDF that
needs the same host, without touching the agent. The RunConfig is a full
declaration and is documented on its own; the two things worth knowing here are:

- **Placement is derived, not declared.** The RunConfig's `host_key` is a hash
  of the subset that affects placement (`runner`, `image`, `resources`,
  [`host_setup`](/lakeshore/host-setup.md), `provider`, `extras`). Two agents with
  the same key can share a warm worker; two that differ in `image` never will,
  however similar they look.
- **The lifetime mismatch is the open design problem.** A RunConfig describes a
  *host for an invocation*. An agent is meant to outlive individual
  invocations, which raises questions a RunConfig does not answer — how long the
  process is kept alive, what happens to its workdir between calls, what
  restarts it. Those are tracked on the lakeshore side under
  [Agent lifecycle](/dev/agent-lifecycle), and
  none of it is built.

> **Note:** Nothing in an agent file says how many machines to keep warm or which instance
>   types are permitted. That is queue-level fleet policy and it belongs to the
>   operator. `run_config.extras.instance_type` is a *request*.

## Arguments — values substituted into the prompt

The second extension. An agent's prompt is usually a procedure with holes in it:
*which* invocation, *which* dataset, *how many* lines. Declaring those holes is
better than string-formatting the prompt at the call site, for a reason that is
about auditability rather than convenience — see
[Why arguments and not string formatting](#why-arguments-and-not-string-formatting).

### Declaring

```yaml
arguments:
  - name: invocation_id
    type: string
    required: true
    description: The failed invocation to triage.
  - name: max_log_lines
    type: integer
    default: 200
    description: How much of the tail to read before giving up on a quote.
  - name: strict
    type: boolean
    default: false
```

| Key | Required | Meaning |
| --- | --- | --- |
| `name` | yes | The placeholder's name. Must be a valid identifier. |
| `type` | yes | `string` \| `integer` \| `number` \| `boolean` \| `enum`. |
| `required` | no | Defaults to `false`. `required: true` and `default:` together is an error — one of them is a lie. |
| `default` | no | Used when the caller omits the argument. Must typecheck. |
| `values` | for `enum` | The permitted values. |
| `description` | no | Documentation. Not sent to the model. |

### Substituting

A `{{ name }}` placeholder in the body is replaced with the argument's value.
Whitespace inside the braces is ignored, so `{{name}}` and `{{ name }}` are the
same placeholder.

```markdown
The invocation is `{{ invocation_id }}`. Read at most `{{ max_log_lines }}`
lines from the end of its log.
```

dispatched with `{ "invocation_id": "01KZCCC93W…" }` renders as:

```markdown
The invocation is `01KZCCC93W…`. Read at most `200` lines from the end of its
log.
```

The rules, in the order they bite:

1. **A placeholder with no matching `arguments` entry is an error at declare
   time**, not a blank at dispatch time. A typo'd `{{ invocaton_id }}` fails the
   declaration.
2. **A declared argument that appears in no placeholder is a warning**, not an
   error — an argument may legitimately be consumed by a tool rather than the
   prompt.
3. **A missing required argument is an error at dispatch**, before the agent
   starts. It is never rendered as an empty string.
4. **Values are inserted as text, never as instructions.** Substitution happens
   once, on the literal placeholder; a value containing `{{ … }}` is not
   re-scanned. This is not a security boundary — an argument that reaches the
   prompt is still untrusted input the model reads — but it does mean an
   argument cannot manufacture *new* placeholders.
5. **Placeholders are body-only.** `{{ … }}` in frontmatter is not substituted.
   An agent does not get to pick its own `run_config` at dispatch.

### Claude compatibility

Claude Code's slash-command argument forms are accepted and lower onto the same
list:

| Written | Equivalent |
| --- | --- |
| `$1` … `$9` | Positional arguments — `arguments[0]` … `arguments[8]` by declaration order. |
| `$ARGUMENTS` | Every positional argument, joined with a single space. |

Named `{{ … }}` is preferred for anything declared, because `$3` does not
survive someone reordering the `arguments:` list. The positional forms exist so
a command file pasted from `.claude/commands/` keeps working.

### Why arguments and not string formatting

You could format the prompt yourself and send the finished text. The reason not
to is that **`markdown` lives on the mutable head and has no version history.**

Pin an agent to a 64-character version and you have pinned the *code* it runs.
Its *instructions* can be rewritten by anyone with write access, at any time,
and every pinned caller silently gets the new text. That is a deliberate storage
decision — folding prose into `sha256(source)` would fork the identity of every
function — but it means a run is not reconstructable from the declaration alone.

Declared arguments are the seam that fixes this: because the platform does the
substitution, it can **record the rendered prompt on the invocation**. The
template can drift afterwards and the run still says what it actually asked.
Hand-formatted prompts have no such record, so the audit trail is only as good
as whatever your call site happened to log.

## What is stored today

Field by field, because the gap between "specified" and "stored" is the whole
risk of reading this page.

| Field | Stored | Validated | Read by anything |
| --- | --- | --- | --- |
| `name` → `module` + `qualname` | ✅ column | ✅ identity is `(namespace, "<module>.<qualname>")` | ✅ |
| `description` | ✅ column | — | ✅ list + search |
| body → `markdown` | ✅ column | — | ✅ rendered in the dashboard |
| `channel` | ✅ column | ✅ shape `{ kind, target?, config? }` **and a full-depth credential walk** | ❌ nothing attaches |
| `kind: agent` | ✅ column | ✅ `markdown`/`channel` on a non-agent is a 400 | ✅ `?kind=agent` filter |
| `tools` | ⚠️ `metadata` only | ❌ | ❌ |
| `permissions` | ⚠️ `metadata` only | ❌ | ❌ |
| `grants` | ⚠️ `metadata` only | ✅ *in workflow `uda` nodes* — not on runnables | ❌ enforcement ships with the engine |
| `run_config` / `queue` | ⚠️ `metadata` only | ❌ as an agent field. The RunConfig **row** is real, and its `host_key` **is** enforced at claim time | partially |
| `arguments` | ⚠️ `metadata` only | ❌ | ❌ no renderer server-side |
| lifecycle | ❌ | ❌ | ❌ undesigned |

`metadata` is a free-form JSON object on the runnable — it round-trips, so a
declaration can carry the full spec today and a client can read it back and act
on it. Two consequences to be clear-eyed about:

- **`metadata` is not walked for credentials.** Only `channel` is. A token in
  `metadata` is stored in plaintext and returned to anyone with read access to
  the namespace. There are deliberately no secret columns on `Runnable`;
  credentials belong on a `Source` or a `LakeshoreProvider` and are referenced
  by name.
- **`metadata` is shallow-merged on update.** Patching `metadata.tools` replaces
  the whole `tools` key; it does not merge into it.

## Not the same as a workflow `uda` node

DreamLake workflows have their own agent concept — **`uda`** nodes: a typed
member of a static workflow graph, with `instructions`, `model`, `tools` and
`permissions` inline. That is a *node in a graph*, not a long-lived worker, and
it is documented at [Workflow node types](/workflows/node-types.md) and
[Agent permissions](/workflows/agent-permissions.md).

The two now share a vocabulary — `tools` is a tool list in both, `permissions`
means grants in the node and tool rules here — which is exactly the collision
worth flagging rather than leaving for someone to trip over. They are separate
declarations today and converging them is future work, not a shipped thing.

## Try it

The one-liner, against whatever you are logged into:

```bash
dreamlake agents create triage --prompt 'You triage failed runs.' --dry-run
dreamlake agents create triage --prompt 'You triage failed runs.'
dreamlake runnable show <the name it printed>
```

`02_agent` in the [starter kit](/lakeshore/quickstart.md) goes further — it declares
the fully-loaded agent against a real server, reads it back, renders its prompt
with real arguments, and earns a real 400 from the credential walk:

```bash
cd dreamlake-starter-kit/02_agent
make check    # offline — validates the file and the argument bindings
make render   # offline — prints the rendered prompt for a set of arguments
make live     # LIVE-REG — declares it, reads it back, renders it server-side-of-the-wire
```

## Execution overview

See [Agent execution overview](/lakeshore/agents/overview.md) for RunConfig
placement, prompt arguments, and the open lifecycle questions.
