DreamLake

Agents

An agent is a name and a prompt. Everything else is optional and layered on when the agent earns it: Claude's tools, model and permissions when it should be constrained, an attached RunConfig when it needs a machine, and arguments when its prompt has holes in it.

bash
dreamlake agents create triage --prompt 'You triage failed runs. Read the
terminal event first, classify the failure, quote the line that decided it.'

That is a complete agent. No tools list, no permission block, no run config, no arguments — and it is the normal case, not a degraded one.

Status. The declaration is real: DreamLake stores, validates and serves it. Execution is not. There is no agent API in the SDK and no agent route on the control plane. The What is stored today table says field-by-field which half you are in.

Nothing dispatches an agent yet

Everything that currently runs is a UDF — invoked, runs to completion, returns. If you need remote compute now, that is the surface. This page declares an agent; nothing on it drives one.

An agent is not a UDF

They share a storage table and nothing else worth knowing, and the difference shows up first in how they are named. A UDF is addressed as <module>.<qualname> because it is code, and that is where code lives. An agent has a name, the way a person does — triage, run-triage, colmap-driver — scoped to the repository it was created in, so two repos can each hold a triage without colliding.

Nothing about authoring an agent asks you to think in module paths. The CLI derives the storage segments and prints them once; after that the name is the name.

UDFAgent
Identity<module>.<qualname> — a code patha name, scoped to a repo
Bodya Python functiona markdown prompt
A RunConfig isrequired to place itoptional, attached when it needs one
Versionedyes — version is sha256(source)the prompt is not; see below

Why Claude's spec

The alternative was to invent a fifth agent format. Claude Code, the Claude Agent SDK, and Claude Managed Agents already agree on one — name, description, tools, model, a system prompt in the body, and an allow/ask/deny permission block — and an agent author almost certainly already writes it. Adopting it means an agent file is portable in both directions: the same file drives a local .claude/agents/ subagent and a declared platform agent, and the fields a reader has to learn are the two that are genuinely ours.

So the rule for this page: where Claude's spec has an answer, it is the answer. The Lakeshore extensions are marked as extensions everywhere they appear, and they are additive — a plain Claude agent file is a valid declaration.

The agent file

The prompt can be a bare markdown file with no frontmatter at all. Add frontmatter when you want to constrain the agent or give it a machine — this is the fully-loaded version, and every field in it is optional:

markdown
---
# ── Claude's agent spec — unchanged ────────────────────────────────
name: run-triage
description: >-
  Triages failed invocations: reads a run's logs and events, classifies the
  failure, and proposes the smallest RunConfig change that would fix it.
tools: [Read, Grep, Glob, Bash]
model: claude-opus-5

permissions:
  allow:
    - Read(./runs/**)
    - Bash(lakeshore logs:*)
    - Bash(lakeshore events:*)
    - WebFetch(domain:docs.dreamlake.ai)
  ask:
    - Edit(./runconfigs/**)
  deny:
    - Read(./.env)
    - Read(./secrets/**)
    - Bash(curl:*)
  additionalDirectories: ["../runconfigs/"]
  defaultMode: default

# ── Lakeshore extensions ───────────────────────────────────────────
run_config: gpu-a10g
queue: gpu-a10g
channel:
  kind: http-sse
  target: /v1/agents/triage/stream
  config: { heartbeatSeconds: 15, maxIdleSeconds: 900 }

grants:
  - dreamlake.nodes.read
  - lakeshore.queues.submit:gpu-a10g

arguments:
  - name: invocation_id
    type: string
    required: true
    description: The failed invocation to triage.
  - name: max_log_lines
    type: integer
    default: 200
    description: How much of the tail to read before giving up on a quote.
---

# Run triage

You triage failed invocations on a Lakeshore control plane. You are given one
invocation id and read-only access to its logs, its events, and the RunConfig
it resolved against.

The invocation is `{{ invocation_id }}`. Read at most `{{ max_log_lines }}`
lines from the end of its log.

## Procedure

1. Read the terminal event first…

The body is the system prompt, exactly as in Claude Code. Everything above the second --- is configuration; everything below it is behaviour.

Creating one

dreamlake agents create [name] [options]

Three ways in, and they compose — frontmatter supplies the defaults, flags win.

Writing the prompt in the terminal

A prompt longer than one line wants a quoted heredoc. Quote the delimiter — <<'AGENT', not <<AGENT — or the shell expands $VARIABLES and backticks inside your prose before the CLI ever sees it:

bash
dreamlake agents create triage --file - <<'AGENT'
# Run triage

You triage failed invocations. Read the terminal event before the logs — it
says whether the worker died, the container exited, or the body raised, and
those three send you to different places.

Quote the single log line that decided your classification. One line.
AGENT

No frontmatter, and none needed. To open $EDITOR on a filled-in scaffold and declare whatever you save:

bash
dreamlake agents create run-triage --edit

To use a file you already have — the same file that works in .claude/agents/:

bash
dreamlake agents create --file agents/run-triage.md
A pipe works too

cat agents/run-triage.md | dreamlake agents create reads stdin without --file -. The explicit form is worth typing anyway: it makes the command read the same whether or not something is piped into it.

Specifying properties

Every field in the frontmatter has a flag. Reach for the flags when you are adding one property to an otherwise simple agent, and for the file when the agent has enough configuration that a diff of it is worth reading.

PropertyFlagExample
Namepositional, or --namedreamlake agents create run-triage
Description--description--description 'Triages a failed invocation.'
Tools--tools (CSV)--tools Read,Grep,Glob,Bash
Model--model--model claude-opus-5
Allowed--allow (repeatable)--allow 'Read(./runs/**)'
Ask first--ask (repeatable)--ask 'Edit(./runconfigs/**)'
Denied--deny (repeatable)--deny 'Bash(rm:*)'
Extra roots--add-dir (repeatable)--add-dir ../runconfigs/
Unmatched calls--permission-mode--permission-mode acceptEdits
Machine--run-config--run-config gpu-a10g
Queue--queue--queue gpu-a10g
Resource grants--grant (repeatable)--grant dreamlake.nodes.read
Prompt arguments--arg (repeatable)--arg 'invocation_id:string!'
Channel--channel--channel http-sse:/v1/agents/triage/stream

Everything on one line:

bash
dreamlake agents create run-triage \
  --description 'Triages a failed invocation and proposes the smallest fix.' \
  --tools Read,Grep,Bash \
  --model claude-opus-5 \
  --allow 'Read(./runs/**)' \
  --allow 'Bash(lakeshore logs:*)' \
  --deny  'Read(./.env)' \
  --deny  'Bash(rm:*)' \
  --run-config gpu-a10g --queue gpu-a10g \
  --arg 'invocation_id:string!' \
  --arg 'max_log_lines:integer=200' \
  --arg 'verdict_detail:enum(terse|normal|forensic)=normal' \
  --prompt 'Triage {{ invocation_id }} at {{ verdict_detail }} detail.
            Read at most {{ max_log_lines }} lines from the end of its log.'
Quote the rules and the args

Bash(npm run test:*) contains a glob and parentheses; enum(a|b|c) contains a pipe. Unquoted, the shell gets there first — | becomes a pipeline and * becomes whatever is in the current directory. Single-quote every --allow, --ask, --deny and --arg.

The --arg mini-syntax

Designed so the common cases fit in one shell word:

WrittenMeans
invocation_id:string!required — !
max_log_lines:integer=200defaulted — =
verdict_detail:enum(terse|normal|forensic)=normala fixed set, with a default
dry_run:boolean=falsea boolean default
scene:string! "the scene id"a trailing description

! and = are mutually exclusive: required plus a default are two statements that contradict each other, and either reading silently disables the other.

Naming

An agent's name is scoped to the repository you are standing in, so dreamlake agents create triage inside dreamlake-starter-kit stores a triage that cannot collide with the triage in another repo. Both the name and the scope are printed before anything is sent:

agent    triage
  remote:      https://api.dreamlake.ai
  namespace:   chengdu
  scope:       dreamlake_starter_kit   (from repo /Users/…/dreamlake-starter-kit)
  prompt:      412 characters, from stdin
  tools:       (omitted — Claude reads that as inherit all)
  permissions: 0 allow / 0 ask / 0 deny   defaultMode=default
  run_config:  —   queue: —
  arguments:   0 declared, 0 substituted

With no name at all one is generated — from the description if there is one, otherwise as agent-<8 hex>. The generated form is deliberately ugly so it invites being renamed rather than being left in place forever. Override any of it with --name, --scope or --qualname.

Checking before writing

--dry-run resolves everything, runs every validation, prints the result and sends nothing:

bash
dreamlake agents create --file agents/run-triage.md --dry-run

That is the cheapest way to see what a placeholder typo or an unquoted rule actually does, because the checks that matter are the ones that only fire once the prompt and the spec are read together:

✗ the prompt uses {{ invocaton_id }} but declares no such argument.
✗ argument 'x' is required AND carries a default. One of those is a lie.
✗ permissions has unknown key(s) allowed — expected only allow, ask, deny, …
✗ frontmatter.auth.api_key looks like a credential.

Fields taken from Claude, unchanged

FieldTypeMeaning
namestringThe agent's identifier. Kebab-case, unique in the namespace.
descriptionstringWhen to use this agent — written for whatever is choosing between agents, not for a human scanning a table.
toolslist | omittedWhich tools the agent may call. Omit to inherit everything available. A list is exhaustive — there is no "all except".
modelstringThe model to run it on. Claude Code's aliases (opus, sonnet, haiku, inherit) are accepted; so is a full model id such as claude-opus-5.

tools names capabilities; it is not a permission list. tools: [Bash] says the agent may call Bash at all; permissions says which commands. Both are required for a Bash call to happen, and they are validated independently.

Permissions — what it may touch

Lifted verbatim from Claude Code's settings.json permissions block, so the rule strings are the ones you already write.

KeyMeaning
allowRules that run without asking.
askRules that pause for confirmation.
denyRules that are refused outright.
additionalDirectoriesExtra roots the agent may read and write outside its workdir.
defaultModedefault | acceptEdits | plan | bypassPermissions — what happens to a call no rule matches.

Precedence is deny > ask > allow, and it is not overridable. A path in both deny and additionalDirectories is denied; that is the point of having deny.

Rule syntax

A rule is Tool or Tool(specifier). A bare tool name matches every use of it.

RuleMatches
Bash(npm run test:*)Any Bash command whose text starts with npm run test. :* is a prefix wildcard, not a glob.
Read(./data/**)Reads under data/, relative to the agent file.
Edit(src/**/*.ts)Edits to TypeScript under src/. Gitignore-style patterns.
Read(//var/log/**)An absolute path — note the leading //.
Read(~/.config/**)A path under the invoking user's home.
WebFetch(domain:docs.dreamlake.ai)Fetches of that host only.
ReadEvery read, anywhere the agent can reach.

Three traps worth naming, because each of them is a rule that looks like it works and does not:

  • Bash rules match the command string, not the effect. Bash(git:*) does not stop sh -c "git push", and it never will. Treat Bash rules as a convenience over an already-trusted command set, not as a sandbox boundary.
  • A deny on a path does not stop a tool that reaches it another way. Read(./.env) denies the Read tool; Bash(cat:*) still prints the file. Deny the route, not only the destination.
  • Relative paths resolve against the agent file, not the invocation's workdir. Under a run_config that mounts a repo somewhere else, an unanchored Read(src/**) is not the path you pictured. Use additionalDirectories and absolute rules when the layout is set by a mount.

Tool rules vs. resource grants

Two lists, deliberately, because they answer different questions:

GovernsVocabulary
permissionsTools and files — what the process may doTool(specifier), from Claude
grantsDreamLake and Lakeshore resources — what the identity may reachdomain.resource.verb[:scope], IAM-style

Read(./data/**) says nothing about whether the agent may publish a dataset version; dreamlake.datasets.release says nothing about whether it may run rm. The grant registry, its closed verb set, and the scope rules are on Agent permissions.

Attaching a run configuration

The first of the two extensions, and the one Claude's spec has no opinion about because Claude Code runs on the machine you are already sitting at.

It is optional. A prompt-only agent needs no machine and most do not have one. Attach a RunConfig when the agent has to run somewhere in particular — a GPU, a specific image, a prepared host — and leave it off otherwise.

FieldMeaning
run_configThe name of a RunConfig — image, resources, host_setup, run_setup, env, timeout_s, mounts, provider.
queueWhich queue to submit on. The queue is the operator's surface; the RunConfig is the author's.

run_config names an existing RunConfig; it does not copy one. That is the point of the indirection: the machine can be re-tuned, or shared with a UDF that needs the same host, without touching the agent. The RunConfig is a full declaration and is documented on its own; the two things worth knowing here are:

  • Placement is derived, not declared. The RunConfig's host_key is a hash of the subset that affects placement (runner, image, resources, host_setup, provider, extras). Two agents with the same key can share a warm worker; two that differ in image never will, however similar they look.
  • The lifetime mismatch is the open design problem. A RunConfig describes a host for an invocation. An agent is meant to outlive individual invocations, which raises questions a RunConfig does not answer — how long the process is kept alive, what happens to its workdir between calls, what restarts it. Those are tracked on the lakeshore side under Agent lifecycle, and none of it is built.
Fleet policy is not in here

Nothing in an agent file says how many machines to keep warm or which instance types are permitted. That is queue-level fleet policy and it belongs to the operator. run_config.extras.instance_type is a request.

Arguments — values substituted into the prompt

The second extension. An agent's prompt is usually a procedure with holes in it: which invocation, which dataset, how many lines. Declaring those holes is better than string-formatting the prompt at the call site, for a reason that is about auditability rather than convenience — see Why arguments and not string formatting.

Declaring

yaml
arguments:
  - name: invocation_id
    type: string
    required: true
    description: The failed invocation to triage.
  - name: max_log_lines
    type: integer
    default: 200
    description: How much of the tail to read before giving up on a quote.
  - name: strict
    type: boolean
    default: false
KeyRequiredMeaning
nameyesThe placeholder's name. Must be a valid identifier.
typeyesstring | integer | number | boolean | enum.
requirednoDefaults to false. required: true and default: together is an error — one of them is a lie.
defaultnoUsed when the caller omits the argument. Must typecheck.
valuesfor enumThe permitted values.
descriptionnoDocumentation. Not sent to the model.

Substituting

A {{ name }} placeholder in the body is replaced with the argument's value. Whitespace inside the braces is ignored, so {{name}} and {{ name }} are the same placeholder.

markdown
The invocation is `{{ invocation_id }}`. Read at most `{{ max_log_lines }}`
lines from the end of its log.

dispatched with { "invocation_id": "01KZCCC93W…" } renders as:

markdown
The invocation is `01KZCCC93W…`. Read at most `200` lines from the end of its
log.

The rules, in the order they bite:

  1. A placeholder with no matching arguments entry is an error at declare time, not a blank at dispatch time. A typo'd {{ invocaton_id }} fails the declaration.
  2. A declared argument that appears in no placeholder is a warning, not an error — an argument may legitimately be consumed by a tool rather than the prompt.
  3. A missing required argument is an error at dispatch, before the agent starts. It is never rendered as an empty string.
  4. Values are inserted as text, never as instructions. Substitution happens once, on the literal placeholder; a value containing {{ … }} is not re-scanned. This is not a security boundary — an argument that reaches the prompt is still untrusted input the model reads — but it does mean an argument cannot manufacture new placeholders.
  5. Placeholders are body-only. {{ … }} in frontmatter is not substituted. An agent does not get to pick its own run_config at dispatch.

Claude compatibility

Claude Code's slash-command argument forms are accepted and lower onto the same list:

WrittenEquivalent
$1 … $9Positional arguments — arguments[0] … arguments[8] by declaration order.
$ARGUMENTSEvery positional argument, joined with a single space.

Named {{ … }} is preferred for anything declared, because $3 does not survive someone reordering the arguments: list. The positional forms exist so a command file pasted from .claude/commands/ keeps working.

Why arguments and not string formatting

You could format the prompt yourself and send the finished text. The reason not to is that markdown lives on the mutable head and has no version history.

Pin an agent to a 64-character version and you have pinned the code it runs. Its instructions can be rewritten by anyone with write access, at any time, and every pinned caller silently gets the new text. That is a deliberate storage decision — folding prose into sha256(source) would fork the identity of every function — but it means a run is not reconstructable from the declaration alone.

Declared arguments are the seam that fixes this: because the platform does the substitution, it can record the rendered prompt on the invocation. The template can drift afterwards and the run still says what it actually asked. Hand-formatted prompts have no such record, so the audit trail is only as good as whatever your call site happened to log.

What is stored today

Field by field, because the gap between "specified" and "stored" is the whole risk of reading this page.

FieldStoredValidatedRead by anything
name → module + qualname✅ column✅ identity is (namespace, "<module>.<qualname>")✅
description✅ column—✅ list + search
body → markdown✅ column—✅ rendered in the dashboard
channel✅ column✅ shape { kind, target?, config? } and a full-depth credential walk❌ nothing attaches
kind: agent✅ column✅ markdown/channel on a non-agent is a 400✅ ?kind=agent filter
tools⚠️ metadata only❌❌
permissions⚠️ metadata only❌❌
grants⚠️ metadata only✅ in workflow uda nodes — not on runnables❌ enforcement ships with the engine
run_config / queue⚠️ metadata only❌ as an agent field. The RunConfig row is real, and its host_key is enforced at claim timepartially
arguments⚠️ metadata only❌❌ no renderer server-side
lifecycle❌❌❌ undesigned

metadata is a free-form JSON object on the runnable — it round-trips, so a declaration can carry the full spec today and a client can read it back and act on it. Two consequences to be clear-eyed about:

  • metadata is not walked for credentials. Only channel is. A token in metadata is stored in plaintext and returned to anyone with read access to the namespace. There are deliberately no secret columns on Runnable; credentials belong on a Source or a LakeshoreProvider and are referenced by name.
  • metadata is shallow-merged on update. Patching metadata.tools replaces the whole tools key; it does not merge into it.

Not the same as a workflow uda node

DreamLake workflows have their own agent concept — uda nodes: a typed member of a static workflow graph, with instructions, model, tools and permissions inline. That is a node in a graph, not a long-lived worker, and it is documented at Workflow node types and Agent permissions.

The two now share a vocabulary — tools is a tool list in both, permissions means grants in the node and tool rules here — which is exactly the collision worth flagging rather than leaving for someone to trip over. They are separate declarations today and converging them is future work, not a shipped thing.

Try it

The one-liner, against whatever you are logged into:

bash
dreamlake agents create triage --prompt 'You triage failed runs.' --dry-run
dreamlake agents create triage --prompt 'You triage failed runs.'
dreamlake runnable show <the name it printed>

02_agent in the starter kit goes further — it declares the fully-loaded agent against a real server, reads it back, renders its prompt with real arguments, and earns a real 400 from the credential walk:

bash
cd dreamlake-starter-kit/02_agent
make check    # offline — validates the file and the argument bindings
make render   # offline — prints the rendered prompt for a set of arguments
make live     # LIVE-REG — declares it, reads it back, renders it server-side-of-the-wire

Execution overview

See Agent execution overview for RunConfig placement, prompt arguments, and the open lifecycle questions.