Workflows
A workflow is a typed graph of data production: ordered stages whose member nodes do the work — deterministic compute (UDFs), judgment-making agents, statistical samplers, and control flow — joined by typed data edges. You push it from the terminal and it renders as a live canvas.
This page is about the WorkflowSpec — the typed stage/node graph you push
as JSON and run on a canvas. Claude Code's JS orchestration scripts and Python
pipelines are also called workflows, share the Workflow server model, and
answer to the same CLI. If that is what you came for, start at
CJS Workflows vs Python Pipelines.
How you use it
One round, from a sentence to a published dataset. Skills live in dreamlake-skills.
-
Install the CLI and the three skills.
Claude Code is where you talk to it — install that first.
$ curl -fsSL https://dl.dreamlake.ai/install.sh | sh$ dreamlake login --env prod$ git clone https://github.com/dreamlake-ai/dreamlake-skills.git$ mkdir -p ~/.claude/skills$ cp -r dreamlake-skills/{video-labeling-workflow,\remote-source-check,workflow-publish} ~/.claude/skills/all three — they call each other by namerestart Claude Code to load themSkill Does video-labeling-workflowfills the template, drives the whole round remote-source-checkproves your source really holds those paths workflow-publishvalidates and pushes, and asks which namespace Install all three — the first invokes the other two by name.
--env prodis worth typing. Without itlogintargets whichever environment is active on that machine, so anyone who has ever logged into staging lands there silently — and the failures that follow (an emptysource list, a push to the wrong place) never point back at the cause.dreamlake auth env listshows which login environment is active.Install the standalone CLI from the CLI guide and check
dreamlake workflow --help. The Pythondreamlakepackage provides the SDK; it is not the supported CLI installation path. -
Say what you want — it asks for the data, verifies it, shows you what it found.
claude code · video-labeling-workflowcollect → check → confirmCreate a video labeling workflowFour things:
1 · source name — a connected S3 / Dropbox / HF store2 · video path inside it3 · reference annotation path (.json, scored against)4 · task description — one line, e.g. mount a wall shelffootage · clips/713488-assembly-shelf.mp4 · gold/713488_annotation.json · mount a wall shelfChecked. Confirm before I create it:
source footage (namespace acme)video clips/713488-assembly-shelf.mp4 · 287 MBreference gold/713488_annotation.json · 64 phasestask mount a wall shelfpublish to acmeConfirmed.the check is not a formality — a path that lists but will not fetch, or a swapped video/annotation pair, costs one line here and a failed run later -
It pushes — validated against the schema and the graph rules first.
$ dreamlake workflow push video-labeling-0807-1501.workflow.json \--namespace acmeworkflow: video-labeling-0807-1501 v1 · 5 stages · 9 nodes · 10 edges✓ pushed video-labeling-0807-1501 v1open: dreamlake.ai/acme/workflows/video-labeling-0807-1501 -
Review on DreamLake — iterate, hot-reload.
-
Run it — nodes light up, output streams.
queuedRuntask on the queue→runningworker claims itnodes light up, logs stream→awaitinggatewaits for a person — no cost→completeddatasetpublished, versioneda run's four states — the gate is the only one that waits on a human -
Review at the gate, then approve.
The run publishes its dataset and stops. Select the gate node: ↗ Review dataset opens what was produced, ✓ Approve & continue completes the run. The link stays after approval — it is the way back to what was signed off.
What a spec is made of
| Part | What it is |
|---|---|
| Stages | phases that group and order — not barriers; execution follows the edges |
| Nodes | the members of a stage, from four families |
| Edges | typed connections between ports (samples, dataset, metrics, …); types must match end to end, and a collect port fans several streams into one input |
| Family | Does | Runs today |
|---|---|---|
compute | deterministic work — filter, transcode, train, publish (any UDF) | ✅ |
uda | agent judgment — labeling, review, curation | ✅ |
control | human approval gates | ✅ approval only |
sampler | statistical subset selection | ⏳ spec-only |
The schema also accepts condition, switch, loop and foreach control
types. They pass validation, then stop a run with not supported yet —
vocabulary to design against, not behaviour to rely on.
A uda node declares the tools and permissions its agent may use, so
judgment never runs on raw credentials. Intermediate results stay in the
worker's run directory; what reaches DreamLake is what a publish node
writes.
Version it
Pushes never overwrite. Each one appends an immutable version, and the spec picker on the workflow page pins the canvas to any of them — pick the latest to resume following pushes.
Editing a node's settings on the canvas and pressing Apply does the same thing: the edited spec is validated and appended as the next version. Nothing is edited in place, so a run always names the version it executed.
Next steps
The spec graph, Claude Code's JS orchestration scripts, and Python pipelines side by side — and the seven places where one of them does not exist yet.
The full vocabulary — every family's inputs, outputs, and configuration, rendered with the real components.
The grant registry uda nodes draw from — domains, verbs, and scopes.
dreamlake workflow push and list — flags, namespaces, versioning.