DreamLake

Workflows

A workflow is a typed graph of data production: ordered stages whose member nodes do the work — deterministic compute (UDFs), judgment-making agents, statistical samplers, and control flow — joined by typed data edges. You push it from the terminal and it renders as a live canvas.

Prepare
stage · 2 members2 done
Annotate
stage · 1 member1 done
Hand pose
stage · 2 members2 done
Evaluate
stage · 2 members2 done
Publish
stage · 2 members1 done
video_source
compute · udf
direct
video_to_lerobot
compute · udf
daemon
subtask_labeler
uda · agent
2 permsgemini-3.5-flashgemini-api
video_frames
compute · udf
direct
hand_pose
compute · udf
EC2 · g5.xlargedaemon
gold_source
compute · udf
direct
subtask_metrics
compute · udf
direct
publish_dataset
compute · udf
direct
publish_gate
control · approvalOpen the published dataset and review it. Approving marks the run complete.
videosvideosdatasetdatasetlabeledvideoframesframeskeypointscomparisongoldlabeledgoldmetricslabeledkeypointsvideodatasetdatasetsamplesdirectoryfiledatasetdatasetsamplesdatasetfiledatasetsamples
video-labeling-0807-1501, replayed — the worker walks the graph in the engine's own order (edges, not stages), then parks on the gate until a person approves. Scroll sideways for the rest.
Three different things are called a workflow

This page is about the WorkflowSpec — the typed stage/node graph you push as JSON and run on a canvas. Claude Code's JS orchestration scripts and Python pipelines are also called workflows, share the Workflow server model, and answer to the same CLI. If that is what you came for, start at CJS Workflows vs Python Pipelines.

How you use it

One round, from a sentence to a published dataset. Skills live in dreamlake-skills.

  1. Install the CLI and the three skills.

    Claude Code is where you talk to it — install that first.

    $ curl -fsSL https://dl.dreamlake.ai/install.sh | sh
    $ dreamlake login --env prod
    $ git clone https://github.com/dreamlake-ai/dreamlake-skills.git
    $ mkdir -p ~/.claude/skills
    $ cp -r dreamlake-skills/{video-labeling-workflow,\
    remote-source-check,workflow-publish} ~/.claude/skills/
    all three — they call each other by name
    restart Claude Code to load them
    SkillDoes
    video-labeling-workflowfills the template, drives the whole round
    remote-source-checkproves your source really holds those paths
    workflow-publishvalidates and pushes, and asks which namespace

    Install all three — the first invokes the other two by name.

    --env prod is worth typing. Without it login targets whichever environment is active on that machine, so anyone who has ever logged into staging lands there silently — and the failures that follow (an empty source list, a push to the wrong place) never point back at the cause. dreamlake auth env list shows which login environment is active.

    Install the standalone CLI from the CLI guide and check dreamlake workflow --help. The Python dreamlake package provides the SDK; it is not the supported CLI installation path.

  2. Say what you want — it asks for the data, verifies it, shows you what it found.

    claude code · video-labeling-workflowcollect → check → confirm
    Create a video labeling workflow

    Four things:

    1 · source name — a connected S3 / Dropbox / HF store
    2 · video path inside it
    3 · reference annotation path (.json, scored against)
    4 · task description — one line, e.g. mount a wall shelf
    footage · clips/713488-assembly-shelf.mp4 · gold/713488_annotation.json · mount a wall shelf

    Checked. Confirm before I create it:

    source footage (namespace acme)
    video clips/713488-assembly-shelf.mp4 · 287 MB
    reference gold/713488_annotation.json · 64 phases
    task mount a wall shelf
    publish to acme
    Confirmed.
    the check is not a formality — a path that lists but will not fetch, or a swapped video/annotation pair, costs one line here and a failed run later
  3. It pushes — validated against the schema and the graph rules first.

    $ dreamlake workflow push video-labeling-0807-1501.workflow.json \
    --namespace acme
    workflow: video-labeling-0807-1501 v1 · 5 stages · 9 nodes · 10 edges
    ✓ pushed video-labeling-0807-1501 v1
    open: dreamlake.ai/acme/workflows/video-labeling-0807-1501
  4. Review on DreamLake — iterate, hot-reload.

    dreamlake.ai/acme/workflows/video-labeling-0807-1501spec v2 ▾
    The canvas renders the spec; click nodes to audit config, permissions, and wiring. Ask for a change — Claude pushes v2 and the open page hot-reloads in place within ~2.5s. No refresh.v1 → v2 → v3 · picker pins any
  5. Run it — nodes light up, output streams.

    queued
    Run
    task on the queue
    →
    running
    worker claims it
    nodes light up, logs stream
    →
    awaiting
    gate
    waits for a person — no cost
    →
    completed
    dataset
    published, versioned
    a run's four states — the gate is the only one that waits on a human
  6. Review at the gate, then approve.

    The run publishes its dataset and stops. Select the gate node: ↗ Review dataset opens what was produced, ✓ Approve & continue completes the run. The link stays after approval — it is the way back to what was signed off.

What a spec is made of

PartWhat it is
Stagesphases that group and order — not barriers; execution follows the edges
Nodesthe members of a stage, from four families
Edgestyped connections between ports (samples, dataset, metrics, …); types must match end to end, and a collect port fans several streams into one input
FamilyDoesRuns today
computedeterministic work — filter, transcode, train, publish (any UDF)✅
udaagent judgment — labeling, review, curation✅
controlhuman approval gates✅ approval only
samplerstatistical subset selection⏳ spec-only

The schema also accepts condition, switch, loop and foreach control types. They pass validation, then stop a run with not supported yet — vocabulary to design against, not behaviour to rely on.

A uda node declares the tools and permissions its agent may use, so judgment never runs on raw credentials. Intermediate results stay in the worker's run directory; what reaches DreamLake is what a publish node writes.

Version it

Pushes never overwrite. Each one appends an immutable version, and the spec picker on the workflow page pins the canvas to any of them — pick the latest to resume following pushes.

Editing a node's settings on the canvas and pressing Apply does the same thing: the edited spec is validated and appended as the next version. Nothing is edited in place, so a run always names the version it executed.

dreamlake workflow push video-labeling-0807-1501 (push again → new version)v1v2v3← the page follows latest; the spec picker pins any version

Next steps

CJS Workflows vs Python Pipelines →

The spec graph, Claude Code's JS orchestration scripts, and Python pipelines side by side — and the seven places where one of them does not exist yet.

Node Types Reference →

The full vocabulary — every family's inputs, outputs, and configuration, rendered with the real components.

Agent Permissions →

The grant registry uda nodes draw from — domains, verbs, and scopes.

CLI Reference →

dreamlake workflow push and list — flags, namespaces, versioning.