DreamLake

Workflow Node Types

A workflow is a spine of ordered stages; typed member nodes fan out from their stage and connect through typed data edges. This page defines every node family and the data types edges can carry — rendered with the real canvas components.

The vocabulary is grounded in established systems: the artifact type system follows Kubeflow Pipelines v2 and Flyte, control flow follows BPMN 2.0 and the workflow control-flow patterns literature, and samplers use statistical naming with Spark/SQL parameter conventions.

Stage

Annotate
stage · 3 members2 done
stage node — a hub in the flow; members fan out from it

A stage groups the work of one phase and is rendered as a node on the spine. Stages are hubs, not barriers: work converges into a stage node and fans out to its members, but data edges may cross stage boundaries freely. Stage order (the array order in the spec) is the spine.

FieldTypeNotes
idstring^[a-z0-9][a-z0-9_-]{0,63}$, unique
titlestringdisplay name
detailstring?one-line description

Compute (UDF)

bimanual_filter
compute · udf
SLURM · xeon-g6-voltadirect
clip_transcode
compute · udf
EC2 · g5.xlargedaemon
compute (UDF) nodes — launcher badge, per-machine fields, dispatch mode

A compute node runs a user-defined function on Lakeshore. The provider block follows the Lakeshore provider schema: a launcher class (SSH · SLURM · EC2 · GCE · Kube) or a named server-stored provider, plus per-machine RunConfig fields.

FieldTypeNotes
compute.udfstringUDF reference, e.g. pipelines.bimanual_filter
compute.paramsobject?static UDF arguments — small inline values are node parameters, never port data (the KFP/Snakemake split)
compute.providerProviderRef?provider (named) XOR launcher + kwargs; per-machine: instance_type, image, partition, resources, runner
compute.dispatchdirect | daemonhow the launch happens; defaults per launcher
executionExecutionPolicy?retry {max_attempts (1 = no retry), retry_on, backoff {initial, factor, max}}, per-attempt timeout, cache {enabled, version}

Secrets in provider kwargs must use {"$secret": "name"} markers — never inline credentials.

UDA (User-Defined Agent)

vlm_annotator
uda · agent
2 permsqwen-vl-72bqueue: gpu-a10g
episode_curator
uda · agent
2 permsclaude-haikuSSH
uda (user-defined agent) nodes — instructions preview, model, permission count, queue or provider target

A uda node runs a remote agent. Field names follow agent-framework conventions (OpenAI Agents SDK, Claude Agent SDK, A2A):

FieldTypeNotes
uda.instructionsstringthe system prompt (fixed persona/behavior — distinct from any per-run input)
uda.descriptionstring?when to delegate to this agent (routing metadata)
uda.modelstring?model id
uda.toolsstring[]?tool grants by name — tools are not permission strings
uda.permissionsstring[]data/resource grants — see Agent Permissions
uda.max_turnsnumber?turn budget
uda.output_schemaobject?JSON Schema for the structured final output
uda.provider XOR uda.queueProviderRef / stringwhere the agent runs: a provider/host, or a Lakeshore queue
executionExecutionPolicy?retry + timeout only — cache is forbidden on agents: agent runs are non-deterministic; caching one replays a stale answer

Sampler

bernoulli_10pct
sampler · 10% · ≥50
random_500
sampler · n=500
per_episode
sampler · by episode · 30%
head_100
sampler · first 100
sampler strategies — bernoulli · random_n · stratified · first_n

Sampler strategies use statistically correct names. The result of bernoulli is probabilistic in count — that is what SQL TABLESAMPLE BERNOULLI and Spark df.sample(fraction) actually do.

StrategyParametersSemantics
bernoullifraction ∈ (0, 1], min_size?, seed?each item kept independently with probability fraction. min_size is a DreamLake extension: if fraction·N < min_size, degrade to an exact-size random sample
random_nsize, with_replacement? = false, seed?exact-size simple random sample (reservoir sampling is the streaming implementation)
stratifiedstratify_by, fraction? or fractions? (per-stratum map), min_size?, seed?per-stratum sampling, Spark sampleBy shape
first_nsizehead/LIMIT — deterministic and order-dependent; not a statistical sample

Signature: samplers take exactly one input, restricted to the collection-like types (samples · table · directory), and are type-preserving — the output port is derived with the same type (a sample of a clip collection is a clip collection).

Determinism: identical input + identical seed ⇒ identical sample (SQL REPEATABLE semantics). Absent seed ⇒ non-reproducible (the run records the effective seed).

Control flow

is_confident
control · conditionconfidence >= 0.95
tier_switch
control · switch3 cases + default
retry_loop
control · loopwhile · until rework < 0.05
qa_gate
control · approvalReview before publish
control-flow nodes — condition · switch · loop (while | foreach) · approval
TypeOut portsSemantics
conditiontrue: T, false: Tbinary exclusive choice (expression)
switchone per case + required default, all Tn-way exclusive choice (WCP-4); first matching case wins; default keeps routing total
loopout: Tmode: 'while' — condition-bounded, until + max_iterations required (no unbounded loops); mode: 'foreach' — collection-driven (over, max_concurrency)
approvalout: Thuman gate with Argo-suspend semantics: pauses until a decision; timeout without decision ⇒ error, never auto-approve

Control nodes are type-preserving pass-throughs: every derived out port carries the input port's type T.

Plain fan-out and fan-in are implicit in edges — there are no AND-gateway nodes. An input port accepts one inbound edge by default; fan-in is either a collect: true port (ordered collection of same-type producers) or the XOR-merge exception (branches of the same condition/switch may share a target port, since at most one fires).

Edge data types

Every port carries a type from a closed artifact lattice. artifact is the root and doubles as "any": any subtype is accepted where artifact is expected.

TypeMeaning
artifactroot — compatible with everything
file / directoryopaque blob, single vs multipart; format is metadata, not more types
tableschema-carrying tabular data (a shape)
dataseta versioned data product — shards + manifest
modeltrained model / checkpoint
metricsevaluation metrics
samplesDreamLake domain type — an addressable collection of media samples / episodes

Incompatible connections are flagged on the canvas before anything runs. Small inline values (strings, numbers, JSON config) are node configuration, not edge data — a type never exists on both sides of the parameter/artifact boundary.

Run states

idle
compute · udf
SLURM · xeon-g6-voltadirect
queued
compute · udf
SLURM · xeon-g6-voltadirect
progress
compute · udf
SLURM · xeon-g6-voltadirect
done
compute · udf
SLURM · xeon-g6-voltadirect
error
compute · udf
SLURM · xeon-g6-voltadirect
run states tint the same card — idle · queued · progress (pulses) · done · error

During a run, each node carries a state — idle · queued · progress · done · error · skipped — tinting the same card. Agent instances fan out under their uda node:

annotator-01
48.2k tok · 182s
annotator-02
12.1k tok
annotator-03
3.4k tok · 41s
agent instances — run-time cards that stack under their uda node

The canvas

Stages are hubs the work flows through — members fan out from their stage node and converge into the next. Two orientations, one card style:

annotator-01
48.2k tok
annotator-02
12.1k tok
Collect
stage · 2 members2 done
Annotate
stage · 1 member0 done
bimanual_filter
compute · udf
SLURM · xeon-g6
take_sample
sampler · 10% · ≥50
vlm_annotator
uda · agent
1 permqwen-vl-72bqueue: gpu-a10g
samplessamples
vertical orientation — stages as hubs, members fanning out, live run overlay
annotator-01
48.2k tok
annotator-02
12.1k tok
Collect
stage · 2 members2 done
Annotate
stage · 1 member0 done
bimanual_filter
compute · udf
SLURM · xeon-g6
take_sample
sampler · 10% · ≥50
vlm_annotator
uda · agent
1 permqwen-vl-72bqueue: gpu-a10g
samplessamples
horizontal orientation — stages as hubs, members fanning out, live run overlay