DreamLake

Transform and Merge

transform is what a plain @ls.udf becomes when nothing else applies — the default, and by count the most common node in any pipeline. merge is the same decorator with a different arity. You do not declare it; the tracer infers it from how many upstream stages you happened to pass in.

Every other node kind on this site is something you ask for. source and sink come from kind=, review comes from kind=, model and filter come from kind=. transform and merge are the two the tracer decides for you, and they are decided by the same rule, which is why they are one page.

The inference rule

Stated once, exactly:

A plain UDF whose inputs include two or more non-source data inputs is typed merge. Anything with fewer is transform.

Three qualifications are load-bearing, and all three come straight from _infer_kinds in the tracer:

  • Only data edges count. A mask edge — the dashed kind a review UDF produces — does not contribute to the count. A stage that takes one table and two masks is a transform, because it has one data input. See Review and masks.
  • Inputs drawn from a source node do not count. This is deliberate: a side input such as a prompt column, loaded by the same source as the data, should not make a stage a merge. Otherwise almost every first-layer stage would be one.
  • An explicit kind= wins. If you write kind="merge" on the decorator, the tracer preserves it verbatim and skips the inference entirely — the same way it skips source, sink and review. Inference only ever fills in a blank.

The consequence that matters more than the rule: you never write merge(...). There is no merge function, no merge node type to reach for, no combinator. A merge is just a UDF you happened to call with several upstream results. If you write a three-argument function and pass it three stage outputs, you have written a merge, whether or not you meant to.

Fan-out is calling the same UDF N times

The tracer gives each call its own node. A UDF called N times yields N distinct nodes with generated ids — semantic_match, semantic_match_2, semantic_match_3 — so both the fan-out and the later fan-in survive in the graph rather than collapsing onto one card with three inbound edges. Three calls to three different UDFs behave identically; the generated-id case matters only when you reuse one function.

Then the caveat that matters, and it is the one that catches people: fan-out width is structural, not dynamic. The tracer does no loop unrolling. It walks the body once. So

python
for i in range(3):
    x = refine(x)

yields one refine node, not three, and the iteration count is lost — there is nowhere in the graph schema to record it. If you want three nodes, write three calls. See Control flow for the full list of what the walk keeps and what it discards.

A comprehension is a map

A comprehension is not a fan-out either, and for a different reason. The tracer binds the loop target to the iterable's provenance and traces the element expression once, so

python
captions = [caption(c) for c in items.clips]

yields one caption node — a map over the stream, which is what a UDF already is. This is the right answer rather than a limitation: a UDF operates on a column, so mapping it over the elements of that column is the identity. List and tuple literals, by contrast, are traced element by element, so [f(a), f(b)] really is two nodes. Again, Control flow has the table.

The worked example

Three captioners with the same signature, and one plain UDF that takes all three of their results:

python
@ls.udf
def caption_model_a(clips, prompts: String["N"]) -> Tuple["caption", "confidence"]: ...

@ls.udf
def caption_model_b(clips, prompts: String["N"]) -> Tuple["caption", "confidence"]: ...

@ls.udf
def caption_model_c(clips, prompts: String["N"]) -> Tuple["caption", "confidence"]: ...

@ls.udf
def merge_captions(a, b, c) -> Tuple["labels"]:
    """Quotient three captionings into one label set."""
    ...


@dl.pipeline
def video_task_annotation():
    src = load_clips()
    for items in batch(src, n=32):
        a = caption_model_a(items.clips, items.prompts)
        b = caption_model_b(items.clips, items.prompts)
        c = caption_model_c(items.clips, items.prompts)
        labels = merge_captions(a, b, c)
clips
clips
clips
a
b
c
load_clips
source · 0→1
caption_model_a
transform · 2→1
caption_model_b
transform · 2→1
caption_model_c
transform · 2→1
merge_captions
merge · 3→1
three calls to a captioning UDF → three nodes → one inferred merge

Look at merge_captions. It was declared with a plain @ls.udf and no kind= — the decorator is byte-identical to the three captioners above it — yet it is painted as a merge, because it received three stage inputs. The three captioners were declared the same way and stayed transform, because each one received a single data input from a source, which the rule does not count. Nothing in the source text distinguishes the four functions. The topology does.

Note also what the source nodes contribute. items.clips and items.prompts are two columns off one upstream stage, not two upstream stages, so they are one input for counting purposes even before the source exemption applies. The .clips attribute access creates no node — which is the next section.

What passes through without becoming a node

Only a call to a decorated function creates a node. Everything else in a pipeline body moves provenance around without appearing on the canvas:

ExpressionTraced as
review.all(axis=0)a method call — provenance of review, passed through
labels[mask]a subscript — labels as data, mask as mask
a & b, a | b, a ^ ba binary operator — the union of both operands
~ma unary operator — the provenance of m, unchanged
items.clipsan attribute — the provenance of items

This is why a merge is always a real function call and never an operator. Writing a & b & c combines three provenance sets into one variable and adds zero nodes to the graph; whatever consumes that variable sees a single input, not three. If you want the reconciliation to be a stage — schedulable, inspectable, versioned, with its own card — it has to be a UDF you call. The mask algebra is documented in full at Review and masks.

Proposal — not importable yet

Everything on this page is real behaviour of the dl_trace tracer, and the example above traces exactly as shown. The imports do not resolve. The shipped dreamlake package exports no pipeline, batch, to_dataset or requeue; lakeshore.udf's real signature is udf(fn=None, *, queue=None, transport="auto", config=None) — it accepts no kind= argument, so kind is tracer-only notation today; and lakeshore.types does not exist at all. These files are a specification format that happens to be Python syntax. The tracer never imports or runs them, which is why they trace anyway.