Transform and Merge
transform is what a plain @ls.udf becomes when nothing else applies —
the default, and by count the most common node in any pipeline. merge is
the same decorator with a different arity. You do not declare it; the
tracer infers it from how many upstream stages you happened to pass in.
Every other node kind on this site is something you ask for. source and
sink come from kind=, review comes from kind=, model and filter
come from kind=. transform and merge are the two the tracer decides for
you, and they are decided by the same rule, which is why they are one page.
The inference rule
Stated once, exactly:
A plain UDF whose inputs include two or more non-source data inputs is typed
merge. Anything with fewer istransform.
Three qualifications are load-bearing, and all three come straight from
_infer_kinds in the tracer:
- Only
dataedges count. Amaskedge — the dashed kind areviewUDF produces — does not contribute to the count. A stage that takes one table and two masks is atransform, because it has one data input. See Review and masks. - Inputs drawn from a
sourcenode do not count. This is deliberate: a side input such as a prompt column, loaded by the same source as the data, should not make a stage a merge. Otherwise almost every first-layer stage would be one. - An explicit
kind=wins. If you writekind="merge"on the decorator, the tracer preserves it verbatim and skips the inference entirely — the same way it skipssource,sinkandreview. Inference only ever fills in a blank.
The consequence that matters more than the rule: you never write
merge(...). There is no merge function, no merge node type to reach for, no
combinator. A merge is just a UDF you happened to call with several upstream
results. If you write a three-argument function and pass it three stage
outputs, you have written a merge, whether or not you meant to.
Fan-out is calling the same UDF N times
The tracer gives each call its own node. A UDF called N times yields N distinct
nodes with generated ids — semantic_match, semantic_match_2,
semantic_match_3 — so both the fan-out and the later fan-in survive in the
graph rather than collapsing onto one card with three inbound edges. Three
calls to three different UDFs behave identically; the generated-id case
matters only when you reuse one function.
Then the caveat that matters, and it is the one that catches people: fan-out width is structural, not dynamic. The tracer does no loop unrolling. It walks the body once. So
yields one refine node, not three, and the iteration count is lost —
there is nowhere in the graph schema to record it. If you want three nodes,
write three calls. See Control flow for the full
list of what the walk keeps and what it discards.
A comprehension is a map
A comprehension is not a fan-out either, and for a different reason. The tracer binds the loop target to the iterable's provenance and traces the element expression once, so
yields one caption node — a map over the stream, which is what a UDF
already is. This is the right answer rather than a limitation: a UDF operates
on a column, so mapping it over the elements of that column is the identity.
List and tuple literals, by contrast, are traced element by element, so
[f(a), f(b)] really is two nodes. Again,
Control flow has the table.
The worked example
Three captioners with the same signature, and one plain UDF that takes all three of their results:
Look at merge_captions. It was declared with a plain @ls.udf and no
kind= — the decorator is byte-identical to the three captioners above it —
yet it is painted as a merge, because it received three stage inputs. The three
captioners were declared the same way and stayed transform, because each one
received a single data input from a source, which the rule does not count.
Nothing in the source text distinguishes the four functions. The topology does.
Note also what the source nodes contribute. items.clips and items.prompts
are two columns off one upstream stage, not two upstream stages, so they are
one input for counting purposes even before the source exemption applies. The
.clips attribute access creates no node — which is the next section.
What passes through without becoming a node
Only a call to a decorated function creates a node. Everything else in a pipeline body moves provenance around without appearing on the canvas:
| Expression | Traced as |
|---|---|
review.all(axis=0) | a method call — provenance of review, passed through |
labels[mask] | a subscript — labels as data, mask as mask |
a & b, a | b, a ^ b | a binary operator — the union of both operands |
~m | a unary operator — the provenance of m, unchanged |
items.clips | an attribute — the provenance of items |
This is why a merge is always a real function call and never an operator.
Writing a & b & c combines three provenance sets into one variable and adds
zero nodes to the graph; whatever consumes that variable sees a single
input, not three. If you want the reconciliation to be a stage — schedulable,
inspectable, versioned, with its own card — it has to be a UDF you call. The
mask algebra is documented in full at
Review and masks.
Everything on this page is real behaviour of the dl_trace tracer, and the
example above traces exactly as shown. The imports do not resolve. The
shipped dreamlake package exports no pipeline, batch, to_dataset or
requeue; lakeshore.udf's real signature is
udf(fn=None, *, queue=None, transport="auto", config=None) — it accepts
no kind= argument, so kind is tracer-only notation today; and
lakeshore.types does not exist at all. These files are a specification
format that happens to be Python syntax. The tracer never imports or runs
them, which is why they trace anyway.