# Transform and Merge

  `transform` is what a plain `@ls.udf` becomes when nothing else applies —
  the default, and by count the most common node in any pipeline. `merge` is
  the **same decorator with a different arity**. You do not declare it; the
  tracer infers it from how many upstream stages you happened to pass in.

Every other node kind on this site is something you ask for. `source` and
`sink` come from `kind=`, `review` comes from `kind=`, `model` and `filter`
come from `kind=`. `transform` and `merge` are the two the tracer decides for
you, and they are decided by the same rule, which is why they are one page.

## The inference rule

Stated once, exactly:

> A plain UDF whose inputs include **two or more non-source data inputs** is
> typed `merge`. Anything with fewer is `transform`.

Three qualifications are load-bearing, and all three come straight from
`_infer_kinds` in the tracer:

- **Only `data` edges count.** A `mask` edge — the dashed kind a `review` UDF
  produces — does not contribute to the count. A stage that takes one table and
  two masks is a `transform`, because it has one data input. See
  [Review and masks](/pipelines/review-and-masks.md).
- **Inputs drawn from a `source` node do not count.** This is deliberate: a
  side input such as a prompt column, loaded by the same source as the data,
  should not make a stage a merge. Otherwise almost every first-layer stage
  would be one.
- **An explicit `kind=` wins.** If you write `kind="merge"` on the decorator,
  the tracer preserves it verbatim and skips the inference entirely — the same
  way it skips `source`, `sink` and `review`. Inference only ever fills in a
  blank.

The consequence that matters more than the rule: **you never write
`merge(...)`.** There is no merge function, no merge node type to reach for, no
combinator. A merge is just a UDF you happened to call with several upstream
results. If you write a three-argument function and pass it three stage
outputs, you have written a merge, whether or not you meant to.

## Fan-out is calling the same UDF N times

The tracer gives each call its own node. A UDF called N times yields N distinct
nodes with generated ids — `semantic_match`, `semantic_match_2`,
`semantic_match_3` — so both the fan-out and the later fan-in survive in the
graph rather than collapsing onto one card with three inbound edges. Three
calls to three *different* UDFs behave identically; the generated-id case
matters only when you reuse one function.

Then the caveat that matters, and it is the one that catches people:
**fan-out width is structural, not dynamic.** The tracer does no loop
unrolling. It walks the body once. So

```python
for i in range(3):
    x = refine(x)
```

yields **one** `refine` node, not three, and the iteration count is lost —
there is nowhere in the graph schema to record it. If you want three nodes,
write three calls. See [Control flow](/pipelines/control-flow.md) for the full
list of what the walk keeps and what it discards.

## A comprehension is a map

A comprehension is not a fan-out either, and for a different reason. The tracer
binds the loop target to the iterable's provenance and traces the **element
expression once**, so

```python
captions = [caption(c) for c in items.clips]
```

yields **one** `caption` node — a map over the stream, which is what a UDF
already is. This is the right answer rather than a limitation: a UDF operates
on a column, so mapping it over the elements of that column is the identity.
List and tuple *literals*, by contrast, are traced element by element, so
`[f(a), f(b)]` really is two nodes. Again,
[Control flow](/pipelines/control-flow.md) has the table.

## The worked example

Three captioners with the same signature, and one plain UDF that takes all
three of their results:

```python
@ls.udf
def caption_model_a(clips, prompts: String["N"]) -> Tuple["caption", "confidence"]: ...

@ls.udf
def caption_model_b(clips, prompts: String["N"]) -> Tuple["caption", "confidence"]: ...

@ls.udf
def caption_model_c(clips, prompts: String["N"]) -> Tuple["caption", "confidence"]: ...

@ls.udf
def merge_captions(a, b, c) -> Tuple["labels"]:
    """Quotient three captionings into one label set."""
    ...


@dl.pipeline
def video_task_annotation():
    src = load_clips()
    for items in batch(src, n=32):
        a = caption_model_a(items.clips, items.prompts)
        b = caption_model_b(items.clips, items.prompts)
        c = caption_model_c(items.clips, items.prompts)
        labels = merge_captions(a, b, c)
```

Look at `merge_captions`. It was declared with a plain `@ls.udf` and **no
`kind=`** — the decorator is byte-identical to the three captioners above it —
yet it is painted as a merge, because it received three stage inputs. The three
captioners were declared the same way and stayed `transform`, because each one
received a single data input from a `source`, which the rule does not count.
Nothing in the source text distinguishes the four functions. The topology does.

Note also what the source nodes contribute. `items.clips` and `items.prompts`
are two columns off one upstream stage, not two upstream stages, so they are
one input for counting purposes even before the source exemption applies. The
`.clips` attribute access creates no node — which is the next section.

## What passes through without becoming a node

Only a call to a decorated function creates a node. Everything else in a
pipeline body moves provenance around without appearing on the canvas:

| Expression | Traced as |
| --- | --- |
| `review.all(axis=0)` | a method call — provenance of `review`, passed through |
| `labels[mask]` | a subscript — `labels` as `data`, `mask` as `mask` |
| `a & b`, `a \| b`, `a ^ b` | a binary operator — the union of both operands |
| `~m` | a unary operator — the provenance of `m`, unchanged |
| `items.clips` | an attribute — the provenance of `items` |

This is why a merge is always a real function call and never an operator.
Writing `a & b & c` combines three provenance sets into one variable and adds
**zero** nodes to the graph; whatever consumes that variable sees a single
input, not three. If you want the reconciliation to be a stage — schedulable,
inspectable, versioned, with its own card — it has to be a UDF you call. The
mask algebra is documented in full at
[Review and masks](/pipelines/review-and-masks.md).

> **Warning:** Everything on this page is real behaviour of the `dl_trace` tracer, and the
>   example above traces exactly as shown. The *imports* do not resolve. The
>   shipped `dreamlake` package exports no `pipeline`, `batch`, `to_dataset` or
>   `requeue`; `lakeshore.udf`'s real signature is
>   `udf(fn=None, *, queue=None, transport="auto", config=None)` — it accepts
>   **no `kind=` argument**, so `kind` is tracer-only notation today; and
>   `lakeshore.types` does not exist at all. These files are a specification
>   format that happens to be Python syntax. The tracer never imports or runs
>   them, which is why they trace anyway.
