# Review and Masks

  A review is a UDF that returns a boolean mask instead of data. The mask is
  not a table you pass along — it is a gate. The tracer marks every edge the
  mask touches, and the renderer draws those edges dashed.

That distinction is the whole page. Every other UDF in a pipeline answers the
question *what is the value?*; a review answers *which rows survive?* The two
answers travel along the same graph, but they are not the same kind of thing,
and a graph that drew them identically would be lying about what the pipeline
does.

## The review node

A review is declared with `kind="review"` and returns a `Mask[…]`:

```python
@ls.udf(kind="review")
def review_boxes(labels) -> Mask["R", "N"]:
    """R reviewers × N rows of accept/reject."""
    ...


@ls.udf(kind="review")
def semantic_match(labels, reference) -> Mask["N"]: ...
```

The `kind` is preserved verbatim. The tracer reads the decorator's keyword
arguments, copies them into the node's `config`, and uses `kind` as the node's
`kind` field — so `kind="review"` is what paints the card's review dot, and any
other keyword you pass survives on `config` untouched. Nothing validates the
string; `kind` is a label the tracer forwards, not an enum it enforces.

Note the annotation rule, because it surprises people. The return annotation
sets the result's **column schema**, and only `Tuple[…]` spreads its arguments
into separate columns. Every other subscripted type collapses to a single
column named after the base type, lowercased. So `Mask["R", "N"]` yields
`columns: ['mask']` — one column, not two. `"R"` and `"N"` are dimensions of
that one mask, not names of two of them. The same rule governs
`Tensor["N", "H", "W", 3]`, which is one `tensor` column. See
[UDF](/pipelines/udf.md) for the full annotation table.

A review node is otherwise an ordinary node: its parameters are its input
ports, and it has exactly one output port carrying that one `mask` column.

## The algebra is edges, not nodes

Here is the part worth reading twice. A pipeline manipulates masks with
ordinary Python operators, and **none of the following creates a node**. Each
one passes its operands' provenance through and tags it:

| Expression | What it means |
| --- | --- |
| `~a` | complement — the rows the mask rejects |
| `a & b` | intersection |
| `a \| b` | union |
| `a ^ b` | symmetric difference |
| `a & ~b` | difference |
| `labels[mask]` | SELECT the subset — this is where a mask **enters** |
| `poses.confidence > 0.5` | threshold — a comparison yields a mask |
| `review.all(axis=0)` | reduce R reviewers to one verdict |
| `review.mean(axis=0) > 0.5` | majority instead of unanimity |

The tracer treats subscripts, unary and binary operators, comparisons, boolean
operators, attribute access and method calls as pure provenance plumbing. They
have no node template to instantiate — there is no `@ls.udf` behind `&` — so
there is nothing to draw. Only a call to a decorated function becomes a node.

Now the precise mechanic, which is worth stating exactly because it is the one
rule that explains every dashed edge you will ever see:

> A mask is introduced at a **subscript's slice** — the thing inside the
> brackets — or at an **`if` test**. From there it taints its whole provenance:
> every node the mask descended from becomes a `mask` contributor to that edge.

Two consequences follow. First, the taint is transitive and total. In
`labels[review.all(axis=0)]` the slice is a method call on `review`, whose
provenance is the `review_boxes` node — so `review_boxes` is tagged `mask`, and
so is anything `review_boxes` itself descended from, all the way back. Second,
a mask only becomes a mask *at the point of use*. `review = review_boxes(labels)`
by itself is a plain data assignment; it is `labels[consensus]` that decides
`consensus` is a selector. Write a review's output into a sink directly and it
is data, because nothing ever put it in gate position.

## Data dominates mask

When the same pair of nodes is reached by both a data path and a mask path, the
edge is `data`. The tracer keys edges on `(from, fromPort, to, toPort)` and
resolves collisions in one direction only: `data` wins, always, no matter which
tag arrived first.

The rule is not arbitrary — it is what makes a dashed edge *mean* something.
Read it as a guarantee about the graph:

**A dashed edge means the source only gates the target, and never carries a
value into it.** A solid edge means at least one value flows. Because `data`
dominates, an edge that is still dashed at render time has had no data path
found on any traced branch — so the dashes are a negative claim you can rely
on, not a stylistic hint.

That is exactly why, in the figure below, the dashed edges come from
`review_boxes` and the solid ones come from `annotate`.

## The worked example

```python
@dl.pipeline
def image_object_annotation():
    src = load_images()
    for items in batch(src, n=64):
        labels = annotate(items.images)
        review = review_boxes(labels)          # -> Mask[R, N]
        consensus = review.all(axis=0)         # R reviewers must agree
        save_dataset(labels[consensus])        # the mask GATES what is written
        rework(labels[~consensus])             # the complement goes back
```

Read the graph against the source and notice what is *missing*. There are four
nodes for four UDF calls. `review.all(axis=0)` and `~consensus` produced no
nodes at all — they are edge tags. The consensus reduction over R reviewers and
the complement that routes the rejects to `rework` are both invisible as
geometry and entirely visible as line style.

Notice also where each sink's rows actually come from. `save_dataset` and
`rework` are each reached by two edges, and in both cases the solid one starts
at `annotate`. That is the honest reading of `labels[consensus]`: the value
being written is `labels`, and `consensus` only chose a subset of it. If the
graph drew a single edge from `review_boxes` to `save_dataset`, it would be
claiming the sink writes the mask — which is not what the code says.

The two sinks are also the same subset split two ways: `consensus` and
`~consensus` partition the rows, so every row lands in exactly one of them. The
graph cannot show that — the complement is not a node — which is a real
fidelity loss and the reason the mask expression stays in the source pane.

## Where masks come from and where they go

- [Sampling](/pipelines/sampling.md) — the subscript is the tap. Taking a subset
  off the main stream is a mask subscript, not a sampler function, and that
  page says what does and does not exist for it.
- [UDF](/pipelines/udf.md) — the annotation rules that turn `Mask[…]` into one
  column and `Tuple[…]` into several.
- [Control flow](/pipelines/control-flow.md) — an `if` test is the other place a
  mask enters. The tracer evaluates the condition in mask position, then
  discards the condition itself and walks both branches.

> **Warning:** Everything on this page is notation the tracer reads statically. None of it
>   runs.
>
>   `lakeshore.types` does not exist as a module — `Mask`, `String`, `Tensor` and
>   `Tuple` are names the tracer matches in an annotation, and
>   `from lakeshore.types import Mask` fails on a real interpreter.
>   `lakeshore.udf` accepts no `kind=` argument; its shipped signature is
>   `udf(fn=None, *, queue=None, transport="auto", config=None)`, so `kind` is
>   notation the tracer reads and the runtime never sees. The shipped `dreamlake`
>   package exports no `pipeline`, `batch`, `to_dataset`, or `requeue`.
>
>   The example files trace because the tracer parses them and never imports or
>   executes anything. Treat this page as a specification of the graph the tracer
>   derives, not as an API you can call today.
