DreamLake

Review and Masks

A review is a UDF that returns a boolean mask instead of data. The mask is not a table you pass along — it is a gate. The tracer marks every edge the mask touches, and the renderer draws those edges dashed.

That distinction is the whole page. Every other UDF in a pipeline answers the question what is the value?; a review answers which rows survive? The two answers travel along the same graph, but they are not the same kind of thing, and a graph that drew them identically would be lying about what the pipeline does.

The review node

A review is declared with kind="review" and returns a Mask[…]:

python
@ls.udf(kind="review")
def review_boxes(labels) -> Mask["R", "N"]:
    """R reviewers × N rows of accept/reject."""
    ...


@ls.udf(kind="review")
def semantic_match(labels, reference) -> Mask["N"]: ...

The kind is preserved verbatim. The tracer reads the decorator's keyword arguments, copies them into the node's config, and uses kind as the node's kind field — so kind="review" is what paints the card's review dot, and any other keyword you pass survives on config untouched. Nothing validates the string; kind is a label the tracer forwards, not an enum it enforces.

Note the annotation rule, because it surprises people. The return annotation sets the result's column schema, and only Tuple[…] spreads its arguments into separate columns. Every other subscripted type collapses to a single column named after the base type, lowercased. So Mask["R", "N"] yields columns: ['mask'] — one column, not two. "R" and "N" are dimensions of that one mask, not names of two of them. The same rule governs Tensor["N", "H", "W", 3], which is one tensor column. See UDF for the full annotation table.

A review node is otherwise an ordinary node: its parameters are its input ports, and it has exactly one output port carrying that one mask column.

The algebra is edges, not nodes

Here is the part worth reading twice. A pipeline manipulates masks with ordinary Python operators, and none of the following creates a node. Each one passes its operands' provenance through and tags it:

ExpressionWhat it means
~acomplement — the rows the mask rejects
a & bintersection
a | bunion
a ^ bsymmetric difference
a & ~bdifference
labels[mask]SELECT the subset — this is where a mask enters
poses.confidence > 0.5threshold — a comparison yields a mask
review.all(axis=0)reduce R reviewers to one verdict
review.mean(axis=0) > 0.5majority instead of unanimity

The tracer treats subscripts, unary and binary operators, comparisons, boolean operators, attribute access and method calls as pure provenance plumbing. They have no node template to instantiate — there is no @ls.udf behind & — so there is nothing to draw. Only a call to a decorated function becomes a node.

Now the precise mechanic, which is worth stating exactly because it is the one rule that explains every dashed edge you will ever see:

A mask is introduced at a subscript's slice — the thing inside the brackets — or at an if test. From there it taints its whole provenance: every node the mask descended from becomes a mask contributor to that edge.

Two consequences follow. First, the taint is transitive and total. In labels[review.all(axis=0)] the slice is a method call on review, whose provenance is the review_boxes node — so review_boxes is tagged mask, and so is anything review_boxes itself descended from, all the way back. Second, a mask only becomes a mask at the point of use. review = review_boxes(labels) by itself is a plain data assignment; it is labels[consensus] that decides consensus is a selector. Write a review's output into a sink directly and it is data, because nothing ever put it in gate position.

Data dominates mask

When the same pair of nodes is reached by both a data path and a mask path, the edge is data. The tracer keys edges on (from, fromPort, to, toPort) and resolves collisions in one direction only: data wins, always, no matter which tag arrived first.

The rule is not arbitrary — it is what makes a dashed edge mean something. Read it as a guarantee about the graph:

A dashed edge means the source only gates the target, and never carries a value into it. A solid edge means at least one value flows. Because data dominates, an edge that is still dashed at render time has had no data path found on any traced branch — so the dashes are a negative claim you can rely on, not a stylistic hint.

That is exactly why, in the figure below, the dashed edges come from review_boxes and the solid ones come from annotate.

The worked example

python
@dl.pipeline
def image_object_annotation():
    src = load_images()
    for items in batch(src, n=64):
        labels = annotate(items.images)
        review = review_boxes(labels)          # -> Mask[R, N]
        consensus = review.all(axis=0)         # R reviewers must agree
        save_dataset(labels[consensus])        # the mask GATES what is written
        rework(labels[~consensus])             # the complement goes back
labels
rows
rows
rows
rows
annotate
transform · 1→1
review_boxes
review · 1→1
save_dataset
sink · 1→0
rework
sink · 1→0
the review node's edge is dashed — it gates, it does not carry

Read the graph against the source and notice what is missing. There are four nodes for four UDF calls. review.all(axis=0) and ~consensus produced no nodes at all — they are edge tags. The consensus reduction over R reviewers and the complement that routes the rejects to rework are both invisible as geometry and entirely visible as line style.

Notice also where each sink's rows actually come from. save_dataset and rework are each reached by two edges, and in both cases the solid one starts at annotate. That is the honest reading of labels[consensus]: the value being written is labels, and consensus only chose a subset of it. If the graph drew a single edge from review_boxes to save_dataset, it would be claiming the sink writes the mask — which is not what the code says.

The two sinks are also the same subset split two ways: consensus and ~consensus partition the rows, so every row lands in exactly one of them. The graph cannot show that — the complement is not a node — which is a real fidelity loss and the reason the mask expression stays in the source pane.

Where masks come from and where they go

  • Sampling — the subscript is the tap. Taking a subset off the main stream is a mask subscript, not a sampler function, and that page says what does and does not exist for it.
  • UDF — the annotation rules that turn Mask[…] into one column and Tuple[…] into several.
  • Control flow — an if test is the other place a mask enters. The tracer evaluates the condition in mask position, then discards the condition itself and walks both branches.
Proposal — not importable yet

Everything on this page is notation the tracer reads statically. None of it runs.

lakeshore.types does not exist as a module — Mask, String, Tensor and Tuple are names the tracer matches in an annotation, and from lakeshore.types import Mask fails on a real interpreter. lakeshore.udf accepts no kind= argument; its shipped signature is udf(fn=None, *, queue=None, transport="auto", config=None), so kind is notation the tracer reads and the runtime never sees. The shipped dreamlake package exports no pipeline, batch, to_dataset, or requeue.

The example files trace because the tracer parses them and never imports or executes anything. Treat this page as a specification of the graph the tracer derives, not as an API you can call today.