DreamLake

Control Flow

A pipeline's graph is derived from its source by a static tracer that never imports and never executes anything. It walks the syntax tree once. That single fact decides everything on this page: the graph over-approximates — it draws every path the code could take — and it cannot draw anything that only exists at run time.

Both halves of that sentence cost something, and the cost is worth stating plainly rather than discovering later. Over-approximation means a graph with four nodes may describe a run that touched two of them. "Nothing that exists only at run time" means the predicate that chose between those two — the thing you most want to see when you are reading the picture — is absent.

There is no control-flow node

The first surprise is that there is nothing to look for. condition, switch, loop and approval are node types, but they belong to the WorkflowSpec model — hand-authored JSON, a different artifact, documented at Workflow Node Types. A Python pipeline has no equivalent and does not want one.

In a pipeline, control flow is ordinary Python. It is not declared, not typed, and not addressable. It shows up in the derived graph as shape — which nodes exist, and which edges join them — and never as a node of its own. If you are looking for the box labelled "if", there isn't one, and its absence is the design rather than a gap in the renderer.

The summary

ConstructWhat the tracer does
if / elseboth branches traced; the resulting provenance is the union; the condition is not drawn
forthe body is traced once; the iteration count is lost
for … in batch(…)the iterator is exempt from the decoration rule; adds no node
[f(x) for x in xs]a map — the element expression is traced once, so one node
[a(), b()] / (a(), b())a literal fan-out — one node per element expression
while / withthe body is traced once, like a for

Each row below.

if / else

python
@dl.pipeline
def image_object_annotation():
    src = load_images()
    for items in batch(src, n=64):
        labels = detect(items.images)
        if fast_mode:
            labels = refine_fast(labels)
        else:
            labels = refine_slow(labels)
        save_dataset(labels)
labels
labels
rows
rows
detect
transform · 1→1
refine_fast
transform · 1→1
refine_slow
transform · 1→1
save_dataset
sink · 1→0
both branches of the if are traced; the condition itself is not drawn

Read the figure literally. Both refine_fast and refine_slow are in the graph, and both feed the sink. That is not a rendering artefact: the tracer walks the if body and the else body from the same pre-if bindings, then merges the two results, so after the branch the name labels carries the union of the two provenances. save_dataset(labels) therefore draws one edge per contributing node. The alternative — trace one branch, drop the other — would leave a real node dangling with no path to the sink, which is worse.

fast_mode appears nowhere. There is no node for it, no port, no label on either edge. The tracer reads the test, uses it (see below), and discards it.

The second-order consequence is the one that bites. The graph shows what the pipeline could do, not what one run did. A four-node picture of a two-node execution is correct as a description of the program and wrong as a description of the run. When runtime status is painted onto the graph, the branch that never fired stays idle — which is the only honest thing an over-approximated graph can say, and is not the same as "this node is not part of the pipeline".

There is a related mechanic worth knowing, because it is the one place the condition leaves a trace. An if test is evaluated in mask role — one of exactly two places a mask enters provenance, the other being a subscript's slice, labels[mask]. So a conditional written on a mask column taints everything downstream of it: the edges that carry that provenance are tagged mask and render dashed. The predicate is still not a node, but its influence is visible in the edge tags. See Review and Masks for the tagging rules and the data-dominates-mask merge.

Loops do not unroll

python
    for i in range(3):
        x = refine(x)          # ONE refine node, not three
x
load_frames
source · 0→1
refine
transform · 1→1
three iterations, one node — the count is lost

Flatly: the iteration count is lost. The body is walked once, refine becomes one node, and nothing in the emitted JSON records that it runs three times. The 3 is not stored as a property, an annotation, or a badge — the tracer never evaluates range(3), because evaluating it would mean running the code.

This is not a bug to route around; it is the direct consequence of a single-pass walk. If you want three stages in the graph, write three calls. Calling one UDF N times is the documented way to get N nodes — the tracer disambiguates them (refine, refine_2, refine_3) and both the fan-out and the fan-in survive. See Transform and Merge.

The loop that is idiomatic in a pipeline body is the batching loop, for items in batch(src, n=64). Its iterator is the single exemption from the rule that every call in a pipeline body must be decorated: batch may be undecorated, adds no node, and the loop variable simply inherits the iterated source's provenance. The body is strict again. See Batching.

A comprehension is a map

python
    captions = [caption(c) for c in items.clips]

The element expression is traced once, with the loop target bound to the iterable's provenance, so this yields exactly one caption node. That is the right reading: a comprehension over a stream is a map, and a map is one stage applied to every row, not one stage per row. The same holds for set comprehensions, dict comprehensions and generator expressions.

A list or tuple literal is a different construct and traces differently:

python
    scored = [model_a(x), model_b(x)]

Here each element is a separate expression, each is traced, and each becomes its own node. This is a genuine fan-out — two nodes, two edges — and it is the compact way to write one. The distinction is syntactic and worth internalising: a comprehension has one element expression and yields one node; a literal has N element expressions and yields N.

while and with bodies are traced once, exactly like a for. A while's test is evaluated in data role rather than mask role, but it is likewise not drawn.

The open gap

Known limitation

Data-dependent routing has no representation. Because the if condition is discarded, a traced graph cannot distinguish "route rows to A or B by a predicate" from "send everything to both A and B". Both produce the same four nodes and the same four edges. This is tracked upstream as an open tracer-fidelity question — DreamLake issue #50 — and there is no annotation, edge tag, or node property that carries the answer today.

One thing does express routing, and it is worth reaching for when the routing matters: the mask subscript. rows[mask] and rows[~mask] split a stream into two disjoint halves, and because a mask is data it survives tracing — the split appears as two dashed mask edges from the node that produced the mask. A predicate held in a Python if is invisible; the same predicate held in a mask column is in the graph. See Sampling for the subscript form and where it stops being enough.

The practical rule that follows from all of the above: keep the pipeline body straight-line dataflow, and move data-dependent control inside a UDF, where the tracer treats the whole thing as one stage and makes no claim about its internals. A branch the graph cannot see is better hidden than half-drawn.

Proposal — not importable yet

The Python on this page is a specification format that happens to be Python syntax. The shipped dreamlake package exports no pipeline, batch, to_dataset or requeue; lakeshore.udf accepts no kind= argument; and lakeshore.types does not exist at all. These files trace precisely because the tracer never imports or executes them. Do not expect them to run.