Control Flow
A pipeline's graph is derived from its source by a static tracer that never imports and never executes anything. It walks the syntax tree once. That single fact decides everything on this page: the graph over-approximates — it draws every path the code could take — and it cannot draw anything that only exists at run time.
Both halves of that sentence cost something, and the cost is worth stating plainly rather than discovering later. Over-approximation means a graph with four nodes may describe a run that touched two of them. "Nothing that exists only at run time" means the predicate that chose between those two — the thing you most want to see when you are reading the picture — is absent.
There is no control-flow node
The first surprise is that there is nothing to look for. condition, switch,
loop and approval are node types, but they belong to the WorkflowSpec
model — hand-authored JSON, a different artifact, documented at
Workflow Node Types. A Python pipeline has no
equivalent and does not want one.
In a pipeline, control flow is ordinary Python. It is not declared, not typed, and not addressable. It shows up in the derived graph as shape — which nodes exist, and which edges join them — and never as a node of its own. If you are looking for the box labelled "if", there isn't one, and its absence is the design rather than a gap in the renderer.
The summary
| Construct | What the tracer does |
|---|---|
if / else | both branches traced; the resulting provenance is the union; the condition is not drawn |
for | the body is traced once; the iteration count is lost |
for … in batch(…) | the iterator is exempt from the decoration rule; adds no node |
[f(x) for x in xs] | a map — the element expression is traced once, so one node |
[a(), b()] / (a(), b()) | a literal fan-out — one node per element expression |
while / with | the body is traced once, like a for |
Each row below.
if / else
Read the figure literally. Both refine_fast and refine_slow are in the
graph, and both feed the sink. That is not a rendering artefact: the tracer
walks the if body and the else body from the same pre-if bindings, then
merges the two results, so after the branch the name labels carries the
union of the two provenances. save_dataset(labels) therefore draws one
edge per contributing node. The alternative — trace one branch, drop the other
— would leave a real node dangling with no path to the sink, which is worse.
fast_mode appears nowhere. There is no node for it, no port, no label on
either edge. The tracer reads the test, uses it (see below), and discards it.
The second-order consequence is the one that bites. The graph shows what the
pipeline could do, not what one run did. A four-node picture of a two-node
execution is correct as a description of the program and wrong as a description
of the run. When runtime status is painted onto the graph, the branch that
never fired stays idle — which is the only honest thing an over-approximated
graph can say, and is not the same as "this node is not part of the pipeline".
There is a related mechanic worth knowing, because it is the one place the
condition leaves a trace. An if test is evaluated in mask role — one of
exactly two places a mask enters provenance, the other being a subscript's
slice, labels[mask]. So a conditional written on a mask column taints
everything downstream of it: the edges that carry that provenance are tagged
mask and render dashed. The predicate is still not a node, but its influence
is visible in the edge tags. See
Review and Masks for the tagging rules and the
data-dominates-mask merge.
Loops do not unroll
Flatly: the iteration count is lost. The body is walked once, refine becomes
one node, and nothing in the emitted JSON records that it runs three times. The
3 is not stored as a property, an annotation, or a badge — the tracer never
evaluates range(3), because evaluating it would mean running the code.
This is not a bug to route around; it is the direct consequence of a
single-pass walk. If you want three stages in the graph, write three calls.
Calling one UDF N times is the documented way to get N nodes — the tracer
disambiguates them (refine, refine_2, refine_3) and both the fan-out and
the fan-in survive. See
Transform and Merge.
The loop that is idiomatic in a pipeline body is the batching loop,
for items in batch(src, n=64). Its iterator is the single exemption from the
rule that every call in a pipeline body must be decorated: batch may be
undecorated, adds no node, and the loop variable simply inherits the iterated
source's provenance. The body is strict again. See
Batching.
A comprehension is a map
The element expression is traced once, with the loop target bound to the
iterable's provenance, so this yields exactly one caption node. That is
the right reading: a comprehension over a stream is a map, and a map is one
stage applied to every row, not one stage per row. The same holds for set
comprehensions, dict comprehensions and generator expressions.
A list or tuple literal is a different construct and traces differently:
Here each element is a separate expression, each is traced, and each becomes its own node. This is a genuine fan-out — two nodes, two edges — and it is the compact way to write one. The distinction is syntactic and worth internalising: a comprehension has one element expression and yields one node; a literal has N element expressions and yields N.
while and with bodies are traced once, exactly like a for. A while's
test is evaluated in data role rather than mask role, but it is likewise not
drawn.
The open gap
Data-dependent routing has no representation. Because the if condition
is discarded, a traced graph cannot distinguish "route rows to A or B by a
predicate" from "send everything to both A and B". Both produce the same four
nodes and the same four edges. This is tracked upstream as an open
tracer-fidelity question — DreamLake issue #50 — and there is no annotation,
edge tag, or node property that carries the answer today.
One thing does express routing, and it is worth reaching for when the routing
matters: the mask subscript. rows[mask] and rows[~mask] split a stream
into two disjoint halves, and because a mask is data it survives tracing —
the split appears as two dashed mask edges from the node that produced the
mask. A predicate held in a Python if is invisible; the same predicate held
in a mask column is in the graph. See Sampling for the
subscript form and where it stops being enough.
The practical rule that follows from all of the above: keep the pipeline body straight-line dataflow, and move data-dependent control inside a UDF, where the tracer treats the whole thing as one stage and makes no claim about its internals. A branch the graph cannot see is better hidden than half-drawn.
The Python on this page is a specification format that happens to be Python
syntax. The shipped dreamlake package exports no pipeline, batch,
to_dataset or requeue; lakeshore.udf accepts no kind= argument; and
lakeshore.types does not exist at all. These files trace precisely because
the tracer never imports or executes them. Do not expect them to run.