Sampling
The pipeline model has no sampler function. There is no sample, no tap,
no take, no head. What it has instead is a subscript: stream[mask]
is the branch you keep and stream[~mask] is the branch you divert. That
single mechanism is the whole of "tapping data off the main stream" in a
Python pipeline, and this page is about what it can and cannot express.
That is the honest answer, and it is worth stating before anything else,
because the obvious guess — that there is a pipeline.sample(0.1) somewhere —
is wrong in a way that costs a reader an afternoon. Real strategy samplers do
exist in DreamLake. They are workflow nodes, not pipeline functions, and
they are hand-authored in a different model entirely. The last two sections of
this page say where they are and why the two models diverge.
The tap
A mask is a boolean selector over rows. Subscripting a table by one selects a subset:
The mechanic worth stating precisely is where in the expression the mask enters. A subscript has two parts, and the tracer treats them differently:
- the value — the thing being subscripted,
traj— keeps the role it already had, which isdata; - the slice — the expression inside the brackets,
ok— is evaluated in mask position, which downgrades its entire provenance tomask.
That asymmetry is the entire reason the renderer draws two different lines into
the same sink. The solid edge is the table whose rows are travelling; the
dashed edge is the table that only decided which rows travel. Neither the
subscript nor the ~ creates a node — both are pure provenance plumbing.
One consequence catches people out: when a node contributes to both the value
and the mask, data wins. An edge carries one role, and a node that is already
shipping rows down an edge does not also get a dashed gate edge alongside it.
The dashed edges in a real trace are therefore only the nodes that touch the
mask and nothing else.
Three ways to build the mask
None of these creates a node. They are expressions over columns, and the tracer evaluates them for provenance only:
A comparison merges the provenance of both sides. &, | and ^ merge the
provenance of both operands. ~ passes provenance straight through — the
complement of a mask descends from exactly the same nodes the mask does, which
is why traj[ok] and traj[~ok] produce identical edge topology and differ
only in the rows that actually flow at run time. The full algebra is at
Review and Masks.
The worked example
The figure is the real <PipelineGraph> component, not a picture of one: drag
a card, scroll to pan, hold ⌘ or ctrl and scroll to zoom. It trims the source
and the batch loop, neither of which adds a tap.
Read what it shows. Two sinks are fed by the same upstream table. Nothing
in the graph distinguishes the kept branch from the diverted one — not the node
kind, not the port, not the column set. fit_trajectory reaches both sinks on
a solid data edge, identically. The only structural difference is the dashed
mask edge from estimate_pose, and even that is symmetric: the complement
gets one too.
That is a real limit, not an oversight. The tracer reads the source and never runs
it, so it cannot know that ok and ~ok partition the stream rather than
duplicating it. The graph is an over-approximation of the dataflow — it records
that the mask reached both sinks and stops there.
estimate_pose is the only node on a dashed edge even though ok reads a
column from traj as well. fit_trajectory is already sending rows down a
solid edge into both sinks, and data dominates mask, so its mask
contribution is absorbed. What survives as a visible gate is precisely the node
whose only contribution was to the predicate.
What the pipeline model cannot express
The mask subscript is a filter, not a sampler, and the difference is not cosmetic. Four things a sampling API would give you have no expression here:
There is no fractional or random sampling. bernoulli(0.1) has no
equivalent. A mask is a predicate over rows — deterministic, evaluated from
column values — not a coin flip, and the tracer has no notion of randomness at
all. Nothing in the graph could record "keep about a tenth of these".
There is no first_n or head. The tracer does not unroll loops and has
no counting construct, so "take the first 100" cannot be expressed in the
graph. A for body traces once regardless of how many times it would run; see
Control Flow.
There is no stratified sampling. There is no notion of a stratify-by key, no per-stratum fraction, and no place in the node or edge schema to put one.
The batch size is not sampling. batch(src, n=32) is a chunking hint for
the run: it says the stream should be processed 32 rows at a time. It adds no
node, it changes no edge, and it does not reduce the stream — every row still
goes through. Reading n=32 as a sample size is an easy mistake to make, and
it is wrong in both directions — it neither selects rows nor bounds them. See
Batching.
Where the real samplers live
DreamLake does have proper samplers, with the statistically correct names and
parameters — bernoulli (fraction, min_size, seed), random_n
(size, with_replacement, seed), stratified (stratify_by, fraction
or per-stratum fractions, min_size, seed) and first_n (size). They
are node types in the hand-authored WorkflowSpec model, not Python, and
they are documented in full at
Workflow Node Types. Do not duplicate their
schema into a pipeline file; there is nothing there to receive it.
The split is a consequence of how each graph comes to exist. A pipeline graph
is derived statically from source — the tracer reads the .py file and can
only describe what the syntax structurally says, and Python syntax has no way
to say "sample 10% with seed 7" that a parser could recognise without running
the program. A workflow spec is authored directly as a typed node graph, so
it can simply name a strategy and its parameters as fields, and the runtime
reads them. One model can only report what the code looks like; the other can
be told what to do.
The authoring API on this page is not implemented in shipped Python.
@dl.pipeline, batch, to_dataset and requeue are not exported by the
dreamlake package; lakeshore.udf accepts no kind= argument; and
lakeshore.types does not exist. These .py files are a specification
format that happens to be Python syntax — they trace because the tracer
parses the source and never imports or executes it.