DreamLake

Sampling

The pipeline model has no sampler function. There is no sample, no tap, no take, no head. What it has instead is a subscript: stream[mask] is the branch you keep and stream[~mask] is the branch you divert. That single mechanism is the whole of "tapping data off the main stream" in a Python pipeline, and this page is about what it can and cannot express.

That is the honest answer, and it is worth stating before anything else, because the obvious guess — that there is a pipeline.sample(0.1) somewhere — is wrong in a way that costs a reader an afternoon. Real strategy samplers do exist in DreamLake. They are workflow nodes, not pipeline functions, and they are hand-authored in a different model entirely. The last two sections of this page say where they are and why the two models diverge.

The tap

A mask is a boolean selector over rows. Subscripting a table by one selects a subset:

python
save_trajectory(traj[ok])      # the rows where ok is true
requeue_bad(traj[~ok])         # the complement

The mechanic worth stating precisely is where in the expression the mask enters. A subscript has two parts, and the tracer treats them differently:

  • the value — the thing being subscripted, traj — keeps the role it already had, which is data;
  • the slice — the expression inside the brackets, ok — is evaluated in mask position, which downgrades its entire provenance to mask.

That asymmetry is the entire reason the renderer draws two different lines into the same sink. The solid edge is the table whose rows are travelling; the dashed edge is the table that only decided which rows travel. Neither the subscript nor the ~ creates a node — both are pure provenance plumbing.

One consequence catches people out: when a node contributes to both the value and the mask, data wins. An edge carries one role, and a node that is already shipping rows down an edge does not also get a dashed gate edge alongside it. The dashed edges in a real trace are therefore only the nodes that touch the mask and nothing else.

Three ways to build the mask

None of these creates a node. They are expressions over columns, and the tracer evaluates them for provenance only:

python
ok        = poses.confidence > 0.5                              # a threshold on a column
ok        = (poses.confidence > 0.5) & (traj.residual < 1.0)    # the conjunction of two
consensus = review.all(axis=0)                                  # a reviewer reduction

A comparison merges the provenance of both sides. &, | and ^ merge the provenance of both operands. ~ passes provenance straight through — the complement of a mask descends from exactly the same nodes the mask does, which is why traj[ok] and traj[~ok] produce identical edge topology and differ only in the rows that actually flow at run time. The full algebra is at Review and Masks.

The worked example

python
@dl.pipeline
def camera_pose_trajectory():
    src = load_frames()
    for items in batch(src, n=8):
        poses = estimate_pose(items.frames)
        traj = fit_trajectory(poses)

        ok = (poses.confidence > 0.5) & (traj.residual < 1.0)
        save_trajectory(traj[ok])      # the kept tap
        requeue_bad(traj[~ok])         # the diverted tap
poses
rows
rows
rows
rows
estimate_pose
transform · 1→1
fit_trajectory
transform · 1→1
save_trajectory
sink · 1→0
requeue_bad
sink · 1→0
one stream, two taps — the dashed edges are the gate, the solid ones carry the rows

The figure is the real <PipelineGraph> component, not a picture of one: drag a card, scroll to pan, hold ⌘ or ctrl and scroll to zoom. It trims the source and the batch loop, neither of which adds a tap.

Read what it shows. Two sinks are fed by the same upstream table. Nothing in the graph distinguishes the kept branch from the diverted one — not the node kind, not the port, not the column set. fit_trajectory reaches both sinks on a solid data edge, identically. The only structural difference is the dashed mask edge from estimate_pose, and even that is symmetric: the complement gets one too.

That is a real limit, not an oversight. The tracer reads the source and never runs it, so it cannot know that ok and ~ok partition the stream rather than duplicating it. The graph is an over-approximation of the dataflow — it records that the mask reached both sinks and stops there.

estimate_pose is the only node on a dashed edge even though ok reads a column from traj as well. fit_trajectory is already sending rows down a solid edge into both sinks, and data dominates mask, so its mask contribution is absorbed. What survives as a visible gate is precisely the node whose only contribution was to the predicate.

What the pipeline model cannot express

The mask subscript is a filter, not a sampler, and the difference is not cosmetic. Four things a sampling API would give you have no expression here:

There is no fractional or random sampling. bernoulli(0.1) has no equivalent. A mask is a predicate over rows — deterministic, evaluated from column values — not a coin flip, and the tracer has no notion of randomness at all. Nothing in the graph could record "keep about a tenth of these".

There is no first_n or head. The tracer does not unroll loops and has no counting construct, so "take the first 100" cannot be expressed in the graph. A for body traces once regardless of how many times it would run; see Control Flow.

There is no stratified sampling. There is no notion of a stratify-by key, no per-stratum fraction, and no place in the node or edge schema to put one.

The batch size is not sampling. batch(src, n=32) is a chunking hint for the run: it says the stream should be processed 32 rows at a time. It adds no node, it changes no edge, and it does not reduce the stream — every row still goes through. Reading n=32 as a sample size is an easy mistake to make, and it is wrong in both directions — it neither selects rows nor bounds them. See Batching.

Where the real samplers live

Samplers are a workflow node, not a pipeline function

DreamLake does have proper samplers, with the statistically correct names and parameters — bernoulli (fraction, min_size, seed), random_n (size, with_replacement, seed), stratified (stratify_by, fraction or per-stratum fractions, min_size, seed) and first_n (size). They are node types in the hand-authored WorkflowSpec model, not Python, and they are documented in full at Workflow Node Types. Do not duplicate their schema into a pipeline file; there is nothing there to receive it.

The split is a consequence of how each graph comes to exist. A pipeline graph is derived statically from source — the tracer reads the .py file and can only describe what the syntax structurally says, and Python syntax has no way to say "sample 10% with seed 7" that a parser could recognise without running the program. A workflow spec is authored directly as a typed node graph, so it can simply name a strategy and its parameters as fields, and the runtime reads them. One model can only report what the code looks like; the other can be told what to do.

Proposal — not importable yet

The authoring API on this page is not implemented in shipped Python. @dl.pipeline, batch, to_dataset and requeue are not exported by the dreamlake package; lakeshore.udf accepts no kind= argument; and lakeshore.types does not exist. These .py files are a specification format that happens to be Python syntax — they trace because the tracer parses the source and never imports or executes it.