# Sampling

  The pipeline model has no sampler function. There is no `sample`, no `tap`,
  no `take`, no `head`. What it has instead is a **subscript**: `stream[mask]`
  is the branch you keep and `stream[~mask]` is the branch you divert. That
  single mechanism is the whole of "tapping data off the main stream" in a
  Python pipeline, and this page is about what it can and cannot express.

That is the honest answer, and it is worth stating before anything else,
because the obvious guess — that there is a `pipeline.sample(0.1)` somewhere —
is wrong in a way that costs a reader an afternoon. Real strategy samplers do
exist in DreamLake. They are **workflow** nodes, not pipeline functions, and
they are hand-authored in a different model entirely. The last two sections of
this page say where they are and why the two models diverge.

## The tap

A mask is a boolean selector over rows. Subscripting a table by one selects a
subset:

```python
save_trajectory(traj[ok])      # the rows where ok is true
requeue_bad(traj[~ok])         # the complement
```

The mechanic worth stating precisely is *where in the expression* the mask
enters. A subscript has two parts, and the tracer treats them differently:

- the **value** — the thing being subscripted, `traj` — keeps the role it
  already had, which is `data`;
- the **slice** — the expression inside the brackets, `ok` — is evaluated in
  mask position, which downgrades its entire provenance to `mask`.

That asymmetry is the entire reason the renderer draws two different lines into
the same sink. The solid edge is the table whose rows are travelling; the
dashed edge is the table that only decided which rows travel. Neither the
subscript nor the `~` creates a node — both are pure provenance plumbing.

One consequence catches people out: when a node contributes to *both* the value
and the mask, `data` wins. An edge carries one role, and a node that is already
shipping rows down an edge does not also get a dashed gate edge alongside it.
The dashed edges in a real trace are therefore only the nodes that touch the
mask **and nothing else**.

## Three ways to build the mask

None of these creates a node. They are expressions over columns, and the tracer
evaluates them for provenance only:

```python
ok        = poses.confidence > 0.5                              # a threshold on a column
ok        = (poses.confidence > 0.5) & (traj.residual < 1.0)    # the conjunction of two
consensus = review.all(axis=0)                                  # a reviewer reduction
```

A comparison merges the provenance of both sides. `&`, `|` and `^` merge the
provenance of both operands. `~` passes provenance straight through — the
complement of a mask descends from exactly the same nodes the mask does, which
is why `traj[ok]` and `traj[~ok]` produce identical edge topology and differ
only in the rows that actually flow at run time. The full algebra is at
[Review and Masks](/pipelines/review-and-masks.md).

## The worked example

```python
@dl.pipeline
def camera_pose_trajectory():
    src = load_frames()
    for items in batch(src, n=8):
        poses = estimate_pose(items.frames)
        traj = fit_trajectory(poses)

        ok = (poses.confidence > 0.5) & (traj.residual < 1.0)
        save_trajectory(traj[ok])      # the kept tap
        requeue_bad(traj[~ok])         # the diverted tap
```

The figure is the real `` component, not a picture of one: drag
a card, scroll to pan, hold ⌘ or ctrl and scroll to zoom. It trims the source
and the batch loop, neither of which adds a tap.

Read what it shows. **Two sinks are fed by the same upstream table.** Nothing
in the graph distinguishes the kept branch from the diverted one — not the node
kind, not the port, not the column set. `fit_trajectory` reaches both sinks on
a solid `data` edge, identically. The only structural difference is the dashed
`mask` edge from `estimate_pose`, and even that is symmetric: the complement
gets one too.

That is a real limit, not an oversight. The tracer reads the source and never runs
it, so it cannot know that `ok` and `~ok` partition the stream rather than
duplicating it. The graph is an over-approximation of the dataflow — it records
that the mask reached both sinks and stops there.

`estimate_pose` is the only node on a dashed edge even though `ok` reads a
column from `traj` as well. `fit_trajectory` is already sending rows down a
solid edge into both sinks, and `data` dominates `mask`, so its mask
contribution is absorbed. What survives as a visible gate is precisely the node
whose *only* contribution was to the predicate.

## What the pipeline model cannot express

The mask subscript is a filter, not a sampler, and the difference is not
cosmetic. Four things a sampling API would give you have no expression here:

**There is no fractional or random sampling.** `bernoulli(0.1)` has no
equivalent. A mask is a predicate over rows — deterministic, evaluated from
column values — not a coin flip, and the tracer has no notion of randomness at
all. Nothing in the graph could record "keep about a tenth of these".

**There is no `first_n` or `head`.** The tracer does not unroll loops and has
no counting construct, so "take the first 100" cannot be expressed in the
graph. A `for` body traces once regardless of how many times it would run; see
[Control Flow](/pipelines/control-flow.md).

**There is no stratified sampling.** There is no notion of a stratify-by key,
no per-stratum fraction, and no place in the node or edge schema to put one.

**The batch size is not sampling.** `batch(src, n=32)` is a chunking hint for
the run: it says the stream should be processed 32 rows at a time. It adds no
node, it changes no edge, and it does not reduce the stream — every row still
goes through. Reading `n=32` as a sample size is an easy mistake to make, and
it is wrong in both directions — it neither selects rows nor bounds them. See
[Batching](/pipelines/batching.md).

## Where the real samplers live

> **Note:** DreamLake does have proper samplers, with the statistically correct names and
>   parameters — `bernoulli` (`fraction`, `min_size`, `seed`), `random_n`
>   (`size`, `with_replacement`, `seed`), `stratified` (`stratify_by`, `fraction`
>   or per-stratum `fractions`, `min_size`, `seed`) and `first_n` (`size`). They
>   are **node types in the hand-authored WorkflowSpec model**, not Python, and
>   they are documented in full at
>   [Workflow Node Types](/workflows/node-types.md#sampler). Do not duplicate their
>   schema into a pipeline file; there is nothing there to receive it.

The split is a consequence of how each graph comes to exist. A pipeline graph
is **derived statically from source** — the tracer reads the `.py` file and can
only describe what the syntax structurally says, and Python syntax has no way
to say "sample 10% with seed 7" that a parser could recognise without running
the program. A workflow spec is **authored directly** as a typed node graph, so
it can simply name a strategy and its parameters as fields, and the runtime
reads them. One model can only report what the code looks like; the other can
be told what to do.

> **Warning:** The authoring API on this page is not implemented in shipped Python.
>   `@dl.pipeline`, `batch`, `to_dataset` and `requeue` are not exported by the
>   `dreamlake` package; `lakeshore.udf` accepts no `kind=` argument; and
>   `lakeshore.types` does not exist. These `.py` files are a specification
>   format that happens to be Python syntax — they trace because the tracer
>   parses the source and never imports or executes it.
