DreamLake

Declaring access

[ draft ] — none of these fields exist yet. udf() takes exactly fn, queue, and transport, and nothing enforces what a body reads or writes. See Providers for what the marker means.

A scope sets where keys resolve. It does not say what a function may reach. These three fields do — as paths relative to the workdir, the same coordinate system the body already writes in.

The three fields

access.pypython
@dls.udf(
    queue="gpu",
    source=["raw/**"],          # may read, relative to workdir
    target=["frames/**"],     # may write, relative to workdir
    mounts=["raw-video"],        # filesystems to attach, by name
)
def encode(src: str) -> dict[str, str]:
    ...
FieldGrantsNames
sourceReadA key pattern, relative to the workdir
targetWriteA key pattern, relative to the workdir
mountsAttachA Mount by name — a resource, not a path

Relative to the workdir, not absolute

This is the part worth getting right. source and target are not Source identifiers. They are patterns in the same coordinate system the body already uses:

python
@dls.udf(queue="gpu", source=["raw/**"], target=["frames/**"])
def encode(src: str) -> dict[str, str]:
    data = dls.run.read(f"raw/{src}")           # matches source=
    out  = dls.run.write("frames/final.png")    # matches target=
    ...
    return {"frame": "frames/final.png"}

The declaration and the body are the same strings. Nothing has to be resolved or translated to check one against the other, and the returned-key contract — which already holds every key the body wrote, relative — can compare them directly.

The scope supplies the absolute half at invocation — see Specifying the prefix below. The declaration only constrains the shape inside it. Two consequences follow, and both are the reason to do it this way:

  • The function is portable. The same encode runs against any project. If the declaration named alice@robotics/raw, it would be pinned to one dataset and a second project would need a second function that differed only in a string.
  • The graph is project-independent. A pipeline drawn from declarations is the shape of the computation, not a picture of one dataset. The same graph describes every run.

The workdir, and where things land

Two fixed paths on every worker, so nothing has to be configured per job:

/workspace              ← the workdir. source= and target= resolve here.
/mounts/<name>          ← each mount in `mounts=`, attached under its own name.

A body never writes either path literally. dls.run.read / dls.run.write take workdir-relative keys and resolve under /workspace; a mount is reached at /mounts/<its name>:

encode.pypython
@dls.udf(queue="gpu", source=["raw/**"], target=["frames/**"], mounts=["raw-video"])
def encode(src: str) -> dict[str, str]:
    data = dls.run.read(f"raw/{src}")          # → /workspace/raw/run01.mp4
    out  = dls.run.write("frames/0000.png")    # → /workspace/frames/0000.png
    ...
    return {"frame": "frames/0000.png"}
/workspace/
  raw/run01.mp4          ← matched by source: raw/**
  frames/0000.png        ← matched by target: frames/**
/mounts/
  raw-video/             ← the mount named in mounts=["raw-video"]
This is the mount target, by convention

Mounting Storage notes that no mount kind carries a target path, so a registered mount cannot say where it attaches. /mounts/<name> answers that without adding a field: the mount's own name is its attachment point. That also makes it unique by construction, since mount names are already unique per namespace.

Specifying the prefix

/workspace is the worker-side path. What it maps to — which project, which episode, which stage — is the prefix, and it is assembled from up to three layers:

run.pypython
with dls.scope("alice@robotics/run-042"):       # 1. the caller, or the pipeline
    encode(src="run01.mp4")
encode.pypython
@dls.udf(queue="gpu", workdir="encode")         # 2. a per-UDF sub-prefix
def encode(src: str) -> dict[str, str]:
    ...

They compose, outermost first:

alice@robotics/run-042  /  encode        → mounted at /workspace
      caller or pipeline      decorator

So the body's frames/0000.png lands at alice@robotics/run-042/encode/frames/0000.png, and the same function run under a different scope writes somewhere else entirely without editing anything.

The absolute root is not repo config

It is tempting to put workdir: alice@robotics in .dreamrc and stop passing it. Do not — it pins the repository to one project and undoes the portability that made these paths relative in the first place. A repo is not a dataset. The caller or the pipeline knows which project a run belongs to; the code does not, and should not have to be edited to move.

.dreamrc still carries the relative defaults below. What it must not carry is the absolute root.

The decorator's workdir= is optional and also relative. With neither, the scope alone is the prefix — which is what scopes already do today, so this adds one inner level rather than replacing the mechanism.

Worked example: two stages over one run

pipeline.pypython
import dreamlake.lakeshore as dls

@dls.udf(queue="gpu", workdir="encode",
         source=["raw/**"], target=["frames/**"])
def encode(src: str) -> dict[str, str]:
    ...

@dls.udf(queue="cpu", workdir="label",
         source=["frames/**"], target=["labels/**"])
def label(frame: str) -> dict[str, str]:
    ...

with dls.scope("alice@robotics/run-042"):     # the caller names the project
    frames = encode(src="run01.mp4")
    label(frame=frames["frame"])
StageEffective prefixReadsWrites
encodealice@robotics/run-042/encoderaw/**frames/**
labelalice@robotics/run-042/labelframes/**labels/**
Sibling stages do not share a workdir

encode writes frames/** under its own prefix, and label reads frames/** under its prefix — two different absolute locations. Declared patterns overlap on the relative name, which is what draws the edge in the graph, but the keys the body returns are what actually carry data between them. A stage sub-prefix isolates; it does not connect. If two stages must share a directory, give them the same workdir= or none at all.

Two spellings, one meaning

Every field takes either a list or a multi-line string, and the two are equivalent:

access.pypython
@dls.udf(queue="gpu", source=["raw/**", "calibration/*.json"])
def encode(src: str) -> dict[str, str]:
    ...

The string form is dedented automatically, so the block indents with the code around it and the leading whitespace never reaches the parser. Lines are split, stripped, and empties dropped — which means the closing """ can sit on its own line without producing a phantom entry.

python
"""
    raw/**
    calibration/*.json
"""                             # → ["raw/**", "calibration/*.json"]

Both spellings normalize to the same list before anything else sees them, so a decorator, a .dreamrc, and a stored record all hold one shape.

The multi-line form exists because these lists get long and are read far more often than they are written. A list literal of six entries wraps awkwardly and diffs badly; six lines in a block diff one line at a time.

Read and write are different lists

Access has a direction, and one list cannot carry both. A function that reads a raw capture and writes derived frames should be able to say so — and be denied the reverse:

access.pypython
@dls.udf(
    queue="gpu",
    source="""
        raw/**
    """,
    target="""
        frames/**
    """,
)
def encode(src: str) -> dict[str, str]:
    ...

Why a separate field rather than a mode on one list. The compact alternative is per-entry modes on source= — ["raw", "frames:rw"] — one field, no duplication. Three things argue against it:

  • Direction is not derivable, and the safe default is wrong. One list means either everything is read-write, which over-permissions the ordinary read-raw-write-derived function, or a mode has to be spelled per entry anyway.
  • The enforcement site already exists, and it is write-shaped. The returned-key contract runs on the way out of the run context and already knows every key the body wrote. Checking those against a declared write list is the same pass with one more predicate — no new interception point. There is no equivalent existing hook on the read side.
  • The write list is the one you audit. Functions read from several places and usually write to one. "Which functions can write to frames?" should be a field lookup, not a substring search for :rw inside a list of strings.

The cost is real but small: a read-modify-write source appears in both fields. That repetition is legible, which is the point.

Why these names

source= and target= are singular because each names a role, not a count — the source of this function, the target of this function — even though both accept several patterns. They mirror the src / dest vocabulary the bodies already use.

mounts= stays plural deliberately: it is a set of distinct resources rather than one role, and the entries are names of things that exist independently.

Mounts

source and target grant access to keys. mounts asks for a filesystem — the S3 bucket or network share a body expects to find already on disk:

access.pypython
@dls.udf(queue="gpu", mounts=["raw-video"])
def encode(src: str) -> dict[str, str]:
    ...

Each entry names a registered Mount in the namespace. The mount already carries its own target — an S3 mount points at a Storage, which holds the bucket, region and credentials — so the UDF names which mount, never how to reach it.

That separation is the reason this is a list of names and not a config block: a credential rotation or an endpoint change happens once on the Storage, and every UDF referencing it follows without an edit.

This is the missing consumer

The Mounting Storage page notes that no queue, mode, or RunConfig field names a mount, so even a daemon able to attach one would not know which. mounts= is that field. Until it exists, a registered mount changes nothing about how a job runs.

Repo-wide defaults in .dreamrc

Most repositories read and write the same handful of shapes everywhere, and repeating them on every decorator is noise. .dreamrc — the YAML config the SDK already resolves — carries the repo-level set:

.dreamrcyaml
source:
  - raw/**
  - calibration/*.json

target:
  - frames/**

mounts:
  - raw-video

The resolution order is the one the SDK already documents: an explicit path, then $DREAMRC, then ./.dreamrc walking up to the filesystem root, then ~/.dreamrc. So a repo's file sits at its root and applies to every UDF beneath it.

How the two combine

The decorator adds to the file; it does not replace it. A UDF's effective access is the union of what .dreamrc grants and what the decorator names:

.dreamrcyaml
source:
  - raw/**
encode.pypython
@dls.udf(queue="gpu", source=["calibration/*.json"])
def encode(src: str) -> dict[str, str]:
    ...
    # effective source: raw/** + calibration/*.json

Union rather than override, because the alternative surprises in the dangerous direction. If a decorator replaced the file, adding one narrow entry to a function would silently revoke everything the repo had granted it, and the failure would appear at runtime in a body that used to work.

Narrowing is a separate act

Union means a decorator cannot reduce access. That is deliberate — a field that sometimes widens and sometimes narrows is impossible to review. Running with less than the repo grants is a real need, and it gets its own form at the call site rather than an overload here: see Overrides.

Why this belongs on the spec, not in the arguments

The obvious objection is that src is already a parameter — encode(src="raw/run01.mp4") says what it reads. It does, but only at call time, and that is the whole problem.

An argument is a value; a declaration is a property of the function. Three things need the second and cannot use the first:

  • The graph is static. The tracer derives a pipeline from source without running it. If the only statement of what a node reads is an argument passed at some future call site, there is nothing to draw until the pipeline has already run. Declared access is what makes a node's ports knowable before execution.
  • The registry cannot be queried otherwise. "Which functions write to frames?" is answerable from a declaration and unanswerable from argument values, which are not stored on the function at all.
  • Checking happens too late. An argument is validated when the body uses it. A declaration can be checked at dispatch, before a worker is paid for.

So the argument stays — it says which raw capture this call reads. The declaration says which family of things this function is ever allowed to read. They are different statements and both are needed.

Wildcards

A declaration that had to name concrete keys would be useless: the scope prefix is per-invocation, fan-out writes shard names that do not exist yet, and a function meant to run over every episode in a project cannot enumerate them. Patterns are the normal case, not an escape hatch.

access.pypython
@dls.udf(
    queue="gpu",
    source="""
        raw/*.mp4                     # one segment, any mp4 directly under raw/
        calibration/**                # any depth under calibration/
    """,
    target="""
        frames/**                     # any depth under frames/
    """,
)
def encode(src: str) -> dict[str, str]:
    ...
PatternMatches
*One segment — raw/*.mp4 matches raw/run01.mp4, not raw/a/b.mp4
run-*Prefix within a segment — frames/run-*
**Any depth — frames/** covers frames/a/b.png

Enforcement is then "does the concrete key match a declared pattern", which is the check the returned-key contract is already positioned to make — it runs on the way out of the run context and already holds every key the body wrote.

Wildcards are not symmetric

A broad source pattern costs you read amplification. A broad target pattern is how a function silently overwrites something it was never meant to touch, and it is the one that cannot be undone. A bare ** in target grants the entire workdir — it reads as convenience and behaves as a blank cheque, and is worth rejecting rather than accepting as shorthand.

Overrides, and the direction they may move

Overrides are wanted — the same function against a dev dataset, a backfill pointed at last year's project, a test run confined to one episode. The rule that keeps them safe is that they may only ever narrow:

access.pypython
inv = encode.submit(
    src="raw/run01.mp4",
    _source=["raw/run01.mp4"],            # intersected with the declaration
)

That gives two layers with deliberately different rules:

LayerCombines byWhat it is
.dreamrc + decoratorUnionThe grant. Static, reviewable, what the graph draws.
Call-site overrideIntersectionThis run's slice of the grant. Never widens it.

The asymmetry is the point. Declaration unions because a decorator that replaced the repo file would silently revoke access and fail at runtime in a body that used to work. Invocation intersects because a call that could widen its own permissions is not a permission system — the declaration would become a default rather than a bound.

An override naming something outside the grant is an error at submit, not a silent no-op. Failing loudly matters more here than convenience: the caller asked for something the function is not allowed to do, and quietly clamping to the intersection would hide that.

In the node graph

Declared access is what lets the pipeline graph draw a node without running it, and it maps onto the existing shape rather than needing a new one:

  • source become inbound ports; target become outbound ports. These are the inputs and outputs the node schema already carries.
  • Edges come from overlap. Where one node's target pattern intersects another's source pattern, there is an edge. The graph stops being an inference from dataflow in the body and becomes a join over declarations — which is both cheaper and harder to get wrong.
  • Mounts are annotations, not ports. A mount is an attachment, not a flow between nodes. It belongs on the node as a badge; drawing it as an edge would imply a producer that does not exist.

Wildcards are where it gets interesting, and they need their own treatment:

A pattern port is a family, not a thing. frames/** is one declaration that may correspond to a thousand keys or none. Drawn as an ordinary port it looks like a single artifact; drawn as nothing it disappears. It wants to render as a port that visibly stands for a set.

A pattern edge is "may flow", not "does flow". Two overlapping patterns prove a possible connection, not an actual one — the concrete keys may never intersect at runtime. That is a real distinction and the graph already has a vocabulary for it, since edges are typed. A speculative edge should not look identical to an observed one.

Runtime resolves it. Once invocations exist, the concrete keys are known, so a run overlay can promote the speculative edges that actually carried data and grey the ones that did not. The static graph shows what could connect; the overlay shows what did.

This needs the lineage wire

Promoting speculative edges requires knowing which invocation produced which key. Invocation already carries correlationId, parentInvocationId and rootInvocationId — and nothing writes them today. The static half of this works without that; the overlay does not.

Sources (DreamDB) →

What a source name resolves to, and how one is created.

Mounting Storage →

The Mount record mounts= names, and the Storage behind it.

Simple functions →

Scopes and the returned-key contract these fields build on.