# Declaring access

**`[ draft ]`** — none of these fields exist yet. `udf()` takes exactly `fn`,
`queue`, and `transport`, and nothing enforces what a body reads or writes. See
[Providers](/lakeshore/providers.md) for what the marker means.

  A [scope](/lakeshore/patterns.md) sets *where* keys resolve. It does not say
  *what* a function may reach. These three fields do — as paths **relative to
  the workdir**, the same coordinate system the body already writes in.

## The three fields

```python file="access.py"
@dls.udf(
    queue="gpu",
    source=["raw/**"],          # may read, relative to workdir
    target=["frames/**"],     # may write, relative to workdir
    mounts=["raw-video"],        # filesystems to attach, by name
)
def encode(src: str) -> dict[str, str]:
    ...
```

| Field | Grants | Names |
| --- | --- | --- |
| `source` | Read | A key pattern, relative to the workdir |
| `target` | Write | A key pattern, relative to the workdir |
| `mounts` | Attach | A [Mount](/lakeshore/mounts.md) by name — a resource, not a path |

### Relative to the workdir, not absolute

This is the part worth getting right. `source` and `target` are **not**
[Source](/lakeshore/dreamdb.md) identifiers. They are patterns in the same
coordinate system the body already uses:

```python
@dls.udf(queue="gpu", source=["raw/**"], target=["frames/**"])
def encode(src: str) -> dict[str, str]:
    data = dls.run.read(f"raw/{src}")           # matches source=
    out  = dls.run.write("frames/final.png")    # matches target=
    ...
    return {"frame": "frames/final.png"}
```

The declaration and the body are the same strings. Nothing has to be resolved or
translated to check one against the other, and the returned-key contract — which
already holds every key the body wrote, relative — can compare them directly.

The **scope supplies the absolute half** at invocation — see
[Specifying the prefix](#specifying-the-prefix) below. The declaration only
constrains the shape inside it. Two consequences follow, and both are the reason to do it
this way:

- **The function is portable.** The same `encode` runs against any project. If
  the declaration named `alice@robotics/raw`, it would be pinned to one dataset
  and a second project would need a second function that differed only in a
  string.
- **The graph is project-independent.** A pipeline drawn from declarations is
  the shape of the computation, not a picture of one dataset. The same graph
  describes every run.

## The workdir, and where things land

Two fixed paths on every worker, so nothing has to be configured per job:

```
/workspace              ← the workdir. source= and target= resolve here.
/mounts/<name>          ← each mount in `mounts=`, attached under its own name.
```

A body never writes either path literally. `dls.run.read` / `dls.run.write` take
workdir-relative keys and resolve under `/workspace`; a mount is reached at
`/mounts/<its name>`:

```python file="encode.py"
@dls.udf(queue="gpu", source=["raw/**"], target=["frames/**"], mounts=["raw-video"])
def encode(src: str) -> dict[str, str]:
    data = dls.run.read(f"raw/{src}")          # → /workspace/raw/run01.mp4
    out  = dls.run.write("frames/0000.png")    # → /workspace/frames/0000.png
    ...
    return {"frame": "frames/0000.png"}
```

```
/workspace/
  raw/run01.mp4          ← matched by source: raw/**
  frames/0000.png        ← matched by target: frames/**
/mounts/
  raw-video/             ← the mount named in mounts=["raw-video"]
```

> **Note:** [Mounting Storage](/lakeshore/mounts.md) notes that no mount kind carries a
>   target path, so a registered mount cannot say where it attaches. `/mounts/<name>`
>   answers that without adding a field: the mount's own name *is* its
>   attachment point. That also makes it unique by construction, since mount names
>   are already unique per namespace.

## Specifying the prefix

`/workspace` is the *worker-side* path. What it maps to — which project, which
episode, which stage — is the prefix, and it is assembled from up to three
layers:

```python file="run.py"
with dls.scope("alice@robotics/run-042"):       # 1. the caller, or the pipeline
    encode(src="run01.mp4")
```

```python file="encode.py"
@dls.udf(queue="gpu", workdir="encode")         # 2. a per-UDF sub-prefix
def encode(src: str) -> dict[str, str]:
    ...
```

They **compose, outermost first**:

```
alice@robotics/run-042  /  encode        → mounted at /workspace
      caller or pipeline      decorator
```

So the body's `frames/0000.png` lands at
`alice@robotics/run-042/encode/frames/0000.png`, and the same function run under
a different scope writes somewhere else entirely without editing anything.

> **Warning:** It is tempting to put `workdir: alice@robotics` in `.dreamrc` and stop passing
>   it. Do not — it pins the repository to one project and undoes the portability
>   that made these paths relative in the first place. A repo is not a dataset. The
>   caller or the pipeline knows which project a run belongs to; the code does not,
>   and should not have to be edited to move.
>
>   `.dreamrc` still carries the *relative* defaults below. What it must not carry
>   is the absolute root.

The decorator's `workdir=` is optional and also relative. With neither, the
scope alone is the prefix — which is [what scopes already do
today](/lakeshore/patterns.md), so this adds one inner level rather than replacing
the mechanism.

### Worked example: two stages over one run

```python file="pipeline.py"
import dreamlake.lakeshore as dls

@dls.udf(queue="gpu", workdir="encode",
         source=["raw/**"], target=["frames/**"])
def encode(src: str) -> dict[str, str]:
    ...

@dls.udf(queue="cpu", workdir="label",
         source=["frames/**"], target=["labels/**"])
def label(frame: str) -> dict[str, str]:
    ...

with dls.scope("alice@robotics/run-042"):     # the caller names the project
    frames = encode(src="run01.mp4")
    label(frame=frames["frame"])
```

| Stage | Effective prefix | Reads | Writes |
| --- | --- | --- | --- |
| `encode` | `alice@robotics/run-042/encode` | `raw/**` | `frames/**` |
| `label` | `alice@robotics/run-042/label` | `frames/**` | `labels/**` |

> **Warning:** `encode` writes `frames/**` under its own prefix, and `label` reads `frames/**`
>   under *its* prefix — two different absolute locations. Declared patterns
>   overlap on the relative name, which is what draws the edge in the graph, but
>   the keys the body returns are what actually carry data between them. A stage
>   sub-prefix isolates; it does not connect. If two stages must share a directory,
>   give them the same `workdir=` or none at all.

## Two spellings, one meaning

Every field takes **either a list or a multi-line string**, and the two are
equivalent:

**List**

```python file="access.py"
@dls.udf(queue="gpu", source=["raw/**", "calibration/*.json"])
def encode(src: str) -> dict[str, str]:
    ...
```

**Multi-line string**

```python file="access.py"
@dls.udf(queue="gpu", source="""
    raw/**
    calibration/*.json
""")
def encode(src: str) -> dict[str, str]:
    ...
```

The string form is **dedented automatically**, so the block indents with the
code around it and the leading whitespace never reaches the parser. Lines are
split, stripped, and empties dropped — which means the closing `"""` can sit on
its own line without producing a phantom entry.

```python
"""
    raw/**
    calibration/*.json
"""                             # → ["raw/**", "calibration/*.json"]
```

Both spellings normalize to the same list before anything else sees them, so a
decorator, a `.dreamrc`, and a stored record all hold one shape.

The multi-line form exists because these lists get long and are read far more
often than they are written. A list literal of six entries wraps awkwardly and
diffs badly; six lines in a block diff one line at a time.

## Read and write are different lists

Access has a direction, and one list cannot carry both. A function that reads a
raw capture and writes derived frames should be able to say so — and be denied
the reverse:

```python file="access.py"
@dls.udf(
    queue="gpu",
    source="""
        raw/**
    """,
    target="""
        frames/**
    """,
)
def encode(src: str) -> dict[str, str]:
    ...
```

**Why a separate field rather than a mode on one list.** The compact alternative
is per-entry modes on `source=` — `["raw", "frames:rw"]` — one field, no
duplication. Three things argue against it:

- **Direction is not derivable, and the safe default is wrong.** One list means
  either everything is read-write, which over-permissions the ordinary
  read-raw-write-derived function, or a mode has to be spelled per entry anyway.
- **The enforcement site already exists, and it is write-shaped.** The
  returned-key contract runs on the way out of the run context and already knows
  every key the body wrote. Checking those against a declared write list is the
  same pass with one more predicate — no new interception point. There is no
  equivalent existing hook on the read side.
- **The write list is the one you audit.** Functions read from several places
  and usually write to one. "Which functions can write to `frames`?" should be a
  field lookup, not a substring search for `:rw` inside a list of strings.

The cost is real but small: a read-modify-write source appears in both fields.
That repetition is legible, which is the point.

> **Note:** `source=` and `target=` are singular because each names a **role**, not a
>   count — the source of this function, the target of this function — even though
>   both accept several patterns. They mirror the `src` / `dest` vocabulary the
>   bodies already use.
>
>   `mounts=` stays plural deliberately: it is a set of distinct resources rather
>   than one role, and the entries are names of things that exist independently.

## Mounts

`source` and `target` grant access to keys. `mounts` asks for a
**filesystem** — the S3 bucket or network share a body expects to find already
on disk:

```python file="access.py"
@dls.udf(queue="gpu", mounts=["raw-video"])
def encode(src: str) -> dict[str, str]:
    ...
```

Each entry names a registered [Mount](/lakeshore/mounts.md) in the namespace. The
mount already carries its own target — an S3 mount points at a
[Storage](/lakeshore/mounts.md#the-type), which holds the bucket, region and
credentials — so the UDF names *which* mount, never how to reach it.

That separation is the reason this is a list of names and not a config block: a
credential rotation or an endpoint change happens once on the Storage, and every
UDF referencing it follows without an edit.

> **Warning:** The [Mounting Storage](/lakeshore/mounts.md) page notes that no queue, mode, or
>   `RunConfig` field names a mount, so even a daemon able to attach one would not
>   know which. `mounts=` is that field. Until it exists, a registered mount
>   changes nothing about how a job runs.

## Repo-wide defaults in `.dreamrc`

Most repositories read and write the same handful of shapes everywhere, and repeating
them on every decorator is noise. `.dreamrc` — the YAML config the SDK already
resolves — carries the repo-level set:

```yaml file=".dreamrc"
source:
  - raw/**
  - calibration/*.json

target:
  - frames/**

mounts:
  - raw-video
```

The resolution order is the one the SDK already documents: an explicit path,
then `$DREAMRC`, then `./.dreamrc` walking up to the filesystem root, then
`~/.dreamrc`. So a repo's file sits at its root and applies to every UDF beneath
it.

### How the two combine

**The decorator adds to the file; it does not replace it.** A UDF's effective
access is the union of what `.dreamrc` grants and what the decorator names:

```yaml file=".dreamrc"
source:
  - raw/**
```

```python file="encode.py"
@dls.udf(queue="gpu", source=["calibration/*.json"])
def encode(src: str) -> dict[str, str]:
    ...
    # effective source: raw/** + calibration/*.json
```

Union rather than override, because the alternative surprises in the dangerous
direction. If a decorator replaced the file, adding one narrow entry to a
function would silently revoke everything the repo had granted it, and the
failure would appear at runtime in a body that used to work.

> **Note:** Union means a decorator cannot *reduce* access. That is deliberate — a field
>   that sometimes widens and sometimes narrows is impossible to review. Running
>   with less than the repo grants is a real need, and it gets its own form at the
>   call site rather than an overload here: see
>   [Overrides](#overrides-and-the-direction-they-may-move).

## Why this belongs on the spec, not in the arguments

The obvious objection is that `src` is already a parameter — `encode(src="raw/run01.mp4")`
says what it reads. It does, but only at call time, and that is the whole
problem.

An argument is a **value**; a declaration is a **property of the function**.
Three things need the second and cannot use the first:

- **The graph is static.** The tracer derives a pipeline from source *without
  running it*. If the only statement of what a node reads is an argument passed
  at some future call site, there is nothing to draw until the pipeline has
  already run. Declared access is what makes a node's ports knowable before
  execution.
- **The registry cannot be queried otherwise.** "Which functions write to
  `frames`?" is answerable from a declaration and unanswerable from argument
  values, which are not stored on the function at all.
- **Checking happens too late.** An argument is validated when the body uses it.
  A declaration can be checked at dispatch, before a worker is paid for.

So the argument stays — it says *which* raw capture this call reads. The
declaration says *which family of things this function is ever allowed to read*.
They are different statements and both are needed.

## Wildcards

A declaration that had to name concrete keys would be useless: the scope prefix
is per-invocation, fan-out writes shard names that do not exist yet, and a
function meant to run over every episode in a project cannot enumerate them.
Patterns are the normal case, not an escape hatch.

```python file="access.py"
@dls.udf(
    queue="gpu",
    source="""
        raw/*.mp4                     # one segment, any mp4 directly under raw/
        calibration/**                # any depth under calibration/
    """,
    target="""
        frames/**                     # any depth under frames/
    """,
)
def encode(src: str) -> dict[str, str]:
    ...
```

| Pattern | Matches |
| --- | --- |
| `*` | One segment — `raw/*.mp4` matches `raw/run01.mp4`, not `raw/a/b.mp4` |
| `run-*` | Prefix within a segment — `frames/run-*` |
| `**` | Any depth — `frames/**` covers `frames/a/b.png` |

Enforcement is then "does the concrete key match a declared pattern", which is
the check the returned-key contract is already positioned to make — it runs on
the way out of the run context and already holds every key the body wrote.

> **Warning:** A broad `source` pattern costs you read amplification. A broad `target`
>   pattern is how a function silently overwrites something it was never meant to
>   touch, and it is the one that cannot be undone. A bare `**` in `target`
>   grants the entire workdir — it reads as convenience and behaves as a blank
>   cheque, and is worth rejecting rather than accepting as shorthand.

## Overrides, and the direction they may move

Overrides are wanted — the same function against a dev dataset, a backfill
pointed at last year's project, a test run confined to one episode. The rule
that keeps them safe is that they may only ever **narrow**:

```python file="access.py"
inv = encode.submit(
    src="raw/run01.mp4",
    _source=["raw/run01.mp4"],            # intersected with the declaration
)
```

That gives two layers with deliberately different rules:

| Layer | Combines by | What it is |
| --- | --- | --- |
| `.dreamrc` + decorator | **Union** | The grant. Static, reviewable, what the graph draws. |
| Call-site override | **Intersection** | This run's slice of the grant. Never widens it. |

The asymmetry is the point. Declaration unions because a decorator that replaced
the repo file would silently revoke access and fail at runtime in a body that
used to work. Invocation intersects because a call that could widen its own
permissions is not a permission system — the declaration would become a default
rather than a bound.

An override naming something outside the grant is an error at submit, not a
silent no-op. Failing loudly matters more here than convenience: the caller
asked for something the function is not allowed to do, and quietly clamping to
the intersection would hide that.

## In the node graph

Declared access is what lets the [pipeline graph](/workflows.md) draw a node
without running it, and it maps onto the existing shape rather than needing a
new one:

- **`source` become inbound ports; `target` become outbound ports.** These
  are the `inputs` and `outputs` the node schema already carries.
- **Edges come from overlap.** Where one node's `target` pattern intersects
  another's `source` pattern, there is an edge. The graph stops being an
  inference from dataflow in the body and becomes a join over declarations —
  which is both cheaper and harder to get wrong.
- **Mounts are annotations, not ports.** A mount is an attachment, not a flow
  between nodes. It belongs on the node as a badge; drawing it as an edge would
  imply a producer that does not exist.

Wildcards are where it gets interesting, and they need their own treatment:

**A pattern port is a family, not a thing.** `frames/**` is one declaration that
may correspond to a thousand keys or none. Drawn as an ordinary port it looks
like a single artifact; drawn as nothing it disappears. It wants to render as a
port that visibly stands for a set.

**A pattern edge is "may flow", not "does flow".** Two overlapping patterns
prove a *possible* connection, not an actual one — the concrete keys may never
intersect at runtime. That is a real distinction and the graph already has a
vocabulary for it, since edges are typed. A speculative edge should not look
identical to an observed one.

**Runtime resolves it.** Once invocations exist, the concrete keys are known, so
a run overlay can promote the speculative edges that actually carried data and
grey the ones that did not. The static graph shows what *could* connect; the
overlay shows what did.

> **Note:** Promoting speculative edges requires knowing which invocation produced which
>   key. `Invocation` already carries `correlationId`, `parentInvocationId` and
>   `rootInvocationId` — and nothing writes them today. The static half of this
>   works without that; the overlay does not.

    What a source name resolves to, and how one is created.

    The Mount record `mounts=` names, and the Storage behind it.

    Scopes and the returned-key contract these fields build on.
