# Mounting Storage

## The type

A mount is not a store. **`Storage` says where bytes live; `Mount` says how they
are attached to a job.** So the S3-family kinds carry a `storage` reference
rather than re-declaring a bucket, and the connection details stay in one place.

```ts file="Mount"
type SecretRef = { $secret: string }

interface MountBase {
  id:           string
  namespaceId:  string        // the prefix — name is unique within it
  name:         string
  createdAt:    Date
  updatedAt:    Date
}

type Mount = MountBase & (
  // ── Object storage: the target is a registered Storage ──────────────
  | { kind: 's3';           storage: string;  prefix?: string }
  | { kind: 's3fs';         storage: string;  prefix?: string; options?: string }

  // ── Everything else carries its own connection ──────────────────────
  | { kind: 'nfs';          server: string;  path: string;   options?: string }
  | { kind: 'samba';        server: string;  share: string;  options?: string; creds?: SecretRef }
  | { kind: 'ftp';          server: string;  path?: string;  port?: number;    creds?: SecretRef }
  | { kind: 'sftp';         server: string;  path?: string;  port?: number;    creds?: SecretRef; key?: SecretRef }
  | { kind: 'google_drive'; folderId: string; creds?: SecretRef }
  | { kind: 'dropbox';      path: string;     creds?: SecretRef }
  | { kind: 'bind';         hostPath: string; readOnly?: boolean }
  | { kind: 'configmap';    kubeNamespace: string; configMapName: string; items?: Record<string, string> }
)
```

`storage` is a **Storage name**, resolved within the same namespace — the same
by-name reference [code snapshots](/lakeshore/code-snapshots.md) already use in
`storageName` to presign a download.

Like every other Lakeshore resource the identity is the pair —
`@@unique([namespaceId, name])` — so `raw-video` in two namespaces is two
mounts. The union is closed on `kind`, so a writer knows which fields are
required without consulting a second table, and a reader can narrow on `kind`
and get the right fields typed.

### Why the S3 kinds delegate

A `Storage` already carries `bucket`, `prefix`, `endpoint`, `region`, and a
`creds` secret — and it is deduped on `(namespaceId, bucket, prefix, endpoint)`,
so one S3 target is one row. Restating those on the mount would create a second
source of truth for the same bucket, with no constraint keeping them agreed.

It also inherits two things a mount would otherwise have to reimplement:
provisioning (`provision: true` on create calls `CreateBucket`, idempotently)
and the immutability rule — `bucket` / `prefix` / `endpoint` cannot be edited
after creation, because changing an S3 target in place would silently break
every reference to it.

`prefix` still appears on the mount, and it is not a duplicate: `Storage.prefix`
scopes the store, while `Mount.prefix` scopes *this attachment* within it.

> **Note:** `Storage` is S3-compatible object storage — kinds `s3` and `s3-prefix`.
>   There is no registered-resource equivalent for NFS, SMB, FTP, Drive, Dropbox,
>   a host path, or a ConfigMap, so those kinds still carry their own connection
>   fields and their own `$secret` refs. If a shared registry for non-S3 targets
>   is wanted later, this is the seam where it would go.

> **Warning:** **`configmap` has to rename its `name`.** The Kubernetes ConfigMap's own name
>   collided with the mount's `name` once both sat at the root, so it appears
>   above as `configMapName`.

**`[ dev ]`** — a mount can be registered, validated and read back, but nothing
attaches it yet, and the record cannot yet say where it would attach. See
[Providers](/lakeshore/providers.md) for what the marker means, and
[What is missing](#what-is-missing) for the specifics.

  A **mount** is a shared filesystem the runner attaches into a job's workdir
  before dispatch. Where a [code snapshot](/lakeshore/code-snapshots.md) ships the
  code, a mount supplies the data the code expects to find already on disk.

## Coming from a jaynes config

Jaynes describes a mount by **role and direction**. Lakeshore describes one by
**protocol**. That difference is why a jaynes config does not translate
field-for-field.

```yaml file=".jaynes.yml"
mounts:
  - !mounts.SSHCode &code_mount
    local_path: .
    host_path: "{secret.JYNS_HOME}/demo/{now:%Y-%m-%d}"
    container_path: /workspace
    pypath: true
    excludes: >-
      --exclude='data' --exclude='*.git' --exclude='*__pycache__'
    compress: true
```

The class picks the role. `SSHCode`, `S3Code` and `GSCode` carry code **in**;
`S3Output` syncs results **out** on an `interval`; `Host` and `TarMount` attach
something already present.

### The paths

Every jaynes mount names `container_path` — where it appears to the job — and
all but `Host` also name a `local_path` to ship from. Lakeshore answers both,
but by **convention rather than by field**:

| jaynes field | What it expresses | Lakeshore |
| --- | --- | --- |
| `container_path` | Where it appears to the job | **`/mounts/<name>`** — the mount's name is its attach point |
| `host_path` | Staging path on the host | Daemon-managed; not something you set |
| `local_path` | Source on your machine | [Code snapshots](/lakeshore/code-snapshots.md) for code; a `Storage` for data |

That is a real difference in kind, not a missing feature: jaynes makes you name
three paths per mount, and Lakeshore fixes two of them so a mount only has to
say *what* it is. See [Declaring access](/lakeshore/access.md#the-workdir-and-where-things-land)
for the `/workspace` and `/mounts/<name>` layout.

### What genuinely has no equivalent

| jaynes | What it expresses | Lakeshore |
| --- | --- | --- |
| `S3Code` vs `S3Output` | Direction — read in, or sync out | **Nothing.** `kind` names a protocol; a mount has no direction |
| `interval`, `sync_s3` | Periodic push-out cadence | **Nothing.** There is no sync-out mount at all |
| `pypath: true` | Add the mount to `PYTHONPATH` | **Nothing** for mounts. Code snapshots do their own `sys.path` insert |
| `docker_mount_type` | Attach *mechanism* — `bind` / `volume` / `tmpfs` | **Nothing.** Lakeshore's `bind` kind names a *source* (a host directory), which is a different axis |
| `volume`, `mount_path`, `sub_path` | Kubernetes volume + subPath wiring | **Nothing.** The `configmap` kind is the only k8s-aware mount |
| `init_image`, `init_image_pull_policy`, `init_image_pull_secret`, `cpu`, `mem` | The init container that stages the data, and its resources | **Nothing.** `S3Code` and `GSCode` emit a full `init_container` and `volume_mount` spec; Lakeshore emits neither |
| `excludes`, `file_mask`, `compress`, `exclude_vcs`, `exclude_from` | Tar shaping on upload | Belongs to [code snapshots](/lakeshore/code-snapshots.md), not to mounts |
| `{secret.X}` interpolation | Credential injection | `{ "$secret": "name" }` — the nearest thing that **is** implemented |

The Kubernetes rows are the largest omission and the easiest to miss. A jaynes
`S3Code` mount does not merely describe storage — it *emits* a Kubernetes init
container that stages the tarball into a volume, plus the `volumeMount` (with
`subPath`) that exposes it to the job, sized by `cpu` and `mem`. Lakeshore's
Kube launcher runs a pod; nothing in the `Mount` model contributes to its spec.

### Two structural differences

**Jaynes mounts are per-run; Lakeshore mounts are registered resources.** A
`.jaynes.yml` builds its mount list fresh for each launch, which is why
`{now:%Y-%m-%d}` templating appears in the paths at all. A Lakeshore `Mount` is a
durable row under `(namespace, name)` that many jobs reference. Anything per-run
in a jaynes path has nowhere to go on the Lakeshore side.

**Jaynes code mounts are Lakeshore code snapshots, not Lakeshore mounts.**
`SSHCode` / `S3Code` tar a local tree, upload it, and unpack it remotely — which
is exactly what a [code snapshot](/lakeshore/code-snapshots.md) does, and it works
today. Porting a jaynes config, the code mount is the part that already has a
home; the data mounts are the part that does not.

## Kinds

| Kind | Fields at root | Notes |
| --- | --- | --- |
| `s3` | `storage`, `prefix?` | Userspace client — **copy-on-read, no FUSE**. Target comes from the named Storage. |
| `s3fs` | `storage`, `prefix?`, `options?` | The s3fs-fuse driver — a real filesystem path |
| `nfs` | `server`, `path`, `options?` | POSIX network filesystem; `path` is the export |
| `samba` | `server`, `share`, `options?`, `creds?` | CIFS; runner uses `mount -t cifs` |
| `ftp` | `server`, `path?`, `port?`, `creds?` | only `server` is structurally required |
| `sftp` | `server`, `path?`, `port?`, `creds?`, `key?` | runner uses sshfs or curlftpfs |
| `google_drive` | `folderId`, `creds?` | rclone-backed; `folderId` may be `"root"` |
| `dropbox` | `path`, `creds?` | rclone-backed |
| `bind` | `hostPath`, `readOnly?` | Host directory; `hostPath` is the source |
| `configmap` | `kubeNamespace`, `configMapName`, `items?` | Kubernetes; renamed to avoid colliding with the mount's `name` |

> **Note:** Same target, different activation. Both name a Storage; **`s3`** pulls objects into the
>   workdir on demand through a userspace client — higher throughput, simple
>   semantics, no privileges. **`s3fs`** mounts the bucket through FUSE so job
>   code sees a real path. Reach for `s3` unless the code genuinely needs
>   filesystem behaviour it cannot get from a copied tree.

## Credentials

For the S3 kinds there is no credential field on the mount at all — the
`creds` secret belongs to the [Storage](https://lakeshore.dreamlake.ai/admin/storages)
it names, which is the point of delegating.

For every other kind, credentials never sit in the row as plaintext. Use a
`$secret` marker in any credential field and the control plane resolves it
through the same plumbing the providers and tunnels routes use:

```json
{ "kind": "sftp", "server": "lab-nas", "creds": { "$secret": "nas-login" } }
```

The referenced secret must exist in the same namespace, or the write is rejected
with a 422.

Credential refs are **not** required on kinds that can attach anonymously — a
public NFS export, an unauthenticated FTP server. That is deliberate: the runner
enforces it at activation time, where it can give a clearer error than "field
required" would at write time.

## The interface

### Python

There is none. No `Mount` type, no mount argument on `@udf`, nothing in
`RunConfig` that names one. A job cannot request a mount from Python today —
mounts are an operator-side resource registered out of band.

### CLI

```bash file="terminal"
lakeshore mounts add raw-video --kind s3 \
    --kwarg storage=training-data \
    --kwarg prefix=2026/

lakeshore mounts list
lakeshore mounts show raw-video
lakeshore mounts update raw-video --kwarg prefix=2027/
lakeshore mounts remove raw-video
```

`--kwarg key=value` is repeatable and dotted keys nest, so
`creds.$secret=nas-login` builds a nested credential ref. For anything larger,
`--config-file foo.yaml` (or `.json`) is posted as-is. The two combine — the
file seeds the config and `--kwarg` entries override on top — and one of them is
required.

### HTTP

```json file="POST /v1/namespaces/:ns/mounts"
{
  "name": "raw-video",
  "kind": "s3",
  "storage": "training-data",
  "prefix": "2026/"
}
```

> **Warning:** The flattened shape above is the intended one. The route as deployed today
>   accepts `{ name, kind, config }` with the kind-specific fields inside
>   `config`, and the `Mount` row stores `config` as an opaque `Json` column. Both
>   spellings appear in this repo's history; the flat one is where it is going.

The rest is ordinary CRUD over `(namespace, name)`:

```
GET    /v1/namespaces/:ns/mounts
GET    /v1/namespaces/:ns/mounts/:name
PATCH  /v1/namespaces/:ns/mounts/:name
DELETE /v1/namespaces/:ns/mounts/:name
```

## What is missing

The control plane's own header says it: this pass is **metadata-only**. It
accepts the kind-specific config, validates `$secret` refs, and persists the row
as-is. The actual `mount -t <kind> …` — or the rclone / aws-cli equivalent — at
launch time is described there as "a daemon-side follow-up."

Concretely, today:

- **A mount attaches at `/mounts/<name>`, by convention.** No kind's config
  carries a target path, and none needs to — the name is the attach point. What
  is still absent is any way to *override* that for a job expecting a specific
  location.
- **Nothing reads a `Mount`.** The daemon does not fetch mounts, and the Python
  SDK has no concept of one.
- **No job can reference one.** No field on a queue, a mode, or a `RunConfig`
  names a mount, so even a daemon that could attach one would not know which.
  [`mounts=` on the decorator](/lakeshore/access.md#mounts) is the proposed field
  that would close this.

If you need data on a worker now, the working path is the one the SDK already
uses: string keys through `dls.run.read` / `dls.run.write` under a
`dls.scope(...)`, with bytes travelling by reference. See
[Simple functions](/lakeshore/patterns.md).

    The other half of what a worker needs — the code, versioned by commit. This
    is where a jaynes `SSHCode` / `S3Code` mount actually lands.

    What runs on a worker before it takes work.
