DreamLake

Mounting Storage

The type

A mount is not a store. Storage says where bytes live; Mount says how they are attached to a job. So the S3-family kinds carry a storage reference rather than re-declaring a bucket, and the connection details stay in one place.

Mountts
type SecretRef = { $secret: string }

interface MountBase {
  id:           string
  namespaceId:  string        // the prefix — name is unique within it
  name:         string
  createdAt:    Date
  updatedAt:    Date
}

type Mount = MountBase & (
  // ── Object storage: the target is a registered Storage ──────────────
  | { kind: 's3';           storage: string;  prefix?: string }
  | { kind: 's3fs';         storage: string;  prefix?: string; options?: string }

  // ── Everything else carries its own connection ──────────────────────
  | { kind: 'nfs';          server: string;  path: string;   options?: string }
  | { kind: 'samba';        server: string;  share: string;  options?: string; creds?: SecretRef }
  | { kind: 'ftp';          server: string;  path?: string;  port?: number;    creds?: SecretRef }
  | { kind: 'sftp';         server: string;  path?: string;  port?: number;    creds?: SecretRef; key?: SecretRef }
  | { kind: 'google_drive'; folderId: string; creds?: SecretRef }
  | { kind: 'dropbox';      path: string;     creds?: SecretRef }
  | { kind: 'bind';         hostPath: string; readOnly?: boolean }
  | { kind: 'configmap';    kubeNamespace: string; configMapName: string; items?: Record<string, string> }
)

storage is a Storage name, resolved within the same namespace — the same by-name reference code snapshots already use in storageName to presign a download.

Like every other Lakeshore resource the identity is the pair — @@unique([namespaceId, name]) — so raw-video in two namespaces is two mounts. The union is closed on kind, so a writer knows which fields are required without consulting a second table, and a reader can narrow on kind and get the right fields typed.

Why the S3 kinds delegate

A Storage already carries bucket, prefix, endpoint, region, and a creds secret — and it is deduped on (namespaceId, bucket, prefix, endpoint), so one S3 target is one row. Restating those on the mount would create a second source of truth for the same bucket, with no constraint keeping them agreed.

It also inherits two things a mount would otherwise have to reimplement: provisioning (provision: true on create calls CreateBucket, idempotently) and the immutability rule — bucket / prefix / endpoint cannot be edited after creation, because changing an S3 target in place would silently break every reference to it.

prefix still appears on the mount, and it is not a duplicate: Storage.prefix scopes the store, while Mount.prefix scopes this attachment within it.

Only the S3 family has a Storage to point at

Storage is S3-compatible object storage — kinds s3 and s3-prefix. There is no registered-resource equivalent for NFS, SMB, FTP, Drive, Dropbox, a host path, or a ConfigMap, so those kinds still carry their own connection fields and their own $secret refs. If a shared registry for non-S3 targets is wanted later, this is the seam where it would go.

One thing flattening forces

configmap has to rename its name. The Kubernetes ConfigMap's own name collided with the mount's name once both sat at the root, so it appears above as configMapName.

[ dev ] — a mount can be registered, validated and read back, but nothing attaches it yet, and the record cannot yet say where it would attach. See Providers for what the marker means, and What is missing for the specifics.

A mount is a shared filesystem the runner attaches into a job's workdir before dispatch. Where a code snapshot ships the code, a mount supplies the data the code expects to find already on disk.

Coming from a jaynes config

Jaynes describes a mount by role and direction. Lakeshore describes one by protocol. That difference is why a jaynes config does not translate field-for-field.

.jaynes.ymlyaml
mounts:
  - !mounts.SSHCode &code_mount
    local_path: .
    host_path: "{secret.JYNS_HOME}/demo/{now:%Y-%m-%d}"
    container_path: /workspace
    pypath: true
    excludes: >-
      --exclude='data' --exclude='*.git' --exclude='*__pycache__'
    compress: true

The class picks the role. SSHCode, S3Code and GSCode carry code in; S3Output syncs results out on an interval; Host and TarMount attach something already present.

The paths

Every jaynes mount names container_path — where it appears to the job — and all but Host also name a local_path to ship from. Lakeshore answers both, but by convention rather than by field:

jaynes fieldWhat it expressesLakeshore
container_pathWhere it appears to the job/mounts/<name> — the mount's name is its attach point
host_pathStaging path on the hostDaemon-managed; not something you set
local_pathSource on your machineCode snapshots for code; a Storage for data

That is a real difference in kind, not a missing feature: jaynes makes you name three paths per mount, and Lakeshore fixes two of them so a mount only has to say what it is. See Declaring access for the /workspace and /mounts/<name> layout.

What genuinely has no equivalent

jaynesWhat it expressesLakeshore
S3Code vs S3OutputDirection — read in, or sync outNothing. kind names a protocol; a mount has no direction
interval, sync_s3Periodic push-out cadenceNothing. There is no sync-out mount at all
pypath: trueAdd the mount to PYTHONPATHNothing for mounts. Code snapshots do their own sys.path insert
docker_mount_typeAttach mechanism — bind / volume / tmpfsNothing. Lakeshore's bind kind names a source (a host directory), which is a different axis
volume, mount_path, sub_pathKubernetes volume + subPath wiringNothing. The configmap kind is the only k8s-aware mount
init_image, init_image_pull_policy, init_image_pull_secret, cpu, memThe init container that stages the data, and its resourcesNothing. S3Code and GSCode emit a full init_container and volume_mount spec; Lakeshore emits neither
excludes, file_mask, compress, exclude_vcs, exclude_fromTar shaping on uploadBelongs to code snapshots, not to mounts
{secret.X} interpolationCredential injection{ "$secret": "name" } — the nearest thing that is implemented

The Kubernetes rows are the largest omission and the easiest to miss. A jaynes S3Code mount does not merely describe storage — it emits a Kubernetes init container that stages the tarball into a volume, plus the volumeMount (with subPath) that exposes it to the job, sized by cpu and mem. Lakeshore's Kube launcher runs a pod; nothing in the Mount model contributes to its spec.

Two structural differences

Jaynes mounts are per-run; Lakeshore mounts are registered resources. A .jaynes.yml builds its mount list fresh for each launch, which is why {now:%Y-%m-%d} templating appears in the paths at all. A Lakeshore Mount is a durable row under (namespace, name) that many jobs reference. Anything per-run in a jaynes path has nowhere to go on the Lakeshore side.

Jaynes code mounts are Lakeshore code snapshots, not Lakeshore mounts. SSHCode / S3Code tar a local tree, upload it, and unpack it remotely — which is exactly what a code snapshot does, and it works today. Porting a jaynes config, the code mount is the part that already has a home; the data mounts are the part that does not.

Kinds

KindFields at rootNotes
s3storage, prefix?Userspace client — copy-on-read, no FUSE. Target comes from the named Storage.
s3fsstorage, prefix?, options?The s3fs-fuse driver — a real filesystem path
nfsserver, path, options?POSIX network filesystem; path is the export
sambaserver, share, options?, creds?CIFS; runner uses mount -t cifs
ftpserver, path?, port?, creds?only server is structurally required
sftpserver, path?, port?, creds?, key?runner uses sshfs or curlftpfs
google_drivefolderId, creds?rclone-backed; folderId may be "root"
dropboxpath, creds?rclone-backed
bindhostPath, readOnly?Host directory; hostPath is the source
configmapkubeNamespace, configMapName, items?Kubernetes; renamed to avoid colliding with the mount's name
s3 versus s3fs

Same target, different activation. Both name a Storage; s3 pulls objects into the workdir on demand through a userspace client — higher throughput, simple semantics, no privileges. s3fs mounts the bucket through FUSE so job code sees a real path. Reach for s3 unless the code genuinely needs filesystem behaviour it cannot get from a copied tree.

Credentials

For the S3 kinds there is no credential field on the mount at all — the creds secret belongs to the Storage it names, which is the point of delegating.

For every other kind, credentials never sit in the row as plaintext. Use a $secret marker in any credential field and the control plane resolves it through the same plumbing the providers and tunnels routes use:

json
{ "kind": "sftp", "server": "lab-nas", "creds": { "$secret": "nas-login" } }

The referenced secret must exist in the same namespace, or the write is rejected with a 422.

Credential refs are not required on kinds that can attach anonymously — a public NFS export, an unauthenticated FTP server. That is deliberate: the runner enforces it at activation time, where it can give a clearer error than "field required" would at write time.

The interface

Python

There is none. No Mount type, no mount argument on @udf, nothing in RunConfig that names one. A job cannot request a mount from Python today — mounts are an operator-side resource registered out of band.

CLI

terminalbash
lakeshore mounts add raw-video --kind s3 \
    --kwarg storage=training-data \
    --kwarg prefix=2026/

lakeshore mounts list
lakeshore mounts show raw-video
lakeshore mounts update raw-video --kwarg prefix=2027/
lakeshore mounts remove raw-video

--kwarg key=value is repeatable and dotted keys nest, so creds.$secret=nas-login builds a nested credential ref. For anything larger, --config-file foo.yaml (or .json) is posted as-is. The two combine — the file seeds the config and --kwarg entries override on top — and one of them is required.

HTTP

POST /v1/namespaces/:ns/mountsjson
{
  "name": "raw-video",
  "kind": "s3",
  "storage": "training-data",
  "prefix": "2026/"
}
The shipped route still nests

The flattened shape above is the intended one. The route as deployed today accepts { name, kind, config } with the kind-specific fields inside config, and the Mount row stores config as an opaque Json column. Both spellings appear in this repo's history; the flat one is where it is going.

The rest is ordinary CRUD over (namespace, name):

GET    /v1/namespaces/:ns/mounts
GET    /v1/namespaces/:ns/mounts/:name
PATCH  /v1/namespaces/:ns/mounts/:name
DELETE /v1/namespaces/:ns/mounts/:name

What is missing

The control plane's own header says it: this pass is metadata-only. It accepts the kind-specific config, validates $secret refs, and persists the row as-is. The actual mount -t <kind> … — or the rclone / aws-cli equivalent — at launch time is described there as "a daemon-side follow-up."

Concretely, today:

  • A mount attaches at /mounts/<name>, by convention. No kind's config carries a target path, and none needs to — the name is the attach point. What is still absent is any way to override that for a job expecting a specific location.
  • Nothing reads a Mount. The daemon does not fetch mounts, and the Python SDK has no concept of one.
  • No job can reference one. No field on a queue, a mode, or a RunConfig names a mount, so even a daemon that could attach one would not know which. mounts= on the decorator is the proposed field that would close this.

If you need data on a worker now, the working path is the one the SDK already uses: string keys through dls.run.read / dls.run.write under a dls.scope(...), with bytes travelling by reference. See Simple functions.

Code snapshots →

The other half of what a worker needs — the code, versioned by commit. This is where a jaynes SSHCode / S3Code mount actually lands.

Host setup →

What runs on a worker before it takes work.