Mounting Storage
The type
A mount is not a store. Storage says where bytes live; Mount says how they
are attached to a job. So the S3-family kinds carry a storage reference
rather than re-declaring a bucket, and the connection details stay in one place.
storage is a Storage name, resolved within the same namespace — the same
by-name reference code snapshots already use in
storageName to presign a download.
Like every other Lakeshore resource the identity is the pair —
@@unique([namespaceId, name]) — so raw-video in two namespaces is two
mounts. The union is closed on kind, so a writer knows which fields are
required without consulting a second table, and a reader can narrow on kind
and get the right fields typed.
Why the S3 kinds delegate
A Storage already carries bucket, prefix, endpoint, region, and a
creds secret — and it is deduped on (namespaceId, bucket, prefix, endpoint),
so one S3 target is one row. Restating those on the mount would create a second
source of truth for the same bucket, with no constraint keeping them agreed.
It also inherits two things a mount would otherwise have to reimplement:
provisioning (provision: true on create calls CreateBucket, idempotently)
and the immutability rule — bucket / prefix / endpoint cannot be edited
after creation, because changing an S3 target in place would silently break
every reference to it.
prefix still appears on the mount, and it is not a duplicate: Storage.prefix
scopes the store, while Mount.prefix scopes this attachment within it.
Storage is S3-compatible object storage — kinds s3 and s3-prefix.
There is no registered-resource equivalent for NFS, SMB, FTP, Drive, Dropbox,
a host path, or a ConfigMap, so those kinds still carry their own connection
fields and their own $secret refs. If a shared registry for non-S3 targets
is wanted later, this is the seam where it would go.
configmap has to rename its name. The Kubernetes ConfigMap's own name
collided with the mount's name once both sat at the root, so it appears
above as configMapName.
[ dev ] — a mount can be registered, validated and read back, but nothing
attaches it yet, and the record cannot yet say where it would attach. See
Providers for what the marker means, and
What is missing for the specifics.
A mount is a shared filesystem the runner attaches into a job's workdir before dispatch. Where a code snapshot ships the code, a mount supplies the data the code expects to find already on disk.
Coming from a jaynes config
Jaynes describes a mount by role and direction. Lakeshore describes one by protocol. That difference is why a jaynes config does not translate field-for-field.
The class picks the role. SSHCode, S3Code and GSCode carry code in;
S3Output syncs results out on an interval; Host and TarMount attach
something already present.
The paths
Every jaynes mount names container_path — where it appears to the job — and
all but Host also name a local_path to ship from. Lakeshore answers both,
but by convention rather than by field:
| jaynes field | What it expresses | Lakeshore |
|---|---|---|
container_path | Where it appears to the job | /mounts/<name> — the mount's name is its attach point |
host_path | Staging path on the host | Daemon-managed; not something you set |
local_path | Source on your machine | Code snapshots for code; a Storage for data |
That is a real difference in kind, not a missing feature: jaynes makes you name
three paths per mount, and Lakeshore fixes two of them so a mount only has to
say what it is. See Declaring access
for the /workspace and /mounts/<name> layout.
What genuinely has no equivalent
| jaynes | What it expresses | Lakeshore |
|---|---|---|
S3Code vs S3Output | Direction — read in, or sync out | Nothing. kind names a protocol; a mount has no direction |
interval, sync_s3 | Periodic push-out cadence | Nothing. There is no sync-out mount at all |
pypath: true | Add the mount to PYTHONPATH | Nothing for mounts. Code snapshots do their own sys.path insert |
docker_mount_type | Attach mechanism — bind / volume / tmpfs | Nothing. Lakeshore's bind kind names a source (a host directory), which is a different axis |
volume, mount_path, sub_path | Kubernetes volume + subPath wiring | Nothing. The configmap kind is the only k8s-aware mount |
init_image, init_image_pull_policy, init_image_pull_secret, cpu, mem | The init container that stages the data, and its resources | Nothing. S3Code and GSCode emit a full init_container and volume_mount spec; Lakeshore emits neither |
excludes, file_mask, compress, exclude_vcs, exclude_from | Tar shaping on upload | Belongs to code snapshots, not to mounts |
{secret.X} interpolation | Credential injection | { "$secret": "name" } — the nearest thing that is implemented |
The Kubernetes rows are the largest omission and the easiest to miss. A jaynes
S3Code mount does not merely describe storage — it emits a Kubernetes init
container that stages the tarball into a volume, plus the volumeMount (with
subPath) that exposes it to the job, sized by cpu and mem. Lakeshore's
Kube launcher runs a pod; nothing in the Mount model contributes to its spec.
Two structural differences
Jaynes mounts are per-run; Lakeshore mounts are registered resources. A
.jaynes.yml builds its mount list fresh for each launch, which is why
{now:%Y-%m-%d} templating appears in the paths at all. A Lakeshore Mount is a
durable row under (namespace, name) that many jobs reference. Anything per-run
in a jaynes path has nowhere to go on the Lakeshore side.
Jaynes code mounts are Lakeshore code snapshots, not Lakeshore mounts.
SSHCode / S3Code tar a local tree, upload it, and unpack it remotely — which
is exactly what a code snapshot does, and it works
today. Porting a jaynes config, the code mount is the part that already has a
home; the data mounts are the part that does not.
Kinds
| Kind | Fields at root | Notes |
|---|---|---|
s3 | storage, prefix? | Userspace client — copy-on-read, no FUSE. Target comes from the named Storage. |
s3fs | storage, prefix?, options? | The s3fs-fuse driver — a real filesystem path |
nfs | server, path, options? | POSIX network filesystem; path is the export |
samba | server, share, options?, creds? | CIFS; runner uses mount -t cifs |
ftp | server, path?, port?, creds? | only server is structurally required |
sftp | server, path?, port?, creds?, key? | runner uses sshfs or curlftpfs |
google_drive | folderId, creds? | rclone-backed; folderId may be "root" |
dropbox | path, creds? | rclone-backed |
bind | hostPath, readOnly? | Host directory; hostPath is the source |
configmap | kubeNamespace, configMapName, items? | Kubernetes; renamed to avoid colliding with the mount's name |
Same target, different activation. Both name a Storage; s3 pulls objects into the
workdir on demand through a userspace client — higher throughput, simple
semantics, no privileges. s3fs mounts the bucket through FUSE so job
code sees a real path. Reach for s3 unless the code genuinely needs
filesystem behaviour it cannot get from a copied tree.
Credentials
For the S3 kinds there is no credential field on the mount at all — the
creds secret belongs to the Storage
it names, which is the point of delegating.
For every other kind, credentials never sit in the row as plaintext. Use a
$secret marker in any credential field and the control plane resolves it
through the same plumbing the providers and tunnels routes use:
The referenced secret must exist in the same namespace, or the write is rejected with a 422.
Credential refs are not required on kinds that can attach anonymously — a public NFS export, an unauthenticated FTP server. That is deliberate: the runner enforces it at activation time, where it can give a clearer error than "field required" would at write time.
The interface
Python
There is none. No Mount type, no mount argument on @udf, nothing in
RunConfig that names one. A job cannot request a mount from Python today —
mounts are an operator-side resource registered out of band.
CLI
--kwarg key=value is repeatable and dotted keys nest, so
creds.$secret=nas-login builds a nested credential ref. For anything larger,
--config-file foo.yaml (or .json) is posted as-is. The two combine — the
file seeds the config and --kwarg entries override on top — and one of them is
required.
HTTP
The flattened shape above is the intended one. The route as deployed today
accepts { name, kind, config } with the kind-specific fields inside
config, and the Mount row stores config as an opaque Json column. Both
spellings appear in this repo's history; the flat one is where it is going.
The rest is ordinary CRUD over (namespace, name):
What is missing
The control plane's own header says it: this pass is metadata-only. It
accepts the kind-specific config, validates $secret refs, and persists the row
as-is. The actual mount -t <kind> … — or the rclone / aws-cli equivalent — at
launch time is described there as "a daemon-side follow-up."
Concretely, today:
- A mount attaches at
/mounts/<name>, by convention. No kind's config carries a target path, and none needs to — the name is the attach point. What is still absent is any way to override that for a job expecting a specific location. - Nothing reads a
Mount. The daemon does not fetch mounts, and the Python SDK has no concept of one. - No job can reference one. No field on a queue, a mode, or a
RunConfignames a mount, so even a daemon that could attach one would not know which.mounts=on the decorator is the proposed field that would close this.
If you need data on a worker now, the working path is the one the SDK already
uses: string keys through dls.run.read / dls.run.write under a
dls.scope(...), with bytes travelling by reference. See
Simple functions.