Code snapshots
The type
's3' | 'github' are not the same kind of thing. s3 is a protocol —
compare Storage.kind, which is s3 / s3-prefix, consistently one axis.
github is a vendor, and choosing it changes something else entirely:
whether an archive exists at all.
The damage shows up in the neighbouring fields. archiveKey means "S3 object
key" under one value and "GitHub ref name" under the other — one field, two
incompatible types, discriminated by a sibling. storageName is documented as
"null for the github backend", so its presence depends on another field's
value. Both are the signature of an enum doing two jobs.
Underneath are two independent questions that backend welds together:
where the bytes live (an object store, addressed by key) and how the
worker obtains them (fetch a presigned URL, or clone from a third party with
its own credentials). The second is not a storage detail — it decides whether
the worker needs GitHub access at all, which is a different security posture,
not a different bucket.
The tarball proposal removes the field
rather than fixing it: if the tar is the payload, it is inline or it is a ref,
exactly like every other payload on the wire, and there is nothing left for
backend to discriminate.
The identity is the triple — @@unique([namespaceId, remote, commit]). That is
the dedup key, and the reason a second submit from the same tree uploads
nothing. branch is deliberately outside it: the same commit reached from two
branches is one snapshot.
A remote worker needs your code before it can run your function. A code
snapshot is one captured tree of one repo — tarred, uploaded once, and
reused by every invocation that shares its commit. You never create one
explicitly; submitting a .remote() call does it for you.
The mechanism is deliberately dull. On submit, the SDK asks git for the current
remote and commit, checks whether that pair is already archived, and uploads a
git archive tarball if it isn't. The commit is the cache key, so the second
submit from the same tree costs one lookup and no bytes.
Coming from a jaynes config
In jaynes, shipping your code is a mount — you declare it, and it re-tars on every launch:
Here it is not a mount and not declared at all. Submitting a .remote() call
creates the snapshot; there is no entry to write.
| jaynes | What it expresses | Lakeshore |
|---|---|---|
!mounts.SSHCode / S3Code / GSCode | Which channel carries the code | backend, loosely — though it also decides whether an archive exists at all |
local_path: . | Which tree to ship | The git repo you submit from; there is no path to choose |
host_path, container_path | Where it lands, where it appears | Worker-managed cache, keyed by commit. Not configurable. |
local_tar, remote_tar | Where the tarball is staged | archiveKey in the Storage named by storageName |
pypath: true | Add it to PYTHONPATH | Automatic — the worker inserts the extracted tree on ImportError |
excludes, exclude_vcs, file_mask | Tar shaping | git archive decides; no exclude knobs are exposed |
compress | Whether to gzip | Not exposed |
{now:…} in paths | Making each launch distinct | commit — the content decides identity, not the clock |
Three differences change how you work, not just how you configure:
Identity is the commit, not the path. Jaynes gives every launch a fresh
timestamped directory, so the same tree uploads again on every run. Here the key
is (remote, commit), so the second submit from an unchanged tree uploads
nothing at all. The {now:…} templating that makes jaynes paths unique has no
equivalent, and needs none.
It has to be a git repo. Jaynes tars whatever local_path points at. A
snapshot is git archive over a real remote and commit, so a directory that is
not a repo has nothing to capture.
Uncommitted work does not travel. This is the one that bites when porting.
A jaynes SSHCode mount cheerfully tars your working tree, edits and all. Here
a dirty tree currently produces no snapshot and no error — see
Uncommitted changes below. A jaynes workflow that
depended on shipping uncommitted edits needs a commit step.
What the worker actually reads
This is the part worth internalizing, because it explains every other behaviour on this page. The snapshot collection is not on the execution path. At submit time the control plane stamps four fields into the invocation envelope:
The worker reads that stamp, and fetches the archive only if the import fails — then caches the extracted tree by commit. So a warm worker running your fifth invocation of the same commit touches neither S3 nor the control plane. The collection exists for deduplication and for inspection, not for dispatch.
Why ImportError is the trigger. A worker built from an image that
already contains your package imports it directly and never downloads
anything. The archive is the fallback for the case where it can't — which
makes baking code into the image and shipping it per-run the same code path,
with no flag to set.
The interface
Python
Tars the working tree, presigns a PUT against storage_name, uploads, and
registers the snapshot. Returns the snapshot dict.
Returns None — meaning no snapshot was made — in two cases: the working
directory is not a git repo, or the tree is dirty and allow_dirty is
false. See Uncommitted changes below, because the
second case is the one that will surprise you.
Repeat calls within one process short-circuit on an in-memory set keyed
remote:commit, so a script submitting a thousand invocations archives once.
CLI
HTTP
Four routes under the config plane, bearer-authenticated. The path segment is
code-repos for historical reasons — the records are snapshots.
| Verb | Path | Purpose |
|---|---|---|
POST | /v1/namespaces/:ns/code-repos | Register an uploaded archive |
GET | /v1/namespaces/:ns/code-repos | List, newest first — ?remote=, ?limit= |
GET | /v1/namespaces/:ns/code-repos/lookup | Dedup probe — ?remote=&commit=, 200 or 404 |
GET | /v1/namespaces/:ns/code-repos/:id | Fetch one |
limit defaults to 50 and caps at 200. remote is matched lowercase. There is
no delete route — snapshots are append-only.
Snapshots are unique on (namespace, remote, commit). The same commit pushed
twice is one record; the same commit in two namespaces is two.
Uncommitted changes
push_if_needed returns None when the working tree has uncommitted
changes, and the submit path treats that the same as an archive-upload
failure — deliberately, so that a storage outage doesn't fail your job. The
consequence is that an uncommitted edit and a broken archive channel look
identical, and neither one fails the submit.
The invocation still runs. It runs against whatever code the worker already has, which is usually the last commit you pushed — not what you are looking at in your editor.
Until this changes, the reliable habit is to commit before a remote submit.
lakeshore code list shows what was actually archived, and the Dirty column
tells you whether a snapshot captured uncommitted work.
The fix is designed and partially built — snapshot_dirty_tree() creates a
detached commit via write-tree / commit-tree without touching your branch,
and the schema already reserves a synthetic-hash convention and a dirty flag
for it. What is missing is the wire between them.
Where this is going
Two changes are planned. Both are visible in the interface, so they are worth knowing before you build against it.
The record splits in two. One model currently carries both repo identity
(remote, backend, storageName) and per-capture state (commit, branch,
dirty, archiveKey, sizeBytes). Splitting them gives a Repo with a
snapshot history, and gives dirty captures an honest commit field — today a
dirty snapshot writes a synthetic hash into a field named commit, so "which
commit did this run use" is unanswerable for exactly the runs where it matters
most. The key becomes contentHash: the commit when clean, a hash of the tar
when dirty.
Snapshots get a lifecycle. Today a record is written after upload and never
revisited — nothing verifies that the archive still exists, so a lifecycle-purged
object surfaces as an ImportError on a worker mid-job. The planned states:
| State | Meaning |
|---|---|
pending | Record written, upload not yet confirmed |
available | Archive uploaded and addressable |
verified | Object confirmed present at verifiedAt |
missing | Object gone — lifecycle rule, bucket change, manual delete |
failed | Upload aborted, or registration raced |
Existing records migrate to available with contentHash = commit. The
envelope keeps its four fields, so nothing you write against the Python or CLI
surface changes shape.
Proposal: the tarball is the payload
The two porting gaps above — it must be a git repo, and uncommitted work does not travel — both dissolve if the archive stops being a consequence of a commit and becomes the thing itself.
Invert it: the tar is the payload; everything else is advisory metadata.
remote and commit move into meta and stop gating anything. A non-repo
directory produces a snapshot with an empty meta; a clean checkout produces
one with all four fields. Both run.
Why the patch is separate, and worth having. If the tar carried your working
tree directly, it would differ on every keystroke and dedup would never hit.
Instead let the tar be a clean base — git archive HEAD, identical across
runs and cached warm on the worker — and let the patch carry only the
uncommitted delta. The tar is big and stable; the patch is small and per-run.
The worker unpacks, applies, imports:
With no repo there is no base to diff against, so the tar holds the tree and
patch is absent. Same path, one step skipped.
Three mechanisms already exist and carry straight over:
- The inline-or-reference split is the payload store's, reused verbatim.
Under
INLINE_THRESHOLD_BYTES(256 KiB default) it rides in the envelope; over it, presign topayloads/. Hard capMAX_PAYLOAD_BYTES, 100 MiB. A patch is almost always inline; a tar is almost always a ref. - The synthetic hash is already reserved — the schema calls
commit"a full commit hash, or a synthetic hash for dirty snapshots," anddirtyexists.contentHashis that idea promoted to the key. - The
ImportErrortrigger and content-keyed cache are untouched.
backend accepts s3 and github and is still load-bearing; there is no
tar, patch, or contentHash field, and snapshot_dirty_tree() — which
would produce the base — is still called by nothing.
What goes in the tar — jaynes has the answer already
A commit defines its own contents. A directory does not, so the tar needs an
include/exclude model, and jaynes already shipped one. It is five knobs feeding a
single tar invocation, identical across SSHCode, S3Code, GSCode, and
TarMount:
| Knob | Role | Default |
|---|---|---|
file_mask | Whitelist — the tar operand naming what to include | "." (everything under the root; may not be empty) |
excludes | Blacklist — a raw string of --exclude= flags | --exclude='*__pycache__' --exclude='*.git' --exclude='*.idea' --exclude='*.egg-info' |
exclude_vcs | Appends --exclude-vcs | True |
exclude_from | Appends --exclude-from=<file>, resolved against RUN.config_root | None |
**tar_options | Arbitrary passthrough — key_name=v becomes --key-name=v | none |
The shape is right and worth adopting: a whitelist selects, blacklists prune, and an ignore file lets a project carry its own rules next to the config. Two details are worth changing rather than copying:
excludes is a raw shell string. It is interpolated straight into the
command, so it is both injection-prone and impossible to validate. A list of
patterns that the client renders into flags gets the same expressiveness and can
be checked.
The defaults overlap, and are still not enough. exclude_vcs=True and the
default --exclude='*.git' both drop VCS metadata. Meanwhile nothing excludes
.venv, node_modules, data/, or checkpoints — which is why every
.jaynes.yml in the starter kit hand-writes a long exclude line, and why several
of them repeat nearly the same list. A sensible default deny-list would delete
most of that configuration.
The cheapest rule that surprises nobody: git archive when there is a repo
— it already honours the repo's own ignore rules — and the tar model above
only when there is not.