DreamLake

Code snapshots

The type

CodeRepots
interface CodeRepo {
  id:           string
  namespaceId:  string        // the prefix — dedup is scoped to it
  remote:       string        // git remote URL, lowercased for dedup
  commit:       string        // 40-hex commit, or a synthetic hash when dirty
  branch?:      string        // informational; not part of the dedup key

  backend:      's3' | 'github'   // default 's3'
  archiveKey:   string        // S3 object key, or the GitHub ref name
  storageName?: string        // Storage entry used to presign the download;
                              // null for the github backend

  sizeBytes?:   number        // informational
  submodules?:  Array<{ path: string; commit: string; remote: string }>
  dirty:        boolean       // snapshot includes uncommitted changes

  createdAt:    Date
}
`backend` mixes two axes and should not exist

's3' | 'github' are not the same kind of thing. s3 is a protocol — compare Storage.kind, which is s3 / s3-prefix, consistently one axis. github is a vendor, and choosing it changes something else entirely: whether an archive exists at all.

The damage shows up in the neighbouring fields. archiveKey means "S3 object key" under one value and "GitHub ref name" under the other — one field, two incompatible types, discriminated by a sibling. storageName is documented as "null for the github backend", so its presence depends on another field's value. Both are the signature of an enum doing two jobs.

Underneath are two independent questions that backend welds together: where the bytes live (an object store, addressed by key) and how the worker obtains them (fetch a presigned URL, or clone from a third party with its own credentials). The second is not a storage detail — it decides whether the worker needs GitHub access at all, which is a different security posture, not a different bucket.

The tarball proposal removes the field rather than fixing it: if the tar is the payload, it is inline or it is a ref, exactly like every other payload on the wire, and there is nothing left for backend to discriminate.

The identity is the triple — @@unique([namespaceId, remote, commit]). That is the dedup key, and the reason a second submit from the same tree uploads nothing. branch is deliberately outside it: the same commit reached from two branches is one snapshot.

A remote worker needs your code before it can run your function. A code snapshot is one captured tree of one repo — tarred, uploaded once, and reused by every invocation that shares its commit. You never create one explicitly; submitting a .remote() call does it for you.

The mechanism is deliberately dull. On submit, the SDK asks git for the current remote and commit, checks whether that pair is already archived, and uploads a git archive tarball if it isn't. The commit is the cache key, so the second submit from the same tree costs one lookup and no bytes.

Coming from a jaynes config

In jaynes, shipping your code is a mount — you declare it, and it re-tars on every launch:

.jaynes.ymlyaml
mounts:
  - !mounts.SSHCode &code_mount
    local_path: .
    host_path: "{secret.JYNS_HOME}/demo/{now:%Y-%m-%d}/{now:%H%M%S.%f}"
    remote_tar: "{secret.JYNS_HOME}/demo/{now:%Y-%m-%d}.tar"
    pypath: true
    excludes: >-
      --exclude='data' --exclude='*.git' --exclude='*__pycache__'
    compress: true

Here it is not a mount and not declared at all. Submitting a .remote() call creates the snapshot; there is no entry to write.

jaynesWhat it expressesLakeshore
!mounts.SSHCode / S3Code / GSCodeWhich channel carries the codebackend, loosely — though it also decides whether an archive exists at all
local_path: .Which tree to shipThe git repo you submit from; there is no path to choose
host_path, container_pathWhere it lands, where it appearsWorker-managed cache, keyed by commit. Not configurable.
local_tar, remote_tarWhere the tarball is stagedarchiveKey in the Storage named by storageName
pypath: trueAdd it to PYTHONPATHAutomatic — the worker inserts the extracted tree on ImportError
excludes, exclude_vcs, file_maskTar shapinggit archive decides; no exclude knobs are exposed
compressWhether to gzipNot exposed
{now:…} in pathsMaking each launch distinctcommit — the content decides identity, not the clock

Three differences change how you work, not just how you configure:

Identity is the commit, not the path. Jaynes gives every launch a fresh timestamped directory, so the same tree uploads again on every run. Here the key is (remote, commit), so the second submit from an unchanged tree uploads nothing at all. The {now:…} templating that makes jaynes paths unique has no equivalent, and needs none.

It has to be a git repo. Jaynes tars whatever local_path points at. A snapshot is git archive over a real remote and commit, so a directory that is not a repo has nothing to capture.

Uncommitted work does not travel. This is the one that bites when porting. A jaynes SSHCode mount cheerfully tars your working tree, edits and all. Here a dirty tree currently produces no snapshot and no error — see Uncommitted changes below. A jaynes workflow that depended on shipping uncommitted edits needs a commit step.

What the worker actually reads

This is the part worth internalizing, because it explains every other behaviour on this page. The snapshot collection is not on the execution path. At submit time the control plane stamps four fields into the invocation envelope:

python
{"remote": ..., "commit": ..., "archiveKey": ..., "storageName": ...}

The worker reads that stamp, and fetches the archive only if the import fails — then caches the extracted tree by commit. So a warm worker running your fifth invocation of the same commit touches neither S3 nor the control plane. The collection exists for deduplication and for inspection, not for dispatch.

Why ImportError is the trigger. A worker built from an image that already contains your package imports it directly and never downloads anything. The archive is the fallback for the case where it can't — which makes baking code into the image and shipping it per-run the same code path, with no flag to set.

The interface

Python

python
from dreamlake.lakeshore import code_archive

code_archive.push_if_needed(
    base_url: str,
    namespace: str,
    token: str | None,
    *,
    storage_name: str = "code-staging",
    allow_dirty: bool = False,
    timeout: float = 60.0,
) -> dict[str, Any] | None

Tars the working tree, presigns a PUT against storage_name, uploads, and registers the snapshot. Returns the snapshot dict.

Returns None — meaning no snapshot was made — in two cases: the working directory is not a git repo, or the tree is dirty and allow_dirty is false. See Uncommitted changes below, because the second case is the one that will surprise you.

Repeat calls within one process short-circuit on an in-memory set keyed remote:commit, so a script submitting a thousand invocations archives once.

CLI

bash
lakeshore code push          # archive and register the current tree
lakeshore code list          # list archived code snapshots

HTTP

Four routes under the config plane, bearer-authenticated. The path segment is code-repos for historical reasons — the records are snapshots.

VerbPathPurpose
POST/v1/namespaces/:ns/code-reposRegister an uploaded archive
GET/v1/namespaces/:ns/code-reposList, newest first — ?remote=, ?limit=
GET/v1/namespaces/:ns/code-repos/lookupDedup probe — ?remote=&commit=, 200 or 404
GET/v1/namespaces/:ns/code-repos/:idFetch one

limit defaults to 50 and caps at 200. remote is matched lowercase. There is no delete route — snapshots are append-only.

ts
// POST body. `remote` and `commit` are required; the rest is optional.
interface CreateCodeSnapshot {
  remote: string          // git remote URL, normalized lowercase
  commit: string          // 40 hex chars, or a synthetic hash for a dirty tree
  branch?: string         // informational — NOT part of the dedup key
  backend?: string        // "s3" | "github", defaults "s3"
  archiveKey?: string     // S3 object key, or a GitHub ref name
  storageName?: string    // which Storage presigns the download; null for GitHub
  sizeBytes?: number
  submodules?: unknown    // [{ path, commit, remote }] at capture time
  dirty?: boolean
}

// Response — the same shape, plus identity and timestamp.
interface CodeSnapshot {
  id: string
  remote: string
  commit: string
  branch: string | null
  backend: string
  archiveKey: string
  storageName: string | null
  sizeBytes: number | null
  submodules: unknown
  dirty: boolean
  createdAt: string
}

Snapshots are unique on (namespace, remote, commit). The same commit pushed twice is one record; the same commit in two namespaces is two.

Uncommitted changes

A dirty tree currently produces no snapshot, and no error

push_if_needed returns None when the working tree has uncommitted changes, and the submit path treats that the same as an archive-upload failure — deliberately, so that a storage outage doesn't fail your job. The consequence is that an uncommitted edit and a broken archive channel look identical, and neither one fails the submit.

The invocation still runs. It runs against whatever code the worker already has, which is usually the last commit you pushed — not what you are looking at in your editor.

Until this changes, the reliable habit is to commit before a remote submit. lakeshore code list shows what was actually archived, and the Dirty column tells you whether a snapshot captured uncommitted work.

The fix is designed and partially built — snapshot_dirty_tree() creates a detached commit via write-tree / commit-tree without touching your branch, and the schema already reserves a synthetic-hash convention and a dirty flag for it. What is missing is the wire between them.

Where this is going

Two changes are planned. Both are visible in the interface, so they are worth knowing before you build against it.

The record splits in two. One model currently carries both repo identity (remote, backend, storageName) and per-capture state (commit, branch, dirty, archiveKey, sizeBytes). Splitting them gives a Repo with a snapshot history, and gives dirty captures an honest commit field — today a dirty snapshot writes a synthetic hash into a field named commit, so "which commit did this run use" is unanswerable for exactly the runs where it matters most. The key becomes contentHash: the commit when clean, a hash of the tar when dirty.

Snapshots get a lifecycle. Today a record is written after upload and never revisited — nothing verifies that the archive still exists, so a lifecycle-purged object surfaces as an ImportError on a worker mid-job. The planned states:

StateMeaning
pendingRecord written, upload not yet confirmed
availableArchive uploaded and addressable
verifiedObject confirmed present at verifiedAt
missingObject gone — lifecycle rule, bucket change, manual delete
failedUpload aborted, or registration raced

Existing records migrate to available with contentHash = commit. The envelope keeps its four fields, so nothing you write against the Python or CLI surface changes shape.

Proposal: the tarball is the payload

The two porting gaps above — it must be a git repo, and uncommitted work does not travel — both dissolve if the archive stops being a consequence of a commit and becomes the thing itself.

Invert it: the tar is the payload; everything else is advisory metadata.

proposedts
interface CodeSnapshot {
  // ── The payload ─────────────────────────────────────────────────────
  tar:      Inline | Ref     // required. base64 under the threshold, else a ref
  patch?:   Inline | Ref     // optional unified diff, applied after unpacking

  contentHash: string        // sha256 of (tar, patch) — the cache key

  // ── Nice-to-have: recorded, never required, never load-bearing ──────
  meta?: {
    remote?:  string
    commit?:  string
    branch?:  string
    dirty?:   boolean
  }
}

remote and commit move into meta and stop gating anything. A non-repo directory produces a snapshot with an empty meta; a clean checkout produces one with all four fields. Both run.

Why the patch is separate, and worth having. If the tar carried your working tree directly, it would differ on every keystroke and dedup would never hit. Instead let the tar be a clean base — git archive HEAD, identical across runs and cached warm on the worker — and let the patch carry only the uncommitted delta. The tar is big and stable; the patch is small and per-run. The worker unpacks, applies, imports:

unpack tar   →  apply patch (if present)  →  sys.path insert  →  import

With no repo there is no base to diff against, so the tar holds the tree and patch is absent. Same path, one step skipped.

Three mechanisms already exist and carry straight over:

  • The inline-or-reference split is the payload store's, reused verbatim. Under INLINE_THRESHOLD_BYTES (256 KiB default) it rides in the envelope; over it, presign to payloads/. Hard cap MAX_PAYLOAD_BYTES, 100 MiB. A patch is almost always inline; a tar is almost always a ref.
  • The synthetic hash is already reserved — the schema calls commit "a full commit hash, or a synthetic hash for dirty snapshots," and dirty exists. contentHash is that idea promoted to the key.
  • The ImportError trigger and content-keyed cache are untouched.
Proposal, not shipped

backend accepts s3 and github and is still load-bearing; there is no tar, patch, or contentHash field, and snapshot_dirty_tree() — which would produce the base — is still called by nothing.

What goes in the tar — jaynes has the answer already

A commit defines its own contents. A directory does not, so the tar needs an include/exclude model, and jaynes already shipped one. It is five knobs feeding a single tar invocation, identical across SSHCode, S3Code, GSCode, and TarMount:

bash
tar {excludes} {tar_options} -c[z]f {local_tar} -C {local_abs} {file_mask}
KnobRoleDefault
file_maskWhitelist — the tar operand naming what to include"." (everything under the root; may not be empty)
excludesBlacklist — a raw string of --exclude= flags--exclude='*__pycache__' --exclude='*.git' --exclude='*.idea' --exclude='*.egg-info'
exclude_vcsAppends --exclude-vcsTrue
exclude_fromAppends --exclude-from=<file>, resolved against RUN.config_rootNone
**tar_optionsArbitrary passthrough — key_name=v becomes --key-name=vnone

The shape is right and worth adopting: a whitelist selects, blacklists prune, and an ignore file lets a project carry its own rules next to the config. Two details are worth changing rather than copying:

excludes is a raw shell string. It is interpolated straight into the command, so it is both injection-prone and impossible to validate. A list of patterns that the client renders into flags gets the same expressiveness and can be checked.

The defaults overlap, and are still not enough. exclude_vcs=True and the default --exclude='*.git' both drop VCS metadata. Meanwhile nothing excludes .venv, node_modules, data/, or checkpoints — which is why every .jaynes.yml in the starter kit hand-writes a long exclude line, and why several of them repeat nearly the same list. A sensible default deny-list would delete most of that configuration.

The cheapest rule that surprises nobody: git archive when there is a repo — it already honours the repo's own ignore rules — and the tar model above only when there is not.