# Code snapshots

## The type

```ts file="CodeRepo"
interface CodeRepo {
  id:           string
  namespaceId:  string        // the prefix — dedup is scoped to it
  remote:       string        // git remote URL, lowercased for dedup
  commit:       string        // 40-hex commit, or a synthetic hash when dirty
  branch?:      string        // informational; not part of the dedup key

  backend:      's3' | 'github'   // default 's3'
  archiveKey:   string        // S3 object key, or the GitHub ref name
  storageName?: string        // Storage entry used to presign the download;
                              // null for the github backend

  sizeBytes?:   number        // informational
  submodules?:  Array<{ path: string; commit: string; remote: string }>
  dirty:        boolean       // snapshot includes uncommitted changes

  createdAt:    Date
}
```

> **Warning:** `'s3' | 'github'` are not the same kind of thing. `s3` is a **protocol** —
>   compare `Storage.kind`, which is `s3` / `s3-prefix`, consistently one axis.
>   `github` is a **vendor**, and choosing it changes something else entirely:
>   whether an archive exists at all.
>
>   The damage shows up in the neighbouring fields. `archiveKey` means "S3 object
>   key" under one value and "GitHub ref name" under the other — one field, two
>   incompatible types, discriminated by a sibling. `storageName` is documented as
>   "null for the github backend", so its presence depends on another field's
>   value. Both are the signature of an enum doing two jobs.
>
>   Underneath are two independent questions that `backend` welds together:
>   **where the bytes live** (an object store, addressed by key) and **how the
>   worker obtains them** (fetch a presigned URL, or clone from a third party with
>   its own credentials). The second is not a storage detail — it decides whether
>   the worker needs GitHub access at all, which is a different security posture,
>   not a different bucket.
>
>   [The tarball proposal](#proposal-the-tarball-is-the-payload) removes the field
>   rather than fixing it: if the tar is the payload, it is inline or it is a ref,
>   exactly like every other payload on the wire, and there is nothing left for
>   `backend` to discriminate.

The identity is the triple — `@@unique([namespaceId, remote, commit])`. That is
the dedup key, and the reason a second submit from the same tree uploads
nothing. `branch` is deliberately outside it: the same commit reached from two
branches is one snapshot.

  A remote worker needs your code before it can run your function. A **code
  snapshot** is one captured tree of one repo — tarred, uploaded once, and
  reused by every invocation that shares its commit. You never create one
  explicitly; submitting a `.remote()` call does it for you.

The mechanism is deliberately dull. On submit, the SDK asks git for the current
remote and commit, checks whether that pair is already archived, and uploads a
`git archive` tarball if it isn't. The commit is the cache key, so the second
submit from the same tree costs one lookup and no bytes.

## Coming from a jaynes config

In jaynes, shipping your code **is a mount** — you declare it, and it re-tars on
every launch:

```yaml file=".jaynes.yml"
mounts:
  - !mounts.SSHCode &code_mount
    local_path: .
    host_path: "{secret.JYNS_HOME}/demo/{now:%Y-%m-%d}/{now:%H%M%S.%f}"
    remote_tar: "{secret.JYNS_HOME}/demo/{now:%Y-%m-%d}.tar"
    pypath: true
    excludes: >-
      --exclude='data' --exclude='*.git' --exclude='*__pycache__'
    compress: true
```

Here it is not a mount and not declared at all. Submitting a `.remote()` call
creates the snapshot; there is no entry to write.

| jaynes | What it expresses | Lakeshore |
| --- | --- | --- |
| `!mounts.SSHCode` / `S3Code` / `GSCode` | Which channel carries the code | `backend`, loosely — though it also decides whether an archive exists at all |
| `local_path: .` | Which tree to ship | The git repo you submit from; there is no path to choose |
| `host_path`, `container_path` | Where it lands, where it appears | Worker-managed cache, keyed by commit. **Not configurable.** |
| `local_tar`, `remote_tar` | Where the tarball is staged | `archiveKey` in the [Storage](/lakeshore/mounts.md) named by `storageName` |
| `pypath: true` | Add it to `PYTHONPATH` | Automatic — the worker inserts the extracted tree on `ImportError` |
| `excludes`, `exclude_vcs`, `file_mask` | Tar shaping | `git archive` decides; no exclude knobs are exposed |
| `compress` | Whether to gzip | Not exposed |
| `{now:…}` in paths | Making each launch distinct | `commit` — the content decides identity, not the clock |

Three differences change how you work, not just how you configure:

**Identity is the commit, not the path.** Jaynes gives every launch a fresh
timestamped directory, so the same tree uploads again on every run. Here the key
is `(remote, commit)`, so the second submit from an unchanged tree uploads
nothing at all. The `{now:…}` templating that makes jaynes paths unique has no
equivalent, and needs none.

**It has to be a git repo.** Jaynes tars whatever `local_path` points at. A
snapshot is `git archive` over a real remote and commit, so a directory that is
not a repo has nothing to capture.

**Uncommitted work does not travel.** This is the one that bites when porting.
A jaynes `SSHCode` mount cheerfully tars your working tree, edits and all. Here
a dirty tree currently produces no snapshot **and no error** — see
[Uncommitted changes](#uncommitted-changes) below. A jaynes workflow that
depended on shipping uncommitted edits needs a commit step.

## What the worker actually reads

This is the part worth internalizing, because it explains every other behaviour
on this page. The snapshot **collection is not on the execution path**. At
submit time the control plane stamps four fields into the invocation envelope:

```python
{"remote": ..., "commit": ..., "archiveKey": ..., "storageName": ...}
```

The worker reads that stamp, and fetches the archive **only if the import
fails** — then caches the extracted tree by commit. So a warm worker running
your fifth invocation of the same commit touches neither S3 nor the control
plane. The collection exists for deduplication and for inspection, not for
dispatch.

> **Note:** **Why `ImportError` is the trigger.** A worker built from an image that
>   already contains your package imports it directly and never downloads
>   anything. The archive is the fallback for the case where it can't — which
>   makes baking code into the image and shipping it per-run the same code path,
>   with no flag to set.

## The interface

### Python

```python
from dreamlake.lakeshore import code_archive

code_archive.push_if_needed(
    base_url: str,
    namespace: str,
    token: str | None,
    *,
    storage_name: str = "code-staging",
    allow_dirty: bool = False,
    timeout: float = 60.0,
) -> dict[str, Any] | None
```

Tars the working tree, presigns a `PUT` against `storage_name`, uploads, and
registers the snapshot. Returns the snapshot dict.

Returns `None` — meaning *no snapshot was made* — in two cases: the working
directory is not a git repo, or **the tree is dirty and `allow_dirty` is
false**. See [Uncommitted changes](#uncommitted-changes) below, because the
second case is the one that will surprise you.

Repeat calls within one process short-circuit on an in-memory set keyed
`remote:commit`, so a script submitting a thousand invocations archives once.

### CLI

```bash
lakeshore code push          # archive and register the current tree
lakeshore code list          # list archived code snapshots
```

### HTTP

Four routes under the config plane, bearer-authenticated. The path segment is
`code-repos` for historical reasons — the records are snapshots.

| Verb     | Path                                      | Purpose                                       |
| -------- | ----------------------------------------- | --------------------------------------------- |
| `POST`   | `/v1/namespaces/:ns/code-repos`           | Register an uploaded archive                  |
| `GET`    | `/v1/namespaces/:ns/code-repos`           | List, newest first — `?remote=`, `?limit=`    |
| `GET`    | `/v1/namespaces/:ns/code-repos/lookup`    | Dedup probe — `?remote=&commit=`, 200 or 404  |
| `GET`    | `/v1/namespaces/:ns/code-repos/:id`       | Fetch one                                     |

`limit` defaults to 50 and caps at 200. `remote` is matched lowercase. There is
no delete route — snapshots are append-only.

```ts
// POST body. `remote` and `commit` are required; the rest is optional.
interface CreateCodeSnapshot {
  remote: string          // git remote URL, normalized lowercase
  commit: string          // 40 hex chars, or a synthetic hash for a dirty tree
  branch?: string         // informational — NOT part of the dedup key
  backend?: string        // "s3" | "github", defaults "s3"
  archiveKey?: string     // S3 object key, or a GitHub ref name
  storageName?: string    // which Storage presigns the download; null for GitHub
  sizeBytes?: number
  submodules?: unknown    // [{ path, commit, remote }] at capture time
  dirty?: boolean
}

// Response — the same shape, plus identity and timestamp.
interface CodeSnapshot {
  id: string
  remote: string
  commit: string
  branch: string | null
  backend: string
  archiveKey: string
  storageName: string | null
  sizeBytes: number | null
  submodules: unknown
  dirty: boolean
  createdAt: string
}
```

Snapshots are unique on `(namespace, remote, commit)`. The same commit pushed
twice is one record; the same commit in two namespaces is two.

## Uncommitted changes

> **Warning:** `push_if_needed` returns `None` when the working tree has uncommitted
>   changes, and the submit path treats that the same as an archive-upload
>   failure — deliberately, so that a storage outage doesn't fail your job. The
>   consequence is that **an uncommitted edit and a broken archive channel look
>   identical**, and neither one fails the submit.
>
>   The invocation still runs. It runs against whatever code the worker already
>   has, which is usually the last commit you pushed — not what you are looking
>   at in your editor.

Until this changes, the reliable habit is to commit before a remote submit.
`lakeshore code list` shows what was actually archived, and the `Dirty` column
tells you whether a snapshot captured uncommitted work.

The fix is designed and partially built — `snapshot_dirty_tree()` creates a
detached commit via `write-tree` / `commit-tree` without touching your branch,
and the schema already reserves a synthetic-hash convention and a `dirty` flag
for it. What is missing is the wire between them.

## Where this is going

Two changes are planned. Both are visible in the interface, so they are worth
knowing before you build against it.

**The record splits in two.** One model currently carries both repo identity
(`remote`, `backend`, `storageName`) and per-capture state (`commit`, `branch`,
`dirty`, `archiveKey`, `sizeBytes`). Splitting them gives a `Repo` with a
snapshot history, and gives dirty captures an honest `commit` field — today a
dirty snapshot writes a synthetic hash into a field named `commit`, so *"which
commit did this run use"* is unanswerable for exactly the runs where it matters
most. The key becomes `contentHash`: the commit when clean, a hash of the tar
when dirty.

**Snapshots get a lifecycle.** Today a record is written after upload and never
revisited — nothing verifies that the archive still exists, so a lifecycle-purged
object surfaces as an `ImportError` on a worker mid-job. The planned states:

| State       | Meaning                                                       |
| ----------- | ------------------------------------------------------------- |
| `pending`   | Record written, upload not yet confirmed                      |
| `available` | Archive uploaded and addressable                              |
| `verified`  | Object confirmed present at `verifiedAt`                      |
| `missing`   | Object gone — lifecycle rule, bucket change, manual delete    |
| `failed`    | Upload aborted, or registration raced                         |

Existing records migrate to `available` with `contentHash = commit`. The
envelope keeps its four fields, so nothing you write against the Python or CLI
surface changes shape.

### Proposal: the tarball is the payload

The two porting gaps above — it must be a git repo, and uncommitted work does
not travel — both dissolve if the archive stops being a *consequence* of a
commit and becomes the thing itself.

Invert it: **the tar is the payload; everything else is advisory metadata.**

```ts file="proposed"
interface CodeSnapshot {
  // ── The payload ─────────────────────────────────────────────────────
  tar:      Inline | Ref     // required. base64 under the threshold, else a ref
  patch?:   Inline | Ref     // optional unified diff, applied after unpacking

  contentHash: string        // sha256 of (tar, patch) — the cache key

  // ── Nice-to-have: recorded, never required, never load-bearing ──────
  meta?: {
    remote?:  string
    commit?:  string
    branch?:  string
    dirty?:   boolean
  }
}
```

`remote` and `commit` move into `meta` and stop gating anything. A non-repo
directory produces a snapshot with an empty `meta`; a clean checkout produces
one with all four fields. Both run.

**Why the patch is separate, and worth having.** If the tar carried your working
tree directly, it would differ on every keystroke and dedup would never hit.
Instead let the tar be a **clean base** — `git archive HEAD`, identical across
runs and cached warm on the worker — and let the patch carry only the
uncommitted delta. The tar is big and stable; the patch is small and per-run.
The worker unpacks, applies, imports:

```
unpack tar   →  apply patch (if present)  →  sys.path insert  →  import
```

With no repo there is no base to diff against, so the tar holds the tree and
`patch` is absent. Same path, one step skipped.

Three mechanisms already exist and carry straight over:

- **The inline-or-reference split** is the payload store's, reused verbatim.
  Under `INLINE_THRESHOLD_BYTES` (256 KiB default) it rides in the envelope;
  over it, presign to `payloads/`. Hard cap `MAX_PAYLOAD_BYTES`, 100 MiB. A
  patch is almost always inline; a tar is almost always a ref.
- **The synthetic hash** is already reserved — the schema calls `commit` "a full
  commit hash, or a synthetic hash for dirty snapshots," and `dirty` exists.
  `contentHash` is that idea promoted to the key.
- **The `ImportError` trigger and content-keyed cache** are untouched.

> **Warning:** `backend` accepts `s3` and `github` and is still load-bearing; there is no
>   `tar`, `patch`, or `contentHash` field, and `snapshot_dirty_tree()` — which
>   would produce the base — is still called by nothing.

### What goes in the tar — jaynes has the answer already

A commit defines its own contents. A directory does not, so the tar needs an
include/exclude model, and jaynes already shipped one. It is five knobs feeding a
single `tar` invocation, identical across `SSHCode`, `S3Code`, `GSCode`, and
`TarMount`:

```bash
tar {excludes} {tar_options} -c[z]f {local_tar} -C {local_abs} {file_mask}
```

| Knob | Role | Default |
| --- | --- | --- |
| `file_mask` | **Whitelist** — the tar operand naming what to include | `"."` (everything under the root; may not be empty) |
| `excludes` | **Blacklist** — a raw string of `--exclude=` flags | `--exclude='*__pycache__' --exclude='*.git' --exclude='*.idea' --exclude='*.egg-info'` |
| `exclude_vcs` | Appends `--exclude-vcs` | `True` |
| `exclude_from` | Appends `--exclude-from=<file>`, resolved against `RUN.config_root` | `None` |
| `**tar_options` | Arbitrary passthrough — `key_name=v` becomes `--key-name=v` | none |

The shape is right and worth adopting: a whitelist selects, blacklists prune, and
an ignore file lets a project carry its own rules next to the config. Two
details are worth changing rather than copying:

**`excludes` is a raw shell string.** It is interpolated straight into the
command, so it is both injection-prone and impossible to validate. A list of
patterns that the client renders into flags gets the same expressiveness and can
be checked.

**The defaults overlap, and are still not enough.** `exclude_vcs=True` and the
default `--exclude='*.git'` both drop VCS metadata. Meanwhile nothing excludes
`.venv`, `node_modules`, `data/`, or checkpoints — which is why every
`.jaynes.yml` in the starter kit hand-writes a long exclude line, and why several
of them repeat nearly the same list. A sensible default deny-list would delete
most of that configuration.

The cheapest rule that surprises nobody: **`git archive` when there is a repo**
— it already honours the repo's own ignore rules — **and the tar model above
only when there is not.**
