DreamLake

Local dev loop

Everything on this page runs on one machine with no AWS account, no cloud credential, and no download from the release bucket. If you can run Docker, Node, Rust and Python, you can bring up a control plane, install a daemon you compiled yourself, submit a job to it, and watch the result land — start to finish, offline.

That constraint is the point, not a convenience. A bring-up that needs somebody else's cloud account is a bring-up only that person can run, which means it is a bring-up nobody reviews.

What you need, and what you deliberately do not

NeededWhy
DockerOne container, Mongo, with a single-node replica set. The control plane needs transactions.
Node 20+ and pnpmThe control plane and the lakeshore CLI.
A Rust toolchainYou build the nymph binary rather than downloading one.
Python 3.11+The SDK, to submit and to read the result back.
Not neededWhy not
An AWS accountNothing here launches an instance. The daemon is a process on your laptop.
A Heroku token, or any deploy credentialThe control plane runs from source, on loopback.
An admin tokenThe control plane runs in open mode locally.
Any download from the release bucketdist-local.sh produces the exact layout the bucket serves, and the installer reads it over file://.

The loop

Six steps. Every URL below is loopback.

1. Start Mongo

terminalbash
cd lakeshore-workspace
docker compose up -d mongo

The compose healthcheck initiates the rs0 replica set on first start, so there is no separate rs.initiate() step.

2. Start the control plane in open mode

terminalbash
cd lakeshore-controlplane
pnpm dev            # listens on http://localhost:8080

Open mode means no client token is required. It is a local-development posture and nothing else — do not run a plane this way on an address anything else can reach.

3. Build a local nymph dist

terminalbash
cd nymph
bash scripts/dist-local.sh          # → ./dist/latest/…

This lays out the same five files the release bucket serves under dreamlake/nymph/<version>/: the tarball, the raw binary, a .sha256 beside each, and a VERSION file. It is the local half of the release workflow, and it exists so the install path is testable without the internet.

4. Install it from a file:// URL

terminalbash
bash nymph/scripts/install.sh \
    --base-url "file://$(cd nymph/dist && pwd)" \
    --version latest \
    --dest ~/.local/bin

The installer detects your platform, fetches the matching tarball, verifies the checksum, and drops the binary at --dest. It is the same script the published one-liner runs; only --base-url differs.

Why file:// and not a bare path

The installer builds its URL as ${BASE_URL}/${VERSION}/nymph-<arch>-<os>.tar.gz and fetches it with curl. A bare absolute path is rejected outright — curl answers URL rejected: No host part in the URL and exits, which set -e turns into an aborted install. Noisy, and intentionally so.

The matching detail on the producing side: the .sha256 sidecar must record the bare archive filename. The installer downloads both files into a fresh temp directory under exactly those basenames, changes into it, and runs sha256sum -c there. A checksum file recording ./nymph-….tar.gz, or an absolute path, fails to verify.

5. Run the daemon

terminalbash
~/.local/bin/nymph --config dev/local-udf-daemon.toml \
    --server http://localhost:8080 --queue '*'

Watch it move through the phases on the worker lifecycle page: hello, then setup while it drains any pending setup commands, then ready.

Don't drop the --config

runtime.workdir defaults to /var/lib/dreamlake/udf, which only root can create. The daemon probes that directory for writability before it asserts ready, so without a config that points somewhere you own, the daemon heartbeats forever at setting_up and never claims anything — lakeshore daemon list shows it alive and step 6 hangs. The committed dev/local-udf-daemon.toml sets a writable workdir (and a 2 s poll window). There is no --workdir flag — runtime.workdir is settable only from a config file, so pass one.

6. Submit, and read the result back

Point the SDK at the same loopback plane — Quick start §2. Point at a control plane owns that environment, and http://localhost:8080 is a perfectly good value for it. Nothing else about the code changes; you submit and collect exactly as you would against a hosted plane.

The fastest confirmation that the daemon is really attached does not involve the SDK at all — ask the control plane what it can see, then run a command on it:

terminalbash
lakeshore daemon list                       # the worker you just started
lakeshore daemon exec <daemon-id> 'echo hello from a locally built nymph'

That is the whole loop closed: a job you submitted, claimed by a daemon you compiled, on a plane you started, with nothing fetched from anywhere.

The pane-by-pane version

This page is the minimum sequence. The run-lakeshore skill in lakeshore-workspace drives the same bring-up through mprocs — one pane per process, with the ordering, the ports and the wait conditions already encoded. Use it for day-to-day work and use this page to understand what it is doing.

Uploading results to your DreamLake scope

A run that produces a file should put it somewhere durable, and the somewhere is your DreamLake scope. The dreamlake module resolves three things from the worker's environment:

VariableWhat it resolvesDefault
DREAMLAKE_REMOTEWhich DreamLake to talk tohttps://api.dreamlake.ai
DREAMLAKE_PROJECTWhich project to write into, as slug@namespaceRequired. No default
DREAMLAKE_API_KEYThe identity doing the writingFalls back to the local token store

You do not set the first two by hand on the worker. RunConfig.dreamlake = {remote, project} is declared next to the function, travels in the run config, and the SDK expands it into those two variables in the child environment. DREAMLAKE_API_KEY is the one the daemon supplies for itself — from its own environment, its systemd unit, or its host_setup.

Then the body publishes:

lifecycle_upload.pypython
import dreamlake.lakeshore as dls

@dls.udf(queue="local")
def train(steps: int = 100) -> str:
    out = "./checkpoint.pt"
    ...                                    # write the file
    return dls.run.publish(out, name="checkpoint", type="model")

publish reads those variables, imports dreamlake lazily, and uploads. If the declaration is missing, the error names the missing piece rather than failing inside the HTTP client.

Proposal

RunConfig.dreamlake, RunConfig.env (applied to the child environment by the process runner) and dls.run.publish all ship. The limit is who the upload is attributed to.

It is worth stating before it surprises anyone: the upload happens as the daemon's own DreamLake identity, not as the submitting user's. On a personal worker those are the same person, which is exactly why the loop on this page is a faithful test of it. On a shared fleet they are not — the artifact lands under whoever's key is in the daemon's environment.

Closing that gap needs the control plane to mint a short-lived DreamLake token scoped to the submitter, which neither side has today.

A token is never carried in the invocation envelope, and that is a decision rather than an omission. Putting one there would write a live user credential into a database document and onto the daemon wire once per invocation, readable by anything that can read the invocation. The narrower blast radius of a per-worker identity is worth the loss of attribution until the exchange exists.

Connecting the DreamLake dashboard

The dashboard can be pointed at a control plane by link, which is how you configure a locally running plane without typing a URL into a form:

http://localhost:3000/<your-namespace>/lakeshore?connect=http%3A%2F%2Flocalhost%3A8080&connect_ns=default

connect is the control plane's base URL, percent-encoded. connect_ns is the Lakeshore namespace to bind to and defaults to default — namespace-as-prefix is the same idea it names here. connect_name optionally suggests a display name.

The safety posture is the interesting half, because a query parameter is attacker-supplied by construction:

  • The link pre-fills and asks you to confirm. It never creates a connection on its own. You land on the create form, populated, with the host shown in isolation so you can read it before agreeing to it.
  • It pre-checks Skip connection test. Creating a connection normally has the server probe the supplied URL, which means a link that auto-applied would be a one-click server-side request forgery primitive — an attacker choosing which address DreamLake's own network reaches. No outbound fetch happens until you ask for one.
  • It will never accept a token as a query parameter. There is no connect_token and there will not be one. Query strings land in history, referrers and server logs; the bearer token is write-only end to end today and this is not the feature that changes that.
  • Anything else is rejected at parse time and simply not pre-filled: a scheme other than http: or https:, or a URL carrying user:pass@ credentials. Plain http: on a non-loopback host is pre-filled but escalates the banner to a warning — http: on localhost is the whole point of this page, so it passes quietly.
  • The parameters are read once and stripped from the address bar before the dialog opens, so a reload does not re-prompt.

The hosts registry

Once you have more than one machine, the list of machines you can install onto lives in a static TOML file in the repo — dev/hosts.toml — which scripts/hosts.sh iterates.

dev/hosts.tomltoml
# EVERY ADDRESS IN THIS FILE IS A PLACEHOLDER — loopback or RFC1918 only.
# Replace them before use. Credentials live in ~/.ssh/config; this file
# names an alias, never a key.

[[host]]
name    = "local"
address = "127.0.0.1"          # loopback — your own laptop
arch    = "aarch64-apple-darwin"
role    = "dev"
queues  = ["*"]
install = "local-dist"         # build here, install over file://

[[host]]
name    = "bench-1"
address = "192.168.10.24"      # RFC1918 placeholder — not a real machine
ssh     = "bench-1"            # an alias in ~/.ssh/config, never a key path
arch    = "x86_64-unknown-linux-gnu"
role    = "cpu"
queues  = ["cpu-lane", "default"]
install = "ssh"                # the remote fetches the published tarball itself

arch is a Rust target triple because that is what picks the tarball name in step 4. install is how the binary gets onto the box: local-dist is the no-network path this page walks, ssh lets the remote fetch a published tarball, and none means something else already manages it — an AMI, a systemd unit, a container image. hosts.sh list reads the file; install is a dry run until you pass --yes, because a file full of placeholders is a file you must not act on by reflex.

It is a file rather than a control-plane table for one reason that outweighs the rest: it has to work before any host has enrolled. The worker table is an output of provisioning, empty at exactly the moment you are bootstrapping a machine that has never run the daemon — and it empties again every time an elastic queue scales to zero. A git-diffable file is also the right shape for a list whose main failure mode is somebody adding a wrong address.

Quick start →

The shortest path from an empty environment to a remote result, and where the client environment is documented.

Worker lifecycle →

What the daemon you just started is doing, phase by phase.

Providers →

When you outgrow the laptop — launching workers on EC2, GCE or Kube.