DreamLake

Preloaded functions

Some functions are cheap to call and expensive to prepare. A generative model needs its weights in GPU memory before it can produce a single sample — and once they are there, every later call is fast. A preloaded function is a UDF written so that the expensive half happens once per worker rather than once per invocation.

Nothing about the decorator changes. This is a simple function whose body is written to reuse state that outlives the call.

What makes it work

run_worker is a claim-and-execute loop in one long-lived process. It claims a job, runs the body, and comes back for the next one without restarting. So anything held at module scope — or memoized behind a cache — survives from one invocation to the next on that worker.

That is the whole mechanism. There is no preload hook to register:

generate.pypython
import functools
import dreamlake.lakeshore as dls


@functools.lru_cache(maxsize=1)
def _model():
    """Load once per worker process. Every later call returns the same object."""
    import torch
    from diffusers import StableDiffusionPipeline

    pipe = StableDiffusionPipeline.from_pretrained(
        "stabilityai/sd-turbo", torch_dtype=torch.float16
    )
    return pipe.to("cuda")


@dls.udf(queue="gpu")
def generate(prompt: str, steps: int = 4) -> dict[str, str]:
    """Sample one image. The first call on a worker pays the model load."""
    image = _model()(prompt, num_inference_steps=steps).images[0]
    out = dls.run.write("samples/0000.png")
    image.save(out)
    return {"image": "samples/0000.png"}

The first invocation on a given worker pays the load. The second through thousandth do not.

Do not load at import time

Putting _model() at module scope rather than behind a cached accessor moves the load into import, which happens while the worker is resolving the function — before it has claimed anything. A slow import can look like a hung worker, and it runs even for an invocation that would not have touched the model. Load lazily, on first use.

What resets it

The cache lives and dies with the worker process. Four things end it:

EventEffect
Scale to zeroThe worker is terminated; the next invocation reloads from cold
Scale upEach new worker starts cold, even while others are warm
A new code commitThe worker imports a different tree; the old module is gone
Worker replacementSpot reclaim, host failure, drain

The first two are the ones that bite, and they are queue settings rather than code. See Elastic queues.

Keeping workers warm

Preloading trades idle cost against latency, and the queue is where you set the exchange rate:

  • fully_elastic with min: 0 throws the weights away every time the backlog clears. Correct for a nightly batch; wrong for anything interactive.
  • pool_with_threshold with a non-zero min keeps some workers resident, so a request arriving after an idle period hits a warm process. This is the setting preloaded functions usually want.
  • A dedicated queue. Routing is caller-side — @udf(queue="sd-turbo") — so a model-serving queue can hold a warm floor while other work stays elastic. Mixing both on one queue means the autoscaler cannot tell the two apart.

Getting the weights onto the host

The load is only fast if the bytes are local. Three places to put them, in descending order of how well they work:

  1. Bake them into the image. daemonTemplate.image_id — paid once at build time, zero per worker. Best for weights that change rarely.
  2. Fetch them in host setup. A setup script pulls them before the worker takes work. Paid per worker, but off the critical path of the first invocation.
  3. Fetch them in the body, behind the cache. Simplest, and the first caller waits.

The fourth option — a shared filesystem holding the weights — is what Mounting Storage is for, and it is not wired yet.

There is still no preload API

Worth being explicit: @udf takes fn, queue, transport, and config — a RunConfig run declaration whose derived host key decides which worker may claim the run. None of the four is a preload hook. There is no preload=, no init callback, and no way to tell the control plane that a worker is warm. Warmth is a property of the process that the platform cannot currently see — which is why the levers above are all indirect.

Not yet shipped — the lifecycle hooks exist, but nothing calls them

dreamlake.lakeshore.lifecycle is a real module, exported as dls.lifecycle. It ships @on_host_setup (once per worker process, before the first body), @on_run_setup (before every body), @on_shutdown (best effort, on the way out), plus phase() and host_key() for introspection, and three driver functions — run_host_setup_hooks(), run_run_setup_hooks(), run_shutdown_hooks() — that a worker loop is meant to call.

No worker loop calls them. Grepping for lifecycle across the SDK outside that one file hits only its own tests: neither run_worker() nor the daemon entrypoint invokes a driver. The module's own docstring admits it — wiring them into worker.py and daemon/entry.py "is a follow-up — those files belong to another workstream in this slice, so today an author drives them explicitly."

So a @lifecycle.on_host_setup hook in your code will simply never run unless you call run_host_setup_hooks() yourself. The working preload idiom today is still the lazy memoised accessor above.

Preload does not work at all on the nymph path

This is the correction that matters most in practice, and it undoes everything on this page for one of the two execution paths.

Preloading depends on the worker being one long-lived process. The nymph daemon is not that. Its process runner spawns a fresh python3 -m dreamlake.lakeshore.daemon.entry child per invocation, and its docker runner is docker run --rm — a brand-new container every time.

The queue-native Python workerThe nymph daemon
Process modelOne process, claim-and-execute loopOne fresh process (or container) per invocation
Module-level global survives?YesNo — reloaded every single invocation
lru_cache accessor survives?YesNo

On the nymph path the only thing that survives between invocations is bytes on the host filesystem, and under the docker runner only if a bind mount is declared for them. A cached model object is reconstructed from scratch on every call, so the "first call pays, the rest are free" property does not hold — every call is the first call.

The practical consequence: bake weights into the image or fetch them in host setup so that each fresh process is at least reading from local disk, and do not expect the in-memory cache to buy you anything. The fourth option named above — a shared filesystem holding the weights — is not available as a fallback either: it is not wired yet, and nothing in the stack reads a Mount today.

Concurrency

One process, one cached model, several in-flight invocations. A daemon reports a maximum invocation count and can run more than one at a time, so a cached object may be shared concurrently. Most inference pipelines are not safe to call re-entrantly on the same device.

If the body is not concurrency-safe, cap the worker rather than the queue — guard the shared object with a lock, or run one invocation per worker. The queue has no per-function concurrency setting.

Simple functions →

The base type — body shapes, scopes, and key-based returns.

Elastic queues →

The scaling policies that decide whether a warm worker survives.

Host setup →

Fetching weights before the worker takes its first job.