Preloaded functions
Some functions are cheap to call and expensive to prepare. A generative model needs its weights in GPU memory before it can produce a single sample — and once they are there, every later call is fast. A preloaded function is a UDF written so that the expensive half happens once per worker rather than once per invocation.
Nothing about the decorator changes. This is a simple function whose body is written to reuse state that outlives the call.
What makes it work
run_worker is a claim-and-execute loop in one long-lived process. It
claims a job, runs the body, and comes back for the next one without restarting.
So anything held at module scope — or memoized behind a cache — survives from
one invocation to the next on that worker.
That is the whole mechanism. There is no preload hook to register:
The first invocation on a given worker pays the load. The second through thousandth do not.
Putting _model() at module scope rather than behind a cached accessor moves
the load into import, which happens while the worker is resolving the
function — before it has claimed anything. A slow import can look like a hung
worker, and it runs even for an invocation that would not have touched the
model. Load lazily, on first use.
What resets it
The cache lives and dies with the worker process. Four things end it:
| Event | Effect |
|---|---|
| Scale to zero | The worker is terminated; the next invocation reloads from cold |
| Scale up | Each new worker starts cold, even while others are warm |
| A new code commit | The worker imports a different tree; the old module is gone |
| Worker replacement | Spot reclaim, host failure, drain |
The first two are the ones that bite, and they are queue settings rather than code. See Elastic queues.
Keeping workers warm
Preloading trades idle cost against latency, and the queue is where you set the exchange rate:
fully_elasticwithmin: 0throws the weights away every time the backlog clears. Correct for a nightly batch; wrong for anything interactive.pool_with_thresholdwith a non-zerominkeeps some workers resident, so a request arriving after an idle period hits a warm process. This is the setting preloaded functions usually want.- A dedicated queue. Routing is caller-side —
@udf(queue="sd-turbo")— so a model-serving queue can hold a warm floor while other work stays elastic. Mixing both on one queue means the autoscaler cannot tell the two apart.
Getting the weights onto the host
The load is only fast if the bytes are local. Three places to put them, in descending order of how well they work:
- Bake them into the image.
daemonTemplate.image_id— paid once at build time, zero per worker. Best for weights that change rarely. - Fetch them in host setup. A setup script pulls them before the worker takes work. Paid per worker, but off the critical path of the first invocation.
- Fetch them in the body, behind the cache. Simplest, and the first caller waits.
The fourth option — a shared filesystem holding the weights — is what Mounting Storage is for, and it is not wired yet.
Worth being explicit: @udf takes fn, queue, transport, and config —
a RunConfig run declaration whose derived host key
decides which worker may claim the run. None of the four is a preload hook.
There is no preload=, no init callback, and no way to tell the control plane
that a worker is warm. Warmth is a property of the process that the platform
cannot currently see — which is why the levers above are all indirect.
dreamlake.lakeshore.lifecycle is a real module, exported as dls.lifecycle.
It ships @on_host_setup (once per worker process, before the first body),
@on_run_setup (before every body), @on_shutdown (best effort, on the way
out), plus phase() and host_key() for introspection, and three driver
functions — run_host_setup_hooks(), run_run_setup_hooks(),
run_shutdown_hooks() — that a worker loop is meant to call.
No worker loop calls them. Grepping for lifecycle across the SDK outside
that one file hits only its own tests: neither run_worker() nor the daemon
entrypoint invokes a driver. The module's own docstring admits it — wiring them
into worker.py and daemon/entry.py "is a follow-up — those files belong to
another workstream in this slice, so today an author drives them explicitly."
So a @lifecycle.on_host_setup hook in your code will simply never run unless
you call run_host_setup_hooks() yourself. The working preload idiom today is
still the lazy memoised accessor above.
Preload does not work at all on the nymph path
This is the correction that matters most in practice, and it undoes everything on this page for one of the two execution paths.
Preloading depends on the worker being one long-lived process. The nymph
daemon is not that. Its process runner spawns a fresh
python3 -m dreamlake.lakeshore.daemon.entry child per invocation, and its
docker runner is docker run --rm — a brand-new container every time.
| The queue-native Python worker | The nymph daemon | |
|---|---|---|
| Process model | One process, claim-and-execute loop | One fresh process (or container) per invocation |
| Module-level global survives? | Yes | No — reloaded every single invocation |
lru_cache accessor survives? | Yes | No |
On the nymph path the only thing that survives between invocations is bytes on
the host filesystem, and under the docker runner only if a bind mount is
declared for them. A cached model object is reconstructed from scratch on every
call, so the "first call pays, the rest are free" property does not hold — every
call is the first call.
The practical consequence: bake weights into the image or fetch them in host
setup so that each fresh process is at least reading from local disk, and do not
expect the in-memory cache to buy you anything. The fourth option named above —
a shared filesystem holding the weights — is not available as a fallback either:
it is not wired yet, and nothing in the stack reads a Mount today.
Concurrency
One process, one cached model, several in-flight invocations. A daemon reports a maximum invocation count and can run more than one at a time, so a cached object may be shared concurrently. Most inference pipelines are not safe to call re-entrantly on the same device.
If the body is not concurrency-safe, cap the worker rather than the queue — guard the shared object with a lock, or run one invocation per worker. The queue has no per-function concurrency setting.