# Preloaded functions

  Some functions are cheap to call and expensive to prepare. A generative model
  needs its weights in GPU memory before it can produce a single sample — and
  once they are there, every later call is fast. A **preloaded function** is a
  UDF written so that the expensive half happens once per worker rather than
  once per invocation.

Nothing about the decorator changes. This is a
[simple function](/lakeshore/patterns.md) whose body is written to reuse state that
outlives the call.

## What makes it work

`run_worker` is a **claim-and-execute loop in one long-lived process**. It
claims a job, runs the body, and comes back for the next one without restarting.
So anything held at module scope — or memoized behind a cache — survives from
one invocation to the next on that worker.

That is the whole mechanism. There is no preload hook to register:

```python file="generate.py"
import functools
import dreamlake.lakeshore as dls


@functools.lru_cache(maxsize=1)
def _model():
    """Load once per worker process. Every later call returns the same object."""
    import torch
    from diffusers import StableDiffusionPipeline

    pipe = StableDiffusionPipeline.from_pretrained(
        "stabilityai/sd-turbo", torch_dtype=torch.float16
    )
    return pipe.to("cuda")


@dls.udf(queue="gpu")
def generate(prompt: str, steps: int = 4) -> dict[str, str]:
    """Sample one image. The first call on a worker pays the model load."""
    image = _model()(prompt, num_inference_steps=steps).images[0]
    out = dls.run.write("samples/0000.png")
    image.save(out)
    return {"image": "samples/0000.png"}
```

The first invocation on a given worker pays the load. The second through
thousandth do not.

> **Warning:** Putting `_model()` at module scope rather than behind a cached accessor moves
>   the load into **import**, which happens while the worker is resolving the
>   function — before it has claimed anything. A slow import can look like a hung
>   worker, and it runs even for an invocation that would not have touched the
>   model. Load lazily, on first use.

## What resets it

The cache lives and dies with the worker process. Four things end it:

| Event | Effect |
| --- | --- |
| Scale to zero | The worker is terminated; the next invocation reloads from cold |
| Scale up | Each **new** worker starts cold, even while others are warm |
| A new code commit | The worker imports a different tree; the old module is gone |
| Worker replacement | Spot reclaim, host failure, drain |

The first two are the ones that bite, and they are queue settings rather than
code. See [Elastic queues](/lakeshore/queues/elastic.md).

## Keeping workers warm

Preloading trades idle cost against latency, and the queue is where you set the
exchange rate:

- **`fully_elastic` with `min: 0`** throws the weights away every time the
  backlog clears. Correct for a nightly batch; wrong for anything interactive.
- **`pool_with_threshold` with a non-zero `min`** keeps some workers resident,
  so a request arriving after an idle period hits a warm process. This is the
  setting preloaded functions usually want.
- **A dedicated queue.** Routing is caller-side — `@udf(queue="sd-turbo")` — so
  a model-serving queue can hold a warm floor while other work stays elastic.
  Mixing both on one queue means the autoscaler cannot tell the two apart.

## Getting the weights onto the host

The load is only fast if the bytes are local. Three places to put them, in
descending order of how well they work:

1. **Bake them into the image.** `daemonTemplate.image_id` — paid once at build
   time, zero per worker. Best for weights that change rarely.
2. **Fetch them in [host setup](/lakeshore/host-setup.md).** A setup script pulls
   them before the worker takes work. Paid per worker, but off the critical path
   of the first invocation.
3. **Fetch them in the body, behind the cache.** Simplest, and the first caller
   waits.

The fourth option — a shared filesystem holding the weights — is what
[Mounting Storage](/lakeshore/mounts.md) is for, and it is not wired yet.

> **Note:** Worth being explicit: `@udf` takes `fn`, `queue`, `transport`, and `config` —
>   a [`RunConfig`](/lakeshore/registry.md) run declaration whose derived host key
>   decides which worker may claim the run. None of the four is a preload hook.
>   There is no `preload=`, no init callback, and no way to tell the control plane
>   that a worker is warm. Warmth is a property of the process that the platform
>   cannot currently see — which is why the levers above are all indirect.

> **Warning:** `dreamlake.lakeshore.lifecycle` is a real module, exported as `dls.lifecycle`.
>   It ships `@on_host_setup` (once per worker process, before the first body),
>   `@on_run_setup` (before every body), `@on_shutdown` (best effort, on the way
>   out), plus `phase()` and `host_key()` for introspection, and three driver
>   functions — `run_host_setup_hooks()`, `run_run_setup_hooks()`,
>   `run_shutdown_hooks()` — that a worker loop is meant to call.
>
>   **No worker loop calls them.** Grepping for `lifecycle` across the SDK outside
>   that one file hits only its own tests: neither `run_worker()` nor the daemon
>   entrypoint invokes a driver. The module's own docstring admits it — wiring them
>   into `worker.py` and `daemon/entry.py` *"is a follow-up — those files belong to
>   another workstream in this slice, so today an author drives them explicitly."*
>
>   So a `@lifecycle.on_host_setup` hook in your code will simply never run unless
>   you call `run_host_setup_hooks()` yourself. The working preload idiom today is
>   still the lazy memoised accessor above.

## Preload does not work at all on the nymph path

This is the correction that matters most in practice, and it undoes everything
on this page for one of the two execution paths.

Preloading depends on the worker being **one long-lived process**. The nymph
daemon is not that. Its `process` runner spawns a **fresh**
`python3 -m dreamlake.lakeshore.daemon.entry` child per invocation, and its
`docker` runner is `docker run --rm` — a brand-new container every time.

| | The queue-native Python worker | The nymph daemon |
| --- | --- | --- |
| Process model | One process, claim-and-execute loop | One fresh process (or container) per invocation |
| Module-level global survives? | **Yes** | **No** — reloaded every single invocation |
| `lru_cache` accessor survives? | **Yes** | **No** |

On the nymph path the only thing that survives between invocations is **bytes on
the host filesystem**, and under the `docker` runner only if a bind mount is
declared for them. A cached model object is reconstructed from scratch on every
call, so the "first call pays, the rest are free" property does not hold — every
call is the first call.

The practical consequence: bake weights into the image or fetch them in host
setup so that each fresh process is at least reading from local disk, and do not
expect the in-memory cache to buy you anything. The fourth option named above —
a shared filesystem holding the weights — is not available as a fallback either:
it is not wired yet, and nothing in the stack reads a `Mount` today.

## Concurrency

One process, one cached model, several in-flight invocations. A daemon reports a
maximum invocation count and can run more than one at a time, so a cached object
may be shared concurrently. Most inference pipelines are not safe to call
re-entrantly on the same device.

If the body is not concurrency-safe, cap the worker rather than the queue —
guard the shared object with a lock, or run one invocation per worker. The queue
has no per-function concurrency setting.

    The base type — body shapes, scopes, and key-based returns.

    The scaling policies that decide whether a warm worker survives.

    Fetching weights before the worker takes its first job.
