DreamLake

Elastic queues

A task queue runs on workers you attached. Set providerRef and the same record becomes a launch queue: the control plane watches its backlog and provisions workers to drain it.

Nothing about the task-queue half changes. Submit, claim, complete, and stream behave identically — the queue simply stops depending on you to supply capacity.

The three fields

FieldWhat it sets
providerRefThe provider to launch through, by name. null means bring your own worker.
elasticityThe scaling rule — fixed, fully_elastic, pool_with_threshold, or max_count, plus min / max / threshold.
daemonTemplateWhat a launched worker looks like: instance_type, image_id, runner, setup_scripts. Arguably derivable from the run declaration rather than written by hand.

The four policies

  • fixed never launches or retires anything. The pool is whatever you attached — this is the task-queue case with a provider recorded but unused.
  • fully_elastic tracks the backlog directly, bounded by min and max. With min: 0 it scales to zero, which is what you want for bursty GPU work that can absorb a cold start.
  • pool_with_threshold scales on utilization with hysteresis — a high and a low watermark, and a run of consecutive ticks before it acts either way. Use it when you want warm capacity and no flapping.
  • max_count is demand-driven with a required ceiling.
Tiering is two queues, not one clever queue

A fast lane and a heavy lane are two Queue rows differing in daemonTemplate and elasticity — fast as pool_with_threshold with a small warm minimum, heavy as fully_elastic from zero on GPU instances. Routing is caller-side: @udf(queue="heavy"). There is no automatic classification.

How the controller decides

The elasticity controller ticks on a fixed interval. On each tick it joins pending invocations to their queue by name, compares the backlog against the policy, and launches or retires daemons through the queue's provider.

Every decision is written as an ElasticityEvent — including the ones where it decided to do nothing. That is the useful part. "Why did my cluster not scale up" is answerable from the event feed, because a noop is a logged decision with a reason rather than an absence of evidence.

terminalbash
lakeshore queues events gpu-train     # the decision feed, noops included
lakeshore queues stats gpu-train      # daemon count, utilization, pending

What a provider must support

providerRef names a provider, but not every provider kind can be launched into. Only EC2, GCE, and Kube are accepted by the launch routes; SSH and SLURM providers can be registered and are rejected at launch, so a queue pointing at one cannot grow workers.

See Providers for the current status of each kind and what the [ dev ] marker means.

Scale to zero, and the cold start it buys

fully_elastic with min: 0 retires the last worker when the backlog clears. The next submit then pays a full provision — instance launch, image pull, setup scripts — before anything runs. That is the trade the policy exists to make. If the latency matters more than the idle cost, pool_with_threshold with a non-zero minimum keeps one worker warm.

daemonTemplate.setup_scripts runs on each launched worker, so a heavy setup is paid per worker, not once per queue. Bake what you can into image_id instead.

Utilization can read as zero

/queues/:name/stats computes utilization from the daemon-reported invocation counts. The daemon reports its maximum capacity but not its current in-flight count, so the utilization figure can be structurally zero even on a busy queue. The controller itself does not depend on this — it counts pending invocations directly — but a dashboard reading the stats endpoint will understate load.

Task queue →

The behaviors that do not change when a provider is attached.

Queue model →

The record behind both halves, field by field.

Elasticity reference →

Scaling policies, pools, and how the controller ticks.

Scaling rules →

Thresholds, hysteresis, and the signals the autoscaler reads.