Elastic queues
A task queue runs on workers you attached. Set
providerRef and the same record becomes a launch queue: the control
plane watches its backlog and provisions workers to drain it.
Nothing about the task-queue half changes. Submit, claim, complete, and stream behave identically — the queue simply stops depending on you to supply capacity.
The three fields
| Field | What it sets |
|---|---|
providerRef | The provider to launch through, by name. null means bring your own worker. |
elasticity | The scaling rule — fixed, fully_elastic, pool_with_threshold, or max_count, plus min / max / threshold. |
daemonTemplate | What a launched worker looks like: instance_type, image_id, runner, setup_scripts. Arguably derivable from the run declaration rather than written by hand. |
The four policies
fixednever launches or retires anything. The pool is whatever you attached — this is the task-queue case with a provider recorded but unused.fully_elastictracks the backlog directly, bounded byminandmax. Withmin: 0it scales to zero, which is what you want for bursty GPU work that can absorb a cold start.pool_with_thresholdscales on utilization with hysteresis — a high and a low watermark, and a run of consecutive ticks before it acts either way. Use it when you want warm capacity and no flapping.max_countis demand-driven with a required ceiling.
A fast lane and a heavy lane are two Queue rows differing in
daemonTemplate and elasticity — fast as pool_with_threshold with a
small warm minimum, heavy as fully_elastic from zero on GPU instances.
Routing is caller-side: @udf(queue="heavy"). There is no automatic
classification.
How the controller decides
The elasticity controller ticks on a fixed interval. On each tick it joins pending invocations to their queue by name, compares the backlog against the policy, and launches or retires daemons through the queue's provider.
Every decision is written as an ElasticityEvent — including the ones where
it decided to do nothing. That is the useful part. "Why did my cluster not
scale up" is answerable from the event feed, because a noop is a logged decision
with a reason rather than an absence of evidence.
What a provider must support
providerRef names a provider, but not every provider kind can be launched
into. Only EC2, GCE, and Kube are accepted by the launch routes;
SSH and SLURM providers can be registered and are rejected at launch, so a queue
pointing at one cannot grow workers.
See Providers for the current status of each kind and
what the [ dev ] marker means.
Scale to zero, and the cold start it buys
fully_elastic with min: 0 retires the last worker when the backlog clears.
The next submit then pays a full provision — instance launch, image pull, setup
scripts — before anything runs. That is the trade the policy exists to make. If
the latency matters more than the idle cost, pool_with_threshold with a
non-zero minimum keeps one worker warm.
daemonTemplate.setup_scripts runs on each launched worker, so a heavy setup is
paid per worker, not once per queue. Bake what you can into image_id instead.
/queues/:name/stats computes utilization from the daemon-reported invocation
counts. The daemon reports its maximum capacity but not its current in-flight
count, so the utilization figure can be structurally zero even on a busy
queue. The controller itself does not depend on this — it counts pending
invocations directly — but a dashboard reading the stats endpoint will
understate load.