Model Functions
dl_pipelines describes itself, in its own module docstring, as "the
canonical node functional API for DreamLake pipelines — the contract." It is
a naming layer over model inference: one function per task, so a pipeline
names what it wants done rather than which checkpoint does it.
The argument for a naming layer is the same one that motivates the graph. A
pipeline that says detect_objects(images) still means the same thing after the
checkpoint behind it is swapped, the serving stack is rewritten, or the work
moves from a local GPU to a queue. A pipeline that says
DetrForObjectDetection.from_pretrained("facebook/detr-resnet-50") means one
particular thing on one particular day. The contract fixes the vocabulary at the
level of the task, and leaves the model choice to whatever resolves it.
That is the whole of the design. What follows is the vocabulary — and an honest account of how much of it runs.
Every one of these thirty-three functions is an async def whose body is a
single raise NotImplementedError. A blocking mirror of each is generated
automatically at import time, so there is a synchronous twin of every entry
point — and it raises too, because it calls the async original. The Hugging
Face adaptor that maps task names onto these functions raises in all three of
its resolve, run and wrap_pipeline paths; only its lookup tables are
real data.
None of these functions appears in any shipped example pipeline. The example
UDF named detect_objects in pipelines/image_object_annotation.py is a
local stage that merely shares a name with the canonical
detect_objects. The two are unrelated: the example's body is ..., and it
imports nothing from dl_pipelines.
The five domains
Thirty-three functions across five modules. Names are given exactly as they are exported; the intent line is the function's own docstring intent, not an invented signature.
Vision — twelve
| Function | Intent |
|---|---|
caption_image | caption each image, optionally prompt-guided |
generate_image | generate images from prompts, or unconditionally |
image_to_image | transform each image, optionally text-guided |
classify_image | classify each image against the model's own label set |
zero_shot_classify_image | classify each image against caller-provided labels |
detect_objects | detect objects in each image |
zero_shot_detect_objects | detect caller-specified object classes |
segment_image | semantic or instance segmentation per image |
generate_masks | promptable mask generation, SAM-style |
estimate_depth | per-pixel depth map per image |
detect_keypoints | keypoints — pose, landmarks — per image |
embed_image | one embedding vector per image |
Text — ten
| Function | Intent |
|---|---|
generate_text | continue or complete each prompt |
classify_text | classify each text against the model's own label set |
zero_shot_classify_text | classify each text against caller-provided labels |
classify_tokens | tag spans — NER, POS — in each text |
summarize | summarize each text |
translate | translate each text |
answer_question | extract the answer span for each question/context pair |
fill_mask | fill the mask token with ranked candidates |
embed_text | one embedding vector per text |
rank_texts | relevance score of each text for one query |
Audio — four
| Function | Intent |
|---|---|
transcribe | speech to text, per clip |
synthesize_speech | text to speech, per text |
classify_audio | classify each clip |
audio_to_audio | transform each clip — denoise, separate, enhance |
Video — four
| Function | Intent |
|---|---|
generate_video | generate video from text and/or a conditioning image |
classify_video | classify each video |
caption_video | describe or answer questions about each video |
video_to_video | transform each video, optionally text-guided |
Multimodal — three
| Function | Intent |
|---|---|
visual_question_answer | answer a question about each image |
document_question_answer | answer a question about each document — a scan, a PDF |
table_question_answer | answer a question about each table |
Why the count is thirty-three
The number is not arbitrary, and it is the strongest evidence that this set was derived rather than brainstormed. The Hugging Face adaptor carries a mapping from every one of the forty-seven Hugging Face task types onto exactly this set of names, and that mapping has two properties worth stating.
Every canonical function is reachable. All thirty-three appear as the target of at least one task. There is no name in the API that the task taxonomy does not justify.
Thirty-nine tasks map to a function; eight map to None. Those eight are
the tasks the contract has decided not to cover yet:
any-to-any · audio-text-to-text · image-to-3d · reinforcement-learning ·
tabular-classification · tabular-regression · text-to-3d ·
visual-document-retrieval
Thirty-nine tasks landing on thirty-three functions means six collapses, and each one is a deliberate judgement that Hugging Face's split is about the input arity rather than the task:
| Canonical function | Hugging Face tasks it absorbs |
|---|---|
generate_video | text-to-video, image-to-video, image-text-to-video |
generate_image | text-to-image, unconditional-image-generation |
image_to_image | image-to-image, image-text-to-image |
embed_text | feature-extraction, sentence-similarity |
visual_question_answer | visual-question-answering, image-text-to-text |
generate_video is the clearest case: text-to-video and image-to-video are one
function with two optional arguments, not two functions. The signature reflects
that — prompts and images are both optional, and passing either or both is
the same call.
The task list was crawled from huggingface.co/api/tasks and is pinned in a
comment as accurate on 2026-07-02. Nothing re-crawls it, and nothing tests
that it is still complete. Treat the forty-seven as a snapshot.
Output records
The return types are small dataclasses in dl_pipelines.types. These are the
only part of the module that is fully implemented, because a dataclass has no
body to leave unwritten.
| Record | Fields | Returned by |
|---|---|---|
Caption | text, confidence? | caption_image, caption_video |
Label | name, score | every classify_*, plus fill_mask and rank_texts |
Box | xyxy, label?, score? | detect_objects, zero_shot_detect_objects |
Mask | mask (HxW boolean array), label?, score? | segment_image, generate_masks |
Keypoints | points (Kx2 array), label?, score? | detect_keypoints |
Span | text, start, end, label?, score? | classify_tokens, answer_question, document_question_answer |
Transcript | text, segments: list[Span] | transcribe |
Embedding is not a record — it is list[float], and embed_image /
embed_text return one per input.
The nesting matters when you read a signature. A per-item classifier returns
list[list[Label]]: the outer list is the batch, the inner list is that item's
ranked labels. answer_question returns a flat list[Span], because there is
exactly one answer per question. Nothing about this is enforced today, but it is
what the annotations say.
The input aliases constrain nothing
Image, Audio, Video, Document and Table are all declared as
TypeAlias = Any, with a comment saying inputs are "intentionally loose for now
— an array, a PIL/AV object, a path, or a URL." So a signature reading
detect_objects(images: list[Image]) is, to a type checker, list[Any].
These aliases are documentation that happens to be written in the type language. They tell a reader what kind of thing belongs in that slot, and they will reject nothing at all if the reader is wrong. Do not mistake a green type check for a validated input.
Async and sync
Every canonical function is defined async def. The blocking mirrors are not
written by hand — dl_pipelines/sync/__init__.py walks dl_pipelines.__all__
at import time, and for each coroutine function installs a wrapper that calls
asyncio.run on it. Both calling conventions therefore exist for every entry
point, and stay in step by construction: a new function added to __all__ gets
its synchronous twin for free.
Two consequences of generating the mirror this way. The wrapper carries
__wrapped_async__, so the original coroutine function is still reachable from
its twin. And because it is asyncio.run underneath, a sync.* call made from
inside a running event loop will fail — the mirrors are for synchronous
programs, not for escaping an await you would rather not write.
The tracer does not care which one you use. It reads the pipeline's source
as an AST and never imports or executes anything, so an async def stage and a
blocking one derive precisely the same graph. The one place the distinction
becomes visible to the graph is the async for batching form, covered at
Batching and streaming.
Everything shipped in dreamlake-pipeline-examples is synchronous, and the
pipeline-generator skill instructs authors to stay that way. Prefer the sync.
mirrors unless you have a reason not to.
How you would use one
The teaching point of this page in one sentence: a canonical model function is called inside a UDF body, and is not itself a node.
The tracer trusts only decorated functions. It walks the @dl.pipeline body and
refuses any undecorated call it finds there; it does not descend into UDF
bodies at all. So a model function called directly in the pipeline body is a
trace error, and the same call one level down — inside a @ls.udf — is
invisible to the graph. That is exactly the treatment to_dataset gets inside a
sink, and it is the same rule for the same reason.
The graph has two nodes and one edge. sync.detect_objects is not one of them.
The stage's ports come entirely from the decorated wrapper: images is an input
port because it is a parameter, and boxes / scores / labels are columns on
the single out port because they are the return annotation's names. See
The UDF for that derivation in full.
Note what the wrapper is responsible for. detect_objects returns
list[list[Box]] — one list of Box records per image — and the UDF above
declares three columns. Flattening the records into those columns is the
author's job, inside the body, where the graph cannot see it. The return
annotation is a declaration of the stage's schema, not something derived
from the canonical function's return type.
Note also what the signature does not have. detect_objects takes exactly one
argument today: images. There is no threshold, no model, no device. If
you want to write a call with tuning parameters, you are writing against a
surface that does not exist yet.
model is a kind the renderer paints and the tracer has never produced
The figure above uses kind: 'model' deliberately, and it is showing an
intended state rather than a graph anyone has traced.
NodeKind in @dreamlake/uikit admits seven values — source, transform,
model, filter, merge, sink, review — and the renderer gives each its
own kind-dot colour. The tracer's behaviour is narrower. It copies an explicit
kind= off the decorator verbatim, and when there is none it infers exactly two
values: merge when a stage has two or more upstream data inputs, transform
otherwise. The kinds that actually appear in shipped example pipelines are
source, sink and review — written by hand — plus inferred merge and
transform. No shipped pipeline writes kind="model" or kind="filter", so
neither has ever appeared in a real traced graph. Both are renderer
vocabulary waiting for an author.
Two things stand between kind="model" and working code. Lakeshore's real
decorator is udf(fn=None, *, queue=None, transport="auto", config=None) — it
accepts no kind= at all, so kind is a tracer-only annotation that the
tracer reads out of the decorator's AST without ever importing Lakeshore. And
dreamlake-server types a node's kind as 'source' | 'transform' | 'sink',
narrower than what the tracer already emits and much narrower than what the
renderer paints. A model node would not survive the round trip today.
lakeshore.types, the source of the Tuple[…] return annotation, does not
exist in shipped Lakeshore either — the example .py files are a
specification format that happens to be Python syntax, and they trace because
the tracer never imports them.
Where this leaves you
This page documents an intended surface, and says so at the top rather than at the bottom. The vocabulary is real and it is well derived — thirty-three names that cover thirty-nine of forty-seven public task types, with a record type for every output and a synchronous twin for every entry point. The inference is not there. Every body raises.
Writing a pipeline against these names today buys you two real things: the
naming, so that the intent survives whatever eventually implements it, and the
graph, because the tracer never runs a line of it and traces a placeholder body
exactly as it traces a real one. It does not buy you a prediction. Until the
bodies land, a pipeline that calls sync.detect_objects renders beautifully and
raises NotImplementedError the moment anything tries to run it.
For the node and edge schema underneath all of this, and for what the tracer can and cannot emit, start at Pipeline Functions.