Annotations Reference
The complete API of dreamlake.annotation: the generic Annotation for
custom-schema data (with Schema and Track), the
VideoAnnotation preset (video.annotation/v2) with its
Episode handle, and the storage rules that explain the constraints.
New here? Start with the step-by-step Annotations guide.
The annotation family
Every DreamLake annotation is a catalog row (name, schemaType,
visibility) plus a dreamdb space (the data). schemaType is the one
dispatch key, applied the same way everywhere:
| consumer | known schemaType | unknown schemaType |
|---|---|---|
Annotation.open | returns the preset subclass (video.annotation/v2 → VideoAnnotation) | returns the generic Annotation |
| web app | rich view | catalog entry (generic track browser: roadmap) |
Annotation.list(schema_type=…) | server-side filter | same |
Unknown never refuses — data written by newer tools stays readable generically.
Names. Every name-taking classmethod accepts "name" (your own
namespace) or "namespace/name" (an organisation you belong to; the server
authorizes per request). A leading @ on the namespace is tolerated. At
most one /.
Annotation (custom schemas)
Lifecycle (classmethods)
| method | meaning |
|---|---|
Annotation.create(name, *, schema=None, schema_type=None, visibility="private") | create. schema=None starts empty — declare tracks as you go (embeddings excepted: create-time only). schema_type is your own dispatch label, default "custom/v1"; a registered preset type errors and points at the preset's own create |
Annotation.ensure(name, *, schema=None, schema_type=None, visibility="private") | open-or-create, rerun-safe. Existing annotations are verified, never widened: an explicitly expected schema_type errors on mismatch, and every field of schema= must already be declared |
Annotation.open(name) | open; dispatches by the catalog's schemaType (see table above) |
Annotation.list(namespace=None, schema_type=None) | list[AnnotationInfo] — name, namespace, schema_type, visibility |
Annotation.delete(name, *, purge=False) | remove the catalog row; purge=True also deletes storage. A classmethod on purpose: no open, no credentials — and a safe distance from row-level ann.db.delete(anchors) tombstones |
Tracks — the schema, in its live form
There is no ann.schema: ann.tracks() IS the schema — each Track handle
carries name / kind / mime / dim.
| method | meaning |
|---|---|
ann.add_track(name, kind, *, mime=None) -> Track | declare a track (evolution is by addition only). Idempotent for a matching re-declaration; a kind change errors. "embedding" refused post-create |
ann.track(name) -> Track | the handle; an unknown name errors here, eagerly |
ann.tracks() -> list[Track] | every declared track |
kind is the dreamdb vocabulary, verbatim: "video" / "image" (with
mime=; JSON documents are kind="image", mime="json") / "scalar_float"
/ "scalar_int" / "scalar_bool" / "scalar_string" /
"scalar_categorical" / "scalar_timestamp". Track names:
^[a-z0-9][a-z0-9_]*$, max 64 chars; anchor, _anchor,
_time_anchors are reserved.
Tracks added through the bare ann.db handle are invisible to tracks()
(the SDK keeps a fields mirror in the space meta, key
dreamdb.dataset.fields); another process's add_track becomes visible
after ann.reload().
Row-wise write / read
| method | meaning |
|---|---|
ann.append_rows(rows) -> {"rows": n} | one commit for the batch. Each row: {"anchor": ns, track: value, ...}. Sparse rows are the norm — omit a field rather than passing None |
ann.rows(start=None, end=None, tracks=None) | rows in [start, end), sorted, sparse (absent field = absent key). Video tracks are excluded by default; naming one in tracks= errors |
Round-trip contract: rows() returns exactly the shape append_rows()
accepts, and t.read() exactly what t.append_range() accepts —
ann2.append_rows(ann1.rows(…)) needs no conversion.
Introspection and recovery
| method | meaning |
|---|---|
ann.anchors(start=None, end=None) -> list[int] | every item anchor, ascending — count (len), span (ends), upload verification |
ann.reload() -> ann | refresh in place: catalog row, fields mirror, credential lease. Track handles survive. The fix for both documented stalenesses (cross-process add_track; the 12 h lease) |
ann.schema_type / ann.namespace / ann.name | identity |
ann.visibility / ann.set_visibility("public") | public annotations get anonymous presigned reads |
ann.db | the live dreamdb.Dataset — the escape hatch. The meta keys dreamdb.schema_type / dreamdb.dataset.* are the SDK's; overwriting them breaks the handle |
Anchors
Absolute int nanoseconds. Everywhere an anchor or range bound is taken,
a tz-aware datetime also works (naive datetimes are refused, not
guessed). Sequential data with no clock of its own uses row indices:
sequence_anchors(n, start=0, step=1); continue an existing annotation from
ann.anchors()[-1] + 1.
Write semantics (read this before scripting)
- Append-only, write-once. Re-writing the same (anchor, track) is undefined — the engine resolves same-anchor duplicates by content order, not write order. The SDK rejects duplicates within one batch; across calls it is on you. There is no update verb in v1.
- Every
append*/ingestcall is one commit (a new manifest, the ref advances). Batch withappend_rows/append_range; a loop of single-point appends is slow and churns the ref history. - One writer per annotation at a time. Concurrent writers lose updates. Readers are unrestricted.
- Credentials are a 12 h lease brokered at open. A stale handle raises
AnnotationError("credentials expired — call ann.reload()"). One active platform annotation per process (the lease rides process-level AWS env vars).
Schema
Mirrors dreamdb.Schema one-to-one — same method names, same parameters,
chainable — with two twists: every declaration is recorded (dreamdb's
own Schema cannot be introspected once built), and required is pinned
False (a required field could never be added later; all-optional is what
keeps add_track available).
to_fields() / from_fields() carry the fields wire format (also
stamped into the catalog's schemaJson and the space meta). Validation is
eager and actionable: bad/reserved/duplicate names, video without
mime, embedding without an int dim, required=True, and add_audio
(unsupported — the engine cannot ingest it yet) all error at declaration.
Track
Column-wise reads and writes. One verb, append, with the shape in the
name: no suffix = one point at a specific anchor, _range = a stretch of
timeline (and ann.append_rows = the row-wise member of the same family).
| method | meaning |
|---|---|
t.append(anchor, value) -> {"items": 1} | one data point at one anchor |
t.append_range(items) -> {"items": n} | (anchor, value) pairs — sorted by the SDK, one commit; a duplicate anchor within the batch errors (write-once) |
t.get(anchor) | the value at exactly anchor, else None |
t.read(start=None, end=None) | [(anchor, value)] in [start, end), sorted. A declared-but-never-written track reads [] |
t.ingest(src, *, anchor, frag_seconds=2.0, height=None) | video tracks only — see below |
t.name / t.kind / t.mime / t.dim | metadata (dim on embeddings) |
Values. dreamdb's native representations pass through untouched:
bytes for blobs, int ns for timestamps, native scalars, float vectors.
On top, one-way input conveniences: file paths (image), tz-aware
datetime (scalar_timestamp and every anchor), dict/list on
mime="json" tracks, .npy paths and ndarrays (embedding,
dim-checked). Scalars are strict — bool into scalar_float errors
rather than coercing. None is always refused: absence is an omitted
field, not a null. Reads return the stored representation as-is, with one
exception: mime="json" tracks decode back to the dict/list that
went in.
Video. t.ingest is the only write path for video tracks
(append on one errors and says so). height=None is a lossless remux —
fast, but a CMAF track accepts exactly one codec configuration, so every
clip on the track must be identically encoded. height=N re-encodes to a
uniform h264 profile so mixed sources can share a track. Each clip
occupies [anchor, anchor + duration); overlapping an existing span is
refused up front. Ranged video reads are not in v1 — playback goes
through the platform, raw bytes via ann.db.
VideoAnnotation (the video.annotation/v2 preset)
An annotation of episodes: one recording each — one or more camera
videos on a shared clock, plus per-frame joint annotations and action
segments. Each episode owns a one-hour timeline slot; every camera,
annotation and search vector of the episode lives inside that slot, which
is why multi-camera playback is time-aligned with zero bookkeeping.
Per-episode operations live on the Episode handle.
The preset subclasses the generic Annotation and registers its schemaType:
Annotation.open on one of these returns this class.
VideoAnnotation.create(name=None, *, backend=None, visibility=None, preview_height=720, preview_fps=30.0, frag_seconds=2.0)
| param | meaning |
|---|---|
name | platform mode: catalog entry + managed bucket (needs dreamlake login / DREAMLAKE_API_KEY). Accepts namespace/name |
backend | self-hosted mode: file:///abs/path or https://… S3 URL — no platform involved |
visibility | "private" (default) / "public"; public annotations omit absolute source paths from metadata |
preview_height / preview_fps / frag_seconds | the encoding profile — playback resolution, frame rate, and fragment length. Chosen once, for the annotation's lifetime, stored in the space meta. Per-episode calls never pass encoding |
Creation is never idempotent: an existing name/path errors instead of
silently forking a second history. (VideoAnnotation.ensure(name)
is the open-or-create form; it takes no schema arguments — the preset owns
its schema.)
VideoAnnotation.open(name=None, *, backend=None, preview_height=None, preview_fps=None, frag_seconds=None)
Open by platform name or backend URI, strictly — a space stamped with a
different schemaType is refused (use Annotation.open or dreamlake.db for
those). The optional encoding kwargs verify, never set: pass the
values your script assumes and open() errors on mismatch — so the
try open / except create idiom cannot silently eat a config edit.
ann.add_episode(videos, *, episode_id=None, joints_pose=None, subtasks=None, recon_mesh=None, recon_pose=None, recon_camera=None, recon_hands=None, recon_gravity=None, meta=None, gid=None, raw=False) -> Episode
The upload. One call transcodes every camera for browser playback and
commits annotations + metadata atomically. Returns the Episode handle;
the ingest report (per-camera fragment counts, annotation counts) is on
epo.report.
| param | meaning |
|---|---|
videos | one path (single camera, stored as main) or {camera: path}. Camera names: [a-z0-9_-], no __ |
joints_pose | a doc (binds to the primary camera — main if present, else the first) or {camera: doc} to annotate several views. Docs may also be JSON file paths |
subtasks | episode-level action segments (doc or path) |
recon_mesh | 3D-reconstruction: object meshes {object: obj_text} (episode-level, no camera) |
recon_pose recon_camera recon_hands recon_gravity | 3D-reconstruction, per camera: object 6-DoF poses / pinhole intrinsics / MANO hands / gravity up-vector. Bare doc → primary camera, or {camera: doc}. All optional — see Annotation formats |
meta | labels: {"task": ..., "scene": ...} — whitelist-only, nothing is inferred |
gid | pin a specific slot (rarely needed) |
raw | True additionally archives a lossless remux (roughly doubles the upload; best-effort on mixed sources) |
Fails fast, before any transcoding: any camera ≥ 3600 s, duplicate
episode_id, unknown meta key, joints naming a camera not in videos,
annotation fps disagreeing with its camera by >10 accumulated frames, or an
aspect ratio differing from that camera's existing track (aspect is per
camera — head 4:3 and wrist 16:9 coexist).
ann.episode(episode_id) / ann.episodes(after_gid=None, limit=None) / ann.episode_count()
Lookup, listing, and count — all scale-safe: id lookups and counts go
through a slim per-episode index track (a scalar_string of bare ids —
one cacheable fetch yields the whole id→slot map; the fat episode_meta
column is only ranged-read for the slots actually needed).
Bare episodes() is the full listing (fine up to a few thousand); past
that, page: episodes(limit=100) then
episodes(after_gid=page[-1].gid, limit=100) — each page is one
slot-window read. [e.meta for e in ann.episodes()] recovers plain dicts.
ann.cameras() -> list[str] / ann.tracks() -> list[Track]
What is in this annotation. cameras() is the union over episodes, main
first. tracks() returns the same Track handles as the generic layer
(name/kind/mime/dim) plus the preset extras role (video_preview,
joints_pose, subtasks, recon_mesh, recon_pose, recon_camera,
recon_hands, recon_gravity, episode_meta, search, user), camera,
and the derived preset flag. Preset reads still go through the named
APIs (read_joints_pose, epo.read_track, …).
ann.add_track(name, kind, *, mime=None) -> Track
Declare a custom column. Same kind vocabulary as the generic layer, with
one extra rule: names must match ^x_[a-z0-9_]+$ — the user namespace the
preset promises never to claim, so your columns can never collide with a
future SDK release. Re-declaring the same name+kind is idempotent; changing
a track's kind is refused; embeddings cannot be added after creation. The
web viewer does not render x_* tracks; they are first-class data for
training loaders. Write and read through the Episode handle.
ann.embed_episodes(*, camera=None, fps=1.0, source_dir=None, batch_size=32) -> dict
The search sweep: for every episode, sample frames at fps from one
camera's source file (primary unless camera=), encode with CLIP, encode
subtask texts with BGE, upload. Episodes commit one at a time — interrupt
and re-run freely. source_dir locates moved files via each camera's
recorded source_rel. Needs pip install "dreamlake[search]". One episode
= epo.embed().
ann.search(query, top_k=10, kind="both") -> list[dict]
Natural-language moments: CLIP text tower against frame vectors and/or BGE
against subtask texts, fused by reciprocal rank. Returns
[{"episode_id", "time_sec", "score", "source", "subtask"?}].
kind: "frames" / "subtasks" / "both".
ann.encoding / ann.db
encoding — the stored profile, read-only: the create-time defaults plus
cameras, each camera track's adopted playback profile
({width, height, fps}, fixed by the aspect ratio of that camera's first
clip; every later clip is validated against it before any transcoding).
db — the live dreamdb handle:
the escape hatch for anything outside the preset (absolute-anchor appends,
columnar bulk reads, branching). Layout invariants are yours to respect.
Episode
Obtained from add_episode / ann.episode(id) / ann.episodes(). The
identity triple (episode_id, gid, anchor) is immutable — handles never
dangle. The meta snapshot is a read convenience; every write re-reads the
stored metadata at call time, so a stale handle can never clobber a newer
revision.
| member | meaning |
|---|---|
episode_id / gid / anchor | identity: your stable id, the slot number, the slot's base time |
report | ingest report (only on the handle add_episode returned) |
meta / cameras / task / scene / duration_s | snapshot reads, zero IO |
refresh() | re-read metadata, returns self |
info() | fresh meta + annotation summaries per camera (does IO) |
Reads
epo.add_cameras(videos, *, joints_pose=None, raw=False) -> dict
Late-arriving cameras — add-only (video fragments occupy their slot and cannot be replaced; re-sending an existing camera errors). Dict form adds N cameras with one metadata write.
epo.revise(*, joints_pose=None, subtasks=None, recon_mesh=None, recon_pose=None, recon_camera=None, recon_hands=None, recon_gravity=None, meta=None) -> dict
The one revision verb: everything passed lands in one committed row.
Storage is append-only — a revision is a new version, readers see the
newest, history stays addressable. meta takes the same whitelist as
add_episode; there is no inference, and passing values identical to what
is stored is reported as an error rather than a silent no-op.
Custom-track values (declare with ann.add_track first)
Values follow the track's kind (dicts/lists round-trip as JSON on
image/json tracks). Timestamps are seconds on the episode's own clock and
are bounds-checked to the slot. Re-writing the same (track, time) is a
revision — the preset's episode clock has revision semantics the generic
layer's absolute anchors do not. For high-rate series and cross-episode
bulk reads, drop to the engine with epo.anchor_at(t_sec) +
ann.db.append_many / ann.db.iter_all_batches(fields=[...]).
Search, single episode
embed resolves the source file video_path → source_dir/source_rel →
recorded absolute path, and refuses a resolved file whose duration
disagrees with the ingested camera's (wrong file); an explicit video_path
overrides the check. Vectors are searchable the moment they land — there is
no index-build step.
Annotation formats
Wire-compatible with the web viewer's overlays. Joint coordinates live in the pixel space of the camera whose track holds the document.
Reconstruction (3D)
An optional third modality: 3D hand–object reconstruction, as five flat
recon_* arguments (one per stored track). Geometry is in the camera's
OpenCV frame (x-right / y-down / z-forward), metres, quaternion wxyz,
frame f ↔ time f/fps; colour is not stored. recon_mesh is episode-level;
the other four are per-camera (bare doc → primary, or {camera: doc}).
Each is optional; at least one present. Re-passing a piece revises it. recon_pose
tells a bare per-frame doc from a {camera: doc} map by whether the keys are
frame indices vs camera names — so don't name a camera a bare number.
Storage semantics worth knowing (preset)
- Slots. Episode
gidoccupies[gid × 1 h, (gid+1) × 1 h)on one timeline; slots are never reused; any single camera clip ≤ 3600 s. - Encoding is annotation-lifetime. Every clip on one camera track shares
one init segment, so height/fps/fragment length cannot vary per episode —
that is why they live on
create, notadd_episode. - Aspect ratio is per camera track, not per annotation.
- Revisions are versioned, bounded, and never in-place. Each episode-level value carries up to 1024 revisions; readers always resolve to the newest; history stays in the timeline.
- Video is add-only. Cameras can join an episode; fragments are never replaced.
x_is yours, everything else is the contract. Preset track names and metadata keys evolve only by addition, in SDK releases; tools render what they recognize and skip the rest.