DreamLake

Sources (DreamDB)

A DreamLake-hosted store for your own videos and structured data. Push from the terminal — each video is sliced into streamable fragments automatically.

Install

bash
pip install "dreamlake[auth,dreamdb]"    # or: uv add "dreamlake[auth,dreamdb]"
brew install ffmpeg                      # apt install ffmpeg on Linux

ffmpeg must be on PATH for any push with video.

Log in

bash
dreamlake login                          # OAuth device flow

Model: source → collection → data

Think database → table → rows.

LayerWhat it isCommand
SourceNamespaced, permissioned storesource create
CollectionA dataset with a schemasource collection create
DataRecords at time anchorssource push

One source holds many collections; access is granted per source.

Quick start — a folder of videos

bash
dreamlake sources create charlie-57/my-videos
dreamlake sources collections create clips --source my-videos --preset video
dreamlake sources push ./videos --source my-videos --collection clips

--preset video = two fields (video + path); push <dir> uploads every video as one record. Each is normalised to H.264 854×480 (--width/--height to change) and sliced into ~2s fragments.

Naming a source

A source name may carry its namespace as a prefix:

[<namespace>/]<name>

charlie-57/my-videos creates the source under the charlie-57 namespace. Omit the prefix and it lands in your own namespace — dreamlake sources create my-videos is the common case and needs nothing extra.

The prefix is how you write into an org you own: dreamlake sources create charlie-57-org/my-videos. Ownership is still required; the syntax only says where, not whether you may.

The shipped CLI still uses the old spelling

This page describes the target surface. dreamlake-py ships dreamlake source (singular), dreamlake source collection (singular), and takes the namespace as --ns <slug> rather than as a namespace/name prefix. Until the rename lands, translate: dreamlake sources create ns/name is dreamlake source create name --ns ns today.

One argument, not two

The namespace travels with the name rather than in a separate --ns flag, so a source is always identifiable from a single string — in a command, in a config file, or pasted into a message. The unprefixed form stays short for the case that dominates.

Custom schema

Write a schema JSON, pass it to collection create:

schema.jsonjson
{
  "fields": [
    { "name": "video",    "type": "video", "mime": "h264" },
    { "name": "path",     "type": "scalar_string" },
    { "name": "category", "type": "scalar_categorical" },
    { "name": "verified", "type": "scalar_bool" }
  ]
}
bash
dreamlake sources collections create labeling --source charlie-57/my-videos --schema schema.json

Field types

TypeValue in a record
videofile path — needs mime (e.g. "h264")
imagefile path (raw bytes)
embedding.npy path or inline list — needs dim
scalar_string · scalar_categoricala string
scalar_int · scalar_timestampan integer (timestamp = ns)
scalar_floata number
scalar_booltrue / false

Push structured data — manifest

A folder can't carry scalar values, so custom schemas push with a manifest: one JSON listing fields + every record. Its fields block is a schema — same file works for both steps.

bash
dreamlake sources collections create labeling --source my-videos --schema manifest.json
dreamlake sources push --manifest manifest.json --source my-videos --collection labeling
manifest.jsonjson
{
  "fields": [
    { "name": "video",       "type": "video",     "mime": "h264" },
    { "name": "thumb",       "type": "image",      "mime": "jpeg" },
    { "name": "clip",        "type": "embedding",  "dim": 512 },
    { "name": "path",        "type": "scalar_string" },
    { "name": "frame_count", "type": "scalar_int" },
    { "name": "duration_s",  "type": "scalar_float" },
    { "name": "verified",    "type": "scalar_bool" },
    { "name": "category",    "type": "scalar_categorical" },
    { "name": "recorded_at", "type": "scalar_timestamp" }
  ],
  "records": [
    {
      "anchor": 0,
      "video": "clips/a.mp4", "thumb": "thumbs/a.jpg", "clip": "embeds/a.npy",
      "path": "a.mp4", "frame_count": 360, "duration_s": 12.0,
      "verified": true, "category": "assembly", "recorded_at": 1750000000000000000
    },
    {
      "anchor": 3600000000000,
      "video": "clips/b.mp4", "thumb": "thumbs/b.jpg", "clip": "embeds/b.npy",
      "path": "b.mp4", "frame_count": 180, "duration_s": 6.0,
      "verified": false, "category": "electronics", "recorded_at": 1750003600000000000
    }
  ]
}
  • video / image — path relative to the manifest.
  • embedding — a .npy path, or an inline list of exactly dim numbers.
  • anchor — required per record, integer ns. For unrelated videos use index × 3_600_000_000_000 (a 1-hour slot each).

Lay files out next to the manifest: clips/ thumbs/ embeds/.

Everything is validated before anything uploads — bad manifest fails fast, never a half-written collection.

Command reference

bash
dreamlake sources create [<namespace>/]<name>

dreamlake sources collections create <name> --source [<namespace>/]<src> \
    [--preset video | --schema <schema.json>]

dreamlake sources push <dir> --source [<namespace>/]<src> --collection <col> \
    [--width <px>] [--height <px>] [--frag-duration <seconds>]
dreamlake sources push --manifest <manifest.json> --source [<namespace>/]<src> --collection <col>