# Sources (DreamDB)

  A DreamLake-hosted store for your own videos and structured data. Push from
  the terminal — each video is sliced into streamable fragments automatically.

## Install

```bash
pip install "dreamlake[auth,dreamdb]"    # or: uv add "dreamlake[auth,dreamdb]"
brew install ffmpeg                      # apt install ffmpeg on Linux
```

`ffmpeg` must be on `PATH` for any push with video.

## Log in

```bash
dreamlake login                          # OAuth device flow
```

## Model: source → collection → data

Think **database → table → rows**.

| Layer | What it is | Command |
|-------|-----------|---------|
| **Source** | Namespaced, permissioned store | `source create` |
| **Collection** | A dataset with a schema | `source collection create` |
| **Data** | Records at time anchors | `source push` |

One source holds many collections; access is granted per source.

## Quick start — a folder of videos

```bash
dreamlake sources create charlie-57/my-videos
dreamlake sources collections create clips --source my-videos --preset video
dreamlake sources push ./videos --source my-videos --collection clips
```

`--preset video` = two fields (`video` + `path`); `push <dir>` uploads every
video as one record. Each is normalised to H.264 854×480 (`--width`/`--height`
to change) and sliced into ~2s fragments.

## Naming a source

A source name may carry its namespace as a prefix:

```
[<namespace>/]<name>
```

`charlie-57/my-videos` creates the source under the `charlie-57` namespace.
**Omit the prefix and it lands in your own namespace** — `dreamlake sources
create my-videos` is the common case and needs nothing extra.

The prefix is how you write into an org you own: `dreamlake sources create
charlie-57-org/my-videos`. Ownership is still required; the syntax only says
where, not whether you may.

> **Warning:** This page describes the target surface. `dreamlake-py` ships
>   `dreamlake source` (singular), `dreamlake source collection` (singular), and
>   takes the namespace as `--ns <slug>` rather than as a `namespace/name` prefix.
>   Until the rename lands, translate: `dreamlake sources create ns/name` is
>   `dreamlake source create name --ns ns` today.

> **Note:** The namespace travels with the name rather than in a separate `--ns` flag, so
>   a source is always identifiable from a single string — in a command, in a
>   config file, or pasted into a message. The unprefixed form stays short for the
>   case that dominates.

## Custom schema

Write a schema JSON, pass it to `collection create`:

```json file="schema.json"
{
  "fields": [
    { "name": "video",    "type": "video", "mime": "h264" },
    { "name": "path",     "type": "scalar_string" },
    { "name": "category", "type": "scalar_categorical" },
    { "name": "verified", "type": "scalar_bool" }
  ]
}
```

```bash
dreamlake sources collections create labeling --source charlie-57/my-videos --schema schema.json
```

### Field types

| Type | Value in a record |
|------|-------------------|
| `video` | file path — needs `mime` (e.g. `"h264"`) |
| `image` | file path (raw bytes) |
| `embedding` | `.npy` path or inline list — needs `dim` |
| `scalar_string` · `scalar_categorical` | a string |
| `scalar_int` · `scalar_timestamp` | an integer (timestamp = ns) |
| `scalar_float` | a number |
| `scalar_bool` | `true` / `false` |

## Push structured data — manifest

A folder can't carry scalar values, so custom schemas push with a **manifest**:
one JSON listing fields + every record. Its `fields` block **is** a schema —
same file works for both steps.

```bash
dreamlake sources collections create labeling --source my-videos --schema manifest.json
dreamlake sources push --manifest manifest.json --source my-videos --collection labeling
```

```json file="manifest.json"
{
  "fields": [
    { "name": "video",       "type": "video",     "mime": "h264" },
    { "name": "thumb",       "type": "image",      "mime": "jpeg" },
    { "name": "clip",        "type": "embedding",  "dim": 512 },
    { "name": "path",        "type": "scalar_string" },
    { "name": "frame_count", "type": "scalar_int" },
    { "name": "duration_s",  "type": "scalar_float" },
    { "name": "verified",    "type": "scalar_bool" },
    { "name": "category",    "type": "scalar_categorical" },
    { "name": "recorded_at", "type": "scalar_timestamp" }
  ],
  "records": [
    {
      "anchor": 0,
      "video": "clips/a.mp4", "thumb": "thumbs/a.jpg", "clip": "embeds/a.npy",
      "path": "a.mp4", "frame_count": 360, "duration_s": 12.0,
      "verified": true, "category": "assembly", "recorded_at": 1750000000000000000
    },
    {
      "anchor": 3600000000000,
      "video": "clips/b.mp4", "thumb": "thumbs/b.jpg", "clip": "embeds/b.npy",
      "path": "b.mp4", "frame_count": 180, "duration_s": 6.0,
      "verified": false, "category": "electronics", "recorded_at": 1750003600000000000
    }
  ]
}
```

- **`video` / `image`** — path relative to the manifest.
- **`embedding`** — a `.npy` path, or an inline list of exactly `dim` numbers.
- **`anchor`** — required per record, integer ns. For unrelated videos use
  `index × 3_600_000_000_000` (a 1-hour slot each).

Lay files out next to the manifest: `clips/  thumbs/  embeds/`.

Everything is validated **before** anything uploads — bad manifest fails fast,
never a half-written collection.

## Command reference

```bash
dreamlake sources create [<namespace>/]<name>

dreamlake sources collections create <name> --source [<namespace>/]<src> \
    [--preset video | --schema <schema.json>]

dreamlake sources push <dir> --source [<namespace>/]<src> --collection <col> \
    [--width <px>] [--height <px>] [--frag-duration <seconds>]
dreamlake sources push --manifest <manifest.json> --source [<namespace>/]<src> --collection <col>
```
