# LLM Relay

  Agents need a model. Handing every user a provider API key does not scale and
  does not revoke. The relay puts the key on the server: you authenticate with
  the DreamLake token you already have, and DreamLake swaps credentials on the
  way out.

```
POST /relay/anthropic/v1/messages          ← what you call
     ↓  DreamLake token verified, dropped, provider key injected
POST https://api.anthropic.com/v1/messages
```

The provider's wire format is preserved **byte-for-byte in both directions**.
That is the whole design: the official SDKs, Claude Code, and the Claude Agent
SDK work against the relay unmodified. You change a base URL — nothing else.

## Quick start

Point any Anthropic client at the relay and give it your DreamLake token:

**Python tab:** The same two variables work for anything built on the Anthropic SDK — your own
agent loop, a script, a CI job.

**curl tab:** Or with no SDK at all:

**Claude Code**

```bash
export ANTHROPIC_BASE_URL=https://api.dreamlake.ai/relay/anthropic
export ANTHROPIC_AUTH_TOKEN=$(dreamlake token)   # your DreamLake token, not a provider key
claude "summarize the last three runs in this namespace"
```

**Python**

```python
import anthropic, os

client = anthropic.Anthropic(
    base_url="https://api.dreamlake.ai/relay/anthropic",
    auth_token=os.environ["DREAMLAKE_API_KEY"],
)

response = client.messages.create(
    model="claude-opus-5",
    max_tokens=16000,
    thinking={"type": "adaptive"},
    messages=[{"role": "user", "content": "What changed in this run?"}],
)
print(next(b.text for b in response.content if b.type == "text"))
```

**curl**

```bash
curl https://api.dreamlake.ai/relay/anthropic/v1/messages \
  -H "authorization: Bearer $DREAMLAKE_TOKEN" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-opus-5",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "ping"}]
  }'
```

Streaming works the same way — set `"stream": true` and read the SSE events.
The relay does not buffer streams; chunks reach you as the provider emits them.

## Which model?

**Anthropic Claude, default `claude-opus-5`.** Three reasons:

1. **The agents are Claude-shaped.** CLI agents, Claude Code, and the Agent SDK
   all speak the Anthropic Messages API. Relaying that format verbatim means
   they need no adapter — just `ANTHROPIC_BASE_URL`.
2. **The stack already runs on it.** Hosted Dream Chat containers already call
   Claude. A second provider family would mean two prompt formats, two token
   accountings, and two sets of model IDs.
3. **Opus 5 is the right default tier** for agentic and coding work. Want
   cheaper turns? Pass `claude-haiku-4-5` in the body — the relay never
   rewrites your model, it only reports and (optionally) restricts it.

The relay does not pick models, inject system prompts, or count turns. It is a
credential boundary and a usage log. Agent behavior belongs in the agent.

An `openai` provider exists behind an optional key, because the pass-through is
provider-agnostic. It is not configured by default.

## What you can call

Everything under `/relay/:provider/` is forwarded unchanged, so the provider's
whole API comes along — not just chat completions:

| You call | Reaches |
| --- | --- |
| `POST /relay/anthropic/v1/messages` | `POST https://api.anthropic.com/v1/messages` |
| `POST /relay/anthropic/v1/messages/count_tokens` | the token-counting endpoint |
| `GET /relay/anthropic/v1/models` | the live model list |

The upstream host is fixed per provider, so the relay cannot be pointed at an
arbitrary destination.

To discover what a deployment actually offers:

```bash
curl https://api.dreamlake.ai/relay/providers -H "authorization: Bearer $DREAMLAKE_TOKEN"
```

```jsonc
{
  "default": "anthropic",
  "allowedModels": null,           // or a list, when the deployment restricts models
  "providers": [
    {
      "id": "anthropic",
      "configured": true,          // false → that provider answers 503
      "baseUrl": "https://api.dreamlake.ai/relay/anthropic",
      "defaultModel": "claude-opus-5"
    }
  ]
}
```

`baseUrl` is exactly what goes in `ANTHROPIC_BASE_URL`.

## Authentication

Send the DreamLake token on **either** header — the relay accepts both:

```
Authorization: Bearer <dreamlake token>
x-api-key: <dreamlake token>
```

Both exist because provider SDKs send the key differently depending on which
variable you set: `ANTHROPIC_AUTH_TOKEN` produces `Authorization`,
`ANTHROPIC_API_KEY` produces `x-api-key`. Either works, so you don't have to
remember which one your client picked.

Your token is verified the same way as every other DreamLake route, and is
**never forwarded to the provider**.

## Errors

| Status | Meaning |
| --- | --- |
| 401 | Missing, invalid, or expired DreamLake token |
| 403 | The deployment restricts models and yours isn't on the list |
| 404 | Unknown provider — call `GET /relay/providers` |
| 503 | That provider has no credentials on this deployment |
| 502 | Upstream unreachable |
| *anything else* | **The provider's own status and body, verbatim** |

That last row matters: a 429 from Anthropic reaches your SDK as a 429 with its
`retry-after` header intact, so your client's normal retry logic works. The
relay never reserializes a provider response, so no field is ever stripped.

## What it does not do

- **No per-user quota.** Any authenticated DreamLake user can spend tokens. A
  deployment can cap *which* models are allowed, not how many calls.
- **No caching, batching, or retries.** Your SDK's own retry logic works.
- **No prompt logging.** Request bodies are relayed, never stored — only
  metadata (user, model, status, duration, token usage) is logged.
- **Not a replacement for hosted chat.** [Hosted pages](/hosted-pages.md) run a
  whole agent container. This is a raw model API for callers running their own
  loop.

## Next steps

    The relay endpoints alongside the rest of the REST surface.

    Where a relayed model call fits in an agentic workflow.
