LLM Relay
Agents need a model. Handing every user a provider API key does not scale and does not revoke. The relay puts the key on the server: you authenticate with the DreamLake token you already have, and DreamLake swaps credentials on the way out.
The provider's wire format is preserved byte-for-byte in both directions. That is the whole design: the official SDKs, Claude Code, and the Claude Agent SDK work against the relay unmodified. You change a base URL — nothing else.
Quick start
Point any Anthropic client at the relay and give it your DreamLake token:
Python tab: The same two variables work for anything built on the Anthropic SDK — your own agent loop, a script, a CI job.
curl tab: Or with no SDK at all:
export ANTHROPIC_BASE_URL=https://api.dreamlake.ai/relay/anthropic
export ANTHROPIC_AUTH_TOKEN=$(dreamlake token) # your DreamLake token, not a provider key
claude "summarize the last three runs in this namespace"import anthropic, os
client = anthropic.Anthropic(
base_url="https://api.dreamlake.ai/relay/anthropic",
auth_token=os.environ["DREAMLAKE_API_KEY"],
)
response = client.messages.create(
model="claude-opus-5",
max_tokens=16000,
thinking={"type": "adaptive"},
messages=[{"role": "user", "content": "What changed in this run?"}],
)
print(next(b.text for b in response.content if b.type == "text"))curl https://api.dreamlake.ai/relay/anthropic/v1/messages \
-H "authorization: Bearer $DREAMLAKE_TOKEN" \
-H "content-type: application/json" \
-d '{
"model": "claude-opus-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "ping"}]
}'Streaming works the same way — set "stream": true and read the SSE events.
The relay does not buffer streams; chunks reach you as the provider emits them.
Which model?
Anthropic Claude, default claude-opus-5. Three reasons:
- The agents are Claude-shaped. CLI agents, Claude Code, and the Agent SDK
all speak the Anthropic Messages API. Relaying that format verbatim means
they need no adapter — just
ANTHROPIC_BASE_URL. - The stack already runs on it. Hosted Dream Chat containers already call Claude. A second provider family would mean two prompt formats, two token accountings, and two sets of model IDs.
- Opus 5 is the right default tier for agentic and coding work. Want
cheaper turns? Pass
claude-haiku-4-5in the body — the relay never rewrites your model, it only reports and (optionally) restricts it.
The relay does not pick models, inject system prompts, or count turns. It is a credential boundary and a usage log. Agent behavior belongs in the agent.
An openai provider exists behind an optional key, because the pass-through is
provider-agnostic. It is not configured by default.
What you can call
Everything under /relay/:provider/ is forwarded unchanged, so the provider's
whole API comes along — not just chat completions:
| You call | Reaches |
|---|---|
POST /relay/anthropic/v1/messages | POST https://api.anthropic.com/v1/messages |
POST /relay/anthropic/v1/messages/count_tokens | the token-counting endpoint |
GET /relay/anthropic/v1/models | the live model list |
The upstream host is fixed per provider, so the relay cannot be pointed at an arbitrary destination.
To discover what a deployment actually offers:
baseUrl is exactly what goes in ANTHROPIC_BASE_URL.
Authentication
Send the DreamLake token on either header — the relay accepts both:
Both exist because provider SDKs send the key differently depending on which
variable you set: ANTHROPIC_AUTH_TOKEN produces Authorization,
ANTHROPIC_API_KEY produces x-api-key. Either works, so you don't have to
remember which one your client picked.
Your token is verified the same way as every other DreamLake route, and is never forwarded to the provider.
Errors
| Status | Meaning |
|---|---|
| 401 | Missing, invalid, or expired DreamLake token |
| 403 | The deployment restricts models and yours isn't on the list |
| 404 | Unknown provider — call GET /relay/providers |
| 503 | That provider has no credentials on this deployment |
| 502 | Upstream unreachable |
| anything else | The provider's own status and body, verbatim |
That last row matters: a 429 from Anthropic reaches your SDK as a 429 with its
retry-after header intact, so your client's normal retry logic works. The
relay never reserializes a provider response, so no field is ever stripped.
What it does not do
- No per-user quota. Any authenticated DreamLake user can spend tokens. A deployment can cap which models are allowed, not how many calls.
- No caching, batching, or retries. Your SDK's own retry logic works.
- No prompt logging. Request bodies are relayed, never stored — only metadata (user, model, status, duration, token usage) is logged.
- Not a replacement for hosted chat. Hosted pages run a whole agent container. This is a raw model API for callers running their own loop.