# Notes audio playback

Unreleased — 2026-09-27. The Notes title has a borderless waveform button.
It opens a player over the document and reads the current note in short passages.
This implementation uses ElevenLabs; no browser speech fallback is silently substituted.

## Credentials

Set `ELEVENLABS_API_KEY` and `ELEVENLABS_VOICE_ID` in the **dreamlake-server** runtime
secret configuration. For local development use the server's ignored `.env` file.
`ELEVENLABS_MODEL_ID` defaults to `eleven_flash_v2_5`.
Create a restricted ElevenLabs key with text-to-speech access and a provider-side
credit ceiling. Use separate keys for development and production. Restart the
server after changing its environment. Never put the key in `VITE_*`, the frontend,
a note, source control, or chat. Users use their existing DreamLake login; they do
not supply individual ElevenLabs keys.

The internal browser transport is authenticated
`POST /namespaces/:slug/notes/:noteId/speech`, with `{ "text": "A short passage." }`.
Unreleased — 2026-09-29: signed-in namespace members and recipients with an
active, previously accepted read or edit share grant can synthesize audio for a
live note. The same access check protects voice options. Membership in the
recipient's own workspace is not required. Anonymous visitors and signed-in
public readers without a share grant cannot synthesize audio. Open the shared
note while signed in to accept its link before using playback. Disabling sharing
or removing the recipient's grant blocks subsequent audio requests; already
buffered audio may continue playing. The existing per-user limits apply equally
to members and share recipients. Voice-options failures display the server's
error without promising that the default voice will work.
Missing key or voice configuration returns 503 with a setup message.

The server forwards to ElevenLabs `stream/with-timestamps` with `pcm_24000` and
streams NDJSON through without buffering a full note. Browser requests cannot
choose arbitrary provider URLs, voices, or models. The client schedules PCM in
Web Audio as chunks arrive. It highlights a passage, rather than attempting
word alignment against Markdown transformed for speech. Provider alignment data
is available in the transport for a future word-highlight mode.

Each request accepts at most 1,500 characters. Per server process, each user has
two concurrent requests and a 60,000-character hourly budget. A request is
aborted on client disconnect or after 45 seconds. These are process-local
safety limits, not globally distributed billing quotas: the provider-side key
credit ceiling must be configured before production enablement. The server
does not forward provider diagnostics to the client.

Server-side caching — 2026-09-30: completed speech responses (PCM audio and alignment
timestamps) are saved permanently and privately in `S3_BUCKET` under
`speech-cache/v1/<namespaceId>/<noteId>/<sha256>.ndjson`. The digest covers the
voice, output format, and exact provider payload (text, model, and surrounding
context). Every replay rechecks current note access; no public or presigned audio
URL is issued. Cache hits avoid provider usage but retain existing playback limits.
New responses stream immediately and are saved only after successful completion;
failed, cancelled, or responses over 16 MiB are not stored. Storage failures fall
back to synthesis. There is no application expiry or automatic deletion for this
prefix; do not apply an expiring S3 lifecycle rule to it.
The response header `x-speech-cache` reports `hit` or `miss` for verification.
No public skill impact: this changes only the internal browser transport, with
no CLI or SDK task contract changes.
ElevenLabs receives the passages requested for playback; its account retention
settings apply. Stopping cancels the current request; text already submitted
may already have incurred provider usage.

## Editing and playback

- The caret and normal selection remain independent from the passive playback
  highlight. Clicking text never seeks audio.
- Typing or scrolling manually suspends follow mode, without pausing audio.
  **Back to audio** reveals the current passage and restores follow mode.
- The currently spoken passage finishes from its original audio. Editing inside
  it replaces the highlight with a side marker. Anchors map through local and
  collaborative edits. Future passages use the latest document text.
- **Read from cursor / selection** explicitly starts at the caret or restricts
  reading to the selection. The selection's end tracks subsequent edits.
- The UIKit slider seeks within the buffered audio of the current passage.
  This first implementation does not show an estimated whole-note duration.
- Pause/resume and speed changes are local. Released — 2026-10-01, UI PR
  [#658](https://github.com/dreamlake-ai/dreamlake-ai/pull/658):
  playback speed preserves voice pitch through client-side time stretching of
  the PCM stream. The 1× path plays the original audio. Faster rates load a
  Signalsmith Stretch AudioWorklet/WASM processor; progress stays in original
  audio seconds while remaining time accounts for the selected speed. Startup
  waits for the processor and buffered audio, and passage completion drains
  its output tail. If pitch-preserving processing is unavailable or fails,
  playback returns to 1× and shows an error rather than raising the voice pitch.
  The last ten completed passages are
  cached in memory per mounted note/player and reused when their spoken text
  matches. Closing clears the cache. Code fences and Markdown destinations are
  omitted; this is a bounded text normalizer, not a complete Markdown narrator.

## Local voice preferences

Unreleased — 2026-09-30: manual speaker and ElevenLabs model selections become
favorites saved per signed-in user in this browser's localStorage. On each player
opening, an epsilon-greedy bandit picks independently from available favorites:
20% uniform exploration, otherwise the best smoothed preference score (random
ties). Changing a selection rewards the chosen option and rejects the replaced
option. Automatic picks never count as positive feedback. Choices stay fixed
while reading, seeking, pausing, or resuming; removed provider options are excluded.

Memory uses a compact score table (12 favorites per dimension) plus the last 24
manual edits. Scores decay on feedback so recent preferences carry more weight.
Account Settings includes a **Reading voice preferences** section with saved
speaker/model names, individual removal, memory usage, and a reset action. These
changes apply on the next player opening and do not interrupt active audio.
The serialized record is capped at 8 KiB in UTF-8, pruning oldest edits and then
least recently selected favorites. It stores IDs, display names, scores and timestamps, never note text
or audio. Invalid or inaccessible storage falls back to the existing defaults.
A table alone loses the recent change trail; an edit list alone loses learning
when pruned. The bounded combination keeps both without replaying a growing log.
No public skill impact: this is browser-only playback preference state.

## Verification

The UI includes `tests/notes-audio-preview`, a real CodeMirror + player harness
with a quiet synthetic PCM stream, explicitly labeled as synthetic. It exercises
browser audio, pause/resume, selection, scroll behavior and provider errors
without credentials. It does not prove ElevenLabs voice output.

Pitch signal tests render the real WASM engine with a 440 Hz tone at changed
speeds, checking pitch and duration across chunk boundaries, rate changes,
underruns, and a 48 kHz context. Worklet lifecycle tests use an empty input list
for idle, pause, and disposal, matching the browser buffer-only graph. A Chrome
preview measurement of its synthetic 220 Hz stream at 2× returned 219.7 Hz.
This establishes pitch preservation for synthetic signals; physical iPad voice
quality still requires device acceptance.

![Synthetic Notes preview at 2× after pause and paragraph replay](/screenshots/notes-audio/pitch-preserving-2x.png)

UI tests cover Markdown ranges, edit mapping, packet boundaries, odd PCM byte
boundaries, pause/resume and network underruns. Server tests cover access checks,
missing configuration, provider error sanitization, request limits and stream
cleanup. Live voice acceptance still requires a configured key and voice.

The generated workspace `dreamlake` reference includes this operations page.
No distributed task-skill impact: this is an internal browser transport and server
operations page. The public Notes document/CLI API and task procedures are
unchanged; no new CLI speech command or SDK contract is introduced.

## Browser authentication transport

The application sends OAuth token and user-info requests to the DreamLake API
(`/api/auth/oauth2/token` and `/api/auth/oauth2/userinfo`). Authorization still
opens the identity provider. The backend forwards only authorization-code PKCE
and refresh-token grants for `VUER_AUTH_PUBLIC_CLIENT_ID` to the fixed
`VUER_AUTH_URL` HTTPS origin. The identity provider continues to validate PKCE,
redirect registration, and tokens.

`VUER_AUTH_BROWSER_ORIGINS` is a comma-separated exact-origin allowlist. A local
frontend must have its origin configured there and its callback registered with
the identity provider. No Vite proxy is used. Token responses use `no-store`;
provider credentials and diagnostics are never logged or returned in error text.

Deployment injects `ELEVENLABS_API_KEY` from the GitHub environment secret and
`ELEVENLABS_VOICE_ID` / `ELEVENLABS_MODEL_ID` from environment variables. Voice
listing permission is not needed when an allowed voice ID is already configured.

## Inspect and control preference learning

Settings shows saved speaker and model favorites, smoothed preference scores,
and up to 24 recent manual changes. The scores count selecting and replacing
options; automatic selections never reward themselves. Exploration defaults
to 20% among available saved favorites, with 0%, 10% and 20% controls.
Keeping recent change records can be turned off independently of learning.
Clear recent changes preserves favorites and aggregate scores. Removing a
favorite also removes records involving it; Reset voice preferences forgets
favorites, scores, records and restores default settings. Storage remains
bounded to 12 speakers, 12 models and 8 KiB per account in this browser.
Changes apply the next time the audio player opens.

## Volume and iPad playback sessions

Released in UI PR [#647](https://github.com/dreamlake-ai/dreamlake-ai/pull/647),
2026-10-01. The Notes player centers Play between previous/next paragraph
controls. Pin stays at the far left. The cursor group (Read from cursor / selection
and conditional Back to audio) precedes the volume/speaker control in both visual
and keyboard order. A 12 px spacer sits immediately before the speaker icon:
after Read from cursor / selection normally, and after Back to audio when it
appears. The speaker icon sits immediately left of the playback controls. Speed and remaining paragraph time sit right of Play. Remaining time accounts
for playback speed and describes the buffered paragraph, not the entire note.
Voice and model selectors have up to 90 px each and shrink to 40 px in narrow
panels; very narrow panels place them on a second row. These layout refinements
ship in UI PR [#651](https://github.com/dreamlake-ai/dreamlake-ai/pull/651).
The Notes Share trigger shows a person-plus icon without presence avatars and
a plus icon alongside visible avatars; sharing permissions are unchanged.
Tap the volume button to open the slider and mute/unmute control. Unmuting
restores the last nonzero level without restarting streamed or cached audio.

Speech requests `navigator.audioSession.type = 'playback'` when available.
iPadOS can otherwise silence default Web Audio in Silent Mode, even while
media in another tab remains audible. See [WebKit issue 237322](https://bugs.webkit.org/show_bug.cgi?id=237322).
The previous session type is restored when the last Notes player closes.
Browsers without the Audio Session API keep the existing audio path; disabling
Silent Mode is a workaround on older iPadOS versions. Physical iPad verification
is still pending; desktop tests do not establish device audibility.

![Actual Notes player with grouped left controls](/screenshots/notes-audio/controls-near-play-wide.jpg)

![Actual player in a narrow panel](/screenshots/notes-audio/controls-near-play-narrow.jpg)

No public skill impact: this changes the browser UI and playback session,
with no SDK, CLI, or API contract changes.
