Skip to main content
Connect to wss://api.fish.audio/v1/tts/live/with-timestamp with a WebSocket client. The endpoint requires a WebSocket upgrade and returns MessagePack binary frames. A normal HTTP GET request does not start synthesis.

Before you start

You need a Fish Audio API key and a WebSocket client that supports custom authentication headers. To run the Python example on this page, install its dependencies:
Replace <token> in the example with your API key and set reference_id to your voice model ID. The start.request object accepts the same parameters as the Text to Speech API. Send text and receive audio concurrently, then keep reading after sending stop until the server sends finish.

Preserve the final timestamps

Process alignment metadata even when an audio event contains empty audio bytes: a trailing event can carry a final alignment correction. Store each non-null alignment by chunk_seq, replacing its previous snapshot. A null alignment does not erase a snapshot you already received. Add chunk_audio_offset_sec to each segment’s start and end to place it on the session’s audio timeline. The optional time field measures elapsed server session time in milliseconds; use the segment timestamps for captions. After finish, you can send another start on the same socket. Reset your audio buffer and stored alignment snapshots for the new session.