Skip to main content
Your agent speaks with a voice model from the Fish Audio Voice Library: the same voices you use for text to speech. Pick one in the Builder, or set it through the API, and choose the language the agent holds conversations in.

Configuration page

Where the voice and language settings live in the Builder.

Voice Library

Browse public voices and manage your own.

Voice Cloning

Create a custom voice, then use it here. See Voice Cloning.

Pick a voice in the Builder

The voice picker on the Configuration page has two levels: a curated list for a fast start, and the full Voice Library when you want something specific.
1

Open the voice selector

In your agent’s Configuration page, open the voice card. The Choose a voice view shows a curated selection of voices.
2

Browse the full library (optional)

Not seeing the right fit? Select More voices to open the Select Voice browser: the full Voice Library, with your own cloned voices under My Voices.
3

Save automatically

Your selection is written to the agent’s draft configuration as soon as you pick it. Start a preview call to hear the voice in a real conversation.
4

Publish

Draft changes don’t affect live sessions until you Publish. See Versions & publishing.

Use any voice model

The agent’s voice is a voice model id (voice_id): the same ids used as reference_id in Text to Speech. Any public voice model from the Voice Library works, including:
  • Library voices: Ready-made public voices. Find ids in the Voice Library.
  • Your cloned voices: Clone a voice once, then use its model id as your agent’s voice.

Set the voice via API

Voice settings live in the voice section of the agent’s configuration. Patches are partial: only the fields you send change, and the result is saved to the draft:
voice.speaking_language lives in the same section and is patched the same way. As in the Builder, API edits land in the draft. Publish to roll them out.

Speaking language

Speaking language sets the language for the agent’s conversations (voice.speaking_language on the wire). Every session converses in this language. Any of the 52 languages below can be chosen.
The voice model and the speaking language are independent settings: picking a voice does not change the language, and vice versa. Choose a voice that sounds natural in the language you configure.
Both settings can also be replaced for a single session: send overrides.voice_id or overrides.language on the session request. See Overrides.

Speech recognition model

Speech recognition picks the model that transcribes what callers say (asr.model on the wire). Latency is the typical wait from the caller’s last word to the transcript. The agent starts its reply after that. Choose it with the Speech recognition selector in the Builder’s voice section, or via the API:
Speech recognition applies to spoken sessions (voice calls and phone). As with every voice setting, the change lands in the draft: publish to roll it out.

Multilingual recognition

Multilingual recognition lets the agent understand callers who switch to another language partway through a conversation. With it off, speech recognition listens only for the speaking language. That is more accurate when every caller speaks the same language, and it keeps short or accented phrases from being transcribed as another language. It is off by default. Toggle it with the Multilingual recognition switch in the Builder’s voice section, or via the API (asr.multilingual, default false):
Which languages a caller can switch between depends on the speech recognition model. With Deepgram Nova-3, a speaking language outside its switching set stays locked even with the setting on. With ElevenLabs Scribe v2, turning the setting off steers recognition toward the speaking language rather than locking it. To lock it, use strict language.

Strict language

Strict language keeps the agent from understanding any language other than its speaking language. When speech recognition detects that the caller spoke another language, the agent receives [unintelligible speech] instead of the transcript and answers as if it did not catch what was said. It applies whether multilingual recognition is on or off. It is off by default. Turn it on via the API (asr.strict_language, default false):
With the ElevenLabs Scribe v2 models, speech in another language reaches the agent as [unintelligible speech]. With Deepgram Nova-3, recognition runs a model for the speaking language only, so speech in another language is usually not transcribed at all and the agent keeps listening.
Language is detected per utterance, so a short or heavily accented phrase in the speaking language can occasionally be detected as another language and replaced with [unintelligible speech] too. Only turn on strict language when you explicitly need to stop the agent from understanding other languages.

Speaking speed

Speaking speed sets how fast the agent talks, as a multiplier from 0.5 (half speed) to 2.0 (double speed). The default is 1.0. Many English-language agents sound more natural at a slightly faster rate, such as 1.2. Set it with the Speaking speed slider in the Builder’s voice section, or via the API (voice.speed, default 1.0):
The rate applies to everything the agent speaks in voice and phone sessions. As with every voice setting, the change lands in the draft: publish to roll it out.

Expressive mode

Expressive mode makes the agent steer its own delivery: it opens sentences with emotion cues, adds natural pauses and emphasis, laughs where it genuinely fits, and speaks the way people talk, with contractions and the occasional “um”. You get lively, emotionally aware speech without writing any delivery rules into your prompt. It is on by default. Toggle it with the Expressive mode switch in the Builder’s voice section, or via the API (voice.expressive, default true):
The delivery cues are rendered by the voice model. They are never spoken and never appear in transcripts or message history. Expressive mode applies to spoken sessions (voice calls and phone); text chat is unaffected.
Keep your system prompt about what the agent says: who it is, what it knows, what it should do. With expressive mode on, how it sounds is handled for you.

Going further

Agent configuration

System prompt, first message, and conversation settings.

System tools

Built-in capabilities like hanging up the call.

Preview calls

Talk to your draft agent and hear the voice live.

Versions & publishing

How drafts become the live agent.