Skip to main content
Configure how your agent sounds — the voice, speaking speed, and speech recognition.

Text-to-Speech (TTS)

The TTS provider converts the agent’s text responses into spoken audio.

Providers

Voice selection

Each provider offers a set of predefined voices. Select from the dropdown in the agent detail page. Custom voice: If you have a custom voice model (e.g., a cloned voice on ElevenLabs), enter the voice UUID in the custom voice ID field.

Speaking rate

Adjust how fast the agent speaks: Note: Speaking rate is not adjustable for Native Realtime providers (Gemini Live, OpenAI Realtime).

Speech-to-Text (STT)

The STT provider transcribes the caller’s speech into text for the LLM. Note: STT is not configurable for Native Realtime LLMs — they have built-in speech recognition.

Who can edit