Text-to-Speech (TTS)
The TTS provider converts the agent’s text responses into spoken audio.Providers
Voice selection
Each provider offers a set of predefined voices. Select from the dropdown in the agent detail page. Custom voice: If you have a custom voice model (e.g., a cloned voice on ElevenLabs), enter the voice UUID in the custom voice ID field.Speaking rate
Adjust how fast the agent speaks:
Note: Speaking rate is not adjustable for Native Realtime providers (Gemini Live, OpenAI Realtime).
Speech-to-Text (STT)
The STT provider transcribes the caller’s speech into text for the LLM.
Note: STT is not configurable for Native Realtime LLMs — they have built-in speech recognition.