Text-to-Speech

Usage Instructions

Generate natural-sounding speech from text using state-of-the-art AI voices from OpenAI, Deepgram, ElevenLabs, Cartesia, Google Cloud, Azure, and PlayHT. Supports multiple voices, languages, and audio formats.

Actions

OpenAI TTS

Convert text to speech using OpenAI TTS models

Input

ParameterTypeRequiredDescription
textstringYesThe text content to convert to speech (e.g., "Hello, welcome to our service!")
apiKeystringYesOpenAI API key
modelstringNoOpenAI TTS model identifier (e.g., "tts-1", "tts-1-hd", "gpt-4o-mini-tts")
voicestringNoOpenAI voice identifier (e.g., "alloy", "ash", "ballad", "coral", "echo", "sage", "shimmer")
responseFormatstringNoAudio format (mp3, opus, aac, flac, wav, pcm)
speednumberNoSpeech speed multiplier from 0.25 to 4.0 (e.g., 0.5 for slower, 1.0 for normal, 2.0 for faster)

Output

ParameterTypeDescription
audioUrlstringURL to the generated audio file
audioFilefileGenerated audio file object
durationnumberAudio duration in seconds
characterCountnumberNumber of characters processed
formatstringAudio format
providerstringTTS provider used

Deepgram TTS

Convert text to speech using Deepgram Aura

Input

ParameterTypeRequiredDescription
textstringYesThe text content to convert to speech (e.g., "Hello, welcome to our service!")
apiKeystringYesDeepgram API key
modelstringNoDeepgram model/voice identifier (e.g., "aura-asteria-en", "aura-luna-en", "aura-2-luna-en")
voicestringNoDeepgram voice identifier, alternative to model param (e.g., "aura-asteria-en", "aura-orion-en")
encodingstringNoAudio encoding (linear16, mp3, opus, aac, flac)
sampleRatenumberNoSample rate (8000, 16000, 24000, 48000)
bitRatenumberNoBit rate for compressed formats
containerstringNoContainer format (none, wav, ogg)

Output

ParameterTypeDescription
audioUrlstringURL to the generated audio file
audioFilefileGenerated audio file object
durationnumberAudio duration in seconds
characterCountnumberNumber of characters processed
formatstringAudio format
providerstringTTS provider used

ElevenLabs TTS

Convert text to speech using ElevenLabs voices

Input

ParameterTypeRequiredDescription
textstringYesThe text content to convert to speech (e.g., "Hello, welcome to our service!")
voiceIdstringYesElevenLabs voice identifier (e.g., "21m00Tcm4TlvDq8ikWAM", "AZnzlk1XvdvUeBnXmlld")
apiKeystringYesElevenLabs API key
modelIdstringNoElevenLabs model identifier (e.g., "eleven_turbo_v2_5", "eleven_flash_v2_5", "eleven_multilingual_v2")
stabilitynumberNoVoice stability (0.0 to 1.0, default: 0.5)
similarityBoostnumberNoSimilarity boost (0.0 to 1.0, default: 0.8)
stylenumberNoStyle exaggeration (0.0 to 1.0)
useSpeakerBoostbooleanNoUse speaker boost (default: true)

Output

ParameterTypeDescription
audioUrlstringURL to the generated audio file
audioFilefileGenerated audio file object
durationnumberAudio duration in seconds
characterCountnumberNumber of characters processed
formatstringAudio format
providerstringTTS provider used

Cartesia TTS

Convert text to speech using Cartesia Sonic (ultra-low latency)

Input

ParameterTypeRequiredDescription
textstringYesThe text content to convert to speech (e.g., "Hello, welcome to our service!")
apiKeystringYesCartesia API key
modelIdstringNoCartesia model identifier (e.g., "sonic", "sonic-2", "sonic-3", "sonic-multilingual")
voicestringNoCartesia voice identifier or embedding (e.g., "a0e99841-438c-4a64-b679-ae501e7d6091")
languagestringNoLanguage code for speech synthesis (e.g., "en", "es", "fr", "de", "it", "pt")
outputFormatjsonNoOutput format configuration (container, encoding, sampleRate)
speednumberNoSpeech speed multiplier (e.g., 0.5 for slower, 1.0 for normal, 2.0 for faster)
emotionarrayNoEmotion tags for Sonic-3 (e.g., ['positivity:high'])

Output

ParameterTypeDescription
audioUrlstringURL to the generated audio file
audioFilefileGenerated audio file object
durationnumberAudio duration in seconds
characterCountnumberNumber of characters processed
formatstringAudio format
providerstringTTS provider used

Google Cloud TTS

Convert text to speech using Google Cloud Text-to-Speech

Input

ParameterTypeRequiredDescription
textstringYesThe text content to convert to speech (e.g., "Hello, welcome to our service!")
apiKeystringYesGoogle Cloud API key
voiceIdstringNoGoogle Cloud voice identifier (e.g., "en-US-Neural2-A", "en-US-Wavenet-D", "en-GB-Neural2-B")
languageCodestringYesBCP-47 language code for speech synthesis (e.g., "en-US", "es-ES", "fr-FR", "de-DE")
genderstringNoVoice gender (MALE, FEMALE, NEUTRAL)
audioEncodingstringNoAudio encoding (LINEAR16, MP3, OGG_OPUS, MULAW, ALAW)
speakingRatenumberNoSpeaking rate multiplier from 0.25 to 2.0 (e.g., 0.5 for slower, 1.0 for normal, 1.5 for faster)
pitchnumberNoVoice pitch (-20.0 to 20.0, default: 0.0)
volumeGainDbnumberNoVolume gain in dB (-96.0 to 16.0)
sampleRateHertznumberNoSample rate in Hz
effectsProfileIdarrayNoEffects profile (e.g., ['headphone-class-device'])

Output

ParameterTypeDescription
audioUrlstringURL to the generated audio file
audioFilefileGenerated audio file object
durationnumberAudio duration in seconds
characterCountnumberNumber of characters processed
formatstringAudio format
providerstringTTS provider used

Azure TTS

Convert text to speech using Azure Cognitive Services

Input

ParameterTypeRequiredDescription
textstringYesThe text content to convert to speech (e.g., "Hello, welcome to our service!")
apiKeystringYesAzure Speech Services API key
voiceIdstringNoAzure voice identifier (e.g., "en-US-JennyNeural", "en-US-GuyNeural", "en-GB-SoniaNeural")
regionstringNoAzure region (e.g., eastus, westus, westeurope)
outputFormatstringNoOutput audio format
ratestringNoSpeaking rate (e.g., +10%, -20%, 1.5)
pitchstringNoVoice pitch (e.g., +5Hz, -2st, low)
stylestringNoSpeaking style (e.g., cheerful, sad, angry - neural voices only)
styleDegreenumberNoStyle intensity (0.01 to 2.0)
rolestringNoRole (e.g., Girl, Boy, YoungAdultFemale)

Output

ParameterTypeDescription
audioUrlstringURL to the generated audio file
audioFilefileGenerated audio file object
durationnumberAudio duration in seconds
characterCountnumberNumber of characters processed
formatstringAudio format
providerstringTTS provider used

PlayHT TTS

Convert text to speech using PlayHT (voice cloning)

Input

ParameterTypeRequiredDescription
textstringYesThe text content to convert to speech (e.g., "Hello, welcome to our service!")
apiKeystringYesPlayHT API key (AUTHORIZATION header)
userIdstringYesPlayHT user ID (X-USER-ID header)
voicestringNoPlayHT voice identifier or manifest URL (e.g., "s3://voice-cloning-zero-shot/...")
qualitystringNoQuality level (draft, standard, premium)
outputFormatstringNoOutput format (mp3, wav, ogg, flac, mulaw)
speednumberNoSpeech speed multiplier from 0.5 to 2.0 (e.g., 0.5 for slower, 1.0 for normal, 1.5 for faster)
temperaturenumberNoCreativity/randomness (0.0 to 2.0)
voiceGuidancenumberNoVoice stability (1.0 to 6.0)
textGuidancenumberNoText adherence (1.0 to 6.0)
sampleRatenumberNoSample rate (8000, 16000, 22050, 24000, 44100, 48000)

Output

ParameterTypeDescription
audioUrlstringURL to the generated audio file
audioFilefileGenerated audio file object
durationnumberAudio duration in seconds
characterCountnumberNumber of characters processed
formatstringAudio format
providerstringTTS provider used