- Inputs
- Text; authorized personal voice
- Text
- 1 to 5,000 characters per generation
- Output
- 1 WAV file (24 kHz, mono) per generation
- Text files
- Optional transcript and SRT captions, 8 words or 52 characters per caption at most
- Languages
- Auto-detect + 23 languages: English, French, Spanish, German, Italian, Brazilian Portuguese, Japanese, Chinese, Korean, Hindi, Arabic, Russian, Turkish, Dutch, Ukrainian, Vietnamese, Indonesian, Thai, Polish, Romanian, Greek, Czech, Finnish
- Voices
- French, English and Spanish narrators, filterable by gender, plus your personal voices
- Speed
- 0.70× to 1.20× (1.00× by default)
- Tags
- 8 emotional tones and 14 audio effects; AI tagging, 3 per sentence at most
- Personal voice
- 10–60 s of clear speech, with confirmation of your permission
- Dialogues
- Up to 8 voices in the Story Audio workflow