Your scenes will appear here
Write a line or two in the panel and cast a voice for each — the finished scene lands on this canvas as one file.
Write the script line by line and cast a voice for each. The whole exchange comes back as one file, timed like a conversation.
Try it freeEvery line carries its own voice and the model renders the exchange in one pass, so the pauses and the overlap between speakers are performed rather than stitched.
Eleven v3 reads inline tags — [laughs], [whispers], [sighs] — as instructions rather than speaking them, which is the difference between two narrations and a conversation.
Up to ten distinct voices in a scene, in any of 74 languages. Podcast openers, ad spots, game barks, language drills — anything written as a back-and-forth.
Three steps, no setup
One line per turn — the script reads the way it will sound.
Pick a voice per line; two is the default, ten is the ceiling.
The whole scene comes back as one audio file on the canvas.
Text to speech reads one block of text in one voice. Dialogue takes a script — each line has its own speaker — and renders the whole exchange as a single audio file, with the timing between the lines handled for you.