Loading
Loading
A script written for a video is not narration text. It carries camera directions, on-screen text, speaker labels and timestamps, and a text-to-speech model reads every one of them aloud. SUMERA exports the spoken half on its own, in the syntax your ElevenLabs model understands.
Write a script to narrateEverything a voice should not say goes. The words stay exactly as written.
## Hook [Cut to wide shot, dramatic music swells] **On-screen text: THE 1975 FILE** Narrator (V.O.): In 1975, a single memo cost $1,299,000, and nobody asked why. [0:12] Pattern interrupt
[serious] In nineteen seventy five, a single memo cost one million two hundred ninety nine thousand dollars, and nobody asked why. [long pause]
Audio tags and SSML breaks are not interchangeable. Send the wrong one and the model reads the markup out loud. SUMERA switches the whole export when you switch the model.
Most expressive. Audio tags for delivery, ellipses and pause tags for timing, no SSML.
5,000 chars per request
Pauses: [pause] tags
Most stable. SSML break tags for pauses, plain punctuation, no audio tags.
10,000 chars per request
Pauses: SSML breaks
Fastest and cheapest. Same syntax as Multilingual v2.
40,000 chars per request
Pauses: SSML breaks
Generate a script in SUMERA's five-stage workflow, or open one you already saved in your library.
Eleven v3 for expressive reads, Multilingual v2 for stable multi-language narration including Hindi, Flash v2.5 for long scripts at low cost. The character limit and pause syntax change with the model.
SUMERA strips section headers, camera directions, on-screen text, timestamps and emoji, keeps only what a voice should say, and converts your written cues into the syntax the model reads.
Open Text to Speech in ElevenLabs, paste the copied text, choose your voice and generate. Long scripts are split into numbered parts that each fit one request.
Numbers, prices and acronyms are spelled out before export so they are not misread. Anything you want delivered differently is one word change in the script, then copy again.
A faceless channel is narration and stock footage, so the voice is the channel. The export exists because that workflow breaks in a specific way: the script and the footage plan live in the same document, and pasting the whole thing into text to speech produces a narrator who announces “cut to wide shot” between sentences.
Most SUMERA creators write from India, and a large share of them publish in Hindi or Hinglish. Multilingual v2 is the model for that: it handles multi-language narration and takes exact pause lengths, which matters when a Hindi line and an English brand name sit in the same sentence. Write the topic in whichever language you publish in.
Long-form works too. A twenty-minute documentary script is past every model’s single-request limit, so the export is split at paragraph boundaries and numbered, and each part is pasted and generated in order.
Eleven v3 is the expressive model: it takes audio tags such as [excited] or [whispers] and reads ellipses as hesitation, but it ignores SSML. Multilingual v2 is the stable one: it takes SSML break tags for exact pause lengths and no audio tags. Pick v3 when performance matters, v2 when consistency across a long series matters. SUMERA formats for whichever you choose.
Eleven v3 accepts 5,000 characters per request, Multilingual v2 accepts 10,000, and Flash v2.5 accepts 40,000. A ten-minute script runs past the v3 limit, so SUMERA splits the export at paragraph boundaries and never mid-sentence, then numbers the parts.
That depends only on the model, not on preference. SSML break tags work on Multilingual v2 and Flash v2.5 and cap at three seconds. Audio tags work on Eleven v3. Sending SSML to v3 gets it read aloud as text, which is the most common reason an export sounds wrong.
Because the model guesses. "₹499" can come out as digits, "1975" as one thousand nine hundred seventy five rather than nineteen seventy five, and "CEO" as a word. SUMERA expands numbers, currency, percentages, ordinals, years and acronyms before you copy, and you can switch that off if you prefer to control it yourself.
Yes. Starter includes copying the script, Eleven v3 delivery tags, and .txt downloads during the 7-day trial and after subscription.
Yes. Give SUMERA the topic in Hindi, Hinglish or English and export with Multilingual v2, which is the model built for multi-language narration. The export removes the English production cues that would otherwise be read aloud in the middle of a Hindi line.
Starter includes copying, delivery tags and .txt downloads. Try it free for 7 days with a card, then continue from €10.80 a month or cancel.