FreeTTS/Studio/Help

Studio · Help

FreeTTS Studio: how the pieces fit together

Two modes — Single voice and Dialogue — plus takes to fix a single line, a Timeline to chain scenes, SSML to direct the delivery, and export to MP3, WAV, captions or a captioned video. Most projects only need Single voice; the rest unlock when your project gets longer, has multiple characters, or needs scene-by-scene control.

Updated 11 September 202615 min readApplies to Studio on the web
In one breath

FreeTTS Studio has two modes. Single voice turns one script into one MP3 with any of 2,250+ voices; Dialogue turns a Speaker: line script of up to 50 lines into one merged multi-voice MP3. PRO adds takes (regenerate any line), the Timeline (chain up to 100 clips with pauses), WAV and OGG export, a caption-animated video export, a music bed, saved projects, and the Signature and Ultra voice tiers. Free accounts render 1,000 characters per Studio render, 5,000 a day and 15,000 a month, and can preview every voice and every emotion at no cost.

Modes & panels

#What each mode actually does

Studio is one screen with two modes and one overlay. Everything you generate in either mode can be chained in the Timeline, and any line of a render can be re-done on its own.

Single voice. Write or paste, select any span and right-click (long-press on touch) for Pause, Emphasis, Emotion or Pronounce. Six speed presets from 0.5× to 2×, three pitch presets. One Generate, one MP3.
Start here
Single voice
Free 1,000 · PRO 10,000 · Creator 25,000 chars per render

When: one block of text, voice testing, quick clips, per-render speed, pitch and style.

The default mode. Type or paste your script, pick any voice, optionally add an emotion, speed or pitch, and Generate. Each render is one MP3. On PRO, Regenerate by paragraph splits the render into paragraph takes you can redo one at a time.

Output: one MP3 per render, with .srt captions when the render has word timings.

Multi-voice
Dialogue
Up to 50 lines per render · free accounts: one scene a day (4 lines, 500 chars)

When: conversations between characters, multi-speaker scenes, anything with line-by-line voice changes.

Write Speaker: line, one line per row, give each speaker a voice, and Generate. The scene is muxed server-side into one merged MP3 with proper turn-taking. Each line has its own pause and its own optional emotion. Ultra voices are single-voice only and are not offered here.

Output: one merged MP3 with all speakers, with per-line takes on PRO.

PRO · overlay
Timeline
Up to 100 clips per merge · pause after each clip 0–3 s

When: chaining several Single voice or Dialogue renders into one continuous file with pauses between scenes.

An overlay, not a mode. After any render, click Add to Timeline under the player. Reorder clips, set the pause after each, preview any clip, and Merge. Intro narration → dialogue scene → narrator break → next scene → outro, in one file.

Output: one continuous MP3 with every clip and pause baked in.

First visit? Studio opens with a pre-rendered three-voice demo scene so you can hear a Dialogue render before typing anything. It costs nothing and is not counted anywhere.

Takes

#Regenerate one line, keep the rest

A render is a set of lines. Any line can be re-done as a fresh take without touching its neighbours; the track re-stitches itself. Takes are kept, so a worse one is a single tap back. This is a PRO feature in both modes.

Dialogue with takes. Mara's second line has three takes; ‹ › steps between them and ↻ records another. The pause after each line is its own control. Hover any line for the same controls in Single voice after switching on Regenerate by paragraph.
  • Dialogue: every line carries ↻ and ‹ › from the first render.
  • Single voice: click Regenerate by paragraph once; Studio re-renders your script as paragraph takes and the same controls appear per paragraph.
  • A take is a new synthesis of that line only — the same voice, emotion and speed unless you change them first — so it costs only that line's characters.

At a glance

#Which one for what

NeedUseWhy
Test a voice on real textSingle voicePer-render control: six speed presets, three pitch presets, per-span emotion, free previews of every voice.
One narrator, 50+ pagesSingle voice + TimelineRender in chunks under your per-render cap, then chain them in the Timeline (PRO).
Multiple characters speakingDialogueServer-side mux into one MP3 with proper turn-taking and per-line pauses.
Fix one weak line without redoing the renderTakes ↻ (PRO)Re-render just that line as a fresh take; it re-stitches automatically.
Chain clips with pausesTimeline (PRO)Reorder, set a 0–3 s pause after each clip, merge into one MP3.
Per-paragraph emotion (sad / cheerful / whispering)Single voice / DialogueSelect text, right-click → Emotion, or wrap in <mstts:express-as style="…">. Standard and multilingual voices only.
Custom pauses between paragraphsSingle voice / DialogueRight-click → Pause, or <break time="1.5s"/>. Works on every tier except Ultra.
A video for Reels, Shorts or YouTubeExport video (PRO)Any render → caption-animated MP4 in 9:16, 1:1 or 16:9.
Music under the narrationMusic bed (PRO)Pick a bed; it is mixed under the voice with automatic ducking.
Programmatic generation from scriptsDeveloper APISee /developers. Free accounts get one API key, paid plans ten.

Common workflows

#"I want to make…"

I have a single narrator and a 50-page chapter.

  1. Open Single voice mode and pick an HD voice (look for the ✨ filter — they read with the most natural cadence at length).
  2. Render your text in chunks that fit your per-render cap (free 1,000 · PRO 10,000 · Creator 25,000). Add <break time="1.5s"/> between paragraphs and <emphasis> on key terms via right-click.
  3. Click 'Add to Timeline' after each chunk (PRO).
  4. Open the Timeline overlay, set the pause after each clip, then Merge into one continuous MP3.

I have a story with 3 characters who all speak.

  1. Write each multi-character SCENE in Dialogue mode as Speaker: line rows (up to 50 lines per render).
  2. Assign each speaker a distinct voice — standard, multilingual, HD or Signature; Ultra stays in Single voice.
  3. Render the scene. Hover a weak line and hit ↻ for a fresh take. Click 'Add to Timeline' below the audio player.
  4. Write narrative bridging text in Single voice mode and add each to the Timeline too.
  5. Open the Timeline overlay, reorder clips, set pauses (e.g. 1.5 s between scenes), then Merge.
  6. Download the single merged MP3, or Export video for a captioned MP4.

I want emotional variation across paragraphs (sad scene then cheerful scene).

  1. In Single voice on a standard or multilingual voice, select the sad paragraph, right-click → Emotion → sad (or type <mstts:express-as style="sad">your text</mstts:express-as>).
  2. Select the cheerful paragraph and apply cheerful the same way.
  3. Available styles depend on the voice: the chip picker shows what your selected voice supports. Free accounts can preview every emotion; rendering with one is PRO.
  4. Tags must be lowercase. Generate — and if one paragraph lands wrong, use 'Regenerate by paragraph' to fix just that line.

I want to test 10 voices to find my favorite for a project.

  1. Open Single voice mode.
  2. Open the voice picker. Click the play button (▶) next to each voice — those previews don't count toward your monthly quota, and they work for Signature and Ultra too.
  3. Once you've narrowed to 2–3 candidates, paste a real paragraph from your project and Generate. Now you'll hear how each handles YOUR text at your target pace.
  4. Try the speed presets (0.5× to 2×) and pitch (Low / Normal / High) to taste, then 'Save project' (PRO) to reuse the whole setup later.

SSML quick reference

#The 4 tags that cover 95% of use cases

SSML is the markup inside the text box that controls pauses, emphasis, pitch and emotion. It works in Single voice and Dialogue on standard, multilingual and HD voices: select text and right-click (or long-press) to insert a tag, or type it. All tags are lowercase.

Pause every tier but Ultra
<break time="1s"/>

Insert a pause. Use 2s, 500ms, up to 10s.

Emphasis
<emphasis level="moderate">word</emphasis>

Stress a word or phrase. Levels: reduced / moderate / strong.

Tempo & pitch
<prosody rate="slow" pitch="-2st">text</prosody>

Rates: x-slow / slow / medium / fast / x-fast or a percentage. Pitch in semitones (+2st, -2st). DragonHD ignores this tag; use the speed presets there.

Emotion render on PRO
<mstts:express-as style="cheerful">text</mstts:express-as>

Emotional style. Which styles exist depends on the voice — the chip picker shows them. Free accounts can preview any emotion; rendering with one is PRO.

Three things trip up almost everyone. Tag names are lowercase (<Break/> is spoken aloud as text). &, < and > must be escaped inside SSML (&amp;, &lt;, &gt;) — Studio does it for you in plain-text mode but not once you add tags. DragonHD, Signature and Ultra voices drop <mstts:express-as>: they read emotion from the words.

SSML deep reference

#Everything else SSML can do

The cheatsheet above handles most projects. Below is the full surface FreeTTS supports, in categories, every example copyable. Tier rules: standard, multilingual and HD voices take everything; DragonHD ignores express-as and prosody rate/pitch; Signature honours <break> only; Ultra reads plain text.

Pauses & silence

Two ways to insert silence. <break> is the W3C standard and the one tag that also works on Signature voices. <mstts:silence> is an mstts extension that gives you finer placement control.

break (time)
<break time="2s"/>

Hard pause for the exact duration. Units: ms (milliseconds) or s (seconds). Cap is 10s; longer values are clamped.

break (strength)
<break strength="strong"/>

Semantic pause length. Options (shortest → longest): "none", "x-weak", "weak", "medium", "strong", "x-strong". Use strength when you want pauses that scale with the voice's natural cadence; use time when you need exact timing.

mstts:silence (leading)
<mstts:silence type="leading" value="500ms"/>

Adds silence at the very start of the audio. Options for type: "leading", "tailing", "sentenceboundary", "leading-exact", "tailing-exact", "sentenceboundary-exact". The -exact variants override the engine's built-in silences instead of adding to them.

mstts:silence (between sentences)
<mstts:silence type="sentenceboundary" value="800ms"/>

Adds 800ms between every sentence in the document. Great for slow audiobook pacing or training content. Place this once at the top of your SSML, not per sentence.

Prosody — rate, pitch, volume

All three attributes accept named keywords, relative percentages, or absolute values. Combine them in a single <prosody> tag. DragonHD voices ignore rate and pitch — use the Studio speed presets or natural writing there.

prosody (rate, named)
<prosody rate="x-slow">slowed text</prosody>

Named rates: "x-slow" (0.5x), "slow" (0.7x), "medium" (1.0x), "fast" (1.3x), "x-fast" (1.5x). Easiest to read in scripts.

prosody (rate, percent)
<prosody rate="85%">precisely controlled</prosody>

Absolute percentage (50%-200%). Use when named rates don't hit the exact pacing you want. 85% is a popular audiobook setting.

prosody (rate, relative)
<prosody rate="+20%">slightly faster than ambient</prosody>

Relative shift from the surrounding rate. Useful when you want one paragraph to be a bit faster than the rest without committing to an absolute speed.

prosody (pitch, semitones)
<prosody pitch="-2st">lower pitch</prosody>

Pitch shift in semitones (st). Range roughly -12st to +12st before voices distort. -2st makes most voices sound a touch more serious; +2st adds energy. Also accepts "+2Hz", "x-low", "low", "medium", "high", "x-high", or percentages.

prosody (volume)
<prosody volume="loud">SHOUTING WITHOUT CAPS</prosody>

Named volumes: "silent", "x-soft", "soft", "medium", "loud", "x-loud". Or use decibels (e.g. volume="+6dB"). Lets you bake quiet/loud passages into a single output without an audio editor.

prosody (combined)
<prosody rate="slow" pitch="-1st" volume="soft">whispered confession</prosody>

All three attributes work in one tag. Combine for distinctive effects — slow + low + soft = intimate; fast + high + loud = panic.

Pronunciation control

Three ways to fix mispronunciations. Phoneme is the most precise; sub is the easiest; the Pronunciation Dictionary (PRO) is permanent across all your generations.

sub (alias)
<sub alias="Doctor">Dr.</sub> Smith

Replaces the displayed text with the alias for speech only. "Dr." gets spoken as "Doctor". Useful for abbreviations the engine misreads (Mr, Mrs, etc.), medical/legal shorthand, or initialisms.

phoneme (IPA)
<phoneme alphabet="ipa" ph="kəˈmɛrəθ">Camarath</phoneme>

Force-pronounce a word using International Phonetic Alphabet symbols. Best for proper nouns, fantasy names, technical terms. The displayed text is still shown in transcripts; only the audio uses your phoneme spelling.

phoneme (SAPI)
<phoneme alphabet="sapi" ph="t aw m ax t ow">tomato</phoneme>

SAPI phonetic alphabet. Easier to read than IPA if you're not a linguist; only works with English voices. Each phoneme is space-separated.

phoneme (UPS)
<phoneme alphabet="x-microsoft-ups" ph="T1 OW0 M EY1 T OW0">tomato</phoneme>

Universal Phone Set. A cross-language phoneme system. Use the x-microsoft-ups alphabet name (the identifier is locked by the SSML standard). Includes optional stress markers (1 = primary, 2 = secondary, 0 = none).

lexicon (PRO)
<lexicon uri="https://your.cdn/pronunciations.xml"/>

Loads an external pronunciation dictionary (W3C PLS XML format). PRO/Creator users typically use the in-app Pronunciation Dictionary instead — same effect, no external hosting needed. Set once at the top of your SSML.

Say-as — interpret strings literally

Forces the engine to read text as a specific kind of data. Without <say-as>, the engine guesses ("1234" might be read as "one thousand two hundred thirty-four" or "one two three four" depending on context).

characters / spell-out
<say-as interpret-as="characters">NASA</say-as>

Reads each character individually: "N-A-S-A". Use for initialisms you want spelled letter by letter. "spell-out" is a synonym.

cardinal
<say-as interpret-as="cardinal">12345</say-as>

Reads numbers as cardinal numbers: "twelve thousand three hundred forty-five". Use when context might otherwise force digit-by-digit reading.

ordinal
The <say-as interpret-as="ordinal">3</say-as> rule

Reads as ordinal numbers: "third". Without this, "3" might be read as "three".

digits / number_digit
<say-as interpret-as="digits">2024</say-as>

Reads each digit separately: "two zero two four". Useful for years pronounced digit-by-digit, account IDs, etc.

fraction
<say-as interpret-as="fraction">1/2</say-as>

Reads as a fraction: "one-half" (English) or the locale equivalent. Works for both "1/2" and "1 1/2" formats.

date (mdy / dmy / ymd)
<say-as interpret-as="date" format="dmy">12-03-2026</say-as>

Reads a date with explicit format. Format strings: "mdy", "dmy", "ymd", "md", "dm", "ym", "my", "d", "m", "y", "yyyymmdd". Engine will say "the twelfth of March, two thousand twenty-six". Avoids the US-vs-rest-of-world confusion entirely.

time
<say-as interpret-as="time" format="hms24">15:30:00</say-as>

Reads a time. Formats: "hms12" (am/pm), "hms24" (24-hour), "ms" (minutes:seconds). The above reads as "fifteen thirty hours" / "three thirty PM" depending on locale.

telephone
<say-as interpret-as="telephone">+1-555-867-5309</say-as>

Reads phone numbers naturally with country code grouping. Strips dashes and spaces, reads digits one at a time in conventional groupings.

currency
<say-as interpret-as="currency" language="en-US">$42.50</say-as>

Reads currency with the unit name: "forty-two dollars and fifty cents". Optional language attribute helps when the currency symbol is ambiguous.

address
<say-as interpret-as="address">221B Baker St</say-as>

Reads street addresses with the right pacing — number before street, expanding common abbreviations (St → Street, Ave → Avenue).

Language switching mid-document

Use <lang> to read a foreign phrase in its native pronunciation without switching voices entirely. The voice has to support the target language for this to sound right.

lang
She said <lang xml:lang="fr-FR">bonjour mon ami</lang> with a smile.

Reads the wrapped text in the specified language using the current voice. Multilingual voices (look for "Multilingual" in the voice name) handle ~12 languages each — best for code-switching scripts. Standard voices may fall back to phonetic approximation.

voice (nested)
<voice name="en-US-AriaNeural">Hello.</voice> <voice name="es-ES-ElviraNeural">Hola.</voice>

Swap voices mid-document. Each <voice> block can be a completely different voice in a different language. Useful when you need an actual native speaker for the foreign passage instead of a multilingual voice's approximation. This is what Dialogue mode does for you automatically.

Expressive styles (mstts)

<mstts:express-as> is the emotional style tag. Not every voice supports every style — see the complete catalog below. Combine with styledegree (0.01–2.0) and role (for certain Chinese voices). Free accounts can preview any emotion; baking one into a render is PRO.

express-as (basic)
<mstts:express-as style="hopeful">tomorrow will be different</mstts:express-as>

Wraps text in an emotional style. The voice must support the style — Aria, Jenny, Davis (US English) have the widest catalogs; British and other locales have fewer.

express-as (style degree)
<mstts:express-as style="excited" styledegree="2">she's here!</mstts:express-as>

Style intensity. 0.01 = barely-there hint of the emotion; 1.0 (default) = normal; 2.0 = exaggerated. Use higher degrees for dramatic moments, lower for subtle inflection.

express-as (role)
<mstts:express-as style="default" role="YoungAdultFemale">line of dialogue</mstts:express-as>

Roleplay attribute — makes the voice imitate a different speaker type. Options: "Boy", "Girl", "YoungAdultFemale", "YoungAdultMale", "OlderAdultFemale", "OlderAdultMale", "SeniorFemale", "SeniorMale". Currently only some Chinese voices (zh-CN-XiaomoNeural, zh-CN-YunxiNeural, zh-CN-YunyeNeural) support this. Pairs well with Dialogue mode for character variety from a single voice.

Structure & markers

Optional but useful for long-form audiobooks, captions, and engines parsing your output.

p (paragraph)
<p>Paragraph one. Two sentences.</p> <p>Paragraph two.</p>

Explicit paragraph boundary. The engine adds a natural pause between <p> blocks (slightly longer than between sentences). Useful when your text has weird line breaks the engine would otherwise misinterpret.

s (sentence)
<s>This is a sentence.</s> <s>This is another.</s>

Explicit sentence boundary. Forces the engine to treat the wrapped text as a complete sentence even without terminal punctuation — useful for fragments like "Yes." that might otherwise blend into the next.

bookmark
And then <bookmark mark="chapter-2-start"/> she opened the door.

Inserts a named position marker. Doesn't affect audio but appears in the boundary metadata exposed via /api/v1/tts. Useful when you're building a player that needs to jump to specific scenes.

Background audio (mstts)

Mix a background audio track (music, ambience, white noise) under the synthesized speech. Audio file must be publicly accessible HTTPS. In Studio, the built-in music bed (PRO) does this without any markup, with automatic ducking.

mstts:backgroundaudio
<mstts:backgroundaudio src="https://example.com/music.mp3" volume="0.4" fadein="2000" fadeout="3000"/>

Plays the source audio under the entire speech track. volume is 0.0–1.0 (0.4 = 40% volume — keeps speech intelligible). fadein/fadeout in milliseconds. Place inside <speak> but outside <voice>. Only allowed once per document.

All expressive styles

#Every mstts:express-as style FreeTTS supports

Not every voice supports every style. The Studio's chip picker (in Single voice, and per line in Dialogue) shows what your selected voice supports. Aria, Jenny and the Chinese voices Xiaoxiao/Yunxi have the widest catalogs; British and most localized voices typically support just cheerful and sad. Everyone can preview a style; rendering with one is PRO.

Emotion

cheerfulBright, upbeat. Default pick for happy moments and positive announcements.
sadSlower, lower pitch, downward inflection. Use for grief, regret, somber news.
angryHarder consonants, elevated volume. Adversarial dialogue, frustration.
excitedFaster, higher energy than cheerful. Big-reveal moments, action scenes.
fearfulTrembling, slightly higher pitch, breathy. Suspense and horror.
terrifiedExtreme of fearful — short bursts, rapid breaths, top of the register.
hopefulWarm, slightly tentative, rising inflection. Bridging sad → cheerful.
disgruntledAnnoyed but restrained. Sarcasm, minor complaints.
embarrassedHushed, slightly halting. Apologies, awkward moments.
seriousSteady, even-paced, low expression. News commentary, formal narration.
calmSmooth, lower energy than serious. Meditation, instructions, ASMR adjacent.

Voice quality

whisperingBreathy, very low volume. Intimate scenes, secrets, ASMR. Best paired with prosody volume='soft'.
shoutingMaximum volume + clipped delivery. Battle scenes, distant calls. Use sparingly — listening fatigue is real.
gentleSoft, mid-pitch, evenly paced. Children's storytelling, comforting.
lyricalSlight musical inflection. Poetry, song-like delivery.

Conversational

friendlyWarm and approachable. Default for tutorials, onboarding, support content.
unfriendlyCold, dismissive. Antagonist dialogue, hostile NPC.
empatheticSoft, slow, validating tone. Customer-service apologies, sensitive subjects.
chatCasual, mid-energy. Podcast-style discussion, informal updates.
assistantPolite, helpful, neutral. Built for AI assistants and voice UIs.
customerserviceProfessional friendly. Phone-system-style helpful answers.

Narration

narration-professionalClean audiobook-style narration. The default 'just read it well' choice for long-form.
narration-relaxedLooser pacing than professional. Personal essays, memoir.
documentary-narrationAuthoritative, evenly paced. Educational video voiceovers, science explainers.
newscastGeneric newscaster tone. Use the more specific casual/formal variants if available on your voice.
newscast-casualApproachable newscaster — feature segments, morning shows.
newscast-formalTraditional newscaster gravitas — breaking news, formal reports.
poetry-readingSlower pacing with deliberate emphasis on cadence. Verse, prose poetry.
advertisement-upbeatPunchy, energetic commercial reading. Product launches, promos.
sports-commentaryFaster pace, dynamic stress. Sports calls, live event narration.
sports-commentary-excitedSports-commentary cranked up — game-winning moments.

Pronunciation

#Fixing words the voice mispronounces

Four escalation paths — pick the smallest one that works for your case.

Quickest: rewrite phonetically

If a word is mispronounced once or twice, rewrite the spelling so the engine reads it correctly. Camarath → kuh-MARE-uth. Ugly in the transcript but fixes one-offs without any SSML, and it is the only method that also works on Signature and Ultra voices.

One-shot fix: <sub alias="…">

Replace what the engine sees with what you want it to say.

Dr. <sub alias="Watson">Watson</sub> said hello.

Use this for abbreviations and initialisms the engine misreads. The transcript still shows Dr.; only the audio uses Doctor.

Precise: <phoneme alphabet="ipa">

Force exact pronunciation with International Phonetic Alphabet symbols.

<phoneme alphabet="ipa" ph="kəˈmɛrəθ">Camarath</phoneme>

The ph attribute holds the IPA. The displayed text is preserved; only the audio uses your spelling. If you don't know IPA, the dashboard's Pronunciation Dictionary has a "Hear it" button that lets you audition different IPA strings until one sounds right.

Permanent: Pronunciation Dictionary PRO

Add Camarath → kəˈmɛrəθ once in the dashboard (Studio settings → Pronunciation). Every future generation across all modes uses it. Much cleaner than wrapping every instance in <phoneme>. Up to 500 entries per account.

IPA cheat sheet for common English sounds

SoundIPAExample word
"a" in catæcæt
"a" in fatherɑfɑther
"e" in bedɛbɛd
"i" in machineimishin
"i" in bitɪbɪt
"o" in noteoʊnoʊt
"u" in moonumun
schwa (the most common vowel)əəbout
"th" in thinθθin
"th" in thisððis
"sh" in shoeʃʃoo
"zh" in measureʒmeaʒure
Primary stress (placed before the stressed syllable)ˈcˈamera

Voice tiers

#Standard vs Multilingual vs HD vs DragonHD vs Signature vs Ultra

One picker, six tiers. Knowing the tier matters because they differ in what SSML they honour, what a free account can do with them, and what they are best at. Every tier can be previewed for free from the picker's ▶ button.

Standard

Free

The everyday neural voices — 2,250+ across 149 languages, fast, expressive on the styles they support. Every language page's free voices are standard voices.

Best forDefault for everything: tutorials, podcasts, video voiceover, casual audiobooks.
SSMLFull SSML, including <mstts:express-as> on supported voices (styles preview free, render on PRO).

Multilingual

Free

One voice identity that speaks about 12 languages — look for "Multilingual" in the name (Andrew, Ava, Brian, Emma, Jenny…). Your French phrase does not suddenly become someone else.

Best forBilingual narration, language-learning content, scripts with embedded foreign phrases.
SSMLFull SSML. Pairs well with <lang xml:lang="…"> for mid-sentence switches.

HD

PRO · free daily taste

Higher-fidelity Azure voices with noticeably more natural prosody. Look for ":DragonHDLatestNeural" or "HD" in the name. Visitors in HD-eligible countries get a small free daily taste; the language pages lead with the flagship HD voice where one exists.

Best forAudiobook narration, professional voiceover, anywhere quality matters more than speed.
SSMLFull SSML.

DragonHD

PRO · free daily taste

The newest Azure generation. Reads context and expresses emotion from the text itself: write "she whispered" and it whispers. The ":DragonHDLatestNeural" suffix marks it.

Best forHighest-quality narration when you want emotion to come from the writing, not from markup.
SSML<break> works. <mstts:express-as> and prosody rate/pitch are ignored — direct it with the words and the Studio speed presets.

Signature

PRO · preview free

Thirty character voices per language — Nova, Maya, Celeste, Atlas, Felix, Theo and 24 more — in 53 languages, with the same persona set everywhere so a Croatian project and a Hungarian one can share a cast. Every language page shows six of them with a free preview.

Best forCharacter work, marketing reads, podcasts, anything that needs a distinct personality per voice in a non-English language.
SSMLPlain text plus <break>. Other tags are stripped before synthesis, so write pacing with punctuation and breaks.

Ultra

PRO · preview free

Our most lifelike line: 21 English voices and five Arabic dialect voices (Syrian, Palestinian, Egyptian, Gulf, Modern Standard). Previews are static samples you can play any time from the English and Arabic pages or the picker.

Best forHero narration, ads, trailers, anything where realism beats control. Single voice only.
SSMLPlain text. Markup is stripped. Not available in Dialogue mode — render the Ultra part in Single voice and chain it in the Timeline.

Where the free daily HD taste applies. Visitors in HD-eligible countries can render a short HD or DragonHD sample each day without a plan; the language pages lead with that voice (Ava HD on English, Seraphina HD on German, Αθηνά HD on Greek…). Elsewhere the same card is badged PRO. Signature and Ultra never have a free render, only free previews.

Voice recommendations

#Best voices by use case

Tested-and-recommended picks for common projects. Preview each in the voice picker before committing — taste is individual.

Long-form audiobook (single narrator)

  • en-US-Andrew:DragonHDLatestNeural
    Top pick. DragonHD reads emotion from context — exactly what you want for fiction.
  • en-US-AvaMultilingualNeural
    Use for mixed-language books. Same voice identity across English, Spanish, French, etc. Free.
  • en-US-JennyNeural
    Solid default neural. Wide expressive style support if you want manual control.
  • en-GB-LibbyNeural
    British narration for UK-set or period-set fiction.

News, current events, factual content

  • en-US-AriaNeural with newscast-formal
    Authority + clarity. Used by many news automations.
  • en-US-BrandonNeural
    Male newscaster cadence.
  • en-US-DavisNeural
    Mid-energy, factual.

Podcast / conversational

  • en-US-GuyNeural with style='chat'
    Casual, mid-energy. Sounds like a real podcaster.
  • en-US-JennyMultilingualNeural
    Approachable warm female voice. Multilingual flexibility.
  • en-US-Davis:DragonHDLatestNeural
    DragonHD male — natural inflection without style hacking.

Educational / explainer

  • en-US-AriaNeural with documentary-narration
    Standard for science explainers and tutorials.
  • en-US-Emma:DragonHDLatestNeural
    DragonHD female. Clean, even-paced explainer voice.

Ads, trailers, hero narration

  • Duke (Ultra)
    Deep, cinematic. The voice the homepage auditions.
  • Scarlett (Ultra)
    Velvety British read for premium brands.
  • Atlas (Signature, any language)
    Informative and clear — the same persona in 53 languages for a multilingual campaign.

Multi-voice dialogue scenes

  • Mix any two contrasting voices in Dialogue mode
    Pick voices with distinctly different pitches and accents — a young female + an older male reads more clearly than two similar voices.
  • Add 'role' attribute for Chinese voices
    zh-CN-XiaomoNeural, zh-CN-YunxiNeural, etc. can switch between Boy/Girl/YoungAdult/OlderAdult roles for variety from one voice.

Export

#Export: audio, captions, video, music

Every render downloads as MP3. PRO unlocks WAV and OGG, a caption-animated video, and a music bed under the narration. Captions are free whenever the render has word timings.

The export row under the player. Format, captions and Export video sit together; the music bed is chosen before you generate and is mixed under the voice with automatic ducking.

Audio

MP3 on every plan. WAV (48 kHz 16-bit) and OGG (Opus) on PRO and Creator. The Timeline merge always produces MP3.

Captions (.srt)

When a render carries word timings, a .srt download appears next to the audio. Long scripts are rendered in chunks and their timings are stitched, so captions stay aligned across the whole file. Drop the file into Premiere, DaVinci Resolve or CapCut.

Export video PRO

Turns the current render into a caption-animated MP4: 9:16 for Reels and Shorts, 1:1 for square posts, 16:9 for YouTube; a wave, bars or circle visualiser on a curated backdrop; and three caption styles — Bold pop, Highlight and Clean. The captions are driven by the same word timings, so they never drift.

Music bed PRO

Pick one of the curated beds and set its level; Studio mixes it under your narration with automatic ducking, so the music dips while the voice speaks and swells between lines. No markup, no audio editor.

Output formats

#MP3 vs WAV vs OGG — which to pick

FormatPlanSpecsSizeWhen to use
MP3All plans96 kbps mono · 24 kHz≈ 0.7 MB per minuteDefault. Universal playback, small files, good quality. Lossy. Re-encoding a final mix costs a little each time.
WAVPRO16-bit PCM mono · 48 kHz≈ 5.8 MB per minuteEditing in Audition, Pro Tools, Reaper; mastering; re-encoding to anything else without loss. Uncompressed. Use this if the file goes into a DAW.
OGGPROOpus mono · 48 kHz≈ 0.7 MB per minuteWeb playback and game engines that prefer Opus. Smallest files at equivalent quality. Open codec. Modern browsers all support it; some older devices do not.

Gotchas

#Common mistakes and fixes

Uppercase SSML tag names

✗ Wrong<Break time="1s"/>
✓ Right<break time="1s"/>

SSML is case-sensitive. Capitalized tags don't parse and get spoken aloud as text — or fail the whole chunk.

Unescaped & in text

✗ WrongTom & Jerry decided to go.
✓ RightTom &amp; Jerry decided to go.

Inside SSML, & starts an XML entity. The fix is &amp; (or & outside any SSML block in plain-text mode). Same goes for < (use &lt;) and > (use &gt;) inside SSML.

<mstts:express-as> on a DragonHD, Signature or Ultra voice

✗ Wrong<mstts:express-as style="cheerful">…</mstts:express-as> (on en-US-Andrew:DragonHDLatestNeural)
✓ RightJust write expressive text. DragonHD reads emotion from context; Signature and Ultra read plain text.

DragonHD drops mstts:express-as (and prosody rate/pitch); Signature and Ultra strip markup before synthesis. Switch to a standard or multilingual voice when you need explicit style control.

Using a style the voice doesn't support

✗ Wrong<mstts:express-as style="poetry-reading">…</mstts:express-as> (on en-US-GuyNeural)
✓ RightCheck the per-voice supported styles. Aria has the most; British voices typically only cheerful + sad.

Unsupported styles are silently ignored — the audio plays neutral instead of poetic. Use the Studio's chip picker to see which styles your selected voice supports.

Forgetting the mstts namespace

✗ Wrong(in raw SSML files outside the Studio)
✓ Rightxmlns:mstts="http://www.w3.org/2001/mstts" on <speak>

If you're writing raw SSML for the API, the <speak> root must declare the mstts namespace before any mstts: tag will parse. The Studio wraps your text automatically — this only matters for direct API users.

Putting mstts:backgroundaudio inside <voice>

✗ Wrong<voice name="..."> <mstts:backgroundaudio .../> ... </voice>
✓ Right<speak> <mstts:backgroundaudio .../> <voice name="...">...</voice> </speak>

Background audio is per-document, not per-voice. Place it directly inside <speak>, before the <voice> block. Only one allowed per document. In Studio, the music bed does this for you.

Expecting <audio src='...'/> to work

✗ Wrong<audio src="https://example.com/clap.mp3"/>
✓ RightInline audio injection isn't supported. Use the Timeline to chain pre-generated clips, or mstts:backgroundaudio / the music bed for a single underlay track.

Arbitrary inline <audio src> clips aren't supported by the synthesis path; use one underlay track, or chain separate clips in the Timeline.

Picking an Ultra voice for a Dialogue line

✗ WrongNarrator: (Duke · Ultra) … inside Dialogue mode
✓ RightRender the Ultra narration in Single voice, then chain it before the Dialogue scene in the Timeline.

Ultra is a single-voice tier; Dialogue refuses it per line so a scene never mixes tiers unpredictably. The Timeline gives you the same result in two clips.

Pro tips

#Power-user shortcuts

  • Voice previews are free. The play button next to each voice in the picker uses a separate quota-free endpoint — Signature and Ultra included. Audition 50 voices before committing to one; none of it counts toward your monthly chars.
  • HD voices cost the same as standard in char usage, but synthesis is slower. For interactive Studio work that's invisible; on a long render it adds a little time. Usually still worth it.
  • Regenerate one line, not the whole thing (PRO). Hover a line (in Dialogue, or in Single voice after 'Regenerate by paragraph') and hit ↻ for a fresh take of just that line — it re-stitches automatically, and you can step between takes with ‹ ›.
  • Save your project (PRO). 'Save project' keeps a named, server-saved copy you can reopen on any device — pick up exactly where you left off.
  • For very long narration, render in chunks under your per-render cap and chain them in the Timeline (PRO). Set the pause after each, then merge to one continuous MP3.
  • Add to Timeline after every render. It's cheaper than regenerating later — if you might want to chain clips, save them as you go instead of hunting for them in History.
  • Pronunciation dictionary is permanent (PRO). Add word → IPA mappings that auto-apply across all generations, up to 500 per account. Much cleaner than wrapping every instance of 'Camarath' in <phoneme> tags.
  • Know your caps. Studio per render: free 1,000 · PRO 10,000 · Creator 25,000 characters. The homepage box allows 5,000 per generation on the free plan. Free accounts get 5,000 a day and 15,000 a month.
  • Captions come free with word timings. When a render has word timings, the .srt download appears under the player — no extra step. Export video (PRO) uses the same timings for its animated captions.

Hard caps

#Plan limits at a glance

Verified against the live service on 11 September 2026. Previews of any voice or emotion never count.

LimitFreePRO ($19/mo)Creator ($39/mo)
Per render in Studio (Single voice)1,00010,00025,000
Per generation on the homepage & API5,00010,00025,000
Characters per day5,000 signed in 2,000 as a guest——
Characters per month15,0001,000,0005,000,000
Standard & Multilingual voices✓✓✓
HD & DragonHD voicesDaily taste eligible countries✓✓
Signature voices (30 per language)Preview only✓✓
Ultra voices (English, Arabic)Preview only✓ single voice✓ single voice
Emotions (express-as)Preview only✓✓
Dialogue mode1 scene / day 4 lines, 500 chars✓ 50 lines per render✓ 50 lines per render
Takes (regenerate one line)—✓✓
Timeline (up to 100 clips)—✓✓
Captions (.srt)✓✓✓
WAV / OGG exportMP3 only✓✓
Export video (MP4)—✓✓
Music bed with ducking—✓✓
Saved projects—✓✓
Generation history—30 days30 days
Pronunciation dictionary—✓ 500 entries✓ 500 entries
Audiobook batch & M4B (dashboard)——✓
Commercial licence—✓✓
API keys11010

Per-render and daily figures count spoken characters, not SSML markup. Ultra has its own monthly allowance inside PRO, shown in Studio when it applies.

FAQ

#Common questions about Studio

Which mode should I start in?

Single voice. It has the gentlest learning curve — type text, pick a voice, hit Generate. Most users only need Single voice to find their preferred voice and pacing. Move on to Dialogue when you need multiple characters, and add renders to the Timeline when you want to chain several clips into one file. Studio per-render caps: free 1,000 · PRO 10,000 · Creator 25,000 characters.

Does Single voice go through my monthly character budget?

Yes. Every Generate in Single voice counts. Voice previews (the small play button next to a voice in the picker) do NOT count — those use a separate quota-free endpoint. Test as many voices as you want via Preview without burning chars. Free accounts have 5,000 characters a day and 15,000 a month; PRO has 1,000,000 a month, Creator 5,000,000.

Dialogue made one MP3 with all speakers. Can I get the per-speaker files separately?

Not directly. Dialogue muxes the conversation server-side and returns one merged file. If you need separate files per speaker, generate each speaker's lines in Single voice (switching voices between them), then add each one to the Timeline for chaining. It's slower but gives you full control over the individual audio files.

What does the Timeline do?

The Timeline is a PRO overlay that assembles audio you already generated. After any Single voice or Dialogue render, click 'Add to Timeline' below the player. Reorder clips, set a pause after each (0–3 s), preview any clip, and merge up to 100 clips into one continuous MP3. Great for multi-scene projects: intro narration → dialogue scene 1 → narrator break → dialogue scene 2 → outro.

Can I fix one bad line without re-rendering everything?

Yes — that's regenerate-one-line, a PRO feature. In Dialogue (and in Single voice after you switch on 'Regenerate by paragraph'), hover any line and hit ↻ to render just that line as a fresh take; the rest of the track is untouched and it re-stitches automatically. Step between takes with ‹ › — it's non-destructive, so a worse take is one tap to undo.

Can I use SSML markup in Single voice too?

Yes — Single voice and Dialogue both accept SSML on standard, multilingual and HD voices. The most useful tags are <break time="1s"/>, <emphasis level="strong">, <prosody rate="slow">, and <mstts:express-as style="cheerful">. Select text and right-click (or long-press on touch) to insert them, or type them directly. All tags are lowercase — <Break/> with a capital B won't work. Signature voices honour <break> only; Ultra voices read plain text.

What's the difference between HD voices and DragonHD voices?

Standard and multilingual voices take explicit direction: you control delivery with mstts:express-as style tags, and the Studio chip picker shows exactly which styles your selected voice supports. DragonHD voices (like en-US-Andrew:DragonHDLatestNeural) are our newest-generation Azure tier and automatically read emotion from the text itself. You write "She gasped in horror" and DragonHD delivers it dramatically without you needing the terrified style. The catch: DragonHD ignores mstts:express-as tags and prosody rate/pitch entirely, so manual delivery control is gone — write the emotion into the text and use <break> for pacing. Use a standard or multilingual voice when you want explicit control. Use DragonHD when you want natural delivery from natural writing.

Why does my style tag get ignored?

Five common reasons. First, your voice doesn't support that style — only certain voices have specific styles (the Studio chip picker shows what your selected voice supports). Second, you're on a DragonHD, Signature or Ultra voice: DragonHD reads emotion from context, and Signature and Ultra voices read plain text, so all three drop explicit style tags. Third, the wrapped span is too short — emotions need at least 3 words to be reliably audible because the voice needs time to transition into and out of the style; single-word emotion wraps often render as neutral. Wrap a longer phrase, or raise styledegree. Fourth, the wrap crosses a sentence boundary — every period, exclamation, or question mark resets the voice's prosody, so a single emotion wrap that spans two sentences usually only renders the style on one of them. Wrap each sentence separately. Fifth, you used the style outside an mstts:express-as wrapper. The Studio's emotion picker wraps the SSML envelope for you, so this only happens if you're writing raw SSML for the API.

Can I use SSML in the homepage box (not Studio)?

Yes. The homepage /api/tts endpoint and the browser extension both build a full SSML envelope from your text. Drop in <break/>, <emphasis>, <prosody>, even mstts:express-as — they all work on the homepage if the underlying voice supports them. The homepage allows up to 5,000 characters per generation on the free plan (Studio's own free cap is 1,000 per render).

What's styledegree and when should I change it?

styledegree controls how intense the emotional style is. Default is 1.0. Range is 0.01 (barely-there hint) to 2.0 (cranked up). For most narration the default works fine. Bump to 1.5 or 2.0 for dramatic moments (a battle cry, a grief explosion). Drop to 0.5 for subtle inflection (a hint of sadness under a brave face). Add it as an attribute: <mstts:express-as style="sad" styledegree="0.5">.

How do I make a pause longer than 10 seconds?

The <break time> tag caps at 10 seconds. For longer silences, stack multiple breaks: <break time="10s"/><break time="10s"/><break time="5s"/> gives you 25 seconds. Or use mstts:silence at the top of your document to set a global between-sentence silence if you just want spacious pacing throughout. Between clips in the Timeline the pause is set per clip, up to 3 seconds each.

What happens if my SSML is invalid?

Studio returns a clear error in the UI telling you what's wrong — usually with the line position — so you can fix it and re-render. A quick sanity check: keep tags lowercase and balanced (every <emphasis> has a matching </emphasis>), and note that the right-click menu inserts well-formed tags for you, so you rarely have to hand-write them.

Can I generate audio in two languages in one file?

Yes, three ways. Easiest: pick a Multilingual voice (look for "Multilingual" in the voice name) which can speak ~12 languages with one identity — these are free voices. Most precise: use <lang xml:lang="fr-FR">phrase</lang> inside any Azure voice to switch languages mid-sentence for that span. Most flexible: use Dialogue mode and assign each language to a different voice — you get separate native speakers for each language.

Why does the same voice sound slightly different on different generations?

Neural voices have a small amount of natural variation — the same text rendered twice won't be byte-identical even with the same voice and settings. Variation is normally subtle (different breath placement, slight rhythm differences). Larger differences usually come from context — the voice reads a question differently from a statement, an exclamation differently from a declaration. This is a feature, not a bug, and is part of what makes neural voices feel human.

What's the Pronunciation Dictionary?

A PRO/Creator feature that stores custom word → pronunciation mappings that auto-apply to ALL your generations. Add 'Camarath' → 'kəˈmɛrəθ' once, and every time you write Camarath in any Studio mode it gets pronounced correctly. Much cleaner than wrapping every instance in <phoneme>. Find it in the dashboard sidebar under Studio settings. Uses International Phonetic Alphabet (IPA) — there's a 'Hear it' button to audition before saving. Up to 500 entries per account.

Can I set per-segment voices in Dialogue mode without writing SSML?

Yes — that's exactly what Dialogue is for. Write your script as `Speaker: line of dialogue` (one per row). Assign each speaker a voice in the picker. Dialogue mode auto-generates the multi-voice SSML for you. Per-line emotion is a small dropdown next to each line, and each line has its own pause. You only need to drop into raw SSML if you want effects beyond voice + style (custom prosody, say-as, etc.).

What's the audio quality difference between MP3 / WAV / OGG?

MP3 (default, all plans) is 24 kHz mono at 96 kbps — universal and small, about 0.7 MB per minute. WAV (PRO) is uncompressed 16-bit PCM at 48 kHz mono, about 5.8 MB per minute — the pick when the file goes into a DAW or gets re-encoded. OGG (PRO) is Opus at 48 kHz — modern, small, but older devices don't all play it. Pick MP3 unless you have a specific reason to use the others.

Can I use Signature or Ultra voices in Studio?

Yes, on PRO and Creator. Signature voices (30 per language across 53 languages) work in Single voice and Dialogue; they read plain text and honour <break> pauses, other SSML tags are stripped. Ultra voices (21 English voices plus Syrian, Palestinian, Egyptian, Gulf and Modern Standard Arabic) work in Single voice only and read plain text. Everyone can preview both tiers for free from the picker; generating with them needs PRO.

Why can't I pick an Ultra voice in Dialogue mode?

Ultra is a single-voice tier. Dialogue mode refuses Ultra voices per line so a multi-speaker scene never mixes tiers unpredictably. Render the Ultra part in Single voice and chain it with the Dialogue scene in the Timeline instead.

Can I turn a Studio render into a video?

Yes, with Export video (PRO). Any Single voice or Dialogue render becomes a caption-animated MP4 in 9:16 (Reels, Shorts), 1:1 or 16:9 (YouTube), with a wave, bars or circle visualiser and three caption styles: Bold pop, Highlight and Clean. The captions come from the render's word timings, so they stay in sync without editing.