Studio · Help
FreeTTS Studio: how the pieces fit together
Two modes — Single voice and Dialogue — plus takes to fix a single line, a Timeline to chain scenes, SSML to direct the delivery, and export to MP3, WAV, captions or a captioned video. Most projects only need Single voice; the rest unlock when your project gets longer, has multiple characters, or needs scene-by-scene control.
FreeTTS Studio has two modes. Single voice turns one script into one MP3 with any of 2,250+ voices; Dialogue turns a Speaker: line script of up to 50 lines into one merged multi-voice MP3. PRO adds takes (regenerate any line), the Timeline (chain up to 100 clips with pauses), WAV and OGG export, a caption-animated video export, a music bed, saved projects, and the Signature and Ultra voice tiers. Free accounts render 1,000 characters per Studio render, 5,000 a day and 15,000 a month, and can preview every voice and every emotion at no cost.
Modes & panels
#What each mode actually does
Studio is one screen with two modes and one overlay. Everything you generate in either mode can be chained in the Timeline, and any line of a render can be re-done on its own.
When: one block of text, voice testing, quick clips, per-render speed, pitch and style.
The default mode. Type or paste your script, pick any voice, optionally add an emotion, speed or pitch, and Generate. Each render is one MP3. On PRO, Regenerate by paragraph splits the render into paragraph takes you can redo one at a time.
Output: one MP3 per render, with .srt captions when the render has word timings.
When: conversations between characters, multi-speaker scenes, anything with line-by-line voice changes.
Write Speaker: line, one line per row, give each speaker a voice, and Generate. The scene is muxed server-side into one merged MP3 with proper turn-taking. Each line has its own pause and its own optional emotion. Ultra voices are single-voice only and are not offered here.
Output: one merged MP3 with all speakers, with per-line takes on PRO.
When: chaining several Single voice or Dialogue renders into one continuous file with pauses between scenes.
An overlay, not a mode. After any render, click Add to Timeline under the player. Reorder clips, set the pause after each, preview any clip, and Merge. Intro narration → dialogue scene → narrator break → next scene → outro, in one file.
Output: one continuous MP3 with every clip and pause baked in.
First visit? Studio opens with a pre-rendered three-voice demo scene so you can hear a Dialogue render before typing anything. It costs nothing and is not counted anywhere.
Takes
#Regenerate one line, keep the rest
A render is a set of lines. Any line can be re-done as a fresh take without touching its neighbours; the track re-stitches itself. Takes are kept, so a worse one is a single tap back. This is a PRO feature in both modes.
- Dialogue: every line carries ↻ and ‹ › from the first render.
- Single voice: click Regenerate by paragraph once; Studio re-renders your script as paragraph takes and the same controls appear per paragraph.
- A take is a new synthesis of that line only — the same voice, emotion and speed unless you change them first — so it costs only that line's characters.
At a glance
#Which one for what
| Need | Use | Why |
|---|---|---|
| Test a voice on real text | Single voice | Per-render control: six speed presets, three pitch presets, per-span emotion, free previews of every voice. |
| One narrator, 50+ pages | Single voice + Timeline | Render in chunks under your per-render cap, then chain them in the Timeline (PRO). |
| Multiple characters speaking | Dialogue | Server-side mux into one MP3 with proper turn-taking and per-line pauses. |
| Fix one weak line without redoing the render | Takes ↻ (PRO) | Re-render just that line as a fresh take; it re-stitches automatically. |
| Chain clips with pauses | Timeline (PRO) | Reorder, set a 0–3 s pause after each clip, merge into one MP3. |
| Per-paragraph emotion (sad / cheerful / whispering) | Single voice / Dialogue | Select text, right-click → Emotion, or wrap in <mstts:express-as style="…">. Standard and multilingual voices only. |
| Custom pauses between paragraphs | Single voice / Dialogue | Right-click → Pause, or <break time="1.5s"/>. Works on every tier except Ultra. |
| A video for Reels, Shorts or YouTube | Export video (PRO) | Any render → caption-animated MP4 in 9:16, 1:1 or 16:9. |
| Music under the narration | Music bed (PRO) | Pick a bed; it is mixed under the voice with automatic ducking. |
| Programmatic generation from scripts | Developer API | See /developers. Free accounts get one API key, paid plans ten. |
Common workflows
#"I want to make…"
I have a single narrator and a 50-page chapter.
- Open Single voice mode and pick an HD voice (look for the ✨ filter — they read with the most natural cadence at length).
- Render your text in chunks that fit your per-render cap (free 1,000 · PRO 10,000 · Creator 25,000). Add <break time="1.5s"/> between paragraphs and <emphasis> on key terms via right-click.
- Click 'Add to Timeline' after each chunk (PRO).
- Open the Timeline overlay, set the pause after each clip, then Merge into one continuous MP3.
I have a story with 3 characters who all speak.
- Write each multi-character SCENE in Dialogue mode as Speaker: line rows (up to 50 lines per render).
- Assign each speaker a distinct voice — standard, multilingual, HD or Signature; Ultra stays in Single voice.
- Render the scene. Hover a weak line and hit ↻ for a fresh take. Click 'Add to Timeline' below the audio player.
- Write narrative bridging text in Single voice mode and add each to the Timeline too.
- Open the Timeline overlay, reorder clips, set pauses (e.g. 1.5 s between scenes), then Merge.
- Download the single merged MP3, or Export video for a captioned MP4.
I want emotional variation across paragraphs (sad scene then cheerful scene).
- In Single voice on a standard or multilingual voice, select the sad paragraph, right-click → Emotion → sad (or type <mstts:express-as style="sad">your text</mstts:express-as>).
- Select the cheerful paragraph and apply cheerful the same way.
- Available styles depend on the voice: the chip picker shows what your selected voice supports. Free accounts can preview every emotion; rendering with one is PRO.
- Tags must be lowercase. Generate — and if one paragraph lands wrong, use 'Regenerate by paragraph' to fix just that line.
I want to test 10 voices to find my favorite for a project.
- Open Single voice mode.
- Open the voice picker. Click the play button (▶) next to each voice — those previews don't count toward your monthly quota, and they work for Signature and Ultra too.
- Once you've narrowed to 2–3 candidates, paste a real paragraph from your project and Generate. Now you'll hear how each handles YOUR text at your target pace.
- Try the speed presets (0.5× to 2×) and pitch (Low / Normal / High) to taste, then 'Save project' (PRO) to reuse the whole setup later.
SSML quick reference
#The 4 tags that cover 95% of use cases
SSML is the markup inside the text box that controls pauses, emphasis, pitch and emotion. It works in Single voice and Dialogue on standard, multilingual and HD voices: select text and right-click (or long-press) to insert a tag, or type it. All tags are lowercase.
<break time="1s"/>Insert a pause. Use 2s, 500ms, up to 10s.
<emphasis level="moderate">word</emphasis>Stress a word or phrase. Levels: reduced / moderate / strong.
<prosody rate="slow" pitch="-2st">text</prosody>Rates: x-slow / slow / medium / fast / x-fast or a percentage. Pitch in semitones (+2st, -2st). DragonHD ignores this tag; use the speed presets there.
<mstts:express-as style="cheerful">text</mstts:express-as>Emotional style. Which styles exist depends on the voice — the chip picker shows them. Free accounts can preview any emotion; rendering with one is PRO.
Three things trip up almost everyone. Tag names are lowercase (<Break/> is spoken aloud as text). &, < and > must be escaped inside SSML (&, <, >) — Studio does it for you in plain-text mode but not once you add tags. DragonHD, Signature and Ultra voices drop <mstts:express-as>: they read emotion from the words.
SSML deep reference
#Everything else SSML can do
The cheatsheet above handles most projects. Below is the full surface FreeTTS supports, in categories, every example copyable. Tier rules: standard, multilingual and HD voices take everything; DragonHD ignores express-as and prosody rate/pitch; Signature honours <break> only; Ultra reads plain text.
Pauses & silence
Two ways to insert silence. <break> is the W3C standard and the one tag that also works on Signature voices. <mstts:silence> is an mstts extension that gives you finer placement control.
Prosody — rate, pitch, volume
All three attributes accept named keywords, relative percentages, or absolute values. Combine them in a single <prosody> tag. DragonHD voices ignore rate and pitch — use the Studio speed presets or natural writing there.
Pronunciation control
Three ways to fix mispronunciations. Phoneme is the most precise; sub is the easiest; the Pronunciation Dictionary (PRO) is permanent across all your generations.
Say-as — interpret strings literally
Forces the engine to read text as a specific kind of data. Without <say-as>, the engine guesses ("1234" might be read as "one thousand two hundred thirty-four" or "one two three four" depending on context).
Language switching mid-document
Use <lang> to read a foreign phrase in its native pronunciation without switching voices entirely. The voice has to support the target language for this to sound right.
Expressive styles (mstts)
<mstts:express-as> is the emotional style tag. Not every voice supports every style — see the complete catalog below. Combine with styledegree (0.01–2.0) and role (for certain Chinese voices). Free accounts can preview any emotion; baking one into a render is PRO.
Structure & markers
Optional but useful for long-form audiobooks, captions, and engines parsing your output.
Background audio (mstts)
Mix a background audio track (music, ambience, white noise) under the synthesized speech. Audio file must be publicly accessible HTTPS. In Studio, the built-in music bed (PRO) does this without any markup, with automatic ducking.
All expressive styles
#Every mstts:express-as style FreeTTS supports
Not every voice supports every style. The Studio's chip picker (in Single voice, and per line in Dialogue) shows what your selected voice supports. Aria, Jenny and the Chinese voices Xiaoxiao/Yunxi have the widest catalogs; British and most localized voices typically support just cheerful and sad. Everyone can preview a style; rendering with one is PRO.
Emotion
cheerfulBright, upbeat. Default pick for happy moments and positive announcements.sadSlower, lower pitch, downward inflection. Use for grief, regret, somber news.angryHarder consonants, elevated volume. Adversarial dialogue, frustration.excitedFaster, higher energy than cheerful. Big-reveal moments, action scenes.fearfulTrembling, slightly higher pitch, breathy. Suspense and horror.terrifiedExtreme of fearful — short bursts, rapid breaths, top of the register.hopefulWarm, slightly tentative, rising inflection. Bridging sad → cheerful.disgruntledAnnoyed but restrained. Sarcasm, minor complaints.embarrassedHushed, slightly halting. Apologies, awkward moments.seriousSteady, even-paced, low expression. News commentary, formal narration.calmSmooth, lower energy than serious. Meditation, instructions, ASMR adjacent.Voice quality
whisperingBreathy, very low volume. Intimate scenes, secrets, ASMR. Best paired with prosody volume='soft'.shoutingMaximum volume + clipped delivery. Battle scenes, distant calls. Use sparingly — listening fatigue is real.gentleSoft, mid-pitch, evenly paced. Children's storytelling, comforting.lyricalSlight musical inflection. Poetry, song-like delivery.Conversational
friendlyWarm and approachable. Default for tutorials, onboarding, support content.unfriendlyCold, dismissive. Antagonist dialogue, hostile NPC.empatheticSoft, slow, validating tone. Customer-service apologies, sensitive subjects.chatCasual, mid-energy. Podcast-style discussion, informal updates.assistantPolite, helpful, neutral. Built for AI assistants and voice UIs.customerserviceProfessional friendly. Phone-system-style helpful answers.Narration
narration-professionalClean audiobook-style narration. The default 'just read it well' choice for long-form.narration-relaxedLooser pacing than professional. Personal essays, memoir.documentary-narrationAuthoritative, evenly paced. Educational video voiceovers, science explainers.newscastGeneric newscaster tone. Use the more specific casual/formal variants if available on your voice.newscast-casualApproachable newscaster — feature segments, morning shows.newscast-formalTraditional newscaster gravitas — breaking news, formal reports.poetry-readingSlower pacing with deliberate emphasis on cadence. Verse, prose poetry.advertisement-upbeatPunchy, energetic commercial reading. Product launches, promos.sports-commentaryFaster pace, dynamic stress. Sports calls, live event narration.sports-commentary-excitedSports-commentary cranked up — game-winning moments.Pronunciation
#Fixing words the voice mispronounces
Four escalation paths — pick the smallest one that works for your case.
Quickest: rewrite phonetically
If a word is mispronounced once or twice, rewrite the spelling so the engine reads it correctly. Camarath → kuh-MARE-uth. Ugly in the transcript but fixes one-offs without any SSML, and it is the only method that also works on Signature and Ultra voices.
One-shot fix: <sub alias="…">
Replace what the engine sees with what you want it to say.
Dr. <sub alias="Watson">Watson</sub> said hello.Use this for abbreviations and initialisms the engine misreads. The transcript still shows Dr.; only the audio uses Doctor.
Precise: <phoneme alphabet="ipa">
Force exact pronunciation with International Phonetic Alphabet symbols.
<phoneme alphabet="ipa" ph="kəˈmɛrəθ">Camarath</phoneme>The ph attribute holds the IPA. The displayed text is preserved; only the audio uses your spelling. If you don't know IPA, the dashboard's Pronunciation Dictionary has a "Hear it" button that lets you audition different IPA strings until one sounds right.
Permanent: Pronunciation Dictionary PRO
Add Camarath → kəˈmɛrəθ once in the dashboard (Studio settings → Pronunciation). Every future generation across all modes uses it. Much cleaner than wrapping every instance in <phoneme>. Up to 500 entries per account.
IPA cheat sheet for common English sounds
| Sound | IPA | Example word |
|---|---|---|
| "a" in cat | æ | cæt |
| "a" in father | ɑ | fɑther |
| "e" in bed | ɛ | bɛd |
| "i" in machine | i | mishin |
| "i" in bit | ɪ | bɪt |
| "o" in note | oʊ | noʊt |
| "u" in moon | u | mun |
| schwa (the most common vowel) | ə | əbout |
| "th" in thin | θ | θin |
| "th" in this | ð | ðis |
| "sh" in shoe | ʃ | ʃoo |
| "zh" in measure | ʒ | meaʒure |
| Primary stress (placed before the stressed syllable) | ˈ | cˈamera |
Voice tiers
#Standard vs Multilingual vs HD vs DragonHD vs Signature vs Ultra
One picker, six tiers. Knowing the tier matters because they differ in what SSML they honour, what a free account can do with them, and what they are best at. Every tier can be previewed for free from the picker's ▶ button.
Standard
FreeThe everyday neural voices — 2,250+ across 149 languages, fast, expressive on the styles they support. Every language page's free voices are standard voices.
Multilingual
FreeOne voice identity that speaks about 12 languages — look for "Multilingual" in the name (Andrew, Ava, Brian, Emma, Jenny…). Your French phrase does not suddenly become someone else.
HD
PRO · free daily tasteHigher-fidelity Azure voices with noticeably more natural prosody. Look for ":DragonHDLatestNeural" or "HD" in the name. Visitors in HD-eligible countries get a small free daily taste; the language pages lead with the flagship HD voice where one exists.
DragonHD
PRO · free daily tasteThe newest Azure generation. Reads context and expresses emotion from the text itself: write "she whispered" and it whispers. The ":DragonHDLatestNeural" suffix marks it.
Signature
PRO · preview freeThirty character voices per language — Nova, Maya, Celeste, Atlas, Felix, Theo and 24 more — in 53 languages, with the same persona set everywhere so a Croatian project and a Hungarian one can share a cast. Every language page shows six of them with a free preview.
Ultra
PRO · preview freeOur most lifelike line: 21 English voices and five Arabic dialect voices (Syrian, Palestinian, Egyptian, Gulf, Modern Standard). Previews are static samples you can play any time from the English and Arabic pages or the picker.
Where the free daily HD taste applies. Visitors in HD-eligible countries can render a short HD or DragonHD sample each day without a plan; the language pages lead with that voice (Ava HD on English, Seraphina HD on German, Αθηνά HD on Greek…). Elsewhere the same card is badged PRO. Signature and Ultra never have a free render, only free previews.
Voice recommendations
#Best voices by use case
Tested-and-recommended picks for common projects. Preview each in the voice picker before committing — taste is individual.
Long-form audiobook (single narrator)
en-US-Andrew:DragonHDLatestNeural
Top pick. DragonHD reads emotion from context — exactly what you want for fiction.en-US-AvaMultilingualNeural
Use for mixed-language books. Same voice identity across English, Spanish, French, etc. Free.en-US-JennyNeural
Solid default neural. Wide expressive style support if you want manual control.en-GB-LibbyNeural
British narration for UK-set or period-set fiction.
News, current events, factual content
en-US-AriaNeural with newscast-formal
Authority + clarity. Used by many news automations.en-US-BrandonNeural
Male newscaster cadence.en-US-DavisNeural
Mid-energy, factual.
Podcast / conversational
en-US-GuyNeural with style='chat'
Casual, mid-energy. Sounds like a real podcaster.en-US-JennyMultilingualNeural
Approachable warm female voice. Multilingual flexibility.en-US-Davis:DragonHDLatestNeural
DragonHD male — natural inflection without style hacking.
Educational / explainer
en-US-AriaNeural with documentary-narration
Standard for science explainers and tutorials.en-US-Emma:DragonHDLatestNeural
DragonHD female. Clean, even-paced explainer voice.
Ads, trailers, hero narration
Duke (Ultra)
Deep, cinematic. The voice the homepage auditions.Scarlett (Ultra)
Velvety British read for premium brands.Atlas (Signature, any language)
Informative and clear — the same persona in 53 languages for a multilingual campaign.
Multi-voice dialogue scenes
Mix any two contrasting voices in Dialogue mode
Pick voices with distinctly different pitches and accents — a young female + an older male reads more clearly than two similar voices.Add 'role' attribute for Chinese voices
zh-CN-XiaomoNeural, zh-CN-YunxiNeural, etc. can switch between Boy/Girl/YoungAdult/OlderAdult roles for variety from one voice.
Export
#Export: audio, captions, video, music
Every render downloads as MP3. PRO unlocks WAV and OGG, a caption-animated video, and a music bed under the narration. Captions are free whenever the render has word timings.
Audio
MP3 on every plan. WAV (48 kHz 16-bit) and OGG (Opus) on PRO and Creator. The Timeline merge always produces MP3.
Captions (.srt)
When a render carries word timings, a .srt download appears next to the audio. Long scripts are rendered in chunks and their timings are stitched, so captions stay aligned across the whole file. Drop the file into Premiere, DaVinci Resolve or CapCut.
Export video PRO
Turns the current render into a caption-animated MP4: 9:16 for Reels and Shorts, 1:1 for square posts, 16:9 for YouTube; a wave, bars or circle visualiser on a curated backdrop; and three caption styles — Bold pop, Highlight and Clean. The captions are driven by the same word timings, so they never drift.
Music bed PRO
Pick one of the curated beds and set its level; Studio mixes it under your narration with automatic ducking, so the music dips while the voice speaks and swells between lines. No markup, no audio editor.
Output formats
#MP3 vs WAV vs OGG — which to pick
| Format | Plan | Specs | Size | When to use |
|---|---|---|---|---|
| MP3 | All plans | 96 kbps mono · 24 kHz | ≈ 0.7 MB per minute | Default. Universal playback, small files, good quality. Lossy. Re-encoding a final mix costs a little each time. |
| WAV | PRO | 16-bit PCM mono · 48 kHz | ≈ 5.8 MB per minute | Editing in Audition, Pro Tools, Reaper; mastering; re-encoding to anything else without loss. Uncompressed. Use this if the file goes into a DAW. |
| OGG | PRO | Opus mono · 48 kHz | ≈ 0.7 MB per minute | Web playback and game engines that prefer Opus. Smallest files at equivalent quality. Open codec. Modern browsers all support it; some older devices do not. |
Gotchas
#Common mistakes and fixes
Uppercase SSML tag names
SSML is case-sensitive. Capitalized tags don't parse and get spoken aloud as text — or fail the whole chunk.
Unescaped & in text
Inside SSML, & starts an XML entity. The fix is & (or & outside any SSML block in plain-text mode). Same goes for < (use <) and > (use >) inside SSML.
<mstts:express-as> on a DragonHD, Signature or Ultra voice
DragonHD drops mstts:express-as (and prosody rate/pitch); Signature and Ultra strip markup before synthesis. Switch to a standard or multilingual voice when you need explicit style control.
Using a style the voice doesn't support
Unsupported styles are silently ignored — the audio plays neutral instead of poetic. Use the Studio's chip picker to see which styles your selected voice supports.
Forgetting the mstts namespace
If you're writing raw SSML for the API, the <speak> root must declare the mstts namespace before any mstts: tag will parse. The Studio wraps your text automatically — this only matters for direct API users.
Putting mstts:backgroundaudio inside <voice>
Background audio is per-document, not per-voice. Place it directly inside <speak>, before the <voice> block. Only one allowed per document. In Studio, the music bed does this for you.
Expecting <audio src='...'/> to work
Arbitrary inline <audio src> clips aren't supported by the synthesis path; use one underlay track, or chain separate clips in the Timeline.
Picking an Ultra voice for a Dialogue line
Ultra is a single-voice tier; Dialogue refuses it per line so a scene never mixes tiers unpredictably. The Timeline gives you the same result in two clips.
Pro tips
#Power-user shortcuts
- Voice previews are free. The play button next to each voice in the picker uses a separate quota-free endpoint — Signature and Ultra included. Audition 50 voices before committing to one; none of it counts toward your monthly chars.
- HD voices cost the same as standard in char usage, but synthesis is slower. For interactive Studio work that's invisible; on a long render it adds a little time. Usually still worth it.
- Regenerate one line, not the whole thing (PRO). Hover a line (in Dialogue, or in Single voice after 'Regenerate by paragraph') and hit ↻ for a fresh take of just that line — it re-stitches automatically, and you can step between takes with ‹ ›.
- Save your project (PRO). 'Save project' keeps a named, server-saved copy you can reopen on any device — pick up exactly where you left off.
- For very long narration, render in chunks under your per-render cap and chain them in the Timeline (PRO). Set the pause after each, then merge to one continuous MP3.
- Add to Timeline after every render. It's cheaper than regenerating later — if you might want to chain clips, save them as you go instead of hunting for them in History.
- Pronunciation dictionary is permanent (PRO). Add word → IPA mappings that auto-apply across all generations, up to 500 per account. Much cleaner than wrapping every instance of 'Camarath' in <phoneme> tags.
- Know your caps. Studio per render: free 1,000 · PRO 10,000 · Creator 25,000 characters. The homepage box allows 5,000 per generation on the free plan. Free accounts get 5,000 a day and 15,000 a month.
- Captions come free with word timings. When a render has word timings, the .srt download appears under the player — no extra step. Export video (PRO) uses the same timings for its animated captions.
Hard caps
#Plan limits at a glance
Verified against the live service on 11 September 2026. Previews of any voice or emotion never count.
| Limit | Free | PRO ($19/mo) | Creator ($39/mo) |
|---|---|---|---|
| Per render in Studio (Single voice) | 1,000 | 10,000 | 25,000 |
| Per generation on the homepage & API | 5,000 | 10,000 | 25,000 |
| Characters per day | 5,000 signed in 2,000 as a guest | — | — |
| Characters per month | 15,000 | 1,000,000 | 5,000,000 |
| Standard & Multilingual voices | ✓ | ✓ | ✓ |
| HD & DragonHD voices | Daily taste eligible countries | ✓ | ✓ |
| Signature voices (30 per language) | Preview only | ✓ | ✓ |
| Ultra voices (English, Arabic) | Preview only | ✓ single voice | ✓ single voice |
| Emotions (express-as) | Preview only | ✓ | ✓ |
| Dialogue mode | 1 scene / day 4 lines, 500 chars | ✓ 50 lines per render | ✓ 50 lines per render |
| Takes (regenerate one line) | — | ✓ | ✓ |
| Timeline (up to 100 clips) | — | ✓ | ✓ |
| Captions (.srt) | ✓ | ✓ | ✓ |
| WAV / OGG export | MP3 only | ✓ | ✓ |
| Export video (MP4) | — | ✓ | ✓ |
| Music bed with ducking | — | ✓ | ✓ |
| Saved projects | — | ✓ | ✓ |
| Generation history | — | 30 days | 30 days |
| Pronunciation dictionary | — | ✓ 500 entries | ✓ 500 entries |
| Audiobook batch & M4B (dashboard) | — | — | ✓ |
| Commercial licence | — | ✓ | ✓ |
| API keys | 1 | 10 | 10 |
Per-render and daily figures count spoken characters, not SSML markup. Ultra has its own monthly allowance inside PRO, shown in Studio when it applies.
FAQ
#Common questions about Studio
Which mode should I start in?
Single voice. It has the gentlest learning curve — type text, pick a voice, hit Generate. Most users only need Single voice to find their preferred voice and pacing. Move on to Dialogue when you need multiple characters, and add renders to the Timeline when you want to chain several clips into one file. Studio per-render caps: free 1,000 · PRO 10,000 · Creator 25,000 characters.
Does Single voice go through my monthly character budget?
Yes. Every Generate in Single voice counts. Voice previews (the small play button next to a voice in the picker) do NOT count — those use a separate quota-free endpoint. Test as many voices as you want via Preview without burning chars. Free accounts have 5,000 characters a day and 15,000 a month; PRO has 1,000,000 a month, Creator 5,000,000.
Dialogue made one MP3 with all speakers. Can I get the per-speaker files separately?
Not directly. Dialogue muxes the conversation server-side and returns one merged file. If you need separate files per speaker, generate each speaker's lines in Single voice (switching voices between them), then add each one to the Timeline for chaining. It's slower but gives you full control over the individual audio files.
What does the Timeline do?
The Timeline is a PRO overlay that assembles audio you already generated. After any Single voice or Dialogue render, click 'Add to Timeline' below the player. Reorder clips, set a pause after each (0–3 s), preview any clip, and merge up to 100 clips into one continuous MP3. Great for multi-scene projects: intro narration → dialogue scene 1 → narrator break → dialogue scene 2 → outro.
Can I fix one bad line without re-rendering everything?
Yes — that's regenerate-one-line, a PRO feature. In Dialogue (and in Single voice after you switch on 'Regenerate by paragraph'), hover any line and hit ↻ to render just that line as a fresh take; the rest of the track is untouched and it re-stitches automatically. Step between takes with ‹ › — it's non-destructive, so a worse take is one tap to undo.
Can I use SSML markup in Single voice too?
Yes — Single voice and Dialogue both accept SSML on standard, multilingual and HD voices. The most useful tags are <break time="1s"/>, <emphasis level="strong">, <prosody rate="slow">, and <mstts:express-as style="cheerful">. Select text and right-click (or long-press on touch) to insert them, or type them directly. All tags are lowercase — <Break/> with a capital B won't work. Signature voices honour <break> only; Ultra voices read plain text.
What's the difference between HD voices and DragonHD voices?
Standard and multilingual voices take explicit direction: you control delivery with mstts:express-as style tags, and the Studio chip picker shows exactly which styles your selected voice supports. DragonHD voices (like en-US-Andrew:DragonHDLatestNeural) are our newest-generation Azure tier and automatically read emotion from the text itself. You write "She gasped in horror" and DragonHD delivers it dramatically without you needing the terrified style. The catch: DragonHD ignores mstts:express-as tags and prosody rate/pitch entirely, so manual delivery control is gone — write the emotion into the text and use <break> for pacing. Use a standard or multilingual voice when you want explicit control. Use DragonHD when you want natural delivery from natural writing.
Why does my style tag get ignored?
Five common reasons. First, your voice doesn't support that style — only certain voices have specific styles (the Studio chip picker shows what your selected voice supports). Second, you're on a DragonHD, Signature or Ultra voice: DragonHD reads emotion from context, and Signature and Ultra voices read plain text, so all three drop explicit style tags. Third, the wrapped span is too short — emotions need at least 3 words to be reliably audible because the voice needs time to transition into and out of the style; single-word emotion wraps often render as neutral. Wrap a longer phrase, or raise styledegree. Fourth, the wrap crosses a sentence boundary — every period, exclamation, or question mark resets the voice's prosody, so a single emotion wrap that spans two sentences usually only renders the style on one of them. Wrap each sentence separately. Fifth, you used the style outside an mstts:express-as wrapper. The Studio's emotion picker wraps the SSML envelope for you, so this only happens if you're writing raw SSML for the API.
Can I use SSML in the homepage box (not Studio)?
Yes. The homepage /api/tts endpoint and the browser extension both build a full SSML envelope from your text. Drop in <break/>, <emphasis>, <prosody>, even mstts:express-as — they all work on the homepage if the underlying voice supports them. The homepage allows up to 5,000 characters per generation on the free plan (Studio's own free cap is 1,000 per render).
What's styledegree and when should I change it?
styledegree controls how intense the emotional style is. Default is 1.0. Range is 0.01 (barely-there hint) to 2.0 (cranked up). For most narration the default works fine. Bump to 1.5 or 2.0 for dramatic moments (a battle cry, a grief explosion). Drop to 0.5 for subtle inflection (a hint of sadness under a brave face). Add it as an attribute: <mstts:express-as style="sad" styledegree="0.5">.
How do I make a pause longer than 10 seconds?
The <break time> tag caps at 10 seconds. For longer silences, stack multiple breaks: <break time="10s"/><break time="10s"/><break time="5s"/> gives you 25 seconds. Or use mstts:silence at the top of your document to set a global between-sentence silence if you just want spacious pacing throughout. Between clips in the Timeline the pause is set per clip, up to 3 seconds each.
What happens if my SSML is invalid?
Studio returns a clear error in the UI telling you what's wrong — usually with the line position — so you can fix it and re-render. A quick sanity check: keep tags lowercase and balanced (every <emphasis> has a matching </emphasis>), and note that the right-click menu inserts well-formed tags for you, so you rarely have to hand-write them.
Can I generate audio in two languages in one file?
Yes, three ways. Easiest: pick a Multilingual voice (look for "Multilingual" in the voice name) which can speak ~12 languages with one identity — these are free voices. Most precise: use <lang xml:lang="fr-FR">phrase</lang> inside any Azure voice to switch languages mid-sentence for that span. Most flexible: use Dialogue mode and assign each language to a different voice — you get separate native speakers for each language.
Why does the same voice sound slightly different on different generations?
Neural voices have a small amount of natural variation — the same text rendered twice won't be byte-identical even with the same voice and settings. Variation is normally subtle (different breath placement, slight rhythm differences). Larger differences usually come from context — the voice reads a question differently from a statement, an exclamation differently from a declaration. This is a feature, not a bug, and is part of what makes neural voices feel human.
What's the Pronunciation Dictionary?
A PRO/Creator feature that stores custom word → pronunciation mappings that auto-apply to ALL your generations. Add 'Camarath' → 'kəˈmɛrəθ' once, and every time you write Camarath in any Studio mode it gets pronounced correctly. Much cleaner than wrapping every instance in <phoneme>. Find it in the dashboard sidebar under Studio settings. Uses International Phonetic Alphabet (IPA) — there's a 'Hear it' button to audition before saving. Up to 500 entries per account.
Can I set per-segment voices in Dialogue mode without writing SSML?
Yes — that's exactly what Dialogue is for. Write your script as `Speaker: line of dialogue` (one per row). Assign each speaker a voice in the picker. Dialogue mode auto-generates the multi-voice SSML for you. Per-line emotion is a small dropdown next to each line, and each line has its own pause. You only need to drop into raw SSML if you want effects beyond voice + style (custom prosody, say-as, etc.).
What's the audio quality difference between MP3 / WAV / OGG?
MP3 (default, all plans) is 24 kHz mono at 96 kbps — universal and small, about 0.7 MB per minute. WAV (PRO) is uncompressed 16-bit PCM at 48 kHz mono, about 5.8 MB per minute — the pick when the file goes into a DAW or gets re-encoded. OGG (PRO) is Opus at 48 kHz — modern, small, but older devices don't all play it. Pick MP3 unless you have a specific reason to use the others.
Can I use Signature or Ultra voices in Studio?
Yes, on PRO and Creator. Signature voices (30 per language across 53 languages) work in Single voice and Dialogue; they read plain text and honour <break> pauses, other SSML tags are stripped. Ultra voices (21 English voices plus Syrian, Palestinian, Egyptian, Gulf and Modern Standard Arabic) work in Single voice only and read plain text. Everyone can preview both tiers for free from the picker; generating with them needs PRO.
Why can't I pick an Ultra voice in Dialogue mode?
Ultra is a single-voice tier. Dialogue mode refuses Ultra voices per line so a multi-speaker scene never mixes tiers unpredictably. Render the Ultra part in Single voice and chain it with the Dialogue scene in the Timeline instead.
Can I turn a Studio render into a video?
Yes, with Export video (PRO). Any Single voice or Dialogue render becomes a caption-animated MP4 in 9:16 (Reels, Shorts), 1:1 or 16:9 (YouTube), with a wave, bars or circle visualiser and three caption styles: Bold pop, Highlight and Clean. The captions come from the render's word timings, so they stay in sync without editing.