Lifetime Deal: $199 one-time — PRO forever33d 12h 30m 44s86/100claimed·14 leftClaim your spot →
Updated August 2026 . No microphone . Free to try
Turn your study notes into a narrated review course
You already wrote the hard part. Somewhere on your drive there is a folder of lessons, question banks, and notes that took months to get right. The reason it is still sitting there as text is that turning it into audio sounds like a project: a microphone, a quiet room, retakes every time you fumble a drug name, and a full re record whenever a guideline changes. None of that is required any more. Here is the whole workflow, start to finish, the way people actually running audio courses do it.
Published by the FreeTTS editorial team . Honest guidance, real limits quoted
One voice across every lessonNo microphone, no studioFix one line, regenerate one fileFree to try, no signup
An audio review course is a series of narrated lessons, one topic per file, read by a single consistent voice and published in syllabus order.
How to make one
Rewrite your notes as spoken sentences, split them one topic per lesson, pick one voice, generate each lesson in the FreeTTS studio, and download the MP3 files.
Equipment needed
None. No microphone, no quiet room, no retakes. Changing a lesson later means regenerating one file in seconds.
Length per lesson
Five to twelve minutes works best. A script of about 8,000 characters produces roughly nine minutes of audio.
Cost
Free to test at 15,000 characters a month. PRO at 19 dollars a month includes 1,000,000 characters, about 23 hours of audio. Creator at 39 dollars a month includes 5,000,000 characters and 25,000 characters per generation.
If you sell it
You need a commercial licence, which the paid plans include. The free plan is personal use only and adds a short watermark tag.
The one thing people get wrong
Changing voices partway through. The same narrator on every lesson is what makes separate files feel like one course.
Why it works
Why review material suits audio better than almost anything else
Not all content works as audio. A reference table does not. A diagram does not. But review material, the kind you go through again and again before an exam, is close to the ideal case, and it is worth understanding why before you spend a weekend producing any of it.
Review works through repetition. You are not learning a concept for the first time, you are reinforcing something you have already met, and reinforcement improves with exposure. Audio is the only format a learner can consume while doing something else. That is not a small advantage, it is the whole advantage. A student who reads your notes gets through them when they sit down with them. A student who listens gets through them on the commute, at the gym, folding laundry, and in the twenty minutes before a shift. You have not made the material better, you have multiplied the number of moments it can reach them.
Exam preparation adds a second fit: the question and answer format is naturally spoken. A practice question read aloud, a pause, then the answer and the reasoning, is close to how a good tutor actually teaches. It works in audio because it was always a spoken pattern that happened to be written down.
The honest limits are worth stating too. Anything visual belongs somewhere else. If your material leans on tables, images, or long lists of values, audio will be a supplement to your written material rather than a replacement for it. Most successful audio courses in this space publish both and let the learner choose the format that fits the moment.
The how-to
The workflow, in six steps
Written lessons to published audio, without a recording setup anywhere in the loop.
1
Rewrite the notes as something spoken
This is the step people skip, and it is the one that decides whether the result sounds like a course or like a machine reading a slide deck. Bullet fragments do not work out loud. Neither do abbreviations a reader would silently expand. Turn the fragments into full sentences, spell out each abbreviation the first time it appears, and read one paragraph aloud to yourself. Anything you stumble over, a listener will stumble over too. Cut it.
2
Split it one topic per lesson
One file, one subject, numbered in the order of your syllabus. Learners navigate by topic and they replay single topics constantly, which is exactly the behaviour review material should encourage. It also limits your maintenance: when a guideline changes you regenerate one lesson instead of re cutting a long recording and trying to match your own voice from three months ago.
3
Pick one house voice and lock it
Audition two or three voices against a real paragraph of your own material, never the demo sentence, because a demo line is written to flatter the voice. Choose the one you could listen to for a full hour without irritation, then use it for every lesson in the series. There are 400 or so voices to browse, which makes this feel like a big decision. It is, but only once. Make it on lesson one and never revisit it.
4
Run a pronunciation test before you commit
Write one short paragraph containing the ten technical terms that appear most in your material and generate it. Listen once. Anything that comes out wrong gets either a phonetic respelling in your script or an HD voice for that lesson. This takes about two minutes and it is the difference between finding a problem now and finding it after you have produced forty lessons.
5
Generate the lesson and listen through once
Paste the script, generate, and listen at normal speed from start to finish. Not skimming. You are listening for two things: lines that read awkwardly out loud, and places where the pace runs on without a breath. Fix them in the script and regenerate. On the free plan you can paste up to 5,000 characters at a time, which is roughly a five minute lesson, so longer lessons either move to a paid plan or get split.
6
Batch the rest and publish
Once one lesson sounds right, the rest is production rather than decision making. Generate them back to back, download the files, and publish. Most people who run these courses write in batches too: several scripts in one sitting, then all the audio in one pass, then a scheduled release. That rhythm is what makes a weekly publishing schedule survive a full time job.
The one big decision
Your house voice matters more than any other setting
If you take one thing from this article, take this. The single most important production decision in an audio course is not the microphone you did not buy, the speed setting, or the platform you publish on. It is picking one voice and never changing it.
Consistency is what turns forty separate MP3 files into something a learner experiences as a course. When the narrator changes between lesson eleven and lesson twelve, the listener notices immediately, and what they notice is the production rather than the content. Nobody articulates it as a voice change. They just feel that something is off, and a course that feels careless is a course people stop finishing.
For English teaching material, Aria is the voice most educators settle on. It stays even over long stretches, which matters far more than sounding impressive for ten seconds. Guy and Emma are sensible alternates if Aria is not right for your material. Slow the pace slightly below default if you are explaining something dense, since the natural instinct of a listener studying a hard topic is to want a beat more time between ideas.
One practical note on quality tiers. Standard voices are free to use and are what most of a course should be produced in. HD voices are a paid feature and are worth reserving for the places they earn their keep, such as an introduction lesson or a chapter with unusually difficult terminology. Producing an entire course in HD is rarely necessary and burns through your character allowance faster.
Be honest about this
Technical terms, drug names, and the ten minute check
Here is the part most articles about synthetic narration quietly skip. Standard voices do not get every technical term right. In medical material specifically, longer drug names are where it shows. Words like metoprolol, furosemide, and carbamazepine are the ones to check first, because they are common enough to appear repeatedly and unusual enough to trip a general purpose voice.
This is very fixable, and it is fixable in minutes, but only if you check before you produce a whole series rather than after. The test is simple. Write one paragraph containing the ten terms that appear most often in your material. Generate it. Listen once. You will know immediately which words are a problem, and in most cases it is a much shorter list than people fear.
For anything that comes out wrong you have two fixes. The first is phonetic respelling: write the word the way it sounds rather than the way it is spelled, so metoprolol becomes meh toe pro lol in the script. Your listener hears the correct word and never sees the spelling. The second is switching that lesson to an HD voice, which handles unusual words noticeably better. Between the two, almost everything can be made to sound right.
What you should not do is assume it works and publish forty lessons without listening. Mispronounced clinical terminology is not a cosmetic problem in a course people are using to prepare for a professional licensing exam. Test first.
Real numbers
What narrating a full course actually costs
Synthetic narration is priced per character rather than per minute, which is unfamiliar at first but works in your favour for this kind of content. Some arithmetic makes it concrete. A script of about 8,000 characters produces roughly nine minutes of audio. If you publish forty lessons in a month at that length, you are using in the region of 320,000 characters. That number is the one to hold on to while reading the table below.
Plan
Per generation
Per month
Commercial use
Realistic fit
Free
5,000 characters
15,000 characters
No, personal use only, watermark tag
Testing the format on a lesson or two
PRO, 19 dollars a month
10,000 characters
1,000,000 characters, about 23 hours of audio
Yes, watermark removed
A weekly publishing schedule with room to spare
Creator, 39 dollars a month
25,000 characters
5,000,000 characters
Yes, watermark removed
A full course produced in batches
Limits quoted from the FreeTTS pricing page at the time of writing. The free plan also has a daily cap of 5,000 characters.
The practical read on that table: a forty lesson month fits comfortably inside PRO, using roughly a third of its allowance. You move to Creator not because you have run out of characters but because you want longer single generations, since 25,000 characters lets a full lesson go through in one pass rather than being split. If you are producing a course in concentrated bursts rather than weekly, that matters more than the monthly ceiling.
Do not skip this
The licence question every course seller has to answer
The moment you charge for a course, put it behind a paywall, or run advertising against it, you are using the audio commercially, and you need a plan that grants a commercial licence. On FreeTTS the free plan is explicitly for personal, non commercial use and adds a short watermark tag to the end of the audio. The PRO and Creator plans include a commercial licence and produce clean audio with no tag.
The mistake worth avoiding is sequence. People produce the whole course on the free plan, then upgrade when they are ready to sell, and discover that every file they have made carries a watermark and was generated under a personal use licence. Regenerating forty lessons is not difficult, but it is an hour you did not need to spend. Switch before the first paid lesson, not after.
If you are comparing providers on this point, the terms vary more than you would expect, and some restrict commercial use even on paid tiers. We keep a licensing matrix comparing commercial rights across 15 TTS providers if you want to check the fine print before committing to a platform for a course you intend to sell for years.
In practice
How one nurse educator runs this
A working nurse educator who uses FreeTTS is building a review course for practical nursing licensing candidates, and the way she does it is a good model because it is unglamorous. She writes the lessons herself, in sequence, the same way she would teach them: a short recap of what the last lesson covered, a statement of what this one adds, then the material, then practice questions with the reasoning spelled out after each answer.
Everything goes through one voice. Every lesson, without exception. She does not use HD voices for most of the material and she does not switch narrators for variety, because variety is precisely what she does not want. The result is a few hundred review lessons that sound like one continuous course, made without a microphone, a studio, or a single retake.
The part worth copying is the batching. Scripts get written when there is time to write, and audio gets generated in passes rather than one lesson at a time. That separation is what makes the schedule survivable alongside actual clinical work. It is also why the production volume looks intimidating from the outside and is not: no individual day involves very much work, the days simply add up.
Distribution
Where to put the finished lessons
You have MP3 files. There are three sensible homes for them and most educators end up using more than one.
A podcast feed is the lowest friction option for the learner. They already have a player, episodes download for offline listening, and the format matches how review material actually gets consumed. It is also the easiest place to publish free sample lessons that lead people toward a paid product.
A video channel reaches people who search rather than subscribe. A static slide with the lesson title behind the audio is enough. This is the same faceless publishing model used across plenty of other niches, and if that route interests you our guide on starting a faceless channel with an AI voice covers the channel side in detail.
A course platform is where you charge. Audio sits alongside your written material, and the sequence you already built maps directly onto modules. If your material started life as PDFs rather than documents, converting a PDF to audio covers that path specifically.
For the broader case of instructional design with narrated audio, our write up on text to speech for e learning goes deeper on structuring modules than we can here.
Avoid these
The mistakes that make a narrated course sound cheap
Reading your notes verbatim. Written notes are compressed for the eye. Fragments, abbreviations, and dense clauses are fine on a page and painful in audio. If you change nothing else, change this.
Switching voices partway through. Covered above, and it is the most common one. Pick once, then live with it.
Publishing without listening. Generating audio is so fast that it is tempting to skip the listen through. One pass at normal speed catches awkward phrasing and mispronounced terms, and it takes exactly as long as the lesson.
Making lessons too long. A 30 minute file feels efficient to produce and is hard to use. Learners lose their place and do not come back. Split by topic.
Leaving no pause after practice questions. If you read a question and answer it immediately, the listener never gets the chance to think, which removes the entire value of a practice question.
Producing the whole thing on a free plan you intend to sell from. Check the licence before, not after.
Try it on one lesson
Take a lesson you have already written, paste it in, and listen to the first two minutes. That is the whole test, and it costs nothing.
Paste your notes into a text to speech tool, pick one voice, and download the MP3. On FreeTTS you can paste up to 5,000 characters per generation on the free plan, which is roughly a five minute lesson, then download the file with no signup. The only real preparation is cleaning up your notes first so they read like something a person would say out loud, because bullet fragments and abbreviations sound wrong when spoken.
What is an audio review course?▼
An audio review course is a series of narrated lessons that a learner listens to instead of reading, usually organised in the same order as a syllabus. Each lesson is a short script covering one topic, read by a single consistent voice, published as MP3 files on a podcast feed, a video channel, or inside a paid course. It suits exam preparation in particular because review material is repetitive by design, and repetition is what audio does well.
Do I need a microphone or a recording studio?▼
No. That is the entire point of narrating with a synthetic voice. There is no room tone to match, no retakes, and no background noise. When you fix a typo in lesson 22 you regenerate that one file in a few seconds rather than setting up a microphone again and trying to match the tone of a recording you made three weeks ago.
How long does one lesson take to produce?▼
Once your script is written, the audio itself takes seconds. A script of about 8,000 characters produces roughly nine minutes of audio. The realistic bottleneck is writing and proofreading, not generating. Most people who do this at volume write several lessons in one sitting, then generate all of them back to back.
Will an AI voice pronounce medical and drug names correctly?▼
Most read fine, some do not. Names like metoprolol, furosemide, and carbamazepine are the ones to check, and the fix takes seconds. Generate a short test clip with your ten most-used terms, listen once, and for anything that comes out wrong either respell it phonetically in the script (writing meh toe pro lol instead of metoprolol) or switch that lesson to an HD voice, which handles unusual words better. Do this test before you record a whole series, not after.
Can I sell a course narrated with an AI voice?▼
Yes, provided your plan grants a commercial licence. On FreeTTS the free plan is for personal, non commercial use and adds a short watermark tag to the audio. The PRO and Creator plans include a commercial licence and remove the watermark. If you charge for the course, put it behind a paywall, or run ads against it, move to a paid plan before you publish the first lesson rather than after.
Which voice should I use for teaching material?▼
Pick a warm, even, unhurried voice and then never change it. For English review material Aria is the one most educators settle on because it stays steady over long stretches without sounding theatrical. Guy and Emma are reasonable alternates. What matters more than the specific voice is consistency: the same narrator on every lesson is what makes a collection of files feel like a course.
How much does it cost to narrate a whole course?▼
Less than most people expect, because the cost is per character and not per minute. The free plan gives you 15,000 characters a month, which is enough to test the format on a couple of lessons. PRO at 19 dollars a month includes 1,000,000 characters, roughly 23 hours of audio, which covers a weekly publishing schedule with room left over. Creator at 39 dollars a month includes 5,000,000 characters and allows 25,000 characters in a single generation, which suits a full length course produced in batches.
Should each lesson be one file or many?▼
One file per lesson, matching one topic. Learners navigate by lesson, so a file that covers exactly one subject is easier to replay, and replaying is how review material actually gets used. It also limits the damage when content changes: update one guideline and you regenerate one file instead of re cutting a 40 minute recording.
How long should each audio lesson be?▼
Somewhere between five and twelve minutes works for most review material. That is long enough to cover a topic properly and short enough to finish on a commute or during a gym session. If a topic genuinely needs more, split it into part one and part two rather than publishing a 30 minute file, because a learner who loses their place in a long file often does not come back to it.
Can I add pauses between a question and its answer?▼
Yes, and for exam style material you should. A short pause after the question gives the listener a moment to answer in their head before you tell them, which is the whole point of practice questions. Write the pause into the script as a line break or a short beat marker, generate, and listen once to check the timing feels natural rather than rushed.
Where do I publish an audio course once it is made?▼
Three common routes. A podcast feed is the lowest friction because learners already have a player and it downloads for offline listening. A video channel with a static slide behind the audio reaches people searching on YouTube. A paid course platform is where you charge for it. Many educators do all three, publishing early lessons publicly as a sample and keeping the full sequence behind a paywall.
Is it better to record my own voice instead?▼
If your personality is the product, record yourself. If the material is the product, narrate it. Review content is usually the second kind. Nobody enrols in an exam preparation course because of the narrator, they enrol because the content is well organised, and a synthetic voice lets you fix and expand that content forever without booking studio time each time a guideline changes.
What is the fastest way to test whether this works for my material?▼
Take one lesson you have already written, paste it into the studio, pick a voice, and listen to the first two minutes. You will know almost immediately whether the format suits your material, and it costs nothing to find out. Most of the doubts people have about synthetic narration disappear or get confirmed within those two minutes.