Text to Speech

Paste English text, choose a voice and style, and press Speak. Download the result as an MP3 or WAV file.

Voices

Style

Speech is created on your own device, so your text is never uploaded. The AI voice files download once from Hugging Face's public file servers and are then kept in your browser.

How to use the text to speech tool

Type or paste up to 5,000 characters of English text and press Speak. Longer texts are split into sentences and read one after another, so the audio starts after the first part is ready instead of waiting for the whole text. When every part has been created, Download MP3 and Download WAV save exactly what you heard, including the style and effects you chose.

The first time you use the AI voices, your browser downloads the voice model, about 90 MB. It's saved in your browser afterwards, so every later visit starts much faster. If you'd rather not wait, switch to Device voices, which use the voices already built into your computer or phone and start instantly.

Voices, styles and effects

The AI voices include American and British English speakers, both female and male. Heart and Bella are the most natural; Nicole is soft and calm, which suits relaxation and ASMR-style audio; Fenrir has a deeper male voice. Then pick a style, or set the sliders yourself:

StyleWhat it does
Slow and clear / FastChanges the speaking speed without changing the pitch. Slow speech is useful for language learners and announcements.
Loud / SoftLoud compresses and boosts the voice so it cuts through background noise without distorting. Soft is quieter and slightly slower.
Deep voice / High voiceLowers or raises the pitch while keeping the same speed.
RobotA metallic, ring-modulated sound for games, videos and fun messages.
Phone call / Radio / MegaphoneFilters the voice to sound like a phone line, a walkie-talkie or a loudhailer.
Echo / StadiumAdds repeating echoes or the reverb of a large hall.
Monster / ChipmunkExtreme low or high pitch for comedy and characters.

Moving any slider switches to a custom setting, so you can start from a style and fine-tune it. Speed goes from half to double, pitch up or down by a full octave, and loudness up to 200 percent.

How it works, and why it's private

Most online text to speech services send your text to a server and send audio back. This tool doesn't. The AI voices come from Kokoro, an open-source speech model released under the Apache 2.0 license, which runs directly inside your browser. The text you enter never leaves your device, and the tool keeps working without an internet connection once the model has been downloaded.

Because the speech is generated on your device, speed depends on your computer or phone. A recent laptop creates speech faster than it can be played; an older phone may take a few seconds per sentence. On computers with a modern graphics card, ticking Use my graphics card makes generation much faster after a larger one-time download. The AI voices work best in Chrome and Edge. If they can't load in your browser, the device voices are a good fallback.

Tips for natural-sounding speech

  • Use punctuation. Commas and full stops create natural pauses. Break long sentences into shorter ones.
  • Write numbers and abbreviations the way they're said when it matters, for example "twenty twenty-six" or "Doctor Khan", to control the pronunciation.
  • Check the length first. The Speaking Time Calculator tells you how long a script will take to say, and the Text Cleaner removes stray symbols and formatting that can trip up the voice.

Using the audio you create

You can use the audio in videos, presentations, e-learning, podcasts and social media posts, and it's free. The Kokoro model's Apache 2.0 license allows commercial use. Don't use any synthetic voice to impersonate a real person or to mislead listeners, and consider mentioning that a voiceover was generated when that matters to your audience.

Frequently asked questions

Is this text to speech tool free?

Yes, completely free with no sign-up, no daily limit and no watermark. You can download as many MP3 and WAV files as you like.

Is my text uploaded anywhere?

No. The speech is created on your own device. Only the voice model itself is downloaded, once, from Hugging Face's public file servers; your text is never sent to any server.

Why does the first use take a while?

The AI voice model, about 90 MB, downloads the first time and is then saved in your browser. After that, speech starts much more quickly. Device voices need no download at all.

Can I download the audio as MP3?

Yes. After the AI voices finish creating the speech, press Download MP3 or Download WAV. The file includes the voice style, loudness and any effect you chose.

Which languages are supported?

This tool is designed for English, with American and British accents. Text in other languages may be pronounced incorrectly.

Why can't I download audio with device voices?

Browsers play device voices directly through the speakers and don't give websites access to the audio, so it can't be saved. Switch to the AI voices to download.

Last updated