Skip to main content
Voice Studio turns any text into natural-sounding speech, entirely inside your browser. Powered by Kokoro.js, every word is synthesized on your device — no audio is sent to any server. Choose from a large library of voices spanning American, British, and European accents, dial in speed and quality, and download a finished MP3 in seconds.

Generating Audio

1

Enter your text

Type or paste the text you want to convert into speech in the main input area. There is no hard character limit, but shorter passages generate faster and are easier to preview.
2

Select a voice

Pick a voice from the dropdown. Each voice has a distinct accent, pitch, and style — see Available Voices below for the full naming guide.
3

Set speed and bitrate

Adjust the Speed slider to control how fast or slow the speech is delivered. Choose your target MP3 bitrate (128, 192, or 320 kbps) based on the quality you need.
4

Click Generate

Hit Generate and Voice Studio synthesizes the audio in your browser. A waveform preview appears when it’s ready.
5

Download your file

Click Download to save the MP3 to your device. The filename reflects your voice and speed selection for easy organization.

Available Voices

Voice Studio ships with a large library of voices. The naming convention tells you the speaker’s region and gender at a glance: Browse all available voices in the Voice dropdown inside the studio. Each entry plays a short preview so you can compare styles before committing to a full generation.

Settings

You can fine-tune Voice Studio behaviour from the controls in the studio and in ⚙️ Settings → TTS Model:

Voice & Speed

Select any voice from the library and adjust the speed multiplier. Values below 1.0 slow the speech down; values above 1.0 speed it up. Your last-used voice and speed are remembered between sessions.

MP3 Bitrate

Choose 128 kbps for small file sizes, 192 kbps for a balance of size and quality, or 320 kbps for the highest-fidelity output. Higher bitrates produce larger files.

TTS Quantization

Found in ⚙️ Settings → TTS Model. q4 generates audio faster using less memory. q8 takes a little longer but produces noticeably cleaner speech — ideal for final deliverables.

Audio Normalization

Audio normalization is on by default to keep volume levels consistent. Disable it in ⚙️ Settings → TTS Model if you need to preserve the original dynamic range.

Downloading Audio

Once generation completes, the audio player appears with playback controls. Click Download to save your MP3. The file is assembled entirely in your browser — no upload, no waiting for a server, no account needed.
Use q4 quantization for quick drafts and listen-throughs — it generates audio significantly faster. Switch to q8 in ⚙️ Settings → TTS Model before your final export for the cleanest, highest-quality result. You can toggle between the two at any time without re-entering your text.