Text to speech — natural voices, privately in your browser
Paste any text, pick a voice, and generate natural-sounding speech entirely on your device. Play it back and download a WAV. Nothing is uploaded — the model runs in your browser, so it's free, unlimited, and private.
Your text
0 characters · about 0:00
Formatting — add pauses & pronunciation
Type these anywhere in your text. They shape the speech but aren't read aloud.
Pause
Insert a real silence.
[pause] · [pause 1.5s] · [pause 800ms]Pronunciation
Respell a tricky name or acronym.
[Nguyen](win) · [GIF](jif) · [Siobhan](shivawn)Voice & speed
Speech
How to turn text into speech
-
Paste your text
Type or paste the text you want spoken. It never leaves your browser.
-
Pick a voice and speed
Choose from 30+ English voices (American and British, female and male) and set the speaking speed.
-
Generate & download
The voice model runs on your device (first run downloads it, then it's cached). Play the audio back and download it as a WAV file.
Why generate speech here
Your text is never uploaded
Speech is generated 100% in your browser — no server, no account, no upload. Ideal for scripts, drafts and anything confidential.
Natural, expressive voices
Powered by Kokoro, a compact but high-quality voice model, with 30+ English voices to choose from and adjustable speed.
Free & unlimited
Because synthesis happens on your device, there's no per-character cost and no cap. Generate as much speech as you like.
Questions
- Is my text uploaded to a server?
- No. The voice model is downloaded to your browser and runs on your own device — your text never leaves it.
- What do I get to download?
- A standard WAV audio file (mono, 24 kHz) that you can play anywhere or drop into a video editor.
- Which voices and languages are supported?
- 30+ English voices at launch — American and British, female and male — powered by the Kokoro-82M model. More languages are planned.
- Is it really free?
- Yes — free and unlimited. There's no account and no metering because the model runs on your device, not our servers.
- Why is the first run slower?
- The first run downloads the voice model once, then it's cached in your browser, so later generations start quickly.
- Does it need a fast computer?
- It uses your GPU (via WebGPU) when available and falls back to your CPU otherwise. The model is small, so it runs on phones too — just a little slower.