Read Aloud
Reads text aloud, switching voices when the language changes.
Free to use · No sign-up needed
0 characters
Ready
Language map
Segments
Segments appear here as you type.
Browsers will speak but will not hand over the audio, so there is no download button here that could work. Paste this command into a terminal instead: it asks Microsoft's voices for the same text and writes one MP3 on your own machine. Gleekit sends your text nowhere — the command runs from your machine, and it is the command that reaches Microsoft. The first time, run the install line as well: it is one of the # comments at the top, marked First time only.
One thing to know before you rely on it: edge-tts is not Microsoft's own program, and the service it calls is the one the Edge browser uses, which Microsoft does not document for anybody else. It can stop working without notice. The reading above does not depend on it — that runs on your browser's own voices.
# Type some text above and the command appears here.
Text that changes language mid-sentence
Most text-to-speech tools assume a page is written in one language. Real writing often is not. A project update names a Tokyo office, a message to colleagues switches to Chinese for a phrase there is no neat English for, a Korean line arrives in the middle of an English paragraph. Hand that to an ordinary reader and it either reads everything with one accent or gives up on the parts it does not recognise.
This tool reads the text first and works out where each language starts and stops. Each stretch then goes to a voice that actually speaks it, so the Chinese is read as Chinese and the English around it is read as English.
How to use it
- Paste your text, or press Load sample to hear four languages in three lines.
- Watch the language map and the chips under it. They redraw as you type, so you can see where the tool thinks each language begins before you hear a word of it.
- Press Read aloud. The marker tracks the reading and the chip being spoken is raised.
- Click any chip to start from there instead of from the top.
- Set a voice and a speed for each language separately in the Voices panel. Chinese read a little slower than English is often easier to follow than both at the same pace.
The hard part is the shared characters
Chinese and Japanese share a writing system. 会議室 is a perfectly good word in both, and nothing about the characters says which one is meant. Guessing wrong is not a small error: a Chinese voice reading Japanese does not sound accented, it sounds like a different sentence.
The rule this tool uses is the one a person would use. Japanese almost always mixes those characters with hiragana or katakana, so if a sentence contains either, the shared characters in it are Japanese. If the sentence has none but the paragraph around it does, it is still Japanese. Otherwise it is read as Chinese. That handles ordinary writing well and fails on a Japanese heading written entirely in kanji, which is why the setting to override it exists.
Where the voices come from
The voices belong to your browser, not to this site. That has three consequences worth knowing. The tool costs nothing to run, so it will not be taken away or put behind a sign-up. Choose a voice installed on the device and it keeps working with the network switched off. And what you hear depends on what you are using: Microsoft Edge offers neural voices that sound close to a person, Safari on a Mac offers a large local set, and a phone may have only a handful. If a language is missing entirely, the tool says so before you press play rather than falling silent halfway through.
A missing voice is worth fixing rather than working around, because it is usually a download away. Windows adds languages under Time & language, Android under the text-to-speech settings of its speech engine, and iOS and macOS under the spoken content settings in Accessibility. Add the language there, reload this page, and the new voice appears in the list.
Where your text goes
Gleekit is not in the path, whichever voice you use. There is no server behind this tool: the detection, the segmenting and the choice of voice are all code running in your browser, and none of that code sends what you pasted to us or to anybody else. We could not read your text if we wanted to, and there is no share link here that could carry it out by accident.
What your browser does next is worth knowing, because it differs by voice. A voice installed on the device is a program already on your computer, so the words never leave it and the reading works with the network switched off. A streamed voice is not installed at all: Edge's Online (Natural) voices belong to Microsoft and Chrome's Google voices belong to Google, and your browser sends your text to that company to be turned into sound, under its privacy terms rather than ours. Streamed voices are usually the best-sounding ones on a machine, so each language list puts them at the top and marks each one Streamed. If you would rather nothing left the device, pick one of the unmarked voices below them, and then switch the network off: a voice that still speaks is one running on your own computer, and that is a test no promise of ours can substitute for.
When you want a file rather than a playback
A browser will speak text but will not give you the audio, so there is no honest download button to put here. Instead the panel at the foot of the tool, below the segments, writes out a command you can run yourself. It uses edge-tts, a small open-source program, and it carries over what you set up above: the same split into languages, the same voices where they exist on Microsoft's service, the same speeds, pitch and volume. Running it leaves one MP3 on your machine, ready for a slide deck, a lesson, or a message to somebody who would rather listen than read.
None of that happens here. Gleekit never contacts Microsoft, or anybody else, on your behalf. The panel only writes a command out as text; through that route your text reaches Microsoft solely if you choose to run it, and it leaves from your own computer rather than from this page.
Good uses
- Checking that a bilingual announcement reads cleanly before you send it to the team.
- Hearing how a name, a place or a phrase sounds in a language you read but do not speak.
- Proofreading your own writing, because the ear catches repetition the eye skips over.
- Listening to a long article or a set of notes while your hands are busy with something else.
- Language practice, at a speed you choose separately for the language you are learning.
Questions people ask
- Why does it need to know Chinese from Japanese?
- Because they share characters. 漢字 is written the same way in both, and a Chinese voice reading Japanese sounds like nonsense rather than an accent. The tool looks for hiragana or katakana in the sentence, and failing that in the paragraph, to decide which one it is.
- It picked the wrong language for my text.
- Use the Han characters setting to force Chinese or Japanese. Automatic detection fails on a Japanese sentence written entirely in kanji with no kana anywhere nearby, which is uncommon in ordinary writing but does happen in headings and lists.
- Why are the voices different on my phone?
- The voices come from your browser, not from us. Windows, macOS, Android and iOS each ship their own set, and Edge adds Microsoft's neural voices on top. That is also why the tool costs nothing to run. Whether it works with the network switched off depends on which voice you pick: the ones installed on the device do, the streamed ones do not.
- Why is there no download button?
- Browsers will speak but will not hand over the audio, so a download button here could not work. The panel below the segments writes out a command for you to run instead. It uses edge-tts, a small open-source program you install once, and it asks Microsoft's voice service for the same text and saves one MP3 on your own machine. Gleekit never contacts Microsoft: the command does, from your computer, and only if you choose to run it.
- Does my text get sent anywhere?
- Never to us. There is no server behind this tool, and no share link that could carry your text, because we decided a link was not worth the risk of leaking what you pasted. Past that it is your browser's decision, and it turns on the voice: one installed on the device speaks without touching the network, while a streamed voice — Edge's Online (Natural) set, Chrome's Google voices — sends your text to Microsoft or Google to be spoken, under their terms rather than ours. Streamed voices sound best, so each list puts them at the top and marks each one Streamed; the rest are ones your browser reports as installed on the device. The surest test of any voice is to switch the network off: if it still speaks, it is running on your own machine.