Speech to Text
Transcribe English speech from an audio file or your microphone using a Whisper model that runs on your own device.
Listen to any text read aloud using the voices built into your browser and operating system, with no download required.
Reading on a screen for a long time is tiring, and sometimes it is easier to listen. Paste an article, a draft or a set of notes into this tool, choose a voice and a speaking speed, and your device reads it aloud. It is also a simple way to proofread your own writing, since awkward sentences are easier to hear than to see.
Unlike the other tools in the lab, this one has no AI model to download. It uses speech capabilities that already come with your browser and operating system, so it starts instantly.
The tool uses the Web Speech API, a standard browser feature, specifically its speechSynthesis interface. When you press play, the page passes your text and settings to the browser, which hands it to a speech engine provided by your operating system or browser vendor. That engine converts the words into sounds and plays them through your speakers. The voices are those installed on your device, which is why the list differs between Windows, macOS, Android, iOS and Chrome OS. Some browsers also offer online voices that are generated by the browser vendor's own service.
The voices come from your operating system and browser, not from this site. You can usually add more voices in your system's language or accessibility settings.
No. The browser speech feature does not provide audio files to web pages, so playback is live only.
Built-in local voices process text on your device. Some browsers, such as Chrome, also list online voices; if you choose one of those, the browser vendor may process the text on its servers.
Transcribe English speech from an audio file or your microphone using a Whisper model that runs on your own device.
Cut the subject out of a photo and save it as a transparent PNG, processed entirely on your own device.
Turn screenshots, scanned pages and photos of printed text into editable text without sending the image anywhere.
Find common objects in a photo and draw labeled boxes around them, using a detection model that runs locally.