Text to Speech
Listen to any text read aloud using the voices built into your browser and operating system, with no download required.
Transcribe English speech from an audio file or your microphone using a Whisper model that runs on your own device.
This tool converts spoken English into written text. You can upload a recording such as a voice memo, a lecture clip or an interview, or you can press record and speak directly into your microphone. A few moments later you get a plain transcript you can copy, edit and save.
Unlike most transcription services, the audio is not sent to a remote server. The speech recognition model is downloaded to your browser and does all of its work locally, which means recordings of meetings, personal notes or private conversations stay on your computer.
The transcription is performed by Whisper tiny (English), an open speech recognition model released by OpenAI, in the converted version Xenova/whisper-tiny.en. It runs through Transformers.js, a JavaScript library that executes machine learning models in the browser. Your audio is resampled to 16 kHz and turned into a spectrogram, a picture of which frequencies are present over time. The model's encoder reads that picture and its decoder writes out the most likely sequence of words, a short section at a time.
No. This version uses the English-only Whisper tiny model, which is tuned for English and will produce poor output for other languages.
No. Audio from a file or your microphone is processed in memory inside this tab and is discarded when you close or reload the page.
Your browser must be given microphone permission, and most browsers only allow it on secure HTTPS pages. Check the permission icon in the address bar and make sure no other app is holding the microphone.
Listen to any text read aloud using the voices built into your browser and operating system, with no download required.
Cut the subject out of a photo and save it as a transparent PNG, processed entirely on your own device.
Turn screenshots, scanned pages and photos of printed text into editable text without sending the image anywhere.
Find common objects in a photo and draw labeled boxes around them, using a detection model that runs locally.