🔒 Your files stay on your device — core audio processing runs locally. Privacy details →
🎧
Free · In-browser · No upload

Speech to Text.

Transcribe live from your microphone, or run Whisper AI fully on-device to turn audio files and recordings into text. Timestamps on every line, TXT and SRT export.

Before-and-after illustration of Speech to Text: Speech recording becomes Timed transcript through Recognize words.
Before → process → afterSee the signal story, then hear it in the workspace.

VOICE DETAIL MAP

Transcript Timeline

Keep speech intelligible while changing only the target detail. The voice path makes phrases, breaths, consonants and tonal character visible as separate decisions instead of treating dialogue as one uniform waveform.

Before-and-after illustration of Speech to Text: Speech recording becomes Timed transcript through Recognize words.
Speech recordingRecognize wordsTimed transcript
A speech-aware signal path: local voice events are identified, processed within bounded spans and returned to a continuous phrase.
Open the full signal breakdownBest source, processing logic, limits and three-pass listening check
  1. Start with the right material

    Draft transcripts, subtitle timing and turning clear speech into editable text.

  2. Understand the transformation

    Live mode uses the browser or operating-system speech service; file mode downloads a Whisper model and runs inference in the browser.

  3. Verify the usable result

    Editable text plus TXT or timestamped SRT export when timing data is available.

Word timestamps

Where this workflow stops

Accuracy depends on language, noise, speaker overlap and device speed. Live speech may be processed by the browser vendor; choose local Whisper when that distinction matters.

THREE-PASS CHECK

Make the final decision by ear

  • Listen to complete phrases so edits retain natural timing.
  • Protect consonants and breaths that carry intelligibility or emotion.
  • Compare on both headphones and ordinary speakers before delivery.

How to convert speech to text online

  1. Choose your language from the selector so the recognizer knows what to listen for.
  2. Press start and speak — the transcript builds live, with a timestamp on every line. Press again to stop.
  3. Have a recording instead? Use the on-device Whisper AI transcription to turn audio files into text without uploading them.
  4. Export the finished transcript as plain TXT or as SRT subtitles with timecodes.

This tool has two distinct data paths. Local Whisper downloads a model, then transcribes a selected audio file on your device. Live microphone dictation uses the Web Speech service exposed by the browser or operating system and may send audio to that vendor. Accuracy, language support and availability vary; review timestamps and names before using the transcript or SRT.

Speech to Text: quick answer and technical limits

Quick answer: A transcription page with two modes: browser live dictation and optional local Whisper transcription for uploaded audio.

Best for
Draft transcripts, subtitle timing and turning clear speech into editable text.
How it works
Live mode uses the browser or operating-system speech service; file mode downloads a Whisper model and runs inference in the browser.
What you get
Editable text plus TXT or timestamped SRT export when timing data is available.

Know before you use it: Accuracy depends on language, noise, speaker overlap and device speed. Live speech may be processed by the browser vendor; choose local Whisper when that distinction matters.

Privacy: Whisper mode downloads model files from third-party hosting and processes the selected audio locally. Live mode can send microphone audio to the browser or OS speech provider. Privacy details →

FAQ

How do I transcribe speech to text for free?
Pick a language and use live browser speech recognition, or choose local Whisper transcription for a recording. Both are free in AudioWrench, but live recognition may use a browser/OS provider and follow its limits; Whisper needs a model download and local device resources.
Is my voice sent to a server for transcription?
Whisper mode processes the selected audio locally and does not upload it to AudioWrench. Live mic mode uses the browser's built-in speech recognition, which in some browsers can send speech to the browser or operating-system vendor's recognition service.
Can I export subtitles for my videos?
Yes — alongside plain TXT, the transcript exports as an SRT subtitle file with proper sequence numbers and timecodes, ready to load into video editors, YouTube or any player that supports subtitles.
How accurate is the transcription?
With a decent microphone, clear speech and little background noise, accuracy is high for everyday language. Heavy accents, crosstalk, technical jargon and noisy rooms lower it — a quick pass with the noise remover before transcribing a file often helps.